Pipeline dumb resource repeated input detection method and system based on multi-feature analysis

By adopting multi-feature analysis and machine learning models in the dumb resource management of communication pipelines, the problem of repeated data entry is solved, accurate data judgment and automated processing are realized, and management efficiency and decision-making accuracy are improved.

CN120128501APending Publication Date: 2025-06-10INSPUR TIANYUAN COMM INFORMATION SYST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510253030.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

In the management of dumb resources of communication pipelines, manual repeated entry is frequent, resulting in serious data repeated entry, occupying a large number of storage resources, and causing chaos in subsequent data query, analysis and decision-making processes, affecting management efficiency and decision-making accuracy.

Method used

The pipeline duplicate input detection method based on multi-feature analysis is adopted. The data collection, preprocessing, extracting space, optical cable and connection features are used, and the machine learning model is trained using a random forest algorithm to determine whether the data is repeated input, and a probability threshold is set to determine repeated input.

Benefits of technology

Through multi-dimensional feature analysis and machine learning algorithms, accurate judgment of pipeline dumb resource data is achieved, the limitations of manual comparison are avoided, work efficiency is improved, and automated and efficient solutions are provided for communication pipeline dumb resource management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120128501A_ABST
    Figure CN120128501A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of communication, in particular to a pipeline dumb resource repeated entry detection method and system based on multi-feature analysis, and the method comprises the following steps: collecting data; the collected data is preprocessed; extracting spatial features, optical cable features and connection features; training a machine learning model; judging whether data input is repeated or not; the method has the beneficial effects that a complex mode can be learned from features of multiple dimensions by utilizing a machine learning algorithm, especially a random forest algorithm, and the limitation of manual comparison and simple rule judgment is avoided, so that whether pipeline dumb resource data is repeatedly input or not is more accurately judged.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communication technologies, and specifically to a method and system for detecting duplicate entry of pipeline dumb resources based on multi-feature analysis. Background Art

[0002] In the field of management of pipeline dumb resources in communication, the accuracy of data is the key to ensuring the efficient operation of the network and scientific decision-making. However, there are currently many pain points. For example, the phenomenon of manual duplicate entry is frequent. Due to factors such as human operation errors, complex system data interaction, and data entry by different personnel at different times, the duplicate entry of pipeline dumb resource data is serious. This not only occupies a large amount of storage resources, but also causes chaos in subsequent data query, analysis, and decision-making processes, greatly affecting the management efficiency and decision-making accuracy of the communication network. Summary of the Invention

[0003] The purpose of the present invention is to provide a method and system for detecting duplicate entry of pipeline dumb resources based on multi-feature analysis to solve the problems raised in the above background art.

[0004] To achieve the above purpose, the present invention provides the following technical solution: A method for detecting duplicate entry of pipeline dumb resources based on multi-feature analysis, the method comprising the following steps:

[0005] Data acquisition;

[0006] Preprocessing the acquired data;

[0007] Extracting spatial features, optical cable features, and connection features;

[0008] Machine learning model training;

[0009] Judging whether the data is entered repeatedly, inputting the data to be judged into the trained machine learning model, and the model judges according to the extracted features and outputs whether the data is repeatedly entered data; setting a probability threshold, and when the probability of repeated entry output by the model is greater than the threshold, it is determined as repeated entry.

[0010] Preferably, the specific operations of data acquisition include:

[0011] Extensively collecting communication pipeline dumb resource data, covering the location information of the pipeline: accurate to longitude and latitude coordinates, length, type: such as pipeline, overhead, connection relationship, and optical cable information in the manhole: including the number and specifications of optical cables.

[0012] Preferably, the specific operations of preprocessing the acquired data include:

[0013] Perform a comprehensive cleaning on the collected data, remove data with obvious errors, such as abnormal positions and negative lengths; perform denoising processing to eliminate interference factors in the data, and standardize data in different formats to ensure data consistency and processability.

[0014] Preferably, the specific operations for extracting spatial features, optical cable features, and connection features include:

[0015] Spatial features: Calculate the distance between manholes. If there are multiple manholes within a certain range and the distance is too close, extract this feature as one of the bases for possible duplicates. The specific calculation formula is where (x 1 , y 1 ) and (x 2 , y 2 ) are the plane coordinates corresponding to the longitude and latitude coordinates of two manholes respectively;

[0016] Optical cable features: Analyze the quantity and specification differences of optical cables in manholes. Define the similarity of optical cable features as where n = |n 1 - n 2 | represents the difference value of the number of optical cables in two manholes, g = |g 1 - g 2 | represents the difference value of optical cable specifications, α and β are weight coefficients, and ∈ is a minimum value to avoid a zero denominator;

[0017] Connection features: Determine the pipelines and equipment conditions connected to manholes. Define the similarity of connection features as S connection , if the types of connected pipelines and equipment are exactly the same, it is 1, and if they are partially the same, it is a value between 0 and 1, which is set according to specific circumstances.

[0018] Preferably, the specific operations for machine learning model training include:

[0019] Use the random forest algorithm for model training. The calculation of information gain is determined by comparing the change in information entropy of the dataset before and after splitting. Its prediction formula is where H(x) is the final prediction result of the random forest for the input data x, Y is the possible output category: duplicate entry or non - duplicate entry, T is the number of decision trees, h i (x) is the prediction result of the i - th decision tree for the input data x, and I is the indicator function. When h i (x) = Y, the value of I is 1, otherwise it is 0;

[0020] When constructing a decision tree, the spatial feature values, the similarity of optical cable features, and the similarity of connection features are used as the splitting basis, and the feature that maximizes the information gain is selected for splitting. The calculation formula for information gain is IG(T, a) = H(T) - H(T|a), where IG(T, a) represents the information gain when splitting by feature a, H(T) represents the information entropy of the data set, and H(T|a) represents the conditional information entropy of the data set T given the feature a. The calculation formula for information entropy is where C is the number of categories: two categories, repeated entry and non-repeated entry, and p(i) is the proportion of samples belonging to category i.

[0021] A pipeline dumb resource duplicate entry detection system based on multi-feature analysis is applied to a pipeline dumb resource duplicate entry detection method based on multi-feature analysis. The system includes:

[0022] A data acquisition module for acquiring communication pipeline dumb resource data;

[0023] A preprocessing module for preprocessing the acquired data;

[0024] A feature extraction module for extracting spatial features, optical cable features, and connection features;

[0025] A model training module for training a machine learning model;

[0026] A duplicate judgment module for judging whether the data is entered repeatedly. The data to be judged is input into the trained machine learning model, and the model makes a judgment based on the extracted features and outputs whether the data is repeatedly entered data; a probability threshold is set, and when the repeated entry probability output by the model is greater than this threshold, it is determined as a repeated entry.

[0027] Preferably, the data acquisition module widely collects communication pipeline dumb resource data, covering the location information of the pipeline: accurate to longitude and latitude coordinates, length, type: such as pipeline, overhead, connection relationship, and the optical cable information in the manhole: including the number and specifications of the optical cables.

[0028] Preferably, the preprocessing module comprehensively cleans the acquired data, removes obviously incorrect data, such as abnormal positions and negative lengths; performs denoising processing to eliminate interference factors in the data, and standardizes data in different formats to ensure the consistency and processability of the data.

[0029] Preferably, the feature extraction module includes: spatial features: calculate the distance between manholes. If multiple manholes appear within a certain range and the distance is too close, then extract this feature as one of the possible duplication bases. The specific calculation formula is where (x 1 , y 1 ) and (x 2, y 2 ) are the plane coordinates corresponding to the longitude and latitude coordinates of two pipe shafts respectively;

[0030] Optical cable characteristics: Analyze the quantity and specification differences of optical cables in the pipe shafts, and define the similarity of optical cable characteristics as where n = |n 1 - n 2 | represents the difference value of the number of optical cables in two pipe shafts, g = |g 1 - g 2 | represents the difference value of optical cable specifications, α and β are weight coefficients, and ∈ is a minimum value to avoid the denominator being zero;

[0031] Connection characteristics: Determine the pipelines and equipment connected by the pipe shafts, and define the similarity of connection characteristics as S connection , if the types of pipelines and equipment connected are exactly the same, it is 1, and if they are partially the same, it is a value between 0 and 1, which is set according to specific circumstances.

[0032] Preferably, the model training module uses the random forest algorithm for model training. The calculation of information gain is determined by comparing the change in information entropy of the data set before and after splitting. Its prediction formula is where H(x) is the final prediction result of the random forest for the input data x, Y is the possible output category: repeated entry or non-repeated entry, T is the number of decision trees, h i (x) is the prediction result of the i-th decision tree for the input data x, and I is the indicator function. When h i (x) = Y, the value of I is 1, otherwise it is 0;

[0033] When constructing the decision tree, use the spatial feature value, the similarity of optical cable characteristics, and the similarity of connection characteristics as the splitting basis, and select the feature that maximizes the information gain for splitting. The calculation formula of information gain is IG(T, a) = H(T) - H(T|a), where IG(T, a) represents the information gain when splitting with feature a, H(T) represents the information entropy of the data set, and H(T|a) represents the conditional information entropy of the data set T when the feature a is known. The calculation formula of information entropy is where c is the number of categories: two categories of repeated entry and non-repeated entry, and p(i) is the proportion of samples belonging to category i.

[0034] Compared with the prior art, the beneficial effects of the present invention are:

[0035] The method and system for detecting duplicate entry of pipeline silent resources based on multi-feature analysis proposed by the present invention can learn complex patterns from features in multiple dimensions by using machine learning algorithms, especially the random forest algorithm, avoiding the limitations of manual comparison and simple rule judgment, and thus more accurately determining whether the pipeline silent resource data is entered repeatedly. It can comprehensively analyze multi-dimensional features and overcome the limitations of single-feature judgment. At the same time, the machine learning algorithm can continuously self-optimize and adjust with the change of data, adapt to different scenarios, realize automated and efficient judgment, greatly improve work efficiency, and bring a new solution for the management of communication pipeline silent resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 It is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] In order to clearly and completely describe the objectives, technical solutions of the present invention and make the advantages more clear, the following further details the embodiments of the present invention with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are part of the embodiments of the present invention, rather than all of the embodiments, and are only used to explain the embodiments of the present invention, not to limit the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.

[0038] Embodiment 1, please refer to Figure 1 , the present invention provides a technical solution: a method for detecting duplicate entry of pipeline silent resources based on multi-feature analysis, the method comprising the following steps:

[0039] (1) Data collection

[0040] Widely collect communication pipeline silent resource data, covering the location information of the pipeline (accurate to longitude and latitude coordinates), length, type (such as pipeline, overhead, etc.), connection relationship, and the optical cable information in the manhole (including the number and specifications of the optical cables).

[0041] (2) Data preprocessing

[0042] Comprehensively clean the collected data, removing obviously incorrect data, such as abnormal locations, negative lengths, etc.

[0043] Perform denoising processing to eliminate interference factors in the data.

[0044] Unify and standardize data in different formats to ensure the consistency and processability of the data.

[0045] (3) Feature extraction

[0046] Spatial feature: Calculate the distance between manholes. If there are multiple manholes within a certain range and the distance is too close, extract this feature as one of the bases for possible duplication. The specific calculation formula is where (x 1 , y 1 ) and (x 2 , y 2 ) are the plane coordinates corresponding to the longitude and latitude coordinates of two manholes respectively.

[0047] Optical cable feature: Analyze the quantity and specification differences of optical cables in manholes. Define the similarity of optical cable features as where n = |n 1 - n 2 | represents the difference value of the number of optical cables in two manholes, g = |g 1 - g 2 | represents the difference value of optical cable specifications, α and β are weight coefficients, and ∈ is a minimum value to avoid the denominator being zero.

[0048] Connection feature: Determine the pipelines and equipment connected to manholes. Define the similarity of connection features as S connection . If the types of connected pipelines and equipment are exactly the same, it is 1; if they are partially the same, it is a value between 0 and 1, which is set according to specific circumstances.

[0049] (4) Machine learning model training

[0050] Use the random forest algorithm for model training. The calculation of information gain is determined by comparing the change in information entropy of the data set before and after splitting. Its prediction formula is where H(x) is the final prediction result of the random forest for the input data x, Y is the possible output category (duplicate entry or non-duplicate entry), T is the number of decision trees, h i (x) is the prediction result of the i-th decision tree for the input data x, and I is the indicator function. When h i (x) = Y, the value of I is 1; otherwise, it is 0.

[0051] When constructing a decision tree, use the spatial feature value, the similarity of optical cable features, and the similarity of connection features as the splitting basis, and select the feature that can maximize the information gain for splitting. The calculation formula of information gain is IG(T, a) = H(T) - H(T|a), where IG(T, a) represents the information gain when splitting with feature a, H(T) represents the information entropy of the data set, and H(T|a) represents the conditional information entropy of the data set T given the feature a. The calculation formula of information entropy is where c is the number of categories (two categories: duplicate entry and non-duplicate entry), and p(i) is the proportion of samples belonging to category i.

[0052] (5) Duplication judgment

[0053] Input the data to be judged into the trained machine learning model. The model makes a judgment based on the extracted features and outputs whether the data is duplicate input data.

[0054] Set a probability threshold. When the duplicate input probability output by the model is greater than this threshold, it is determined as duplicate input.

[0055] Example 2: On the basis of Example 1, a pipeline dumb resource duplicate input detection system based on multi-feature analysis is proposed, which is applied to the pipeline dumb resource duplicate input detection method based on multi-feature analysis. The system includes:

[0056] A data acquisition module, which is used to collect communication pipeline dumb resource data, widely collect communication pipeline dumb resource data, covering the location information of the pipeline: accurate to longitude and latitude coordinates, length, type: such as pipeline, overhead, connection relationship, and the information of optical cables in the manhole: including the number and specifications of optical cables.

[0057] A preprocessing module, which preprocesses the collected data; comprehensively cleans the collected data, removes obviously incorrect data, such as abnormal positions and negative lengths; performs denoising processing to eliminate interference factors in the data, and standardizes different formats of data to ensure the consistency and processability of the data.

[0058] A feature extraction module, which extracts spatial features, optical cable features, and connection features; including: Spatial features: Calculate the distance between manholes. If there are multiple manholes within a certain range and the distance is too close, extract this feature as one of the possible duplicates. The specific calculation formula is where (x 1 , y 1 ) and (x 2 , y 2 ) are the plane coordinates corresponding to the longitude and latitude coordinates of two manholes respectively;

[0059] Optical cable features: Analyze the differences in the number and specifications of optical cables in the manhole. Define the optical cable feature similarity as where n = |n 1 - n 2 | represents the difference value of the number of optical cables in two manholes, g = |g 1 - g 2 | represents the difference value of optical cable specifications, α and β are weight coefficients, and ∈ is a minimum value to avoid the denominator being zero;

[0060] Connection features: Determine the pipelines and equipment connected to the manhole. Define the connection feature similarity as S connection . If the types of pipelines and equipment connected are exactly the same, it is 1. If they are partially the same, it is a value between 0 and 1, which is set according to specific circumstances.

[0061] Model training module, for training a machine learning model; the random forest algorithm is used for model training, and the calculation of information gain is determined by comparing the change in information entropy of the data set before and after splitting. Its prediction formula is where H(x) is the final prediction result of the random forest for the input data x, Y is the possible output category: duplicate entry or non-duplicate entry, T is the number of decision trees, and h i (x) is the prediction result of the i-th decision tree for the input data x, and I is the indicator function. When h i (x) = Y, the value of I is 1, otherwise it is 0;

[0062] When constructing the decision tree, the spatial feature value, the similarity of optical cable features, and the similarity of connection features are used as the splitting basis, and the feature that maximizes the information gain is selected for splitting. The calculation formula of information gain is IG(T, a) = H(T) - H(T|a), where IG(T, a) represents the information gain when splitting with feature a, H(T) represents the information entropy of the data set, and H(T|a) represents the conditional information entropy of the data set T given the feature a. The calculation formula of information entropy is where c is the number of categories: two categories of duplicate entry and non-duplicate entry, and p(i) is the proportion of samples belonging to category i.

[0063] Duplicate judgment module, for judging whether the data is entered repeatedly. The data to be judged is input into the trained machine learning model, and the model makes a judgment based on the extracted features and outputs whether the data is duplicate entry data; a probability threshold is set, and when the probability of duplicate entry output by the model is greater than this threshold, it is determined as duplicate entry.

[0064] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A pipeline dumb resource duplicate entry detection method based on multi-feature analysis, characterized by: The method comprises the following steps: Data collection; Preprocess the collected data; Extract spatial features, optical cable features, and connection features; Machine learning model training; To determine whether the data is entered repeatedly, the data to be judged is input into the trained machine learning model. The model makes a judgment based on the extracted features and outputs whether the data is duplicated. A probability threshold is set. When the probability of duplicate entry output by the model is greater than the threshold, it is determined to be a duplicate entry.

2. The pipeline dumb resource duplicate entry detection method based on multi-feature analysis according to claim 1 is characterized in that: The specific operations of data collection include: Extensive collection of communication pipeline dumb resource data, covering the location information of the pipeline: accurate to the longitude and latitude coordinates, length, type: such as pipeline, overhead, connection relationship and optical cable information in the pipe well: including the number and specifications of optical cables.

3. The pipeline dumb resource duplicate entry detection method based on multi-feature analysis according to claim 1 is characterized in that: The specific operations for preprocessing the collected data include: The collected data is cleaned comprehensively to remove obviously erroneous data, such as abnormal positions and negative lengths; denoising is performed to eliminate interference factors in the data, and data in different formats are standardized to ensure data consistency and processability.

4. The pipeline dumb resource duplicate entry detection method based on multi-feature analysis according to claim 1 is characterized in that: The specific operations for extracting spatial features, optical cable features, and connection features include: Spatial feature: Calculate the distance between pipe wells. If multiple pipe wells appear within a certain range and the distance is too close, extract this feature as one of the possible duplication bases. The specific calculation formula is: Where (x1, y1) and (x2, y2) are the plane coordinates corresponding to the longitude and latitude coordinates of the two tube wells respectively; Cable characteristics: Analyze the number and specification differences of the optical cables in the pipe well, and define the similarity of the optical cable characteristics as Where n = |n1-n2| represents the difference in the number of optical cables in the two shafts, g = |g1-g2| represents the difference in the specifications of the optical cables, α and β are weight coefficients, and ∈ is a minimum value to avoid the denominator being zero; Connection characteristics: Determine the pipelines and equipment connected to the pipe well, and define the connection feature similarity as S connection If the connected pipeline type and equipment are exactly the same, it is 1. If they are partially the same, it is a value between 0 and 1, which is set according to the specific situation.

5. The pipeline dumb resource duplicate entry detection method based on multi-feature analysis according to claim 1 is characterized in that: The specific operations of machine learning model training include: The random forest algorithm is used for model training. The information gain is calculated by comparing the changes in information entropy before and after the split of the data set. The prediction formula is: Where H(x) is the final prediction result of the random forest for the input data x, Y is the possible output category: duplicate or non-duplicate, T is the number of decision trees, and h i (x) is the prediction result of the i-th decision tree for the input data x, I is the indicator function, when h i When (x) = Y, the value of I is 1, otherwise it is 0; When constructing a decision tree, the spatial feature value, the similarity of the optical cable feature and the similarity of the connection feature are used as the basis for splitting, and the feature with the largest information gain is selected for splitting. The calculation formula of information gain is IG(T, a) = H(T) - H(T|a), where IG(T, a) represents the information gain when splitting with feature a, H(T) represents the information entropy of the data set, and H(T|a) represents the conditional information entropy of the data set T when feature a is known. The calculation formula of information entropy is Where c is the number of categories: duplicate entries and non-duplicate entries, and p(i) is the proportion of samples belonging to category i.

6. A pipeline dumb resource duplicate entry detection system based on multi-feature analysis, applied to the pipeline dumb resource duplicate entry detection method based on multi-feature analysis as described in any one of claims 1 to 5, characterized in that: The system comprises: Data acquisition module, used to collect communication pipeline dumb resource data; A preprocessing module preprocesses the collected data; Feature extraction module, extracting spatial features, optical cable features and connection features; Model training module, machine learning model training; The duplicate judgment module judges whether the data is entered repeatedly. The data to be judged is input into the trained machine learning model. The model judges based on the extracted features and outputs whether the data is duplicated. A probability threshold is set. When the duplicate entry probability output by the model is greater than the threshold, it is judged as a duplicate entry.

7. The pipeline dumb resource duplicate entry detection system based on multi-feature analysis according to claim 6 is characterized in that: The data acquisition module widely collects communication pipeline dumb resource data, covering the location information of the pipeline: accurate to the longitude and latitude coordinates, length, type: such as pipeline, overhead, connection relationship and optical cable information in the pipe well: including the number and specifications of the optical cables.

8. The pipeline dumb resource duplicate entry detection system based on multi-feature analysis according to claim 6 is characterized by: The preprocessing module performs a comprehensive cleaning of the collected data to remove data with obvious errors, such as abnormal positions and negative lengths; Perform denoising to eliminate interference factors in the data, standardize data in different formats, and ensure data consistency and processability.

9. The pipeline dumb resource duplicate entry detection system based on multi-feature analysis according to claim 6 is characterized in that: The feature extraction module includes: spatial feature: calculating the distance between pipe wells. If multiple pipe wells appear within a certain range and the distance is too close, the feature is extracted as one of the possible duplication bases. The specific calculation formula is: Where (x1, y1) and (x2, y2) are the plane coordinates corresponding to the longitude and latitude coordinates of the two tube wells respectively; Cable characteristics: Analyze the number and specification differences of the optical cables in the pipe well, and define the similarity of the optical cable characteristics as Where n = |n1-n2| represents the difference in the number of optical cables in the two shafts, g = |g1-g2| represents the difference in the specifications of the optical cables, α and β are weight coefficients, and ∈ is a minimum value to avoid the denominator being zero; Connection characteristics: Determine the pipelines and equipment connected to the pipe well, and define the connection feature similarity as S connection If the connected pipeline type and equipment are exactly the same, it is 1. If they are partially the same, it is a value between 0 and 1, which is set according to the specific situation.

10. The pipeline dumb resource duplicate entry detection system based on multi-feature analysis according to claim 6, characterized in that: The model training module uses the random forest algorithm for model training. The information gain is calculated by comparing the information entropy changes of the data set before and after the split. The prediction formula is: Where H(x) is the final prediction result of the random forest for the input data x, Y is the possible output category: duplicate or non-duplicate, T is the number of decision trees, and h i (x) is the prediction result of the i-th decision tree for the input data x, I is the indicator function, when h i When (x) = Y, the value of I is 1, otherwise it is 0; When constructing a decision tree, the spatial feature value, the similarity of the optical cable feature, and the similarity of the connection feature are used as the basis for splitting, and the feature with the largest information gain is selected for splitting. The calculation formula of information gain is IG(T,a)=H(T)-H(T|a), where IG(T,a) represents the information gain when splitting with feature a, H(T) represents the information entropy of the data set, and H(T|a) represents the conditional information entropy of the data set T when feature a is known. The calculation formula of information entropy is Where c is the number of categories: duplicate entries and non-duplicate entries, and p(i) is the proportion of samples belonging to category i.