Multi-source big data fusion access method and system for power distribution communication network

By identifying and classifying data sources in the distribution communication network and collecting and fusion data in real time, the complex problems of data dispersion and processing in the distribution communication network are solved, and data fusion with high accuracy and low noise is achieved, data silos are reduced, and decision-making is improved.

CN120067972APending Publication Date: 2025-05-30STATE GRID LIAONING ELECTRIC POWER CO LTD +4
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510054960.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Data acquisition in the distribution communication network relies on traditional sensors and equipment, resulting in data dispersion and lack of effective integration. Data processing is complex and has poor real-time performance, making it difficult to meet the needs of dynamic monitoring and management, and there is data silos, which affects the scientificity and timeliness of decision-making.

Method used

A method for fusion access of multi-source big data in the power distribution communication network is proposed. By identifying and classifying data sources, data is collected and converted into a unified format in real time, outliers are removed, and matrix is ​​formed, and data fusion is fused using Kalman filtering, matrix weighting, linear summing and random forest algorithms, and finally the fused data is transmitted to the external system.

Benefits of technology

It reduces the noise of temperature sensor data, improves the accuracy of data fusion, reduces the amount of calculation, reduces the data island phenomenon, and improves the scientificity and timeliness of decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067972A_ABST
    Figure CN120067972A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-source big data fusion access method for a power distribution communication network. The method comprises the following steps: identifying types of various data sources in the power distribution communication network and service types of data acquired by the data sources; various data sources are divided into noisy core data sources, noiseless core data sources, auxiliary data sources and non-core data sources according to types and business types of data collected by the data sources; converting the data collected at each moment into a uniform format, and forming a matrix by the data collected by the same type of data sources; carrying out data fusion or not carrying out data fusion on a matrix formed by data collected by different types of data sources in manners of Kalman filtering and matrix weighting, linear addition of weighting based on the types of the data sources and a random forest algorithm; and transmitting the fused data to an external system related to the power distribution communication network. According to the method, the calculation amount is reduced while the accuracy of data fusion is improved, and the phenomenon of data islands is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data fusion, and more specifically, relates to a method and system for multi-source big data fusion access in a distribution communication network. Background Art

[0002] The distribution communication network has the characteristics of wide distribution, numerous terminal nodes, rich and diverse bearing service types, mixed application of multiple access technologies, and harsh on-site operating environment, resulting in problems such as diverse technical systems, numerous equipment manufacturers, and complex networking structures in the distribution communication network, which are not convenient for management and operation and maintenance. This also increases the difficulty of communication monitoring and control in the distribution network, and at the same time brings many management loopholes and defects to operation and maintenance management, fault handling, and service guarantee.

[0003] The data acquisition of the distribution communication network mainly relies on traditional sensors and devices. The data is scattered and lacks effective integration, which not only leads to the complexity of data processing caused by data source diversity, but also has poor real-time performance and is difficult to meet the requirements of dynamic monitoring and management. In addition, the phenomenon of data islands is serious, affecting the scientificity and timeliness of decision-making. Summary of the Invention

[0004] To solve the deficiencies in the prior art, the present invention provides a method and system for multi-source big data fusion access in a distribution communication network.

[0005] The present invention adopts the following technical solutions.

[0006] The first aspect of the present invention proposes a method for multi-source big data fusion access in a distribution communication network, which is characterized by including:

[0007] Identifying the types of various data sources in the distribution communication network and the service types of the data collected by the data sources;

[0008] Classifying various data sources into noisy core data sources, noiseless core data sources, auxiliary data sources, and non-core data sources according to the types and the service types of the data collected by the data sources;

[0009] Real-time collecting various data sources in the distribution communication network, converting the data collected at each moment into a unified format, and detecting and removing outliers; forming a matrix with the data collected by the same type of data sources;

[0010] For the matrices formed by the data collected by noisy core data sources, noiseless core data sources, and auxiliary data sources, data fusion is respectively performed by means of Kalman filtering combined with matrix weighting, linear summation with weighting based on the type of data source, and random forest algorithm. For the data collected by non-core data sources, no data fusion is performed;

[0011] Convert the fused data and the data collected from non-core data sources into a set data exchange format, and use the set data exchange standard to transmit the fused data to external systems related to the distribution communication network through the API.

[0012] Preferably, the types of various data sources in the distribution communication network and the business types of the data collected by the data sources are specifically as follows:

[0013] The business types include real-time monitoring data, sensor status monitoring data, user power consumption data, market transaction data, fault diagnosis data, maintenance record data, user feedback data, and environmental monitoring data;

[0014] The types include temperature sensors, humidity sensors, current sensors, monitoring systems, market research, external third parties, and public databases.

[0015] Preferably, the various data sources are classified into noisy core data sources, noiseless core data sources, auxiliary data sources, and non-core data sources according to the type and the business type of the data collected by the data source, specifically as follows:

[0016] If the business type is real-time monitoring data and the type is a temperature sensor, then the data source is a noisy core data source;

[0017] If the business type is real-time monitoring data and the type is a humidity sensor or a current sensor, then the data source is a noisy core data source;

[0018] If the business type is user power consumption data, sensor status monitoring data, or market transaction data, and the type is a monitoring system or market research, then the data source is an auxiliary data source;

[0019] If the business type is fault diagnosis data, maintenance record data, or user feedback data, or the type is an external third party or a public database, then the data source is classified as a non-core data source.

[0020] Preferably, the various data sources in the distribution communication network are collected in real time, the data collected at each moment is converted into a unified format, and outliers are detected and removed; the data collected by data sources of the same type are combined into a matrix, specifically as follows:

[0021] The formats of the data collected from various data sources include structured data formats and unstructured formats. The structured data formats include CSV and JSON, and the unstructured formats include documents and images. Extract or transform the data collected from various data sources into numerical values, dates, or image feature values. Perform outlier detection on the numerical values, and after deleting the outliers, transform the numerical values, dates, and image feature values extracted or transformed from the data collected from the same type of data source into a matrix format. The formula is;

[0022] Z ki =[z ki1 z ki2 ... z kin

[0023] Among them, Z ki represents the matrix composed of the numerical values, dates, or image feature values extracted or transformed from the data collected from the i-th type of data source at time k. z ki1 、z ki2 、…、z kin respectively represent the numerical values, dates, or image feature values extracted or transformed from the 1st, 2nd, …, n-th data sources of the i-th type of data source at time k. n is the number of data sources of the i-th type.

[0024] Preferably, the extraction or transformation of the data collected from various data sources into numerical values, dates, or image feature values is specifically as follows:

[0025] For the data in CSV format, first read the CSV file where the data is located, and extract the numerical values or dates in each column of the CSV file;

[0026] For the data in JSON format, read the numerical values or dates in each layer from the JSON nested structure where the data is located;

[0027] For documents, use Tesseract for text recognition to extract the numerical values and dates in them,

[0028] The images include the display screen photos of temperature sensors, humidity sensors, and current sensors, and the fault photos taken by external third parties. For images, use OpenCV to extract image feature values. For the display screen photos of temperature sensors, humidity sensors, and current sensors, the image feature value is the number on the display screen. For the fault photos taken by external third parties, the image feature value is the number of faults.

[0029] Preferably, the outlier detection of the numerical values and the deletion of the outliers are specifically as follows:

[0030] Calculate the outlier judgment value. The formula is: ​

[0031]

[0032] Wherein: Z′ is an anomaly judgment value, z is the value for anomaly detection, μ is the mean of the historical data of the data source corresponding to this value, and σ is the standard deviation of the mean of the historical data of the data source corresponding to this value;

[0033] When the absolute value of the anomaly judgment value is greater than or equal to the set threshold, then z is judged as abnormal data and z is deleted.

[0034] Preferably, for the matrices composed of the data collected from the noisy core data source, the noise-free core data source, and the auxiliary data source, data fusion is performed by using Kalman filtering combined with matrix weighting, linear summation with weighting based on the type of data source, and the random forest algorithm, specifically:

[0035] Perform data fusion on the matrix composed of the data of the noisy core data source by using Kalman filtering combined with matrix weighting;

[0036] Perform data fusion on the matrix composed of the data of the noise-free core data source by using linear summation with weighting based on the type of data source;

[0037] Perform data fusion on the matrix composed of the data of the auxiliary data source by using the random forest algorithm.

[0038] Preferably, performing data fusion on the matrix composed of the data of the noisy core data source by using Kalman filtering combined with matrix weighting is specifically:

[0039] For matrix Z ki =[z ki1 z ki2 ... z kin , perform Kalman filtering on each value to obtain each filtered value and the covariance matrix P ki1 , P ki2 , …, P kin

[0040] The formula for Kalman filtering combined with matrix weighting is:

[0041]

[0042] Wherein, is the matrix Z composed of the data of the i-th type of data source at time k ki The value after data fusion, tr(P kij ) is the trace of matrix P kij , P kij is matrix Z ki The j-th value z inkij The covariance matrix is the matrix Z ki The j-th value z in kij The value after filtering, and n is the number of data sources of the i-th type.

[0043] Preferably, for the matrix composed of the data of the noise-free core data sources, data fusion is performed using linear summation weighted based on the type of data source, specifically:

[0044] The linear weighting formula weighted based on the type of data source is:

[0045]

[0046] Where is the matrix Z composed of the data of the data source of the i-th type at time k ki The value after data fusion, z kij is the matrix Z ki The j-th value in, RMSE kij is the root mean square error of the historical data of the j-th data source of the i-th type, w kij is the set weight of the j-th value in the matrix Z ki n is the number of data sources of the i-th type.

[0047] Preferably, the set weight w ki of the j-th value in the matrix Z kij , specifically:

[0048] When the ratio of the accuracy to the resolution of the j-th data source of the i-th type is less than or equal to 5, w kij is 0.5;

[0049] When the ratio of the accuracy to the resolution of the j-th data source of the i-th type is greater than 5 and less than 15, w kij is 0.3;

[0050] When the ratio of the accuracy to the resolution of the j-th data source of the i-th type is greater than or equal to 15, w kij is 0.2.

[0051] The second aspect of the present invention proposes a system for multi-source big data fusion access of a distribution communication network using the method described in the first aspect of the present invention, including an identification module, a classification module, a data acquisition module, a format conversion module, a data fusion module, and a data transmission module, characterized in that:

[0052] Identification module: Identify the types of various data sources in the distribution communication network and the service types of the data collected by the data source;

[0053] Classification module: used to classify various data sources into noisy core data sources, noiseless core data sources, auxiliary data sources, and non-core data sources according to their types and the business types of the data collected by the data sources;

[0054] Data acquisition module: used to control various data sources in the distribution communication network for real-time acquisition;

[0055] Format conversion module: used to convert the data collected at each moment into a unified format, detect and remove outliers; form the data collected by data sources of the same type into a matrix;

[0056] Data fusion module: respectively perform data fusion on the matrices composed of the data collected by noisy core data sources, noiseless core data sources, and auxiliary data sources by means of Kalman filtering combined with matrix weighting, linear summation with weighting based on the type of data source, and random forest algorithm. No data fusion is performed on the data collected by non-core data sources;

[0057] Data transmission module: used to convert the fused data and the data collected by non-core data sources into a set data exchange format, and transmit the fused data to an external system related to the distribution communication network through an API using a set data exchange standard.

[0058] The beneficial effects of the present invention are as follows. Compared with the prior art, various data sources are classified into four types according to different business types and types, including noisy core data sources, noiseless core data sources, auxiliary data sources, and non-core data sources; for noisy core data sources, Kalman filtering combined with matrix weighting is used for data fusion, reducing the noise of temperature sensor data and improving the accuracy of data fusion. For noiseless core data sources, linear summation with weighting based on the type of data source is used, taking into account the influence of different sensor accuracies on data fusion. For auxiliary data sources and non-core data sources, simple random forest algorithm is used for data fusion or no data fusion is performed, reducing the calculation amount; and the fused data is converted into a set data exchange format, and the fused data is transmitted to an external system related to the distribution communication network through an API using a set data exchange standard, reducing the data island phenomenon. Description of the Drawings

[0059] Figure 1 It is a flowchart of the method of the present invention. Detailed Embodiments

[0060] To make the objectives, technical solutions and advantages of the present invention more clear, the following will clearly and completely describe the technical solutions of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. The embodiments described in this application are only a part of the embodiments of the present invention, rather than all embodiments. Based on the spirit of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative work belong to the protection scope of the present invention.

[0061] As Figure 1 shown, Embodiment 1 of the present invention proposes a method for multi-source big data fusion access in a distribution communication network, which is characterized by including:

[0062] Identifying the types of various data sources in the distribution communication network and the service types of the data collected by the data sources;

[0063] Classifying various data sources into noisy core data sources, noiseless core data sources, auxiliary data sources, and non-core data sources according to the types and the service types of the data collected by the data sources;

[0064] Performing real-time collection on various data sources in the distribution communication network, converting the data collected at each moment into a unified format, and detecting and removing outliers; forming a matrix with the data collected by the same type of data sources;

[0065] For the matrices formed by the data collected by the noisy core data sources, noiseless core data sources, and auxiliary data sources, perform data fusion by means of Kalman filtering combined with matrix weighting, linear summation with weighting based on the type of data source, and random forest algorithm respectively. For the data collected by non-core data sources, no data fusion is performed;

[0066] Converting the fused data and the data collected by non-core data sources into a set data exchange format, and transmitting the fused data to an external system related to the distribution communication network through an API using a set data exchange standard.

[0067] Preferably, the types of various data sources in the distribution communication network and the service types of the data collected by the data sources are specifically:

[0068] The service types include real-time monitoring data, sensor status monitoring data, user power consumption data, market transaction data, fault diagnosis data, maintenance record data, user feedback data, and environmental monitoring data;

[0069] The types include temperature sensors, humidity sensors, current sensors, monitoring systems, market research, external third parties, and public databases.

[0070] Preferably, the various data sources are classified into noisy core data sources, noise-free core data sources, auxiliary data sources, and non-core data sources according to their types and the business types of the data collected by the data sources, specifically as follows:

[0071] If the business type is real-time monitoring data and the type is a temperature sensor, then the data source is a noisy core data source;

[0072] If the business type is real-time monitoring data and the type is a humidity sensor or a current sensor, then the data source is a noisy core data source;

[0073] If the business type is user power consumption data, sensor status monitoring data, or market transaction data, and the type is a monitoring system or market research, then the data source is an auxiliary data source;

[0074] If the business type is fault diagnosis data, maintenance record data, or user feedback data, or the type is from an external third party or a public database, then the data source is classified as a non-core data source.

[0075] Preferably, the various data sources in the distribution communication network are collected in real time, the data collected at each moment is converted into a unified format, and outliers are detected and removed; the data collected by the data sources of the same type are combined into a matrix, specifically as follows:

[0076] The formats of the data collected by the various data sources include structured data formats and unstructured formats. The structured data formats include CSV and JSON, and the unstructured formats include documents and images; the data collected by the various data sources are extracted or converted, extracted or converted into numerical values, dates, or image feature values. Among them, the numerical values are subjected to outlier detection, and after removing the outliers, the numerical values, dates, and image feature values extracted or converted from the data collected by the data sources of the same type are subjected to format conversion and converted into matrix format, and the formula is;

[0077] Z ki =[z ki1 z ki2 ... z kin

[0078] where Z ki represents a matrix composed of the numerical values, dates, or image feature values extracted or converted from the data collected by the data source of the i-th type at the k-th moment, and z ki1 、z ki2 、…、z kin represent the numerical values, dates, or image feature values extracted or converted from the data collected by the 1st, 2nd,..., n-th data sources of the i-th type at the k-th moment respectively, and n is the number of data sources of the i-th type.

[0079] ​Preferably, the data collected from various data sources is extracted or transformed into numerical values, dates, or image feature values, specifically as follows:

[0080] For data in CSV format, first read the CSV file where the data is located and extract the numerical values or dates in each column of the CSV file;

[0081] For data in JSON format, read the numerical values or dates in each layer from the JSON nested structure where the data is located;

[0082] For documents, use Tesseract for text recognition to extract the numerical values and dates therein,

[0083] The images include photos of the display screens of temperature sensors, humidity sensors, and current sensors, as well as fault photos taken by external third parties; for images, use OpenCV to extract image feature values; for photos of the display screens of temperature sensors, humidity sensors, and current sensors, the image feature value is the number on the display screen; for fault photos taken by external third parties, the image feature value is the number of faults.

[0084] Preferably, the abnormal value detection of the numerical values and the deletion of abnormal values are specifically as follows:

[0085] Calculate the abnormal judgment value, and the formula is:

[0086]

[0087] Where: Z′ is the abnormal judgment value, z is the numerical value for abnormal value detection, μ is the mean of the historical data of the data source corresponding to this numerical value, and σ is the standard deviation of the mean of the historical data of the data source corresponding to this numerical value;

[0088] When the absolute value of the abnormal judgment value is greater than or equal to the set threshold, then judge z as abnormal data and delete z.

[0089] Preferably, for the matrices composed of the data collected from the noisy core data source, the noiseless core data source, and the auxiliary data source, data fusion is performed by using Kalman filtering combined with matrix weighting, linear summation based on weighting according to the type of data source, and the random forest algorithm, specifically as follows:

[0090] Use Kalman filtering combined with matrix weighting to perform data fusion on the matrix composed of the data of the noisy core data source;

[0091] Use linear summation based on weighting according to the type of data source to perform data fusion on the matrix composed of the data of the noiseless core data source;

[0092] The random forest algorithm is used for data fusion on the matrix composed of the data of the auxiliary data source.

[0093] Preferably, Kalman filtering combined with matrix weighting is used for data fusion on the matrix composed of the data of the noisy core data source. Specifically:

[0094] For matrix Z ki =[z ki1 z ki2 ... z kin , Kalman filtering is performed on each value to obtain each filtered value and the covariance matrix P ki1 、P ki2 、…、P kin

[0095] The formula for Kalman filtering combined with matrix weighting is:

[0096]

[0097] Where is the matrix Z composed of the data of the data source of the i-th type at time k ki The value after data fusion, tr(P kij ) is the trace of matrix P kij , P kij is the covariance matrix of the j-th value z ki in matrix Z kij , is the j-th value z ki in matrix Z kij The filtered value, and n is the number of data sources of the i-th type.

[0098] Preferably, linear summation with weighting based on the type of data source is used for data fusion on the matrix composed of the data of the noise-free core data source. Specifically:

[0099] The formula for linear weighting with weighting based on the type of data source is:

[0100]

[0101] Where is the matrix Z composed of the data of the data source of the i-th type at time k ki The value after data fusion, z kij is the j-th value in matrix Z ki , RMSE kij is the root mean square error of the historical data of the j-th data source of the i-th type, w kij is the set matrix Z kiThe weight of the j-th value in it, where n is the number of data sources of the i-th type.

[0102] Preferably, the set matrix Z ki The weight w of the j-th value in it kij , specifically:

[0103] When the ratio of the accuracy to the resolution of the j-th data source of the i-th type is less than or equal to 5, w kij is 0.5;

[0104] When the ratio of the accuracy to the resolution of the j-th data source of the i-th type is greater than 5 and less than 15, w kij is 0.3;

[0105] When the ratio of the accuracy to the resolution of the j-th data source of the i-th type is greater than or equal to 15, w kij is 0.2.

[0106] Embodiment 2 of the present invention proposes a system for multi-source big data fusion access of a distribution communication network using the method described in Embodiment 1 of the present invention, including an identification module, a classification module, a data acquisition module, a format conversion module, a data fusion module, and a data transmission module, characterized in that:

[0107] Identification module: Identify the types of various data sources in the distribution communication network and the service types of the data collected by the data source;

[0108] Classification module: Used to classify various data sources into noisy core data sources, noiseless core data sources, auxiliary data sources, and non-core data sources according to the type and the service type of the data collected by the data source;

[0109] Data acquisition module: Used to control various data sources in the distribution communication network for real-time acquisition;

[0110] Format conversion module: Used to convert the data collected at each moment into a unified format and detect and remove outliers; form the data collected by the data sources of the same type into a matrix;

[0111] Data fusion module: For the matrices composed of the data collected by the noisy core data sources, noiseless core data sources, and auxiliary data sources, perform data fusion by means of Kalman filtering combined with matrix weighting, linear summation based on the type of data source for weighting, and random forest algorithm respectively. For the data collected by non-core data sources, no data fusion is performed;

[0112] Data transmission module: Used to convert the fused data and the data collected by non-core data sources into a set data exchange format, and transmit the fused data to an external system related to the distribution communication network through the API using the set data exchange standard.

[0113] The present disclosure may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions thereon for causing a processor to implement aspects of the present disclosure.

[0114] The computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. The computer-readable storage medium may be, for example—but not limited to—an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punch card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not to be construed as a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0115] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to a respective computing / processing device, or may be downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in the respective computing / processing device.

[0116] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via an Internet service provider through the Internet). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present disclosure.

[0117] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent substitutions can still be made to the specific embodiments of the present invention. Any modification or equivalent substitution that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.

Claims

1. A method for fusion access of multi-source big data in a power distribution communication network, characterized in that: include: Identify the types of various data sources in the power distribution communication network and the business types of data collected by the data sources; The various data sources are divided into noisy core data sources, noise-free core data sources, auxiliary data sources and non-core data sources according to the types and business types of the data collected by the data sources; Various data sources in the power distribution communication network are collected in real time, the data collected at each moment is converted into a unified format, and abnormal values ​​are detected and removed; the data collected from the same type of data source are formed into a matrix; For the matrices composed of the data collected by the noisy core data source, the noise-free core data source, and the auxiliary data source, Kalman filtering combined with matrix weighting, linear addition based on the type of data source weighting, or random forest algorithm are selected for data fusion. No data fusion is performed on the data collected by the non-core data source; The fused data and data collected from non-core data sources are converted into the set data exchange format, and the fused data is transmitted to external systems related to the distribution communication network through the API using the set data exchange standard.

2. The method for fusion access of multi-source big data in a power distribution communication network according to claim 1, characterized in that: The types of various data sources in the power distribution communication network and the service types of data collected by the data sources are specifically: The business types include real-time monitoring data, sensor status monitoring data, user electricity consumption data, market transaction data, fault diagnosis data, maintenance record data, user feedback data, and environmental monitoring data; The categories include temperature sensors, humidity sensors, current sensors, monitoring systems, market research, external third parties, and public databases.

3. The method for fusion access of multi-source big data in a power distribution communication network according to claim 2, characterized in that: The various data sources are divided into noisy core data sources, noise-free core data sources, auxiliary data sources and non-core data sources according to the types and business types of the data collected by the data sources, specifically: If the service type is real-time monitoring data and the type is temperature sensor, then the data source is a noisy core data source; If the service type is real-time monitoring data and the type is humidity sensor or current sensor, the data source is a noisy core data source; If the business type is user electricity consumption data, sensor status monitoring data or market transaction data, and the type is monitoring system or market research, then the data source is an auxiliary data source; If the business type is fault diagnosis data, maintenance record data or user feedback data, or the type is in an external third party or public database, the data source is classified as a non-core data source.

4. The method for fusion access of multi-source big data in a power distribution communication network according to claim 1, characterized in that: Various data sources in the power distribution communication network are collected in real time, the data collected at each moment is converted into a unified format, and abnormal values ​​are detected and removed; the data collected from the same type of data sources are formed into a matrix, specifically: The formats of the data collected by the various data sources include structured data formats and unstructured formats. The structured data formats include CSV and JSON, and the unstructured formats include documents and images. The data collected from various data sources are extracted or converted into numerical values, dates or image feature values, and the numerical values ​​are detected for outliers. After the outliers are deleted, the numerical values, dates and image feature values ​​extracted or converted from the data collected from the same type of data sources are formatted and converted into a matrix format. The formula is: WITH ki =[of ki1 With ki2 ...With kin ] Among them, Z ki represents the matrix composed of numerical values, dates or image feature values ​​extracted or transformed from the data collected by the i-th type of data source at time k, z ki1 、z ki2 ,…,z kin They respectively represent the numerical values, dates or image feature values ​​extracted or converted from the data collected by the 1st, 2nd, ..., nth data sources of the i-th type at time k, and n is the number of data sources of the i-th type.

5. The method for fusion access of multi-source big data in a power distribution communication network according to claim 4, characterized in that: The data collected from various data sources are extracted or converted into numerical values, dates or image feature values, specifically: For data in CSV format, first read the CSV file where the data is located, and extract the value or date of each column of the CSV file; For data in JSON format, read the value or date of each layer from the JSON nested structure where the data is located; Tesseract is used to perform text recognition on documents to extract values ​​and dates. The images include display screen photos of the temperature sensor, humidity sensor and current sensor, and fault photos taken by an external third party; OpenCV is used to extract image feature values ​​for the images; for display screen photos of the temperature sensor, humidity sensor and current sensor, the image feature values ​​are numbers on the display screen; for fault photos taken by an external third party, the image feature values ​​are the number of faults.

6. The method for fusion access of multi-source big data in a power distribution communication network according to claim 4, characterized in that: The numerical value is subjected to outlier detection, and after the outliers are deleted, specifically: Calculate the abnormal judgment value, the formula is: Z′=(z-μ) σ Where: Z′ is the abnormal judgment value, z is the value for abnormal value detection, μ is the mean of the historical data of the data source corresponding to the value, and σ is the standard deviation of the mean of the historical data of the data source corresponding to the value; When the absolute value of the abnormal judgment value is greater than or equal to the set threshold, z is judged as abnormal data and z is deleted.

7. A method for fusion access of multi-source big data in a power distribution communication network according to any one of claims 4 to 6, characterized in that: The matrix composed of the data collected by the noisy core data source, the noise-free core data source, and the auxiliary data source respectively selects Kalman filtering combined with matrix weighting, linear addition based on the type of data source weighted, or random forest algorithm to perform data fusion, specifically: The matrix composed of data from the noisy core data source is fused using Kalman filtering combined with matrix weighting; The matrix composed of the data of the noise-free core data source is fused by using weighted linear addition based on the type of data source; The random forest algorithm is used to perform data fusion on the matrix composed of data from the auxiliary data source.

8. The method for fusion access of multi-source big data in a power distribution communication network according to claim 7, characterized in that: The matrix composed of data from the noisy core data source is fused using Kalman filtering combined with matrix weighting, specifically: For the matrix Z ki =[z ki1 z ki2 ...z kin ] is Kalman filtered to obtain each value after filtering. And the covariance matrix P of each value ki1 , P ki2 ,…,P kin The Kalman filter combined with the matrix weighted formula is: in, The matrix Z is composed of the data of the i-th type of data source at time k ki The value after data fusion, tr(P kij ) is the matrix P kij The trace of P kij is the matrix Z ki The jth value z in kij The covariance matrix of is the matrix Z ki The jth value z in kij The filtered value, n is the number of data sources of the i-th type.

9. The method for fusion access of multi-source big data in a power distribution communication network according to claim 7, characterized in that: The matrix composed of the data of the noise-free core data source is fused by weighted linear addition based on the type of the data source, specifically: The linear weighting formula based on the type of data source is: in, The matrix Z is composed of the data of the i-th type of data source at time k ki The value after data fusion, z kij is the matrix Z ki The jth value in , RMSE kij is the RMS error of the historical data of the jth data source of the i-th category, w kij is the set matrix Z ki is the weight of the jth value in , and n is the number of data sources of the ith type.

10. The method for fusion access of multi-source big data in a power distribution communication network according to claim 9, characterized in that: The set matrix Z ki The weight w of the jth value in kij , specifically: When the ratio of the precision to resolution of the jth data source of the i-th type is less than or equal to 5, w kij is 0.5; When the ratio of the precision to resolution of the jth data source of the i-th type is greater than 5 and less than 15, w kij is 0.3; When the ratio of the precision to resolution of the jth data source of the i-th type is greater than or equal to 15, w kij is 0.

2.

11. A system for fusion access of multi-source big data in a power distribution communication network using the method according to any one of claims 1 to 10, comprising an identification module, a classification module, a data acquisition module, a format conversion module, a data fusion module, and a data transmission module, characterized in that: Identification module: identifies the types of various data sources in the power distribution communication network and the business types of the data collected by the data sources; Classification module: used to classify various data sources into noisy core data sources, noise-free core data sources, auxiliary data sources and non-core data sources according to the types and business types of the data collected by the data sources; Data acquisition module: used to control various data sources in the power distribution communication network for real-time acquisition; Format conversion module: used to convert the data collected at each moment into a unified format, detect and remove outliers; and form a matrix of data collected from the same type of data source; Data fusion module: For the matrices composed of data collected from noisy core data sources, noise-free core data sources, and auxiliary data sources, data fusion is performed using Kalman filtering combined with matrix weighting, linear addition based on the type of data source, and random forest algorithm. Data collected from non-core data sources are not fused. Data transmission module: used to convert the fused data and data collected from non-core data sources into the set data exchange format, and use the set data exchange standard to transmit the fused data to external systems related to the distribution communication network through API.

Citation Information

Cited By

  • Data retrieval method based on multivariate matrix fusion

    CN121117280A