A Multi-dimensional Data Fusion and Intelligent Analysis System for Smart Park
By dividing the multidimensional data in the smart park into time series, spatial and attribute features, and using corresponding similarity calculation methods to fuse, the heterogeneity problem of multi-source heterogeneous data is solved, and high-precision and high-real-time data processing and analysis are achieved.
Patent Information
- Application Number
- CN202510074083.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-01-17
AI Technical Summary
The heterogeneity of multi-source heterogeneous data in smart parks leads to difficulties in data fusion and analysis, and it is difficult for the existing technology to achieve high-precision, high-reality and intelligent data processing.
A smart park multi-dimensional data fusion and intelligent analysis system is proposed. By dividing the data into time series features, spatial features and attribute features, and using dynamic time regularization, similarity measurement based on graph structure and information entropy algorithm for similarity calculation, data fusion and analysis are realized.
Through multi-dimensional data fusion, information in different data sources can be more accurately integrated, information loss or error fusion can be avoided, data fusion accuracy and complex data relationship capture capabilities, so that the fused data can more truly reflect the actual operating status of the park.
Smart Images

Figure SMS_2 
Figure SMS_54 
Figure QLYQS_1
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data analysis, and particularly to a multi-dimensional data fusion and intelligent analysis system for smart parks. Background Art
[0002] The construction of smart parks involves many fields and complex systems, generating a large amount of multi-source heterogeneous data, such as real-time sensor data generated by Internet of Things devices, image data of video surveillance systems, and structured data in park management business systems. In order to achieve efficient management, intelligent decision-making, and optimized operation of the park, it is necessary to build a system that can effectively fuse these multi-dimensional data and perform in-depth intelligent analysis. This technical solution aims to propose an innovative solution to overcome the limitations of existing technologies in data fusion and analysis and meet the high-precision, high-real-time, and intelligent requirements of smart parks for data processing.
[0003] There are significant differences in data formats, data types, semantic expressions, etc. of different data sources. For example, sensor data may be a continuous numerical stream, with strong real-time but relatively simple semantics; while business data is usually in a structured table form, containing rich semantic information but with a low update frequency. How to convert these heterogeneous data into a unified representation form for effective fusion and analysis is the primary problem to be solved. Therefore, this application proposes a multi-dimensional data fusion and intelligent analysis system for smart parks that can fuse various types of data generated in the park and then perform analysis. Summary of the Invention
[0004] The technical problem to be solved by the present invention is: to provide a multi-dimensional data fusion and intelligent analysis system for smart parks that can fuse various types of data generated in the park and then perform analysis.
[0005] To solve the above technical problem, the technical solution adopted by the present invention is:
[0006] A multi-dimensional data fusion and intelligent analysis system for smart parks, including a controller and a memory. The memory stores various types of data. After artificially dividing various types of data into time series features, spatial features, and attribute features, the controller calculates the similarity of these three types of features respectively and then performs fusion, including:
[0007] For the similarity calculation of time series features, dynamic time warping algorithm:
[0008]
[0009] Wherein, are respectively the time series feature vectors of two data sources, and min represents finding the minimum value is the warping path of the time series, denotes the element value corresponding to the k-th element in the time series under the regular path in the time series ; denotes the k-th element value in the time series ; the minimum cumulative distance obtained by dynamic time warping is used as the similarity of the time series ;
[0010] For the similarity calculation of spatial features, a graph structure-based similarity measurement method is adopted: first, construct a graph model of spatial features and , and by calculating the eigenvalues and eigenvectors of the Laplacian matrix of the graph, the similarity of the eigenvectors is compared to obtain the spatial feature similarity ;
[0011] For the similarity calculation of attribute features, an algorithm based on information entropy is used:
[0012] = 1 -
[0013] where are the time attribute column feature vectors of two data sources respectively, r is the dimension of the attribute feature. For example, includes the power of the device, the brand of the device, and the installation date, similarly; r is the dimension of the attribute feature, then r = 3 (power, brand, installation date), and are respectively the k-th elements in; for example, = , = , then = 200W, = 180W, where k ranges from 1 to r, and is used to traverse each element in the attribute feature vector. The information entropy represents the similarity of the feature, with a value ranging from 0 to 1. The closer it is to 1, the higher the similarity;
[0014] The global similarity matrix M = is obtained by integrating the similarities of each modality:
[0015] M = = + +
[0016] where , , is the weight coefficient of the similarity of each modality, and + + = 1; , , The actual values of can be inversely deduced through the calculation model or can be set manually.
[0017] Calculate the credibility weight of the data source based on the global similarity matrix :
[0018] =
[0019] where n is the total number of data sources, represents the sum of the global similarities between the i-th data source and all other data sources; represents the double sum of the global similarities between all pairs of data sources;
[0020] Finally, perform deep fusion, splice the multi-modal features of each data source to obtain the fused feature vector = , and the fused data F is:
[0021] F =
[0022] The controller outputs in real time according to the fused data F.
[0023] Preferably, there are more than three controllers;
[0024] Each controller calculates the fused data F respectively. If the output data F of each controller is the same value, then the data F is adopted; if there are different data F, then the data F with the highest repeatability is selected; if the output data F of each controller is different, a warning is given and the average value is taken.
[0025] Preferably, construct a generative adversarial network model, and the generative adversarial network model includes a generator and a discriminator;
[0026] The task of the generator is to learn the distribution of normal data and generate data samples similar to the normal data; the discriminator is responsible for distinguishing between real data and the data generated by the generator;
[0027] During the training process, the generator tries to minimize the discrimination accuracy of the discriminator, while the discriminator tries to maximize its discrimination accuracy;
[0028] Input the fused data F into the trained generative adversarial network model, and the discriminator outputs the probability that the data is abnormal. When the probability exceeds the preset value, the data is determined to be abnormal.
[0029] Preferably, the memory stores various types of data for preliminary data cleaning and adjusts to a predetermined format.
[0030] Preferably, the predetermined format includes a date format and a time format. The date format is: year - month - day; the time format is: hour - minute - second; after the adjustment is completed, the formats generated by the park devices are uniformly adjusted to the same date format and time format.
[0031] Preferably, if the time counting methods used by the park devices are inconsistent, the data is converted into the same counting scheme and the same numerical representation.
[0032] Preferably, if the various types of data contain text information, the case is unified, and all the device names and area names in the park are converted into a unified uppercase or lowercase form; special characters in the text information are cleaned or converted, and non - standard line breaks, tab characters, etc. are uniformly processed into standard character formats to ensure the integrity and standardization of the text data.
[0033] Preferably, if there are missing values in the various types of data, the missing parts are filled. The filling methods include filling with the mean, median or mode; if the missing values do not affect the use of the data, a placeholder is selected for filling.
[0034] Preferably, the predetermined format includes a spatial format, and the spatial format includes: uniformly converting the different coordinate systems used by various types of data into the WGS84 coordinate system.
[0035] Preferably, the predetermined format includes an attribute format and a semantic format. The attribute format includes: uniformly converting the different data expression forms used by various types of data into the floating - point number expression form; the semantic format includes: standardizing the attribute vocabulary with the same meaning but different expressions.
[0036] The beneficial effects of the present invention are as follows: By artificially classifying various types of data into time series features, spatial features, and attribute features, the time series features and spatial features can cover most of the data, while the attribute features are used to capture data that is not suitable for time and space distribution. Moreover, by introducing a regular path in the time series, it is possible to use a unified path planning to avoid data reference errors caused by path asynchronization and subsequent fusion failures. By fully considering the time, space, and attribute features of the data for fusion, when processing the equipment operation data and environmental monitoring data in the park, this algorithm can more accurately integrate information about the change of equipment status over time, the spatial distribution of equipment, and the equipment's own attributes (such as model, power, working status, etc.) in different data sources, avoiding information loss or incorrect fusion caused by data heterogeneity, thereby improving the accuracy of data fusion. Its ability to capture complex data relationships is stronger, making the fused data more capable of truly reflecting the actual operation status of the park. Specific Embodiments
[0037] To illustrate the technical content, achieved objectives, and effects of the present invention in detail, the following is described in conjunction with the embodiments.
[0038] A multi-dimensional data fusion and intelligent analysis system for an intelligent park includes a controller and a memory. The memory stores various types of data. After artificially classifying the various types of data into time series features, spatial features, and attribute features, the controller calculates the similarity of these three types of features respectively and then performs fusion, including:
[0039] For the similarity calculation of time series features, dynamic time warping is used algorithm:
[0040]
[0041] Among them, are the time series feature vectors of two data sources respectively, and min represents finding the minimum value is the regular path of the time series, represents under the regular path the element value in the time series corresponding to the k-th element in the time series ; represents the k-th element value in the time series ; The minimum cumulative distance obtained by solving through dynamic time warping is used as the similarity of the time series ;
[0042] For the similarity calculation of spatial features, a similarity measurement method based on a graph structure is used: First, construct a graph model of the spatial features and , the spatial feature similarity is obtained by calculating the eigenvalues and eigenvectors of the Laplacian matrix of the graph and comparing the similarity of the eigenvectors. ;
[0043] For the similarity calculation of attribute features, an algorithm based on information entropy is used: Algorithm:
[0044] = 1 -
[0045] where are the time attribute column feature vectors of two data sources respectively, r is the dimension of the attribute feature, and are respectively the k-th element in;
[0046] The global similarity matrix M is obtained by integrating the similarities of each modality: :
[0047] M = = + +
[0048] where , , are the weight coefficients of the similarities of each modality, and + + = 1;
[0049] The credibility weight of the data source is calculated based on the global similarity matrix: :
[0050] =
[0051] where n is the total number of data sources, represents the sum of the global similarities between the i-th data source and all other data sources; represents the double sum of the global similarities between all data sources pairwise;
[0052] Finally, deep fusion is performed. The multi-modal features of each data source are concatenated to obtain the fused feature vector = , and the fused data F is:
[0053] F =
[0054] The controller outputs in real time according to the fused data F.
[0055] From the above description, various types of data are artificially divided into time-series features, spatial features, and attribute features. Time-series features and spatial features can cover most of the data, while attribute features are used to capture data that is not suitable for time and space distribution. By introducing a regular path in the time series, it is possible to use a unified path planning to avoid data reference errors caused by path asynchronization, which may lead to fusion failure. By fully considering the time, space, and attribute features of the data for fusion, when processing the equipment operation data and environmental monitoring data in the park, the algorithm can more accurately integrate information on the changes in equipment status over time, the spatial distribution of equipment, and the equipment's own attributes (such as model, power, working status, etc.) from different data sources, avoiding information loss or incorrect fusion caused by data heterogeneity, thereby improving the accuracy of data fusion. Its ability to capture complex data relationships is stronger, making the fused data more accurately reflect the actual operation status of the park.
[0056] Further, there are more than three controllers;
[0057] Each controller calculates the fused data F respectively. If the data F output by each controller is the same value, then the data F is adopted; if there are different data F, then the data F with the highest repeatability is selected; if the data F output by each controller is different, a warning is given and the average value is taken.
[0058] From the above description, since the types, scopes, and objects of data involved in the data fusion process vary greatly, it is very easy to make mistakes during the calling process. Therefore, independent calling calculations are required to prevent errors. And if the values obtained from parallel calculations are different, it proves that there is a problem during the data calling process, and a warning is issued. If there is no calculation error, the average value is taken for use first, and then the problem of different outputs is investigated later.
[0059] Further, a generative adversarial network model is constructed. The generative adversarial network model includes a generator and a discriminator;
[0060] The task of the generator is to learn the distribution of normal data and generate data samples similar to the normal data; the discriminator is responsible for distinguishing between real data and the data generated by the generator;
[0061] During the training process, the generator attempts to minimize the discriminative accuracy of the discriminator, while the discriminator attempts to maximize its discriminative accuracy;
[0062] The fused data F is input into the trained generative adversarial network model, and the discriminator outputs the probability that the data is abnormal. When the probability exceeds the preset value, the data is determined to be abnormal.
[0063] From the above description, it can be seen that by detecting abnormal situations in the park, timely alarms can be issued.
[0064] Further, the memory stores various types of data for preliminary data cleaning and adjustment to a predetermined format.
[0065] As can be seen from the above description, by adjusting the data to a predetermined format, it is convenient for reading and subsequent extraction and integration.
[0066] Further, the predetermined format includes a date format and a time format. The date format is: year - month - day; the time format is: hour - minute - second; after the adjustment is completed, the formats generated by the park equipment are uniformly adjusted to the same date format and time format.
[0067] Further, if the time counting methods adopted by the park equipment are inconsistent, the data will be converted into the same counting scheme and the same numerical representation method.
[0068] Further, if the various types of data contain text information, the case will be unified, and all the equipment names and area names in the park will be converted into a unified uppercase or lowercase form; special characters in the text information will be cleaned or converted, and non - standard line breaks, tab characters, etc. will be uniformly processed into standard character formats to ensure the integrity and standardization of the text data.
[0069] Further, if there are missing values in the various types of data, the missing parts will be filled. The filling methods include filling with the mean, median or mode; if the missing values do not affect the use of the data, placeholders will be selected for filling.
[0070] Further, the predetermined format includes a spatial format, and the spatial format includes: uniformly converting the different coordinate systems used by various types of data into the WGS84 coordinate system.
[0071] Further, the predetermined format includes an attribute format and a semantic format. The attribute format includes: uniformly converting the different data representation forms used by various types of data into floating - point representation forms; the semantic format includes: standardizing attribute vocabulary with the same meaning but different expressions.
[0072] Embodiment
[0073] A multi - dimensional data fusion and intelligent analysis system for an intelligent park, including a controller and a memory. The memory stores various types of data, and the various types of data include data D from n different data sources, D = , for each data source, extract its multi - modal features, and artificially divide them into time - series features = , spatial features = and attribute features = ( , , where the k included respectively represents the dimension of each modal feature, which are independent of each other and do not interfere with each other).
[0074] S1. Standardize each modal feature separately so that its mean is 0 and variance is 1.
[0075] S2. The controller calculates the similarity of these three features separately and then fuses them, including:
[0076] For the similarity calculation of time series features, the dynamic time warping algorithm:
[0077]
[0078] where are the time series feature vectors of two data sources respectively, and min represents finding the minimum value is the warping path of the time series, which is a mapping from the time series ; represents that under the warping path , the element value in the time series corresponding to the k-th element in the time series ; represents the k-th element value in the time series ; for example, has m elements, has n elements, then can be represented as a sequence: = , where ) represents the element index in aligned with the k-th element in ; the minimum cumulative distance obtained by solving through dynamic time warping can find the warping path that minimizes the sum of squared differences among all possible alignment methods , and obtain the similarity
[0079] between the two time series; the purpose of this formula is still to be able to capture the similarity between data more accurately and provide a more reliable basis for subsequent data fusion and analysis; = and = and the warping path found through the dynamic time warping algorithm, = This is just a simple example for the purpose of easy understanding, and the actual calculation will be more complex. Then indicates corresponding to in the second element (i.e., 3) of, that is, 2. Such a matching method can handle situations such as the offset and scaling of two time series on the time axis, and measure their similarity more accurately.
[0080] For the similarity calculation of spatial features, a graph-structure-based similarity measurement method is adopted: first, construct a graph model of spatial features and , by calculating the eigenvalues and eigenvectors of the Laplacian matrix of the graph, compare the similarity of the eigenvectors to obtain the spatial feature similarity ; The Laplacian matrix feature is a relatively conventional algorithm in the industry and will not be explained in detail here;
[0081] For the similarity calculation of attribute features, use the algorithm based on information entropy :
[0082] =1 -
[0083] where are the time attribute column feature vectors of two data sources respectively, r is the dimension of the attribute feature, and are respectively the k-th element in; for example, includes the power of the device, the brand of the device, and the installation date, similarly; r is the dimension of the attribute feature, then r = 3 (power, brand, installation date), and are respectively the k-th element in; for example, = , = , then = 200W, = 180W, where k ranges from 1 to r, used to traverse each element in the attribute feature vector. The information entropy represents the similarity of features, with a value ranging from 0 to 1. The closer to 1, the higher the similarity;
[0084] The entire formula is based on the principle of information entropy. By calculating the information entropy-related values of each element in two attribute feature vectors, and performing summation and normalization processing, the similarity of two data sources in terms of attribute features is obtained ; Information entropy is an index to measure information uncertainty. Here, logarithmic operations are used to quantify the uncertainty degree of the information carried by each element in the attribute feature. For example, for , That is, a logarithmic operation is performed on the attribute value, so as to participate in the calculation of information entropy to determine its contribution degree when measuring the similarity of attribute features.
[0085] The global similarity matrix M is obtained by integrating the similarities of each modality: :
[0086] M = = + +
[0087] Among them, , , are the weight coefficients of the similarities of each modality, and + + = 1;
[0088] Calculate the credibility weight of the data source based on the global similarity matrix :
[0089] =
[0090] Among them, n is the total number of data sources; n determines the range of data sources participating in the calculation. For example, if there are 5 different types of data sources in the park (such as environmental monitoring, equipment operation, personnel flow, energy consumption, and security data, etc.), then n = 5; this means that when calculating the credibility weight of each data source, the similarity situation between these 5 data sources in pairs needs to be comprehensively considered.
[0091] represents the sum of the global similarities between the i-th data source and all other data sources; the purpose of this step is to measure the overall similarity degree of the i-th data source and the entire data source set. If this sum value is large, it indicates that the data source has a high similarity with other data sources as a whole, and further implies that its credibility in data fusion may be relatively high, because data sources with high similarity to other data sources often conform more to the overall distribution and rules of the data, and their data quality and reliability are relatively more guaranteed.
[0092] represents the double sum of the global similarities between all data sources in pairs; it calculates the total similarity between all data sources in the entire data source set, and serves as a normalized denominator to ensure that the calculated credibility weight It can be reasonably distributed within the range of 0 to 1, so that the sum of the credibility weights of all data sources is 1. In the subsequent data fusion process, the data of each data source can be weightedly fused (that is, the final deep fusion) according to the relative credibility of each data source, so as to obtain a more accurate and reliable fusion result.
[0093] S3: Perform deep fusion to combine the multimodal features of each data source to obtain a fused feature vector = , the fused data F is:
[0094] F=
[0095] The controller performs real-time output based on the fused data F.
[0096] There are more than three controllers;
[0097] Each controller calculates the fused data F separately. If the output data F of each controller is the same value, the data F is used; if there are different data F, the data F with the highest repeatability is selected; if the output data F of each controller is different, a warning is issued and the average value is taken.
[0098] S4. Build a generative adversarial network model (GAN), which includes a generator and a discriminator;
[0099] The task of the generator is to learn the distribution of normal data and generate data samples similar to normal data; the discriminator is responsible for distinguishing between real data and data generated by the generator;
[0100] During training, the generator tries to minimize the discriminator's accuracy, while the discriminator tries to maximize its accuracy.
[0101] The fused data F is input into the trained generative adversarial network model, and the discriminator outputs the probability that the data is abnormal. When the probability exceeds the preset value, the data is judged to be abnormal.
[0102] The most difficult part of park data integration is actually the data. The data generated by each device is different, and each company's products have a specific recording method. Therefore, the data needs to be processed in a unified manner, and the storage device stores various types of data for preliminary data cleaning and adjustment of the predetermined format.
[0103] If the various types of data contain text information, unify the case, converting all device names and area names in the park to a unified uppercase or lowercase form; for example, for the brand, model, etc. of the device, unify the character encoding format (such as GTX-4080 and gtx-4080, unified as GTX-4080); clean or convert special characters in the text information, and uniformly process non-standard line breaks, tab characters, etc. into standard character formats to ensure the integrity and standardization of the text data.
[0104] If there are missing values in the various types of data, fill in the missing part. The filling methods include filling with the mean, median, or mode; if the missing values do not affect the use of the data, select a placeholder (such as "NA" or "none") for filling.
[0105] The predefined formats include:
[0106] The date format is: year-month-day;
[0107] The time format is: hour-minute-second; after the adjustment is completed, the formats generated by the park devices are uniformly adjusted to the same date format and time format. If the time counting methods used by the park devices are inconsistent, convert the data to the same counting scheme and the same numerical representation method.
[0108] The spatial format is: uniformly convert the different coordinate systems (such as the geodetic coordinate system, the plane rectangular coordinate system, etc.) used by the various types of data to the WGS84 coordinate system. This can be achieved through coordinate conversion algorithms to ensure that all information related to spatial positions (such as device positions, personnel positions, building positions, etc.) is represented in the same coordinate framework, facilitating subsequent spatial analysis and calculations.
[0109] The attribute format is: uniformly convert the different data expression forms used by the various types of data to the floating-point number expression form; for example, if the power attribute of a device in one data source is represented as an integer, and in another data source it is represented as a floating-point number, it is necessary to unify it to one data type, such as a floating-point number, for accurate numerical calculations and analyses.
[0110] The semantic format is: standardize the attribute vocabulary with the same meaning but different expressions. For example, for the operating state of a device, different data sources may use different words such as "running", "working", "on" to describe it. It is necessary to unify these expressions to a standard term, such as "operating state", and define a clear value range to ensure the semantic consistency of the data.
[0111] The above are only embodiments of the present invention, and thus do not limit the patent scope of the present invention. Any equivalent transformation made using the content of the specification of the present invention, directly or indirectly applied in the relevant technical fields, shall similarly be included within the patent protection scope of the present invention.
Claims
1. A smart park multi-dimensional data fusion and intelligent analysis system, characterized in that: The system includes a controller and a memory, wherein the memory stores various types of data. After artificially dividing the various types of data into time series features, spatial features, and attribute features, the controller calculates the similarity of the three types of features and then fuses them, including: For the similarity calculation of time series features, dynamic time warping is used algorithm: in, They are the time series feature vectors of two data sources, and min means finding the minimum value is the regular path of the time series, Indicates that the path is regular Next, time series Corresponding to the time series The element value of the kth element in ; Representing time series The kth element value in ; the minimum cumulative distance obtained by dynamic time warping is used as the similarity of the time series ; For the similarity calculation of spatial features, a similarity measurement method based on graph structure is adopted: first, a graph model of spatial features is constructed and , by calculating the Laplacian matrix eigenvalues and eigenvectors of the graph and comparing the similarity of the eigenvectors, we can get the spatial feature similarity ; For the similarity calculation of attribute features, we use the information entropy-based algorithm: =1- in, are the attribute feature vectors of the two data sources, r is the dimension of the attribute feature, and They are The kth element in ; The global similarity matrix M is obtained by combining the similarities of each modality. : M= = + + in, , , is the weight coefficient of each modality similarity, and + + =1; Calculate the credibility weight of data sources based on the global similarity matrix : = Where n is the total number of data sources, Indicates the sum of the global similarities between the i-th data source and all other data sources; It means to perform double summation of the global similarities between all data sources; Finally, deep fusion is performed to splice the multimodal features of each data source to obtain a fused feature vector = , the fused data F is: F= The controller performs real-time output based on the fused data F.
2. The smart park multi-dimensional data fusion and intelligent analysis system according to claim 1 is characterized in that: There are more than three controllers; Each controller calculates the fused data F separately. If the output data F of each controller is the same value, the data F is used; if there are different data F, the data F with the highest repeatability is selected; if the output data F of each controller is different, a warning is issued and the average value is taken.
3. The smart park multi-dimensional data fusion and intelligent analysis system according to claim 1 is characterized in that: Build a generative adversarial network model, which includes a generator and a discriminator; The task of the generator is to learn the distribution of normal data and generate data samples similar to normal data; the discriminator is responsible for distinguishing between real data and data generated by the generator; During training, the generator tries to minimize the discriminator's accuracy, while the discriminator tries to maximize its accuracy. The fused data F is input into the trained generative adversarial network model, and the discriminator outputs the probability that the data is abnormal. When the probability exceeds the preset value, the data is judged to be abnormal.
4. The smart park multi-dimensional data fusion and intelligent analysis system according to claim 1 is characterized in that: The memory stores various types of data for preliminary data cleaning and adjustment of the predetermined format.
5. The smart park multi-dimensional data fusion and intelligent analysis system according to claim 4 is characterized in that: The preset format includes date format and time format. The date format is: year-month-day; the time format is: hour-minute-second. After the adjustment is completed, the formats generated by the park equipment are uniformly adjusted to the same date format and time format.
6. The smart park multi-dimensional data fusion and intelligent analysis system according to claim 5 is characterized in that: If the time counting methods used by the park equipment are inconsistent, the data should be converted into the same counting scheme and the same numerical representation.
7. The smart park multi-dimensional data fusion and intelligent analysis system according to claim 4 is characterized in that: If various types of data contain text information, the case will be unified and all equipment names and area names in the park will be converted to unified uppercase or lowercase. Special characters in the text information will be cleaned or converted, and non-standard line breaks and tabs will be uniformly processed into standard character formats to ensure the integrity and standardization of the text data.
8. The smart park multi-dimensional data fusion and intelligent analysis system according to claim 4 is characterized in that: If there are missing values in various types of data, the missing parts will be filled using the mean, median or mode. If the missing values do not affect the use of the data, placeholders will be selected to fill them.
9. The smart park multi-dimensional data fusion and intelligent analysis system according to claim 4 is characterized in that: The predetermined format includes a spatial format, and the spatial format includes: converting different coordinate systems used by various types of data into a WGS84 coordinate system.
10. The smart park multi-dimensional data fusion and intelligent analysis system according to claim 4, characterized in that: The predetermined formats include attribute formats and semantic formats. The attribute formats include: uniformly converting different data expressions used by various types of data into floating-point expressions; the semantic formats include: standardizing attribute terms with the same meaning but different expressions.
Citation Information
Patent Citations
Data fusion method and device of Internet of Things equipment, intelligent terminal and storage medium
CN113961622A
Smart park multi-source data dynamic monitoring and real-time analysis system and method
CN116665001A