Smart city multi-modal data acquisition and fusion method and system
By constructing a semantic mapping rule library and using a naive Bayes model for data cleaning and standardization, combined with principal component analysis and preset fusion strategies, the semantic inconsistency problem during multimodal data fusion in smart cities is solved, and data availability and integration are improved.
Patent Information
- Application Number
- CN202510267411.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-03-07
AI Technical Summary
The prior art faces semantic inconsistency when processing multimodal data in smart cities, resulting in confusion and errors during data fusion and reducing the availability of data.
Semantic mapping technology is used to build a semantic mapping rule library, convert heterogeneous data into machine-readable information, and data cleaning and standardization are carried out through pre-trained naive Bayesian models, combining principal component analysis and preset fusion strategies to achieve effective data fusion.
It effectively solves the problem of semantic inconsistency, improves data availability and integration, and ensures the accuracy and reliability of the data fusion process in smart city systems.
Smart Images

Figure CN120197130A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of smart cities, and particularly to a method and system for collecting and fusing multi-modal data in a smart city. Background Art
[0002] A smart city is an urban development model that uses information and communication technologies to improve urban services and management, so as to enhance urban work efficiency and the quality of life of residents. Multi-modal data refers to information data from multiple sources, formats, and types. With the expansion of modern urban scale and the growth of population, a huge amount and variety of multi-modal data have been generated in smart cities, such as video surveillance data, meteorological sensor data, traffic sensor data, social media data, etc. These data can be generated and collected through different types of sensors, devices, and stations. The multi-modal data generated in smart cities can provide useful information about urban operations, population characteristics, and event activities. However, due to different data sources, formats, and types, these data face many difficulties in effective utilization. Collecting and fusing multi-modal data from different sources, formats, and qualities is a key step in building a smart city system and an important foundation for the sustainable development of smart cities.
[0003] In an existing technology, the method for collecting and fusing multi-modal data in a smart city includes establishing an urban information model, generating building models, traffic models, enterprise models, and population flow models; formulating data collection indicators corresponding to the models according to the models, and fusing the data collection indicators; generating data collection tasks according to the data collection indicators, and collecting and integrating the collected data. However, different data sources may describe the same concept differently. For example, "temperature" in a meteorological sensor may be expressed in degrees Celsius, while "temperature" in another data source may be expressed in degrees Fahrenheit; "vehicle speed" in a traffic sensor may be expressed in kilometers per hour, while "vehicle speed" in another data source may be expressed in miles per hour. Therefore, the existing technology only collects and integrates the collected data, and in the case of dealing with semantic inconsistencies, it will lead to confusion and errors in data fusion.
[0004] In summary, in the data collection process of the existing technology, only the collected data is collected and integrated, and in the case of dealing with semantic inconsistencies, it will lead to confusion and errors in data fusion, resulting in poor usability of the data. Summary of the Invention
[0005] The present invention provides a method for collecting and fusing multi-modal data in a smart city, so as to realize constructing a semantic mapping rule library by using semantic mapping technology, converting various heterogeneous data into machine-understandable information, effectively performing data fusion, and improving the usability of data.
[0006] In a first aspect, to solve the above technical problems, the present invention provides a method for collecting and fusing multi-modal data in a smart city, including:
[0007] Obtain smart city system data; wherein, the smart city system data includes meteorological sensor data, traffic sensor data, and social media data;
[0008] Perform data cleaning on the smart city system data according to a pre-trained Naive Bayes model, and extract metadata of data items; the metadata of the data items includes structured data, semi-structured data, and unstructured data;
[0009] Perform data standardization according to the metadata of the data items to obtain standardized data information;
[0010] Perform data conversion on the standardized data information according to a pre-formed semantic mapping rule library to obtain machine-readable data; wherein, the semantic mapping rule library is obtained by mining learning parameters from historical metadata using the Expectation-Maximization algorithm and constructing based on the learning parameters;
[0011] Perform merging and compression on the machine-readable data based on principal component analysis to obtain dimension-reduced readable data;
[0012] Perform data fusion on the dimension-reduced readable data according to a preset fusion strategy to obtain fusion data information and store it.
[0013] As an optional implementation manner, the obtaining of the smart city system data includes:
[0014] Use a data transmission protocol to transmit the internal system data of the smart city to the smart city data platform;
[0015] Use an integration method to import the external system data of the smart city into the smart city data platform;
[0016] Obtain the smart city data platform information through an open API interface to obtain the smart city system data;
[0017] Among them, the data transmission protocol includes HTTP, MQTT, CoAP, DICOM, and FTP;
[0018] The integration method includes ETL tools and APIs.
[0019] As an optional implementation manner, the performing of data cleaning on the smart city system data according to a pre-trained Naive Bayes model and extracting the metadata of the data items includes:
[0020] Perform data inspection and repair on the smart city system data to obtain repaired system data;
[0021] Input the repair system data into a pre-trained Naive Bayes model for data classification to obtain classified system data;
[0022] Perform data deduplication operations based on the classified system data to obtain metadata for data items;
[0023] Among them, the training process of the Naive Bayes model includes:
[0024] Obtain historical smart city system data;
[0025] For the historical smart city system data, use the bag-of-words model for feature extraction to obtain feature encodings;
[0026] Based on the historical smart city system data, perform probability statistical calculations to obtain class prior probabilities;
[0027] Train the initial model based on the feature encodings and the class prior probabilities, and obtain the Naive Bayes model after training is completed.
[0028] As an alternative implementation, the performing data standardization based on the metadata of the data items to obtain standardized data information includes:
[0029] Perform a unified data source operation on the metadata of the data items to obtain first standardized information;
[0030] Based on the first standardized information, use a preset sequence-to-sequence model and a knowledge graph conversion algorithm for format conversion and unit conversion to obtain second standardized information;
[0031] Based on the second standardized information, use standardized semantic tags to unify the data names to obtain standardized data information.
[0032] As an alternative implementation, the using a preset sequence-to-sequence model and a knowledge graph conversion algorithm for format conversion and unit conversion based on the first standardized information to obtain second standardized information includes:
[0033] Based on the first standardized information, use the sequence-to-sequence model for format conversion to obtain standard format information;
[0034] Based on the first standardized information, use the knowledge graph conversion algorithm for unit conversion to obtain standard unit information;
[0035] Based on the standard format information and the standard unit information, perform information splicing to obtain second standardized information.
[0036] As an alternative implementation, the configuration process of the semantic mapping rule library includes:
[0037] Preprocess the historical metadata to obtain preprocessed data; the preprocessing includes removing stop words, stemming, lemmatization, and word form reduction;
[0038] Perform natural language processing on the preprocessed data to obtain natural language data. The natural language processing includes word segmentation, part-of-speech tagging, and named entity recognition;
[0039] Use Word2Vec to extract features from the natural language data to obtain natural language features;
[0040] Use the natural language features as input and train using a multi-layer perceptron to generate text feature vectors;
[0041] Use the expectation-maximization algorithm to cluster each text feature vector and mine the probability distribution of the categories to obtain the semantic mapping rule library.
[0042] As an alternative implementation, the use of the expectation-maximization algorithm to cluster each text feature vector and mine the probability distribution of the categories to obtain the semantic mapping rule library includes:
[0043] Initialize the initial parameters of the expectation-maximization algorithm and start the iteration;
[0044] In the i-th iteration, based on the current parameters, calculate the conditional probability distribution of the current observed data and the current parameters, and sum all possible conditional probability distributions to obtain the expected value of the log-likelihood function;
[0045] Find the parameter value that maximizes the expected value of the log-likelihood function as the parameter estimate value for the (i + 1)-th iteration;
[0046] Repeat the iteration process until the difference between the parameter estimates of two consecutive iterations is less than a preset convergence threshold, then end the iteration and output the parameter value at this time as the learning parameter;
[0047] Mine the probability distribution of the categories based on the learning parameter to obtain the semantic mapping rule library.
[0048] As an alternative implementation, the merging and compression of the machine-readable data based on principal component analysis to obtain the dimension-reduced readable data includes:
[0049] According to the machine-readable data, perform data merging to obtain non-duplicate machine-readable data;
[0050] According to the non-duplicate machine-readable data, perform range data standardization to obtain the first standardized information;
[0051] Perform covariance calculation according to the first standardized information to obtain a covariance matrix;
[0052] Perform eigenvalue decomposition according to the covariance matrix to obtain eigenvalues and eigenvectors;
[0053] Arrange the eigenvalues from largest to smallest, and select the eigenvector corresponding to the first sorted eigenvalue as the second eigenvector;
[0054] Successively calculate the variance contribution rates corresponding to the second eigenvector, and add up the variance contribution rates to obtain the sum of the variance contribution rates. When the sum of the variance contribution rates is greater than or equal to a preset contribution rate threshold, the accumulated second eigenvectors are used as the third eigenvector;
[0055] Extract the principal component data information that constitutes the third eigenvector, and use the principal component data information as the dimension-reduced readable data.
[0056] As an optional implementation manner, the data fusion of the dimension-reduced readable data according to a preset fusion strategy to obtain fusion data information and store it includes:
[0057] Match the dimension-reduced readable data according to a preset fusion strategy. When it is determined that there are conflicts between data sources, an approximation operation is performed on the two conflicting data sources;
[0058] Integrate the matched dimension-reduced readable data in combination with the weighted average to obtain fusion data information;
[0059] Perform data storage on the fusion data information by using a symmetric encryption algorithm;
[0060] Among them, the symmetric encryption algorithm includes AES and DES encryption algorithms.
[0061] In a second aspect, the present invention provides a smart city multi-modal data acquisition and fusion system, including:
[0062] A data acquisition module for acquiring smart city system data;
[0063] A data cleaning module for cleaning the smart city system data according to a pre-trained Naive Bayes model to extract metadata of data items; the metadata of the data items includes structured data, semi-structured data, and unstructured data;
[0064] A data standardization module for performing data standardization according to the metadata of the data items to obtain standardized data information;
[0065] A data conversion module, configured to convert the standardized data information according to a pre-formed semantic mapping rule library to obtain machine-readable data; wherein, the semantic mapping rule library is constructed by mining historical metadata using the expectation maximization algorithm to obtain learning parameters and based on the learning parameters.
[0066] A data compression module, configured to merge and compress the machine-readable data based on principal component analysis to obtain dimension-reduced readable data.
[0067] A data fusion module, configured to perform data fusion on the dimension-reduced readable data according to a preset fusion strategy to obtain fusion data information and store it.
[0068] In a third aspect, the present invention further provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the smart city multi-modal data acquisition and fusion method described in any one of the above.
[0069] Compared with the prior art, the present invention has the following beneficial effects:
[0070] This method cleans the data using the naive Bayes model according to the data of the urban system, thereby extracting the metadata of the data items. These metadata include structured data, semi-structured data, and unstructured data. Next, the data is standardized according to these metadata to obtain standardized data information. Then, according to the pre-established semantic mapping rule library, the standardized data information is converted into data that can be understood by the machine. The semantic mapping rule library is constructed by mining historical metadata using the expectation maximization algorithm to obtain learning parameters. After that, the principal component analysis method is used to merge and compress the machine-readable data to obtain dimension-reduced data. Finally, according to the preset data fusion strategy, the dimension-reduced data is fused to obtain fused data information and stored.
[0071] The method of the present invention uses semantic mapping technology to construct a semantic mapping rule library, converts various heterogeneous data into machine-understandable information, effectively performs data fusion, provides strong support for the construction of smart cities, and improves the usability of data. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] Figure 1 is a schematic flowchart of a smart city multi-modal data acquisition and fusion method provided by the first embodiment of the present invention;
[0073] Figure 2 is a schematic structural diagram of a smart city multi-modal data acquisition and fusion system provided by the second embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0074] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0075] Referring to Figure 1 , the first embodiment of the present invention provides a method for collecting and fusing multi-modal data in a smart city, including the following steps:
[0076] S11, obtaining smart city system data.
[0077] S12, performing data cleaning on the smart city system data according to a pre-trained Naive Bayes model, and extracting metadata of data items; the metadata of the data items includes structured data, semi-structured data, and unstructured data.
[0078] S13, performing data standardization according to the metadata of the data items to obtain standardized data information.
[0079] S14, performing data conversion on the standardized data information according to a pre-formed semantic mapping rule library to obtain machine-readable data; wherein, the semantic mapping rule library is obtained by mining learning parameters from historical metadata using the Expectation-Maximization algorithm and constructing based on the learning parameters.
[0080] S15, merging and compressing the machine-readable data based on principal component analysis to obtain dimension-reduced readable data.
[0081] S16, performing data fusion on the dimension-reduced readable data according to a preset fusion strategy to obtain fusion data information and storing it.
[0082] In step S11, the obtaining of the smart city system data includes:
[0083] Using a data transfer protocol to transfer the internal system data of the smart city to the smart city data platform;
[0084] Using an integration method to import the external system data of the smart city into the smart city data platform;
[0085] Obtaining the smart city data platform information through an open API interface to obtain the smart city system data;
[0086] Among them, the data transfer protocol includes HTTP, MQTT, CoAP, DICOM, and FTP;
[0087] The integration methods include ETL tools and APIs.
[0088] It should be noted that opening an API interface means that a company makes its application programming interface public, enabling external developers or partners to access its software applications and data platforms.
[0089] In this embodiment, the smart city system data includes meteorological sensor data, traffic sensor data, and social media data. The data formats and types of different data sources may be different, and the descriptions of the same concept by different data sources may also be different. For example, "temperature" in meteorological sensors may be expressed in degrees Celsius, while "temperature" in another data source may be expressed in degrees Fahrenheit; "vehicle speed" in traffic sensors may be expressed in kilometers per hour, while "vehicle speed" in another data source may be expressed in miles per hour.
[0090] Exemplarily, the representation form of meteorological sensor data is temperature (degrees Celsius), humidity (percentage); the representation form of traffic sensor data is vehicle speed (kilometers per hour), traffic flow (vehicles per hour); the representation form of social media data is user comments (text).
[0091] In step S12, the data cleaning of the smart city system data is performed according to the pre-trained Naive Bayes model, and the metadata of the data items is extracted, including:
[0092] Performing data inspection and repair on the smart city system data to obtain repaired system data;
[0093] Inputting the repaired system data into the pre-trained Naive Bayes model for data classification to obtain classified system data;
[0094] Performing data deduplication operation according to the classified system data to obtain the metadata of the data items;
[0095] Among them, the training process of the Naive Bayes model includes:
[0096] Obtaining historical smart city system data;
[0097] Performing feature extraction on the historical smart city system data using the bag-of-words model to obtain feature encodings;
[0098] Performing probability statistical calculation according to the historical smart city system data to obtain the class prior probability;
[0099] Training the initial model according to the feature encodings and the class prior probability, and obtaining the Naive Bayes model after the training is completed.
[0100] It should be noted that the data inspection and repair of the smart city system data to obtain the repaired system data includes: checking whether there are missing values in the data and detecting outliers in the data using statistical methods; when there are missing values in the data, filling the missing values using the methods of mean and median according to the distribution characteristics or correlations of the data; if the outliers are caused by data entry errors or equipment failures, searching for the original data for correction.
[0101] It should be further noted that the bag-of-words model is a model that treats text as an unordered set of words, ignores word order and grammatical structure, and only focuses on the frequency of each word appearing in the text.
[0102] It should be further noted that the data deduplication operation on the classification system data to obtain the metadata of data items includes: searching using the bubble method according to the classification system data to find duplicate data, and deleting the duplicate data to obtain the metadata of data items.
[0103] Exemplarily, assume there is the following smart city system data, including meteorological data, traffic flow data, and social media data: Meteorological data: Temperature: 25°C (valid), Humidity: 60% (valid), Precipitation: -5 mm (invalid, negative value is unreasonable); Traffic flow data: Vehicle speed: 80 km / h (valid), Traffic volume: 120 vehicles / hour (valid), Vehicle speed: -20 km / h (invalid, negative value is unreasonable); Social media data: User comment: "The weather is really nice today!" (valid), User comment: "#" (invalid, meaningless).
[0104] Assume the following historical data: Meteorological data: Temperature: [20°C, 22°C, 24°C, 26°C, 28°C], Humidity: [50%, 55%, 60%, 65%, 70%], Precipitation: [0 mm, 2 mm, 5 mm, 10 mm, 15 mm]; Traffic flow data: Vehicle speed: [60 km / h, 70 km / h, 80 km / h, 90 km / h, 100 km / h], Traffic volume: [100 vehicles / hour, 120 vehicles / hour, 140 vehicles / hour, 160 vehicles / hour, 180 vehicles / hour]; Social media data: User comment: ["The weather is really nice today!", "Traffic congestion!", "In a good mood!", "It's raining!", "Traffic jam is so annoying!"]
[0105] Before training the Naive Bayes model, it is necessary to extract features from historical data and calculate prior probabilities. The bag-of-words model is used to extract features from text data to obtain feature encodings. For numerical data, the numerical values themselves are directly used as features. Then, the prior probabilities for each category are calculated. For example, for meteorological data, the prior probabilities of temperature, humidity, and precipitation are respectively: Temperature: P(Temperature) = 1; Humidity: P(Humidity) = 1; Precipitation: P(Precipitation) = 1. Then, the Naive Bayes model is trained using the feature encodings and prior probabilities. The trained model can identify valid and invalid data. When in use, new smart city system data is input into the trained Naive Bayes model for data cleaning. The cleaned data is classified and de-duplicated, and metadata is extracted from the classified and de-duplicated data, including structured data, semi-structured data, and unstructured data. Exemplarily, structured data: Temperature 25°C, Humidity 60%; Vehicle speed 80 km / h, Traffic flow 120 vehicles / hour. Unstructured data such as user comments "The weather is really nice today!" etc.
[0106] In step S13, performing data standardization according to the metadata of the data item to obtain standardized data information, including:
[0107] Performing a unified data source operation on the metadata of the data item to obtain the first standardized information;
[0108] According to the first standardized information, using a preset sequence-to-sequence model and knowledge graph conversion algorithm to perform format conversion and unit conversion to obtain the second standardized information;
[0109] According to the second standardized information, using standardized semantic tags to unify the data names to obtain the standardized data information.
[0110] It should be noted that performing a unified data source operation on the metadata of the data item to obtain the first standardized information includes: defining a naming standard and a unified format specification, mapping the data source name to the standard name, and unifying the data in different regions of the city to obtain the first standardized information.
[0111] It should be noted that according to the second standardized information, using standardized semantic tags to unify the data names to obtain the standardized data information includes: The semantic tags are from a tag system preset by the user, such as roads, bridges, and are not specifically limited here.
[0112] It should be noted that according to the first standardized information, using a preset sequence-to-sequence model and knowledge graph conversion algorithm to perform format conversion and unit conversion to obtain the second standardized information includes:
[0113] According to the first standardized information, perform format conversion using a sequence-to-sequence model to obtain standard format information;
[0114] According to the first standardized information, perform unit conversion using a knowledge graph conversion algorithm to obtain standard unit information;
[0115] According to the standard format information and the standard unit information, perform information splicing to obtain second standardized information.
[0116] It should be further noted that the sequence-to-sequence model is used to convert an input sequence into a context vector of a fixed length, and then an output sequence is gradually generated according to the context vector.
[0117] It is worth noting that the unit conversion using the knowledge graph conversion algorithm includes: constructing a unit knowledge graph using various unit names, mapping the unit knowledge graph to a vector space using a knowledge graph embedding algorithm (such as TransE), and finding corresponding matrix information according to the vector space to obtain standard unit information.
[0118] In step S14, the configuration process of the semantic mapping rule library includes:
[0119] S14, perform data conversion on the standardized data information according to a pre-formed semantic mapping rule library to obtain machine-readable data; wherein, the semantic mapping rule library is constructed by mining learning parameters from historical metadata using the expectation maximization algorithm and based on the learning parameters.
[0120] Preprocess the historical metadata to obtain preprocessed data; the preprocessing includes removing stop words, stemming, lemmatization, and word form reduction;
[0121] Perform natural language processing on the preprocessed data to obtain natural language data, and the natural language processing includes word segmentation, part-of-speech tagging, and named entity recognition;
[0122] Use Word2Vec to extract features from the natural language data to obtain natural language features;
[0123] Take the natural language features as input and use a multi-layer perceptron for training to generate text feature vectors;
[0124] Use the expectation maximization algorithm to cluster each text feature vector and mine the probability distribution of the categories to obtain the semantic mapping rule library.
[0125] It should be noted that the feature extraction of the natural language data using Word2Vec to obtain natural language features includes: obtaining the vector representation form of the natural language using the wv attribute of the Word2Vec model, calculating the average value of the occurrence of vector words based on the vector representation form of the natural language, and obtaining natural language features based on the average value of the occurrence of the vector values.
[0126] It should be further noted that using the natural language features as input and training with a multi-layer perceptron to generate text feature vectors includes: inputting the natural language features into a pre-constructed multi-layer perceptron model, adjusting the weights through the backpropagation algorithm to minimize the loss function, obtaining the adjusted weights, obtaining the output result of the natural language through the adjusted weight model, and obtaining the text feature vector from the output result.
[0127] It should be further noted that the pre-constructed multi-layer perceptron model includes determining the dimension of the input layer according to the dimension of the natural language features, setting the number of preset activation functions, and determining the number of output layers according to the number of output categories; among them, the number of preset activation functions can be 10, and no specific limitation is made here.
[0128] It should be noted that using the expectation-maximization algorithm to cluster each text feature vector and mining the probability distribution of the categories to obtain a semantic mapping rule library includes:
[0129] Initializing the initial parameters of the expectation-maximization algorithm and starting the iteration;
[0130] In the i-th iteration, based on the current parameters, calculating the conditional probability distribution of the current observed data and the current parameters, and summing all possible conditional probability distributions to obtain the expected value of the log-likelihood function;
[0131] Finding the parameter value that maximizes the expected value of the log-likelihood function as the parameter estimate value for the (i + 1)-th iteration;
[0132] Repeating the iteration process until the difference between the parameter estimates of two consecutive iterations is less than the preset convergence threshold, ending the iteration, and outputting the parameter value at this time as the learning parameter;
[0133] Mining the probability distribution of the categories based on the learning parameter to obtain a semantic mapping rule library.
[0134] It should be further noted that the method for initializing the initial parameters of the expectation-maximization algorithm is the random initialization method. Among them, the initial parameters include: mean, covariance matrix, and mixing weights.
[0135] It should be noted that in the i-th iteration, based on the current parameters, the conditional probability distribution of the current observed data and the current parameters is calculated, and the expected value of the log-likelihood function is obtained by summing over all possible conditional probability distributions. The current observed data is obtained by subtracting the learning rate multiplied by the gradient from the previous observed data. The conditional probability distribution of the current parameters is calculated by the following formula:
[0136]
[0137] In the formula, x n is the feature vector of the n-th text feature, θ (i) is the parameter set of the i-th iteration, is the mixing weight of the i-th iteration, is the mean vector of the i-th iteration, is the covariance matrix of the i-th iteration, k represents the category as k, z n represents the latent variable used to indicate the category, K represents the total number of categories in the Gaussian distribution, represents the multivariate normal distribution.
[0138] The expected value of the log-likelihood function is calculated by the following formula:
[0139]
[0140] In the formula, N represents the total number of observed text samples, π k represents the current mixing weight, μ k represents the current iteration mean vector, and θ represents the current parameter set.
[0141] It should be noted that the probability distribution of the categories is mined based on the learning parameters to obtain a semantic mapping rule base, including:
[0142] According to the learning parameters, the probability of each corresponding category is obtained; according to the probability distribution of each category, a semantic mapping rule base is obtained.
[0143] Exemplarily, first, the metadata of the data items is standardized to obtain the following standardized data information: meteorological sensor data: temperature (degrees Celsius), humidity (percentage); traffic sensor data: vehicle speed (km / h), traffic flow (vehicles / h); social media data: user comments (text). Then, the historical metadata is mined through the expectation-maximization algorithm to construct the following semantic mapping rule base: temperature: conversion of degrees Celsius to Fahrenheit; humidity: conversion of percentage to decimal; vehicle speed: conversion of km / h to m / s; traffic flow: conversion of vehicles / h to vehicles / min; user comments: conversion of text to sentiment score. Then, according to the semantic mapping rule base, the standardized data information is converted into machine-readable data:
[0144] Meteorological sensor data: temperature (in degrees Celsius) is converted to temperature (in degrees Fahrenheit); humidity (in percentage) is converted to humidity (in decimal); traffic sensor data: vehicle speed (in km / h) is converted to vehicle speed (in m / s), traffic flow (in vehicles per hour) is converted to traffic flow (in vehicles per minute); social media data: user comments (text) are converted to sentiment scores (numerical values). After these data conversions, they have a unified format and semantic representation, making it easier to fuse them.
[0145] In step S15, the machine-readable data is merged and compressed based on principal component analysis to obtain dimension-reduced readable data, including:
[0146] Based on the machine-readable data, data merging is performed to obtain non-duplicate machine-readable data;
[0147] Based on the non-duplicate machine-readable data, range data standardization is performed to obtain first standardized information;
[0148] Based on the first standardized information, covariance calculation is performed to obtain a covariance matrix;
[0149] Based on the covariance matrix, eigenvalue decomposition is performed to obtain eigenvalues and eigenvectors;
[0150] Sorted in descending order of eigenvalues, the eigenvector corresponding to the first-ranked eigenvalue is selected as the second eigenvector;
[0151] The variance contribution rate corresponding to the second eigenvector is calculated in sequence, and the variance contribution rates are added up to obtain the sum of the variance contribution rates. When the sum of the variance contribution rates is greater than or equal to a preset contribution rate threshold, the accumulated second eigenvectors are used as the third eigenvector;
[0152] The principal component data information constituting the third eigenvector is extracted, and the principal component data information is used as the dimension-reduced readable data.
[0153] It should be noted that the range data standardization based on the non-duplicate machine-readable data to obtain the first standardized information includes:
[0154] Data standardization is achieved according to the following formula:
[0155]
[0156] In the formula, x′ represents the first standardized information, x represents the non-duplicate machine-readable data, min(x) represents the minimum value of the non-duplicate machine-readable data, and max(x) represents the maximum value of the non-duplicate machine-readable data.
[0157] It should be noted that the preset contribution rate threshold is preset by the user, which can be 98%, as long as it is reasonable, and no specific limitation is made here.
[0158] In step S16, the data fusion of the dimensionality-reduced readable data according to the preset fusion strategy to obtain fusion data information and store it includes:
[0159] According to the preset fusion strategy, match the dimensionality-reduced readable data. When it is determined that there is a conflict between data sources, an approximation operation is performed on the two conflicting data sources;
[0160] Combine the weighted average to integrate the matched dimensionality-reduced readable data to obtain fusion data information;
[0161] According to the fusion data information, use a symmetric encryption algorithm for data storage;
[0162] Among them, the symmetric encryption algorithm includes AES and DES encryption algorithms.
[0163] It should be noted that the operation of taking the average value of the data generated by the same data source as an approximation operation is included in the step of matching the dimensionality-reduced readable data according to the preset fusion strategy and performing an approximation operation on the two conflicting data sources when it is determined that there is a conflict between data sources.
[0164] It should be further noted that the operation of combining the weighted average to integrate the matched dimensionality-reduced readable data to obtain fusion data information includes: the sum of the weight values in the weighting process is 1, and the weight values are preset by the user, which can be 0.1, 0.2, etc., and are determined according to the number of the matched dimensionality-reduced readable data, as long as it is reasonable, and no limitation is made here.
[0165] Compared with the prior art, the present invention has the following beneficial effects:
[0166] The present invention provides a method and system for collecting and fusing multi-modal data in a smart city. The method is executed by a smart city data platform and includes: obtaining smart city system data; performing data cleaning on the smart city system data according to a pre-trained Naive Bayes model to extract metadata of data items; the metadata of the data items includes structured data, semi-structured data, and unstructured data; performing data standardization according to the metadata of the data items to obtain standardized data information; performing data conversion on the standardized data information according to a pre-formed semantic mapping rule library to obtain machine-readable data; wherein, the semantic mapping rule library is obtained by mining historical metadata using the Expectation-Maximization algorithm to obtain learning parameters and constructing based on the learning parameters; performing merging and compression on the machine-readable data based on principal component analysis to obtain dimension-reduced readable data; performing data fusion on the dimension-reduced readable data according to a preset fusion strategy to obtain fused data information and storing it. This method cleans the data using the Naive Bayes model according to the data of the city system, thereby extracting the metadata of the data items. These metadata include structured data, semi-structured data, and unstructured data. Next, the data is standardized according to these metadata to obtain standardized data information. Then, according to the pre-established semantic mapping rule library, the standardized data information is converted into data that can be understood by the machine. The semantic mapping rule library is constructed by mining historical metadata using the Expectation-Maximization algorithm to obtain learning parameters. After that, the principal component analysis method is used to merge and compress the machine-readable data to obtain dimension-reduced data. Finally, according to the preset data fusion strategy, the dimension-reduced data is fused to obtain fused data information and stored. The present invention provides a method for collecting and fusing multi-modal data in a smart city to achieve efficient collection and fusion, eliminate data fragmentation, and improve data availability.
[0167] For ease of understanding of the present invention, some preferred embodiments of the present invention will be further described below.
[0168] In this embodiment, a multi-modal data acquisition and fusion device for a smart city is proposed. The device consists of the following key components: a data acquisition device, a data cleaning device, a data standardization device, a data conversion device, a data compression device, and a data fusion device. Among them, the data acquisition device is used to capture raw data from various sources in the city (such as traffic monitoring cameras, environmental monitoring stations, public facilities, etc.). The data cleaning device is used to identify and remove incorrect or incomplete data. This device uses the Naive Bayes model to improve the accuracy and reliability of the data. The data standardization device is used to convert the cleaned data into a unified format and range. This device ensures that all data is measured according to the same standard, facilitating subsequent processing. The data conversion device is used to convert the standardized data into a machine-readable format. This device mines historical metadata through the Expectation-Maximization algorithm to build a semantic mapping rule library. The data compression device is used to reduce the dimensionality of the data while retaining key information. This device improves the efficiency of data storage and processing. The data fusion device is used to integrate the dimensionality-reduced data and resolve conflicts between data sources. This device uses approximation operations such as mean, mode, median, or weighted average to fuse the data. The present invention provides a multi-modal data acquisition and fusion device for a smart city to achieve efficient acquisition and fusion, eliminate data fragmentation, and improve the usability of the data.
[0169] The solution is executed by the following steps:
[0170] Step 1: Transmit internal data using data transfer protocols (HTTP, MQTT, CoAP, DICOM, FTP); import external data through ETL tools and API integration methods; obtain data platform information using open API interfaces.
[0171] Step 2: Use the Naive Bayes model to perform data inspection and repair, data classification, and deduplication operations to obtain metadata for structured, semi-structured, and unstructured data.
[0172] Step 3: Perform data source unification operations, use sequence-to-sequence models and knowledge graph conversion algorithms for format conversion and unit conversion, and use standardized semantic tags to unify data names.
[0173] Step 4: Apply the semantic mapping rule library, which is constructed based on learning parameters mined from historical metadata through the Expectation-Maximization algorithm.
[0174] Step 5: Merge and compress the machine-readable data based on principal component analysis to obtain dimensionality-reduced readable data.
[0175] Step 6: Perform data matching according to the preset fusion strategy, perform an approximation operation on the conflicting data sources, integrate the data by combining the weighted average, and finally use a symmetric encryption algorithm (such as AES, DES) for data storage.
[0176] Compared with the prior art, the present invention has the following beneficial effects:
[0177] The present invention provides a method and system for collecting and fusing multi-modal data in a smart city. The method is executed by a smart city data platform and includes: obtaining smart city system data; performing data cleaning on the smart city system data according to a pre-trained Naive Bayes model to extract metadata of data items; the metadata of the data items includes structured data, semi-structured data, and unstructured data; performing data standardization on the basis of the metadata of the data items to obtain standardized data information; performing data conversion on the standardized data information according to a pre-formed semantic mapping rule library to obtain machine-readable data; wherein, the semantic mapping rule library is obtained by mining historical metadata using the Expectation-Maximization algorithm to obtain learning parameters and constructing based on the learning parameters; performing merging and compression on the machine-readable data based on principal component analysis to obtain dimension-reduced readable data; performing data fusion on the dimension-reduced readable data according to a preset fusion strategy to obtain fusion data information and storing it. This method cleans the data using the Naive Bayes model based on the data of the city system, thereby extracting the metadata of the data items. These metadata include structured data, semi-structured data, and unstructured data. Next, the data is standardized according to these metadata to obtain standardized data information. Then, according to the pre-established semantic mapping rule library, the standardized data information is converted into data that can be understood by the machine. The semantic mapping rule library is constructed by mining historical metadata using the Expectation-Maximization algorithm to obtain learning parameters. After that, the principal component analysis method is used to merge and compress the machine-readable data to obtain the dimension-reduced data. Finally, according to the preset data fusion strategy, the dimension-reduced data is fused to obtain the fused data information and stored. The present invention provides a method for collecting and fusing multi-modal data in a smart city to achieve efficient collection and fusion, eliminate data fragmentation, and improve the usability of data.
[0178] For the convenience of understanding the present invention, some preferred embodiments of the present invention will be further described below.
[0179] Refer to Figure 2 , the second embodiment of the present invention provides a system for collecting and fusing multi-modal data in a smart city, including:
[0180] A data acquisition module, configured to acquire smart city system data;
[0181] A data cleaning module, configured to perform data cleaning on the smart city system data according to a pre-trained Naive Bayes model, and extract metadata of data items; the metadata of the data items includes structured data, semi-structured data, and unstructured data;
[0182] A data standardization module, configured to perform data standardization according to the metadata of the data items to obtain standardized data information;
[0183] A data conversion module, configured to perform data conversion on the standardized data information according to a pre-formed semantic mapping rule library to obtain machine-readable data; wherein, the semantic mapping rule library is obtained by mining historical metadata using the Expectation-Maximization algorithm to obtain learning parameters, and is constructed based on the learning parameters;
[0184] A data compression module, configured to perform merging and compression on the machine-readable data based on principal component analysis to obtain dimension-reduced readable data;
[0185] A data fusion module, configured to perform data fusion on the dimension-reduced readable data according to a preset fusion strategy to obtain fusion data information and store it.
[0186] In one implementation, the obtaining of the smart city system data includes:
[0187] Using a data transmission protocol to transmit the internal system data of the smart city to the smart city data platform;
[0188] Using an integration method to import the external system data of the smart city into the smart city data platform;
[0189] Obtaining the smart city data platform information through an open API interface to obtain the smart city system data;
[0190] Wherein, the data transmission protocol includes HTTP, MQTT, CoAP, DICOM, and FTP;
[0191] The integration method includes ETL tools and APIs.
[0192] In one implementation, the data cleaning module, configured to perform data cleaning on the smart city system data according to a pre-trained Naive Bayes model, and extract metadata of data items; the metadata of the data items includes structured data, semi-structured data, and unstructured data, includes:
[0193] Performing data inspection and repair on the smart city system data to obtain repaired system data;
[0194] Inputting the repaired system data into a pre-trained Naive Bayes model for data classification to obtain classified system data;
[0195] Perform data deduplication on the data according to the classification system data to obtain the metadata of the data items;
[0196] Among them, the training process of the Naive Bayes model includes:
[0197] Obtain historical smart city system data;
[0198] Perform feature extraction on the historical smart city system data using the bag-of-words model to obtain feature encodings;
[0199] Perform probability statistical calculations based on the historical smart city system data to obtain the prior probability of the category;
[0200] Train the initial model based on the feature encodings and the prior probability of the category, and obtain the Naive Bayes model after training is completed.
[0201] In one implementation, the data standardization module is used to perform data standardization according to the metadata of the data items to obtain standardized data information, including:
[0202] Perform a unified operation on the data sources of the metadata of the data items to obtain the first standardized information;
[0203] According to the first standardized information, use a preset sequence-to-sequence model and knowledge graph conversion algorithm to perform format conversion and unit conversion to obtain the second standardized information;
[0204] According to the second standardized information, use standardized semantic tags to unify the data names to obtain the standardized data information.
[0205] In one implementation, the data conversion module is used to convert the standardized data information according to a pre-formed semantic mapping rule library to obtain machine-readable data; among them, the semantic mapping rule library is obtained by mining historical metadata using the expectation maximization algorithm to obtain learning parameters and constructing based on the learning parameters. Its characteristics are that the configuration process of the semantic mapping rule library includes:
[0206] Preprocess the historical metadata to obtain preprocessed data; the preprocessing includes removing stop words, stemming, lemmatization, and word form reduction;
[0207] Perform natural language processing on the preprocessed data to obtain natural language data. The natural language processing includes word segmentation, part-of-speech tagging, and named entity recognition;
[0208] Use Word2Vec to perform feature extraction on the natural language data to obtain natural language features;
[0209] Using the natural language features as input, training is performed using a multi-layer perceptron to generate text feature vectors;
[0210] Using the expectation-maximization algorithm to cluster each text feature vector and mining the probability distribution of the categories to obtain a semantic mapping rule base.
[0211] In one implementation, the data compression module is used to merge and compress the machine-readable data based on principal component analysis to obtain dimension-reduced readable data, including: performing data merging on the machine-readable data to obtain non-duplicate machine-readable data;
[0212] Performing range data standardization on the non-duplicate machine-readable data to obtain first standardized information;
[0213] Performing covariance operation on the first standardized information to obtain a covariance matrix;
[0214] Performing eigenvalue decomposition on the covariance matrix to obtain eigenvalues and eigenvectors;
[0215] Sorting according to the eigenvalues from largest to smallest, and selecting the eigenvector corresponding to the first-ranked eigenvalue as the second eigenvector;
[0216] Successively calculating the variance contribution rate corresponding to the second eigenvector and adding the variance contribution rates to obtain the sum of the variance contribution rates. When the sum of the variance contribution rates is greater than or equal to a preset contribution rate threshold, the accumulated second eigenvectors are used as the third eigenvector;
[0217] Extracting the principal component data information constituting the third eigenvector and using the principal component data information as the dimension-reduced readable data.
[0218] In one implementation, the data fusion module is used to perform data fusion on the dimension-reduced readable data according to a preset fusion strategy to obtain fusion data information and store it. It includes:
[0219] According to the preset fusion strategy, matching the dimension-reduced readable data. When it is determined that there is a conflict between data sources, an approximation operation is performed on the two conflicting data sources;
[0220] Combining with the weighted average, integrating the matched dimension-reduced readable data to obtain fusion data information;
[0221] Performing data storage on the fusion data information using a symmetric encryption algorithm;
[0222] Among them, the symmetric encryption algorithm includes AES and DES encryption algorithms.
[0223] Compared with the prior art, the present invention has the following beneficial effects:
[0224] The present invention provides a method and system for collecting and fusing multi-modal data in a smart city. The method is executed by a smart city data platform and includes: obtaining smart city system data; performing data cleaning on the smart city system data according to a pre-trained Naive Bayes model to extract metadata of data items; the metadata of the data items includes structured data, semi-structured data, and unstructured data; performing data standardization according to the metadata of the data items to obtain standardized data information; performing data conversion on the standardized data information according to a pre-formed semantic mapping rule library to obtain machine-readable data; wherein, the semantic mapping rule library is obtained by mining historical metadata using the Expectation Maximization algorithm to obtain learning parameters and constructing based on the learning parameters; performing merging and compression on the machine-readable data based on principal component analysis to obtain dimension-reduced readable data; performing data fusion on the dimension-reduced readable data according to a preset fusion strategy to obtain fusion data information and storing it. This method cleans the data using the Naive Bayes model based on the data of the city system, thereby extracting the metadata of the data items. These metadata include structured data, semi-structured data, and unstructured data. Next, the data is standardized according to these metadata to obtain standardized data information. Then, according to the pre-established semantic mapping rule library, the standardized data information is converted into data that can be understood by the machine. The semantic mapping rule library is constructed by mining historical metadata using the Expectation Maximization algorithm to obtain learning parameters. After that, the principal component analysis method is used to merge and compress the machine-readable data to obtain dimension-reduced data. Finally, according to the preset data fusion strategy, the dimension-reduced data is fused to obtain fused data information and stored. The present invention provides a method for collecting and fusing multi-modal data in a smart city to achieve efficient collection and fusion, eliminate data fragmentation, and improve data availability.
[0225] It should be noted that a device for collecting and fusing multi-modal data in a smart city provided in an embodiment of the present invention is used to execute all the process steps of a method for collecting and fusing multi-modal data in a smart city in the above embodiment. Their working principles and beneficial effects correspond one by one, so they will not be elaborated here.
[0226] An embodiment of the present invention also provides an electronic device. The electronic device includes: a processor, a memory, and a computer program stored in the memory and executable on the processor, such as a data preprocessing program. When the processor executes the computer program, it implements the steps in each of the above embodiments of the method for collecting and fusing multi-modal data in a smart city, such as Figure 1The step S11 shown. Alternatively, when the processor executes the computer program, it implements the functions of each module / unit in the above-described device embodiments, such as the data acquisition module.
[0227] Exemplarily, the computer program may be divided into one or more modules / units, and the one or more modules / units are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the electronic device.
[0228] The electronic device may be a desktop computer, a notebook, a palm computer, a smart tablet, or other computing devices. The electronic device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the above components are only examples of the electronic device and do not constitute a limitation on the electronic device. It may include more or fewer components than the above, or combine certain components, or different components. For example, the electronic device may further include input / output devices, network access devices, a bus, etc.
[0229] The so-called processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the electronic device and connects various parts of the entire electronic device through various interfaces and lines.
[0230] The memory can be used to store the computer programs and / or modules. By running or executing the computer programs and / or modules stored in the memory, and by invoking the data stored in the memory, the processor can implement various functions of the electronic device. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory can include high-speed random access memory, and can also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices.
[0231] Among them, if the modules / units integrated in the electronic device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present invention, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of the above-mentioned various method embodiments. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0232] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided by the present invention, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0233] The specific embodiments described above have further elaborated on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. In particular, for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A smart city multimodal data collection and fusion method, characterized in that: Executed by the Smart City Data Platform, including: Acquire smart city system data; wherein the smart city system data includes meteorological sensor data, traffic sensor data and social media data; Performing data cleaning on the smart city system data according to a pre-trained naive Bayes model to extract metadata of the data item; the metadata of the data item includes structured data, semi-structured data and unstructured data; Performing data standardization according to the metadata of the data item to obtain standardized data information; The standardized data information is converted according to a pre-formed semantic mapping rule base to obtain machine-readable data; wherein the semantic mapping rule base is obtained by mining historical metadata using an expectation maximization algorithm to obtain learning parameters, and is constructed based on the learning parameters; Combining and compressing the machine-readable data based on principal component analysis to obtain dimension-reduced readable data; The dimension-reduced readable data is fused according to a preset fusion strategy to obtain fused data information, which is then stored.
2. The smart city multimodal data collection and fusion method according to claim 1 is characterized in that: The obtaining of smart city system data includes: Using data transmission protocols, the data of the smart city internal system is transmitted to the smart city data platform; Import data from external smart city systems into the smart city data platform using an integrated approach; Through the open API interface, obtain the smart city data platform information and obtain the smart city system data; Among them, data transmission protocols include HTTP, MQTT, CoAP, DICOM, and FTP; Integration methods include ETL tools and APIs.
3. The smart city multimodal data collection and fusion method according to claim 1 is characterized in that: The data cleaning of the smart city system data according to the pre-trained naive Bayes model to extract metadata of the data items includes: Performing data inspection and repair on the smart city system data to obtain repaired system data; Inputting the repair system data into a pre-trained naive Bayes model for data classification to obtain classified system data; Performing a data deduplication operation according to the classification system data to obtain metadata of the data item; The naive Bayes model training process includes: Access historical smart city system data; For the historical smart city system data, a bag-of-words model is used to perform feature extraction to obtain feature coding; Performing probability statistical calculations based on the historical smart city system data to obtain category prior probabilities; The initial model is trained according to the feature coding and the category prior probability, and a naive Bayes model is obtained after the training is completed.
4. The smart city multimodal data collection and fusion method according to claim 1 is characterized in that: The step of performing data standardization according to the metadata of the data item to obtain standardized data information includes: Performing a data source unification operation on the metadata of the data item to obtain first standardized information; According to the first standardized information, a preset sequence-to-sequence model and a knowledge graph conversion algorithm are used to perform format conversion and unit conversion to obtain second standardized information; According to the second standardized information, the data names are unified using standardized semantic tags to obtain standardized data information.
5. The smart city multimodal data collection and fusion method according to claim 4 is characterized in that: The method of performing format conversion and unit conversion according to the first standardized information by using a preset sequence-to-sequence model and a knowledge graph conversion algorithm to obtain second standardized information includes: According to the first standardized information, a sequence-to-sequence model is used to perform format conversion to obtain standard format information; According to the first standardized information, a knowledge graph conversion algorithm is used to perform unit conversion to obtain standard unit information; Information splicing is performed according to the standard format information and the standard unit information to obtain second standardized information.
6. The smart city multimodal data collection and fusion method according to claim 1 is characterized in that: The configuration process of the semantic mapping rule base includes: Preprocess the historical metadata to obtain preprocessed data; the preprocessing includes removing stop words, root restoration, stem extraction and word form restoration; Performing natural language processing on the preprocessed data to obtain natural language data, wherein the natural language processing includes word segmentation, part-of-speech tagging, and named entity recognition; Using Word2Vec to extract features from the natural language data to obtain natural language features; Taking the natural language features as input, training with a multi-layer perceptron to generate a text feature vector; The expectation maximization algorithm is used to cluster each text feature vector, and the probability distribution of the categories is mined to obtain a semantic mapping rule base.
7. The smart city multimodal data collection and fusion method according to claim 6 is characterized in that: The expectation maximization algorithm is used to cluster each text feature vector and mine the probability distribution of the category to obtain a semantic mapping rule base, including: Initialize the initial parameters of the expectation maximization algorithm and start iteration; In the i-th iteration, based on the current parameters, the conditional probability distribution of the current observation data and the current parameters is calculated, and the expected value of the log-likelihood function is obtained by summing all possible conditional probability distributions; Finding a parameter value that can maximize the expected value of the log-likelihood function as the parameter estimate for the (i+1)th iteration; The iteration process is repeated until the difference between the parameter estimation values of two consecutive iterations is less than the preset convergence threshold, and then the iteration is terminated and the parameter value at this time is output as the learning parameter; The probability distribution of the categories is mined based on the learning parameters to obtain a semantic mapping rule base.
8. The smart city multimodal data collection and fusion method according to claim 1 is characterized in that: The merging and compressing the machine-readable data based on principal component analysis to obtain dimension-reduced readable data includes: Merging the machine-readable data to obtain non-repetitive machine-readable data; Performing range data normalization according to the non-repetitive machine-readable data to obtain first normalized information; Performing a covariance operation according to the first standardized information to obtain a covariance matrix; According to the covariance matrix, eigenvalue decomposition is performed to obtain eigenvalues and eigenvectors; Arrange the eigenvalues from large to small, and select the eigenvector corresponding to the first eigenvalue as the second eigenvector; Calculating the variance contribution rates corresponding to the second eigenvectors in sequence, and adding the variance contribution rates to obtain the sum of the variance contribution rates, and when the sum of the variance contribution rates is greater than or equal to a preset contribution rate threshold, using the accumulated second eigenvectors as the third eigenvector; The principal component data information constituting the third eigenvector is extracted, and the principal component data information is used as dimensionally reduced readable data.
9. The smart city multimodal data collection and fusion method according to claim 1 is characterized in that: The step of fusing the dimension-reduced readable data according to a preset fusion strategy to obtain fused data information and storing the information includes: According to a preset fusion strategy, the dimension-reduced readable data are matched, and when it is determined that there is a conflict between data sources, an approximate value operation is performed on the two conflicting data sources; Combined with the weighted average, the matched dimensionality-reduced readable data is integrated to obtain fused data information; According to the fused data information, a symmetric encryption algorithm is used for data storage; The symmetric encryption algorithms include AES and DES encryption algorithms.
10. A smart city multimodal data collection and fusion system, characterized in that: include: Data acquisition module, used to obtain smart city system data; A data cleaning module, used to clean the smart city system data according to a pre-trained naive Bayes model, and extract metadata of the data items; the metadata of the data items includes structured data, semi-structured data and unstructured data; A data standardization module, used to perform data standardization according to the metadata of the data item to obtain standardized data information; A data conversion module is used to convert the standardized data information according to a pre-formed semantic mapping rule base to obtain machine-readable data; wherein the semantic mapping rule base is obtained by mining historical metadata using an expectation maximization algorithm to obtain learning parameters, and is constructed based on the learning parameters; A data compression module, used for merging and compressing the machine-readable data based on principal component analysis to obtain dimension-reduced readable data; The data fusion module is used to perform data fusion on the dimension-reduced readable data according to a preset fusion strategy to obtain fused data information and store it.
Citation Information
Patent Citations
Smart city traffic planning method and system based on big data
CN118917025A
Methods and systems for reuse of data item fingerprints in generation of semantic maps
US20220156303A1
Cited By
Highway area risk intelligent early warning method and system based on space-time coupling model
CN120472694A
Chemical material data processing method and system based on large language model
CN120930651A