Fusion method and system for multi-source heterogeneous data

By constructing a spatiotemporal index table and using technologies such as Kalman filtering and the national secret SM4 algorithm, the problem of spatiotemporal granularity differences in multi-source heterogeneous data was solved, efficient data fusion and analysis were achieved, and the performance of the 5G network and the accuracy of urban management were improved.

CN120602958APending Publication Date: 2025-09-05GUANGXI ACAD OF SCI +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510668557.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing technologies have alignment errors of up to 15% caused by differences in temporal and spatial granularity when processing multi-source heterogeneous data. Traditional fixed-time window aggregation methods are difficult to adapt dynamically, and the query delay exceeds 200ms, which cannot meet the millisecond-level optimization requirements of 5G.

Method used

By constructing the GPS trajectory spatiotemporal index table and the base station MR spatiotemporal index table, using Kalman filtering for spatiotemporal calibration, and combining the national secret SM4 algorithm and Intel SGX technology, a fusion method for multi-source heterogeneous data is calculated, including obtaining the adjusted raster, GPS trajectory data, base station MR data, sensor data and user complaint logs, constructing the spatiotemporal index table and calculating the eigenvector, ultimately achieving efficient data fusion.

Benefits of technology

It achieves efficient organization and analysis of multi-source heterogeneous data, avoids data dislocation, improves analysis efficiency, reduces query latency, and meets the optimization needs of 5G networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120602958A_ABST
    Figure CN120602958A_ABST
Patent Text Reader

Abstract

The invention discloses a fusion method and system for multi-source heterogeneous data. The fusion method for the multi-source heterogeneous data comprises the steps that adjusted grids, GPS track data, base station MR data, sensor data and user complaint logs are obtained; constructing a GPS track space-time index table based on the adjusted grids and the GPS track data; constructing a base station MR spatial-temporal index table based on the adjusted grids and the base station MR data; based on the GPS trajectory spatial-temporal index table and the base station MR spatial-temporal index table, performing spatial-temporal calibration through Kalman filtering to obtain a first feature vector; according to the method, the space-time grids are dynamically calculated, and the index table is constructed, so that efficient data organization is realized; calculating a second feature vector based on the sensor data and the user complaint log; on the basis of the first feature vector and the second feature vector, the fused multi-source heterogeneous data is calculated through the SM4 cryptographic algorithm, data dislocation is avoided, and the analysis efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field related to the fusion of multi-source heterogeneous data, and in particular to a method and system for the fusion of multi-source heterogeneous data. Background Art

[0002] With the widespread adoption of 5G communication technology, multi-source, heterogeneous data generated by mobile terminals has become a core analytical target for scenarios such as network optimization. This data exhibits strong spatiotemporal dynamics, diverse modalities, and privacy-sensitive characteristics. Efficiently integrating and securely utilizing this data has become a key challenge in improving 5G network performance and the precision of urban management.

[0003] Currently, existing technologies face the following bottlenecks when addressing the above requirements: the differences in spatiotemporal granularity of data from different sources cause alignment errors of up to 15%, making it difficult for traditional fixed-time window aggregation methods to adapt dynamically; and existing indexing technologies have query delays exceeding 200ms with an average daily data volume of 1 billion items, which cannot meet the millisecond-level optimization requirements of 5G. Summary of the Invention

[0004] This application aims to at least solve the technical problems existing in the prior art. To this end, this application proposes a method and system for fusing multi-source heterogeneous data, which can avoid data misalignment and improve analysis efficiency.

[0005] In a first aspect of the present application, a method for fusing multi-source heterogeneous data is provided, comprising the following steps:

[0006] Obtain adjusted grid, GPS trajectory data, base station MR data, sensor data, and user complaint logs;

[0007] Constructing a GPS track spatiotemporal index table based on the adjusted grid and the GPS track data; constructing a base station MR spatiotemporal index table based on the adjusted grid and the base station MR data;

[0008] Based on the GPS trajectory spatiotemporal index table and the base station MR spatiotemporal index table, performing spatiotemporal calibration through Kalman filtering to obtain a first eigenvector;

[0009] calculating a second feature vector based on the sensor data and the user complaint log;

[0010] Based on the first eigenvector and the second eigenvector, the fused multi-source heterogeneous data is calculated using the national secret SM4 algorithm.

[0011] The control method according to the embodiment of the present application has at least the following beneficial effects:

[0012] This method obtains the adjusted grid, GPS trajectory data, base station MR data, sensor data and user complaint log; constructs a GPS trajectory spatiotemporal index table based on the adjusted grid and GPS trajectory data; constructs a base station MR spatiotemporal index table based on the adjusted grid and base station MR data; performs spatiotemporal calibration through Kalman filtering based on the GPS trajectory spatiotemporal index table and the base station MR spatiotemporal index table to obtain a first eigenvector; this application achieves efficient data organization by dynamically calculating the spatiotemporal grid and constructing the index table; calculates the second eigenvector based on the sensor data and user complaint log; and calculates the fused multi-source heterogeneous data based on the first eigenvector and the second eigenvector through the national secret SM4 algorithm, thereby avoiding data dislocation and improving analysis efficiency.

[0013] According to some embodiments of the present application, obtaining the adjusted grid includes:

[0014] Obtain the total number of users, the total number of grids, and the number of users in each grid within a preset time period;

[0015] Calculating a user ratio of a corresponding grid based on the total number of grids and the number of users in each grid;

[0016] The user distribution entropy value is calculated based on the total number of users and the corresponding grid user ratio using the following formula:

[0017]

[0018] Among them, H is the user distribution entropy value, N is the total number of users in the preset time period, and p i is the proportion of grid users corresponding to the i-th grid;

[0019] The adjusted grid side length is calculated based on the user distribution entropy value using the following formula:

[0020] L grid =L max ·e -kH ;

[0021] Among them, L grid To adjust the grid side length, L max is the preset maximum grid side length, k is the preset adjustment coefficient;

[0022] The grid within a preset time period is adjusted based on the adjusted grid side length to obtain the adjusted grid.

[0023] According to some embodiments of the present application, constructing a GPS track spatiotemporal index table based on the adjusted grid and the GPS track data includes:

[0024] Obtain the corresponding stay time and movement speed of each user in the adjusted grid;

[0025] The grid weight is calculated based on the preset time period, the corresponding moving speed and the corresponding stay time using the following formula:

[0026]

[0027] α+β=1

[0028] Where W is the grid weight, T stay is the corresponding stay time, T total is the preset time period, α is the first preset weight coefficient, β is the second preset weight coefficient, and v is the corresponding moving speed;

[0029] The GPS track spatiotemporal index table is constructed based on the adjusted grid, the GPS track data and the grid weight, wherein the GPS track spatiotemporal index table includes the grid number, timestamp, data source type and grid weight of the adjusted grid.

[0030] According to some embodiments of the present application, constructing a base station MR spatiotemporal index table based on the adjusted grid and the base station MR data includes:

[0031] The base station MR spatiotemporal index table is constructed based on the adjusted grid, the base station MR data and the grid weight.

[0032] According to some embodiments of the present application, performing spatiotemporal calibration based on the GPS trajectory spatiotemporal index table and the base station MR spatiotemporal index table through Kalman filtering to obtain a first eigenvector includes:

[0033] Based on the GPS trajectory spatiotemporal index table and the base station MR spatiotemporal index table, the spatiotemporal deviation is corrected by Kalman filtering to obtain calibrated spatiotemporal trajectory data;

[0034] Feature extraction is performed on the calibrated spatiotemporal trajectory data to obtain the first feature vector.

[0035] According to some embodiments of the present application, the calculating the second feature vector based on the sensor data and the user complaint log includes:

[0036] Based on the sensor data, extracting behavioral pattern features through a convolutional neural network model;

[0037] Based on the user complaint log, extracting complaint semantic features through natural language processing;

[0038] The implicit weight value is calculated based on the behavior pattern characteristics and the complaint semantic characteristics using the following formula:

[0039]

[0040] Among them, s i is the behavioral pattern characteristic, s j is the complaint semantic feature, w ij is the implicit weight value;

[0041] Based on the implicit weight value, the high-weight value features are mapped to the corresponding adjusted grid to obtain the second feature vector.

[0042] According to some embodiments of the present application, the calculating of the fused multi-source heterogeneous data by the national secret SM4 algorithm based on the first eigenvector and the second eigenvector includes:

[0043] splicing the first feature vector and the second feature vector based on a timestamp and a grid number of the adjusted grid to obtain a spliced ​​feature vector;

[0044] The concatenated feature vectors are privacy-enhanced using the national secret SM4 algorithm and Intel SGX technology to obtain the fused multi-source heterogeneous data.

[0045] In a second aspect of the present application, a multi-source heterogeneous data fusion system is provided. The multi-source heterogeneous data fusion system includes:

[0046] Data acquisition module, used to obtain adjusted grid, GPS trajectory data, base station MR data, sensor data and user complaint log;

[0047] A spatiotemporal index table construction module, configured to construct a GPS trajectory spatiotemporal index table based on the adjusted grid and the GPS trajectory data; and to construct a base station MR spatiotemporal index table based on the adjusted grid and the base station MR data;

[0048] A first eigenvector calculation module is configured to perform spatiotemporal calibration by Kalman filtering based on the GPS trajectory spatiotemporal index table and the base station MR spatiotemporal index table to obtain a first eigenvector;

[0049] A second eigenvector calculation module, configured to calculate a second eigenvector based on the sensor data and the user complaint log;

[0050] The fused multi-source heterogeneous data calculation module is used to calculate the fused multi-source heterogeneous data based on the first eigenvector and the second eigenvector using the national secret SM4 algorithm.

[0051] This system obtains the adjusted grid, GPS trajectory data, base station MR data, sensor data and user complaint logs; constructs a GPS trajectory spatiotemporal index table based on the adjusted grid and GPS trajectory data; constructs a base station MR spatiotemporal index table based on the adjusted grid and base station MR data; performs spatiotemporal calibration through Kalman filtering based on the GPS trajectory spatiotemporal index table and the base station MR spatiotemporal index table to obtain the first eigenvector; this application achieves efficient data organization by dynamically calculating the spatiotemporal grid and constructing the index table; calculates the second eigenvector based on the sensor data and user complaint logs; and calculates the fused multi-source heterogeneous data based on the first and second eigenvectors through the national secret SM4 algorithm, avoiding data dislocation and improving analysis efficiency.

[0052] In a third aspect of the present application, an electronic device for fusing multi-source heterogeneous data is provided, comprising at least one control processor and a memory for communicating with the at least one control processor; the memory stores instructions that can be executed by the at least one control processor, and the instructions are executed by the at least one control processor so that the at least one control processor can execute the above-mentioned multi-source heterogeneous data fusion method.

[0053] In a fourth aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the above-mentioned multi-source heterogeneous data fusion method.

[0054] It should be noted that the beneficial effects between the second to fourth aspects of the present application and the prior art are the same as the beneficial effects between the above-mentioned multi-source heterogeneous data fusion system and the prior art, and will not be described in detail here.

[0055] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become obvious from the description below, or will be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:

[0057] Figure 1 This is a flow chart of a method for fusing multi-source heterogeneous data according to an embodiment of the present application;

[0058] Figure 2 This is a schematic diagram of the structure of an embodiment of a multi-source heterogeneous data fusion system provided by the present application;

[0059] Figure 3 It is a structural diagram of an embodiment of the electronic device provided by this application. DETAILED DESCRIPTION

[0060] The following describes in detail embodiments of the present application. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application and are not to be construed as limiting the present application.

[0061] In the description of this application, if there is a description of first, second, etc., it is only for the purpose of distinguishing technical features, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features or implicitly indicating the order of the indicated technical features.

[0062] In the description of this application, it should be understood that descriptions involving orientation, such as the orientation or positional relationship indicated by up, down, etc., are based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing this application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on this application.

[0063] In the description of this application, it should be noted that, unless otherwise clearly defined, terms such as setting, installing, and connecting should be understood in a broad sense, and technical personnel in the relevant technical field can reasonably determine the specific meaning of the above terms in this application based on the specific content of the technical solution.

[0064] With the widespread adoption of 5G communication technology, multi-source, heterogeneous data generated by mobile terminals has become a core analytical target for scenarios such as network optimization. This data exhibits strong spatiotemporal dynamics, diverse modalities, and privacy-sensitive characteristics. Efficiently integrating and securely utilizing this data has become a key challenge in improving 5G network performance and the precision of urban management.

[0065] Currently, existing technologies face the following bottlenecks when addressing the above requirements: the differences in spatiotemporal granularity of data from different sources cause alignment errors of up to 15%, making it difficult for traditional fixed-time window aggregation methods to adapt dynamically; and existing indexing technologies have query delays exceeding 200ms with an average daily data volume of 1 billion items, which cannot meet the millisecond-level optimization requirements of 5G.

[0066] In order to solve the above technical defects, the embodiments of the present application provide a method and system for fusing multi-source heterogeneous data.

[0067] See Figure 1 , is a flow chart of a method for fusing multi-source heterogeneous data provided by an embodiment of the present application, the method is applied to an electronic device, which may be a server, etc. Figure 1 As shown, the fusion method of multi-source heterogeneous data includes:

[0068] Step S101: Obtain adjusted grid, GPS trajectory data, base station MR data, sensor data, and user complaint logs;

[0069] Step S102: constructing a GPS trajectory spatiotemporal index table based on the adjusted grid and GPS trajectory data; constructing a base station MR spatiotemporal index table based on the adjusted grid and base station MR data;

[0070] Step S103: Based on the GPS trajectory spatiotemporal index table and the base station MR spatiotemporal index table, perform spatiotemporal calibration through Kalman filtering to obtain a first eigenvector;

[0071] Step S104: Calculate a second eigenvector based on the sensor data and the user complaint log;

[0072] Step S105: Based on the first eigenvector and the second eigenvector, the fused multi-source heterogeneous data is calculated using the national secret SM4 algorithm.

[0073] Specifically, the above-mentioned GPS trajectory data, base station MR data, sensor data and user complaint logs may include but are not limited to GPS trajectory, base station MR data, sensor data, Wi-Fi signaling, AP connection records, geographic information, network working parameters and user complaint logs. The data is then preprocessed, including unifying the timestamp to the UTC standard and converting the spatial coordinates to the WGS84 coordinate system; filtering GPS drift points based on the 3σ principle; and using a spatiotemporal interpolation algorithm to supplement the missing period data of the base station MR to obtain GPS trajectory data, base station MR data, sensor data and user complaint logs.

[0074] This method obtains the adjusted grid, GPS trajectory data, base station MR data, sensor data and user complaint log; constructs a GPS trajectory spatiotemporal index table based on the adjusted grid and GPS trajectory data; constructs a base station MR spatiotemporal index table based on the adjusted grid and base station MR data; performs spatiotemporal calibration through Kalman filtering based on the GPS trajectory spatiotemporal index table and the base station MR spatiotemporal index table to obtain a first eigenvector; this application achieves efficient data organization by dynamically calculating the spatiotemporal grid and constructing the index table; calculates the second eigenvector based on the sensor data and user complaint log; and calculates the fused multi-source heterogeneous data based on the first eigenvector and the second eigenvector through the national secret SM4 algorithm, thereby avoiding data dislocation and improving analysis efficiency.

[0075] In some embodiments, obtaining the adjusted grid in step S101 includes:

[0076] Step S201: Obtain the total number of users, the total number of grids, and the number of users in each grid within a preset time period;

[0077] Step S202: Calculate the user ratio of the corresponding grid based on the total number of grids and the number of users in each grid;

[0078] Step S203: Calculate the user distribution entropy value based on the total number of users and the corresponding grid user ratio using the following formula:

[0079]

[0080] Among them, H is the user distribution entropy value, N is the total number of users in the preset time period, and p i is the proportion of grid users corresponding to the i-th grid;

[0081] Step S204: Calculate the adjusted grid side length based on the user distribution entropy value using the following formula:

[0082] L grid =L max ·e -kH ;

[0083] Among them, L grid To adjust the grid side length, L max is the preset maximum grid side length, k is the preset adjustment coefficient;

[0084] Step S205 : adjusting the grid within the preset time period based on the adjusted grid side length to obtain an adjusted grid.

[0085] This application provides more accurate data for subsequent data calculations and improves the accuracy by dynamically adjusting the grid in real time.

[0086] In some embodiments, in step S102, a GPS track spatiotemporal index table is constructed based on the adjusted grid and the GPS track data; the step includes:

[0087] Step S301: Obtain the corresponding stay time and corresponding movement speed of each user in the adjusted grid;

[0088] Step S302: Calculate the grid weight based on the preset time period, the corresponding moving speed, and the corresponding stay time using the following formula:

[0089]

[0090] α+β=1

[0091] Where W is the grid weight, T stay is the corresponding stay time, T total is the preset time period, α is the first preset weight coefficient, β is the second preset weight coefficient, and v is the corresponding moving speed;

[0092] Step S303: constructing a GPS track spatiotemporal index table based on the adjusted grid, GPS track data and grid weight, wherein the GPS track spatiotemporal index table includes the grid number, timestamp, data source type and grid weight of the adjusted grid.

[0093] Specifically, in some embodiments, the above steps include data encoding: encoding the data into (grid code, timestamp, data source type and weight).

[0094] Storage structure:

[0095] B+ tree index: used for time range queries.

[0096] Inverted index: quickly locate data through grid ID (generation logic: establish an inverted index table based on the mapping relationship between grid ID and data records).

[0097] The grid weight is calculated based on the preset time period, the corresponding movement speed, and the corresponding stay time using the following formula:

[0098]

[0099] α+β=1

[0100] Where W is the grid weight, T stay is the corresponding stay time, T total is the preset time period, α is the first preset weight coefficient, β is the second preset weight coefficient, and v is the corresponding moving speed.

[0101] Output: A structured spatiotemporal index table (raster code, timestamp, data source type, and weight), supporting millisecond-level queries and used for priority analysis. The structured spatiotemporal index table consists of the encoded fields (raster code, timestamp, data source type, and weight) and is stored using a B+ tree and inverted index (default α = 0.7, β = 0.3).

[0102] Output: A structured spatiotemporal index table (raster code, timestamp, data source type, and weight), supporting millisecond-level queries for priority analysis. The structured spatiotemporal index table consists of the encoded fields (raster code, timestamp, data source type, and weight) and is stored using a B+ tree and inverted index.

[0103] In some embodiments, in step S102, constructing a base station MR spatiotemporal index table based on the adjusted grid and the base station MR data includes:

[0104] Step S401: construct a base station MR spatiotemporal index table based on the adjusted grid, base station MR data, and grid weights.

[0105] This application provides a more accurate data source for subsequent data fusion by constructing an index table, thereby improving the accuracy of data fusion.

[0106] In some embodiments, in step S103, based on the GPS trajectory spatiotemporal index table and the base station MR spatiotemporal index table, a spatiotemporal calibration is performed by Kalman filtering to obtain a first eigenvector, including:

[0107] Step S501: Based on the GPS trajectory spatiotemporal index table and the base station MR spatiotemporal index table, the spatiotemporal deviation is corrected by Kalman filtering to obtain calibrated spatiotemporal trajectory data;

[0108] Step S502: extract features from the calibrated spatiotemporal trajectory data to obtain a first feature vector.

[0109] Specifically, in some embodiments, the above steps include performing spatiotemporal calibration using Kalman filtering based on the GPS trajectory spatiotemporal index table and the base station MR spatiotemporal index table, and dynamically correcting the spatiotemporal deviation between GPS and base station MR using the following state equation to eliminate trajectory jitter caused by different sampling frequencies:

[0110]

[0111] in, represents the state estimate at the kth moment, A represents the state transfer matrix, B represents the control input matrix, u k Represents external control input, w k represents process noise.

[0112] Dynamic Time Warping (DTW) alignment: The time axis of the calibrated trajectory is stretched / compressed to ensure that the GPS trajectory is strictly aligned with the timestamp of the base station MR data, with an error rate of ε < 3%.

[0113] Generate the calibrated spatiotemporal trajectory features and obtain the quasi-spatiotemporal trajectory feature vector, including:

[0114] Position sequence (GID, t);

[0115] Movement state (speed v, direction θ);

[0116] Base station signal quality (RSRP, SINR).

[0117] In some embodiments, calculating the second feature vector based on the sensor data and the user complaint log in step S104 includes:

[0118] Step S601: extracting behavioral pattern features through a convolutional neural network model based on sensor data;

[0119] Step S602: Extracting complaint semantic features through natural language processing based on user complaint logs;

[0120] Step S603: Calculate the implicit weight value based on the behavior pattern characteristics and complaint semantic characteristics using the following formula:

[0121]

[0122] Among them, s i is the behavioral pattern characteristic, s j is the complaint semantic feature, w ij is the implicit weight value;

[0123] Step S604: Based on the implicit weight value, map the high-weight value feature to the corresponding adjusted grid to obtain a second feature vector.

[0124] Specifically, in some embodiments, the above steps include edge layer lightweighting processing:

[0125] Sensor data compression: Deploy the TensorFlow Lite model (structure: 1D CNN + quantization layer), with a compression rate of >80%:

[0126] X compressed =ReLU(W conw *X raw +b);

[0127] Among them, X compressed Represents the compressed sensor data, X ray represents the raw sensor data, W conv represents the one-dimensional convolution kernel weight matrix, and b represents the bias term.

[0128] Perform natural language processing on user complaint logs to extract keywords (such as "lag" and "poor signal") and associated location tags.

[0129] Calculate the implicit weights of sensor features and complaint semantics:

[0130]

[0131] Among them, s i is the behavioral pattern characteristic, s j is the complaint semantic feature, w ij is the implicit weight value;

[0132] Map high-weight feature pairs (such as "walking status" and "complaint at the mall entrance") to the corresponding grid, generate scenario-based feature labels, and obtain the second feature vector.

[0133] This application improves data accuracy by mining implicit associations.

[0134] In some embodiments, in step S105, the fused multi-source heterogeneous data is calculated based on the first eigenvector and the second eigenvector using the national secret SM4 algorithm, including:

[0135] Step S701: based on the timestamp and the grid number of the adjusted grid, concatenate the first eigenvector and the second eigenvector to obtain a concatenated eigenvector;

[0136] Step S702: The concatenated feature vectors are privacy-enhanced using the national secret SM4 algorithm and Intel SGX technology to obtain fused multi-source heterogeneous data.

[0137] Specifically, in some embodiments, the above step may be to concatenate the first feature vector and the second feature vector based on the timestamp and the grid number of the adjusted grid to obtain a concatenated feature vector;

[0138] Inject Laplace noise into GPS coordinates (privacy budget ε = 0.1):

[0139]

[0140] in, represents the noisy private data, x represents the original sensitive data, Lap represents the Laplace distribution generating function, Δf represents the data sensitivity, and ε represents the privacy budget.

[0141] Effect: The single point position is irreversibly blurred to a radius of 50m.

[0142] Using the national secret SM4 algorithm, the key is updated every 10 minutes, the encryption package structure is:

[0143] Ciphertext = SM4(K t , plaintext||timestamp||key version)

[0144] Paillier encryption is performed on sensitive fields (such as IMSI) and ciphertext addition is supported:

[0145] E(m1+m2)=E(m1)·E(m2)mod n 2

[0146] Where E(m) represents the homomorphic encryption result of plaintext m, and n represents the public key parameter of the Paillier algorithm.

[0147] Based on Intel SGX technology, data operations are verified with zero trust, audit log hashes are uploaded to the chain, and fused multi-source heterogeneous data is output.

[0148] Specifically, to facilitate understanding by those skilled in the art, a set of best embodiments are provided below:

[0149] 1. Data Acquisition

[0150] Obtain the adjusted grid, GPS trajectory data, base station MR data, sensor data, and user complaint logs, specifically:

[0151] Obtain the total number of users, the total number of grids, and the number of users in each grid within a preset time period;

[0152] Calculate the user ratio of the corresponding grid based on the total number of grids and the number of users in each grid;

[0153] The user distribution entropy value is calculated based on the total number of users and the corresponding grid user ratio using the following formula:

[0154]

[0155] Among them, H is the user distribution entropy value, N is the total number of users in the preset time period, and p i is the proportion of grid users corresponding to the i-th grid;

[0156] The adjusted grid side length is calculated based on the user distribution entropy value using the following formula:

[0157] L grid =L max ·e -kH ;

[0158] Among them, L grid To adjust the grid side length, L max is the preset maximum grid side length, k is the preset adjustment coefficient;

[0159] The grid within a preset time period is adjusted based on the adjusted grid side length to obtain the adjusted grid.

[0160] 2. Index table construction:

[0161] The GPS trajectory spatiotemporal index table is constructed based on the adjusted grid and GPS trajectory data; the base station MR spatiotemporal index table is constructed based on the adjusted grid and base station MR data, specifically:

[0162] Obtain the corresponding stay time and movement speed of each user in the adjusted grid;

[0163] The grid weight is calculated based on the preset time period, the corresponding movement speed, and the corresponding stay time using the following formula:

[0164]

[0165] α+β=1

[0166] Where W is the grid weight, T stay is the corresponding stay time, Ttotal is the preset time period, α is the first preset weight coefficient, β is the second preset weight coefficient, and v is the corresponding moving speed;

[0167] A GPS track spatiotemporal index table is constructed based on the adjusted grid, GPS track data and grid weight, wherein the GPS track spatiotemporal index table includes the grid number, timestamp, data source type and grid weight of the adjusted grid.

[0168] A base station MR spatiotemporal index table is constructed based on the adjusted grid, base station MR data, and grid weights.

[0169] 3. Calculation of the first eigenvector:

[0170] Based on the GPS trajectory spatiotemporal index table and the base station MR spatiotemporal index table, the spatiotemporal calibration is performed through Kalman filtering to obtain the first eigenvector, which is specifically:

[0171] Based on the GPS trajectory spatiotemporal index table and the base station MR spatiotemporal index table, the spatiotemporal deviation is corrected by Kalman filtering to obtain the calibrated spatiotemporal trajectory data;

[0172] Feature extraction is performed on the calibrated spatiotemporal trajectory data to obtain the first eigenvector.

[0173] 4. Calculation of the second eigenvector:

[0174] The second eigenvector is calculated based on sensor data and user complaint logs, specifically:

[0175] Based on sensor data, behavioral pattern features are extracted through a convolutional neural network model;

[0176] Based on user complaint logs, semantic features of complaints are extracted through natural language processing;

[0177] The implicit weight value is calculated based on the behavior pattern characteristics and complaint semantic characteristics using the following formula:

[0178]

[0179] Among them, s i is the behavioral pattern characteristic, s j is the complaint semantic feature, w ij is the implicit weight value;

[0180] Based on the implicit weight value, the high-weight value features are mapped to the corresponding adjusted grid to obtain the second feature vector.

[0181] 5. Data Fusion

[0182] Based on the first and second eigenvectors, the fused multi-source heterogeneous data is calculated using the national secret SM4 algorithm, specifically:

[0183] splicing the first eigenvector and the second eigenvector based on the timestamp and the grid number of the adjusted grid to obtain a spliced ​​eigenvector;

[0184] The privacy of the concatenated feature vectors is enhanced using the national secret SM4 algorithm and Intel SGX technology to obtain fused multi-source heterogeneous data.

[0185] In addition, refer to Figure 2 One embodiment of the present application provides a multi-source heterogeneous data fusion system, including a data acquisition module 1100, a spatiotemporal index table construction module 1200, a first eigenvector calculation module 1300, a second eigenvector calculation module 1400, and a fused multi-source heterogeneous data calculation module 1500, wherein:

[0186] The data acquisition module 1100 is used to obtain the adjusted grid, GPS trajectory data, base station MR data, sensor data and user complaint logs;

[0187] The spatiotemporal index table construction module 1200 is used to construct a GPS trajectory spatiotemporal index table based on the adjusted grid and GPS trajectory data; and to construct a base station MR spatiotemporal index table based on the adjusted grid and base station MR data;

[0188] The first eigenvector calculation module 1300 is used to perform spatiotemporal calibration through Kalman filtering based on the GPS trajectory spatiotemporal index table and the base station MR spatiotemporal index table to obtain a first eigenvector;

[0189] The second feature vector calculation module 1400 is used to calculate the second feature vector based on the sensor data and the user complaint log;

[0190] The fused multi-source heterogeneous data calculation module 1500 is used to calculate the fused multi-source heterogeneous data based on the first eigenvector and the second eigenvector using the national secret SM4 algorithm.

[0191] This system obtains the adjusted grid, GPS trajectory data, base station MR data, sensor data and user complaint logs; constructs a GPS trajectory spatiotemporal index table based on the adjusted grid and GPS trajectory data; constructs a base station MR spatiotemporal index table based on the adjusted grid and base station MR data; performs spatiotemporal calibration through Kalman filtering based on the GPS trajectory spatiotemporal index table and the base station MR spatiotemporal index table to obtain the first eigenvector; this application achieves efficient data organization by dynamically calculating the spatiotemporal grid and constructing the index table; calculates the second eigenvector based on the sensor data and user complaint logs; and calculates the fused multi-source heterogeneous data based on the first and second eigenvectors through the national secret SM4 algorithm, avoiding data dislocation and improving analysis efficiency.

[0192] It should be noted that this system embodiment and the above-mentioned method embodiment are based on the same inventive concept, so the relevant content of the above-mentioned method embodiment is also applicable to this system embodiment and will not be repeated here.

[0193] Figure 3 A schematic diagram of the hardware structure for the fusion of multi-source heterogeneous data provided by an embodiment of the present application is shown.

[0194] The multi-source heterogeneous data fusion device may include a processor 301 and a memory 302 storing computer program instructions.

[0195] Specifically, the processor 301 may include a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.

[0196] The memory 302 may include a large capacity memory for data or instructions. By way of example and not limitation, the memory 302 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory 302 may include removable or non-removable (or fixed) media. Where appropriate, the memory 302 may be inside or outside the integrated gateway disaster recovery device. In a specific embodiment, the memory 302 is a non-volatile solid-state memory.

[0197] In some embodiments, the memory 302 may include read-only memory (ROM), random access memory (RAM), magnetic disk storage media devices, optical storage media devices, flash memory devices, electrical, optical, or other physical / tangible memory storage devices. Thus, generally, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to an aspect of the present disclosure.

[0198] The processor 301 reads and executes computer program instructions stored in the memory 302 to implement any one of the multi-source heterogeneous data fusion methods in the above embodiments.

[0199] In one example, the multi-source heterogeneous data fusion device may further include a communication interface 303 and a bus 310. Figure 3As shown, the processor 301 , the memory 302 , and the communication interface 303 are connected via a bus 310 and communicate with each other.

[0200] The communication interface 303 is mainly used to implement communication between various modules, devices, units and / or equipment in the embodiments of the present application.

[0201] Bus 310 includes hardware, software or both, and the parts of the fusion equipment of multi-source heterogeneous data are coupled to each other.For example, but not limitation, bus may include accelerated graphics port (AGP) or other graphics bus, enhanced industry standard architecture (EISA) bus, front side bus (FSB), hypertransport (HT) interconnection, industry standard architecture (ISA) bus, infinite bandwidth interconnection, low pin count (LPC) bus, memory bus, micro channel architecture (MCA) bus, peripheral component interconnection (PCI) bus, PCI-Express (PCI-X) bus, serial advanced technology attachment (SATA) bus, video electronics standard association local (VLB) bus or other suitable bus or two or more of these combinations.In suitable cases, bus 310 may include one or more buses.Although the present application embodiment describes and shows specific bus, the application considers any suitable bus or interconnection.

[0202] The multi-source heterogeneous data fusion device can execute the multi-source heterogeneous data fusion method in the embodiment of the present application based on the three-dimensional design model, thereby realizing the fusion of Figure 1 and Figure 2 Described is a method and system for fusing multi-source heterogeneous data.

[0203] In addition, in conjunction with the multi-source heterogeneous data fusion method in the above embodiments, embodiments of the present application may provide a computer storage medium for implementation. The computer storage medium stores computer program instructions; when the computer program instructions are executed by a processor, any of the multi-source heterogeneous data fusion methods in the above embodiments is implemented.

[0204] It should be understood that the present application is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, a detailed description of known methods is omitted here. In the above embodiments, several specific steps are described and illustrated as examples. However, the method process of the present application is not limited to the specific steps described and illustrated. Those skilled in the art can make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present application.

[0205] The functional blocks shown in the above block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of the present application are programs or code segments that are used to perform the required tasks. Programs or code segments can be stored in machine-readable media, or transmitted on a transmission medium or a communication link by a data signal carried in a carrier wave. "Machine-readable media" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROMs, flash memories, erasable ROMs (EROMs), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segments can be downloaded via computer networks such as the Internet, intranets, etc.

[0206] It should also be noted that the exemplary embodiments mentioned in this application describe some methods or systems based on a series of steps or devices. However, this application is not limited to the order of the above steps. In other words, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.

[0207] Aspects of the present disclosure have been described above with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer or other programmable data processing device to produce a machine so that these instructions executed by the processor of the computer or other programmable data processing device enable the implementation of the function / action specified in one or more boxes of the flowchart and / or block diagram. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor or a field programmable logic circuit. It is also understood that each box in the block diagram and / or flowchart and the combination of the boxes in the block diagram and / or flowchart can also be implemented by dedicated hardware that performs the specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions.

[0208] The above description is only a specific embodiment of the present application. Those skilled in the art will clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present application, and these modifications or replacements should be included in the scope of protection of the present application.

Claims

1. A method for fusing multi-source heterogeneous data, characterized in that: The multi-source heterogeneous data fusion method includes: Obtain adjusted grid, GPS trajectory data, base station MR data, sensor data, and user complaint logs; Constructing a GPS track spatiotemporal index table based on the adjusted grid and the GPS track data; constructing a base station MR spatiotemporal index table based on the adjusted grid and the base station MR data; Based on the GPS trajectory spatiotemporal index table and the base station MR spatiotemporal index table, performing spatiotemporal calibration through Kalman filtering to obtain a first eigenvector; calculating a second feature vector based on the sensor data and the user complaint log; Based on the first eigenvector and the second eigenvector, the fused multi-source heterogeneous data is calculated using the national secret SM4 algorithm.

2. The method for fusing multi-source heterogeneous data according to claim 1, characterized in that: The step of obtaining the adjusted grid includes: Obtain the total number of users, the total number of grids, and the number of users in each grid within a preset time period; Calculating a user ratio of a corresponding grid based on the total number of grids and the number of users in each grid; The user distribution entropy value is calculated based on the total number of users and the corresponding grid user ratio using the following formula: Among them, H is the user distribution entropy value, N is the total number of users in the preset time period, and p i is the proportion of grid users corresponding to the i-th grid; The adjusted grid side length is calculated based on the user distribution entropy value using the following formula: THE grid =L max ·And -kH ; Among them, L grid To adjust the grid side length, L max is the preset maximum grid side length, k is the preset adjustment coefficient; The grid within a preset time period is adjusted based on the adjusted grid side length to obtain the adjusted grid.

3. The method for fusing multi-source heterogeneous data according to claim 2, characterized in that: The step of constructing a GPS track spatiotemporal index table based on the adjusted grid and the GPS track data comprises: Obtain the corresponding stay time and movement speed of each user in the adjusted grid; The grid weight is calculated based on the preset time period, the corresponding moving speed and the corresponding stay time using the following formula: α+β=1 Where W is the grid weight, T stay is the corresponding stay time, T total is the preset time period, α is the first preset weight coefficient, β is the second preset weight coefficient, and v is the corresponding moving speed; The GPS track spatiotemporal index table is constructed based on the adjusted grid, the GPS track data and the grid weight, wherein the GPS track spatiotemporal index table includes the grid number, timestamp, data source type and grid weight of the adjusted grid.

4. The method for fusing multi-source heterogeneous data according to claim 3, characterized in that: The constructing a base station MR spatiotemporal index table based on the adjusted grid and the base station MR data includes: The base station MR spatiotemporal index table is constructed based on the adjusted grid, the base station MR data and the grid weight.

5. The method for fusing multi-source heterogeneous data according to claim 1, characterized in that: The method of performing spatiotemporal calibration based on the GPS trajectory spatiotemporal index table and the base station MR spatiotemporal index table through Kalman filtering to obtain a first eigenvector includes: Based on the GPS trajectory spatiotemporal index table and the base station MR spatiotemporal index table, the spatiotemporal deviation is corrected by Kalman filtering to obtain calibrated spatiotemporal trajectory data; Feature extraction is performed on the calibrated spatiotemporal trajectory data to obtain the first feature vector.

6. The method for fusing multi-source heterogeneous data according to claim 1, characterized in that: The calculating a second feature vector based on the sensor data and the user complaint log includes: Based on the sensor data, extracting behavioral pattern features through a convolutional neural network model; Based on the user complaint log, extracting complaint semantic features through natural language processing; The implicit weight value is calculated based on the behavior pattern characteristics and the complaint semantic characteristics using the following formula: Among them, s i is the behavioral pattern characteristic, s j is the complaint semantic feature, w ij is the implicit weight value; Based on the implicit weight value, the high-weight value features are mapped to the corresponding adjusted grid to obtain the second feature vector.

7. The method for fusing multi-source heterogeneous data according to claim 1, characterized in that: The calculating of the fused multi-source heterogeneous data based on the first eigenvector and the second eigenvector by using the national secret SM4 algorithm includes: splicing the first feature vector and the second feature vector based on a timestamp and a grid number of the adjusted grid to obtain a spliced ​​feature vector; The concatenated feature vectors are privacy-enhanced using the national secret SM4 algorithm and Intel SGX technology to obtain the fused multi-source heterogeneous data.

8. A multi-source heterogeneous data fusion system, characterized by: The multi-source heterogeneous data fusion system includes: Data acquisition module, used to obtain adjusted grid, GPS trajectory data, base station MR data, sensor data and user complaint log; A spatiotemporal index table construction module, configured to construct a GPS trajectory spatiotemporal index table based on the adjusted grid and the GPS trajectory data; and to construct a base station MR spatiotemporal index table based on the adjusted grid and the base station MR data; A first eigenvector calculation module is configured to perform spatiotemporal calibration by Kalman filtering based on the GPS trajectory spatiotemporal index table and the base station MR spatiotemporal index table to obtain a first eigenvector; A second eigenvector calculation module, configured to calculate a second eigenvector based on the sensor data and the user complaint log; The fused multi-source heterogeneous data calculation module is used to calculate the fused multi-source heterogeneous data based on the first eigenvector and the second eigenvector using the national secret SM4 algorithm.

9. A device for fusion of multi-source heterogeneous data, characterized in that: It includes at least one control processor and a memory for communicating with the at least one control processor; the memory stores instructions that can be executed by the at least one control processor, and the instructions are executed by the at least one control processor to enable the at least one control processor to execute a multi-source heterogeneous data fusion method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the multi-source heterogeneous data fusion method according to any one of claims 1 to 7.