Multi-source data fusion method and system based on cloud computing
By using health assessment models and semantic consistency methods, this approach addresses existing technical issues. It generates health scores and labels, and, based on semantic consistency, resolves these issues, enabling scientific adjudication of conflicting data. This improves the reliability and semantic consistency of the fusion results, enhances the fusion effect, and adapts to the reliability and semantic consistency of the fusion results. It also strengthens rule adaptability and scenario flexibility.
Patent Information
- Application Number
- CN202510068063.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2045-01-16
AI Technical Summary
Existing technologies lack a dynamic data health assessment mechanism in multi-source data fusion, cannot label and process low-quality data in real time, and have insufficient semantic matching processing accuracy, making it difficult to guarantee the accuracy and reliability of the fusion results.
Health scores and labels are generated through a health assessment model. Fusion rules are constructed based on semantic enhancement and conflicting data are adjudicated using a semantic consistency index. The fusion process is optimized by combining knowledge graph reasoning and semantic similarity calculation to improve accuracy and reliability.
It enables scientific adjudication of conflicting data, improves the reliability and semantic consistency of fusion results, enhances rule adaptability and scenario flexibility, and is suitable for complex data processing needs.
Smart Images

Figure CN119989267B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multi-source data fusion technology, and more specifically, to a multi-source data fusion method and system based on cloud computing. Background Technology
[0002] With the rapid development of cloud computing technology, multi-source data fusion has been widely applied in fields such as smart cities, traffic management, and healthcare. Multi-source data fusion supports data-driven decision-making by integrating heterogeneous data from different sources, formats, and semantics. However, the diversity and dynamism of heterogeneous data place extremely high demands on fusion methods, especially in areas such as real-time data processing, semantic consistency maintenance, and conflict resolution. Traditional fusion methods often struggle to cope with complex and ever-changing scenarios.
[0003] However, in practical use, it still has the disadvantage of poor fusion effect. For example, it lacks a dynamic data health assessment mechanism and cannot mark and process low-quality data in real time. In the semantic matching of heterogeneous data, the traditional method is not accurate enough in handling missing fields and inconsistent data, making it difficult to guarantee the accuracy of fusion. In the face of conflicts between multi-source data, it lacks a fusion rule adjudication mechanism based on confidence and semantic similarity, which can easily lead to a decrease in the reliability of fusion results. Summary of the Invention
[0004] To overcome the aforementioned deficiencies of the prior art, this invention provides a cloud computing-based multi-source data fusion method and system, which generates health scores and labels through a health assessment model, constructs fusion rules based on semantic enhancement, and uses a semantic consistency index to adjudicate conflicting data, thereby solving the problems mentioned in the background art.
[0005] To achieve the above objectives, the present invention provides the following technical solution: a multi-source data fusion method based on cloud computing, comprising:
[0006] Step S001: Obtain real-time data from multiple data sources through cloud computing interfaces, construct multi-dimensional feature vectors of the real-time data, generate health scores using a health assessment model, and generate health labels for each data point based on the health scores.
[0007] Step S002: Perform format conversion and semantic enhancement on the data marked as healthy to generate semantically enhanced structured data;
[0008] Step S003: Based on the semantically enhanced structured data, construct fusion rules through knowledge graph reasoning and calculate the confidence of the fusion rules; construct a semantic consistency index based on the confidence of the fusion rules and semantic similarity, and apply the semantic consistency index to adjudicate the fusion result for conflicting data;
[0009] Step S004: Verify the quality of the fusion results using historical data and business rules. If the fusion results do not meet the preset requirements, trigger the feedback optimization process.
[0010] Through the above steps, dynamic multi-source data fusion and full-process optimization are achieved.
[0011] Preferably, the health assessment model includes a time-based dynamic correction term, which refers to dynamically adjusting the health score based on the time characteristics of the data; the health label is used to mark the status and abnormality category of the data.
[0012] Preferably, in step S002, format conversion refers to structural standardization processing through pattern matching algorithm; semantic enhancement refers to constructing semantic associations between data based on a semantic embedding model optimized for scene features. The semantic embedding model dynamically adjusts the semantic matching strategy according to the abnormal category of health label and dynamically completes fields that are missing or inconsistent.
[0013] Preferably, constructing fusion rules based on semantically enhanced structured data includes the following steps:
[0014] Knowledge graphs are generated based on semantic relationships between data using knowledge graph construction technology.
[0015] Input the defined logical relationships, use logical reasoning techniques to infer the logical connections between data from the knowledge graph, and generate fusion rules;
[0016] Calculation of weight and confidence of fusion rules: Combine health labels, data source confidence and semantic features to calculate the weight and confidence of fusion rules.
[0017] Preferably, the method for identifying conflicting data is as follows:
[0018] Step S301, Data Matching and Semantic Alignment: Use a semantic embedding model to perform semantic representation on multi-source data, perform semantic alignment on data with the same topic, and output semantically aligned data;
[0019] Step 302, Difference Detection: For semantically aligned data, determine whether there are descriptive differences among multi-source data under the same topic, and output the difference index; Step 303, If the difference index exceeds the preset threshold, it is determined to be conflicting data.
[0020] Preferably, the quality verification of the fusion results is used to evaluate the accuracy and reliability of the fusion results. Based on historical data, business rules, and verification scoring indicators, the fusion results are quantitatively evaluated. The core of the verification model is to calculate the verification score Qk by analyzing the deviation between the fusion results and the expected target, and thereby determine the quality and optimization direction of the fusion process. This includes the following steps:
[0021] Step S401: Load historical data and business rules as benchmarks. Historical data is used to provide a reference standard for the fusion results, and business rules are used to define the target accuracy and deviation range. Input historical fusion result data and business rules, analyze the statistical characteristics of historical data, and extract the requirements for result accuracy and deviation in the business rules. Output the benchmark standard data and accuracy target for the validation model. The accuracy target defines the tolerance range of scoring deviation and is applied to the threshold comparison of the validation score.
[0022] Step S402: Format the fusion result to be verified to ensure comparability with the benchmark standard data and rules; clean outliers and redundant entries in the current fusion result data, unify the format, ensure that the data dimensions and features are consistent with the benchmark standard data, and output the cleaned fusion result data.
[0023] Step S403: Calculate the verification score: Calculate the verification score based on the cleaned fusion result data to quantify the reliability of the fusion result; output the comprehensive verification score of the fusion result through the following model:
[0024]
[0025] Among them, Acc t The accuracy of record t This represents the fusion result data; Indicates benchmark data; Q k w represents the validation score of the current fusion result. t This represents the weight of the t-th record;
[0026] Step S404: Compare the comprehensive verification score with the preset threshold to determine whether the fusion result meets the requirements. If the comprehensive verification score is higher than or equal to the threshold, the fusion result is considered qualified and the verification is passed. If the comprehensive verification score is lower than the threshold, the fusion result is considered unqualified and the feedback optimization process is triggered.
[0027] Preferably, the feedback optimization process includes health assessment model parameter optimization, semantic embedding model adjustment, and fusion rule optimization, specifically:
[0028] Optimize the time dynamic correction term of the health assessment model to improve its adaptability to dynamic fluctuations in data; reassess abnormal situations in data health scores (such as the proportion of low-scoring data or missing fields), adjust the dynamic correction term and abnormal compensation coefficient in the scoring formula, and optimize feature weight allocation to improve the reliability of data health scores and ensure accurate assessment of data quality in subsequent fusion processes.
[0029] The scene feature priority rules of the semantic embedding model are adjusted to enhance the semantic association ability of specific scenes. By introducing domain-specific corpora to improve the training data of the semantic embedding model, adjusting the feature priority weights of the embedding vectors, and adding contextual features, the semantic embedding model is adjusted to a form that is more suitable for specific application scenarios, thereby improving the accuracy of semantic similarity calculation.
[0030] Improve the confidence calculation logic of the fusion rules to enhance the accuracy and consistency of conflict data adjudication;
[0031] The confidence formula for fusion rules is redefined, the weights of fusion rules are dynamically adjusted, and the rule logic is expanded to improve the ability of fusion rules to adjudicate conflicting data.
[0032] Using the optimized health assessment model, semantic embedding model, and fusion rules, the semantic consistency index is recalculated and compared with a threshold. If it is higher than the threshold, the optimization is confirmed to be effective; otherwise, the optimization scheme is adjusted to form a dynamic closed-loop optimization mechanism.
[0033] Preferably, the application semantic consistency index adjudication fusion result includes:
[0034] If the semantic consistency index is not lower than the set threshold, the current fusion rule is applied to the conflicting data, and the corresponding fusion result is selected as the final output; if the semantic consistency index is lower than the set threshold, it indicates that the current fusion rule is invalid, and the fusion rule needs to be optimized or a new fusion rule needs to be selected.
[0035] Preferably, step S001 includes the following:
[0036] Let i represent the sequential number of the data, m be the number of data features included in each data entry, and j represent the sequential number of the data features; let n be the number of fields included in each data entry, and k represent the sequential number of the fields. The health score is calculated using the following health assessment model.
[0037]
[0038] φ t =δ*|t a -t0|
[0039] Among them, S i For the health score of the i-th data item; v ij Let α be the j-th data feature of the i-th data; j φ represents the feature weight corresponding to the j-th data feature, indicating the degree of influence of the data feature on the score; t This is the time-based dynamic correction term, where δ is the time correction coefficient, and t at0 is the current time, t0 is the data collection time; ψ is the anomaly compensation coefficient, used to adjust for the impact of anomalies in the scoring; Var(f k ) represents the variance of the k-th field, measuring the field's volatility.
[0040] Preferably, the semantic consistency index is obtained in the following way:
[0041] Let r represent the sequential number of the fusion rule, and Ng represent the number of fusion rules; C r Let W be the confidence level of the r-th fusion rule. r To integrate rule weights, the semantic consistency index is calculated using the following formula:
[0042]
[0043] The confidence level of the fusion rule refers to the ratio of the number of valid fusions to the total number of fusions.
[0044] Preferably, the semantic similarity is obtained in the following way:
[0045] Let a set of conflicting data be denoted as (D) i D s ), measuring conflict data through semantic similarity (D i D s The semantic similarity of data is used to assist in determining conflicting data. The semantic similarity Sim(D) is calculated using the following formula. i D s ):
[0046]
[0047] Among them, E ′ (D i ) and E′(D s () are data D i and D s semantic embedding vector; ∥E ′ (D i )∥ represents the vector magnitude, used for normalization; μ is the distance weighting coefficient; Dist(D i D s ) represents the feature distance, f ij and f sj Data D i and D s The j-th feature value, where m is the total number of features.
[0048] To achieve the above objectives, the present invention provides the following technical solution: a cloud computing-based multi-source data fusion system, comprising:
[0049] The data acquisition module acquires real-time data from multiple data sources through a cloud computing interface, constructs a multi-dimensional feature vector of the real-time data, generates a health score using a health assessment model, and generates a health label for each data point based on the health score.
[0050] The data processing module is used to perform format conversion and semantic enhancement on data marked as healthy, generating semantically enhanced structured data;
[0051] The rule construction and adjudication module constructs fusion rules and calculates the confidence of the fusion rules based on the semantically enhanced structured data through knowledge graph reasoning; it constructs a semantic consistency index based on the confidence of the fusion rules and semantic similarity, and applies the semantic consistency index to adjudicate the fusion result for conflicting data;
[0052] The verification and optimization module verifies the quality of the fusion results using historical data and business rules. If the fusion results do not meet the preset requirements, a feedback optimization process is triggered.
[0053] The technical effects and advantages of this invention are as follows:
[0054] (1) This invention provides a cloud computing-based multi-source data fusion method. By introducing a semantic consistency index and using the weighted calculation of rule confidence, rule weight and semantic similarity, it realizes the scientific adjudication of conflicting data. The semantic consistency index can dynamically evaluate the effectiveness of the fusion rules and ensure that the selected or fused data results have higher reliability and semantic consistency when there are conflicts between multi-source data. Compared with traditional methods, it not only improves the accuracy of conflicting data adjudication, but also enhances rule adaptability and scenario flexibility, and is suitable for complex data processing needs in multiple scenarios such as traffic flow prediction and medical data analysis.
[0055] (2) This invention provides a cloud computing-based multi-source data fusion method. By designing a closed-loop optimization process that optimizes health assessment model parameters, adjusts semantic embedding model, and improves fusion rules, the key parameters of each module can be dynamically adjusted according to the results of verification scoring. When the quality of the fusion result does not meet the requirements, the optimization process can specifically improve the accuracy of data health assessment, the semantic association capability of semantic embedding model, and the adjudication efficiency of fusion rules, thereby ensuring continuous improvement of the quality of the fusion process. The closed-loop mechanism provides adaptive optimization capability for multi-source data fusion, which greatly improves the system's adaptability and processing efficiency to dynamic scenarios. Attached Figure Description
[0056] Figure 1 This is a flowchart of the multi-source data fusion method of the present invention.
[0057] Figure 2 This is a block diagram of the multi-source data fusion system of the present invention. Detailed Implementation
[0058] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0059] At the same time, it should be understood that, for ease of description, the dimensions of the various parts shown in the accompanying drawings are not drawn according to actual scale.
[0060] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit the scope of this application and its application or use.
[0061] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered part of the specification.
[0062] Example 1, see Figure 1 The flowchart of a multi-source data fusion method is provided in this embodiment of the invention. The method includes the following steps:
[0063] Step S001: Obtain real-time data from multiple data sources through a cloud computing interface (the real-time data includes each data item, transmission rate, data integrity, and timestamp information), construct a multi-dimensional feature vector of the real-time data, and generate a health score using a health assessment model (the health score is used to reflect the comprehensive quality index of each data item), and generate a health label for each data item based on the health score.
[0064] Step S002: Perform format conversion and semantic enhancement on the data marked as healthy to generate semantically enhanced structured data;
[0065] Step S003: Based on the semantically enhanced structured data, construct fusion rules through knowledge graph reasoning and calculate the confidence of the fusion rules; construct a semantic consistency index based on the confidence of the fusion rules and semantic similarity, and apply the semantic consistency index to adjudicate the fusion result for conflicting data;
[0066] To explain, conflicting data refers to data from different data sources that provide inconsistent descriptions of the same event or object; for example, two data sources may provide different values for traffic flow on a certain road segment, or conflicting descriptions of weather conditions at a certain point in time.
[0067] Step S004: Verify the quality of the fusion results using historical data and business rules. If the fusion results do not meet the preset requirements, trigger the feedback optimization process.
[0068] Through the above steps, dynamic multi-source data fusion and full-process optimization are achieved.
[0069] In this embodiment of the invention, it is necessary to further explain that the health assessment model includes a time dynamic correction term, which refers to dynamically adjusting the health score according to the time characteristics of the data; the health label is used to mark the status of the data (such as healthy data and abnormal data) and the abnormality category (such as transmission delay, field missing).
[0070] In this embodiment of the invention, it is necessary to further explain that the format conversion in step S002 refers to the structural standardization process performed by the pattern matching algorithm; the semantic enhancement refers to the construction of semantic associations between data based on the semantic embedding model optimized by the scene features. The semantic embedding model dynamically adjusts the semantic matching strategy according to the abnormal category of the health label and dynamically completes the field in case of missing or inconsistent fields.
[0071] In this embodiment of the invention, it needs to be further explained that, based on the semantically enhanced structured data, constructing the fusion rules includes the following steps:
[0072] Knowledge graphs are generated based on semantic relationships between data using knowledge graph construction techniques (such as RDF and SPARQL).
[0073] Input the defined logical relationships, and use logical reasoning techniques (such as rule reasoning and OWL inferencer) to infer the logical connections between data from the knowledge graph to generate fusion rules;
[0074] Calculation of weight and confidence of fusion rules: Combine health labels, data source confidence and semantic features to calculate the weight and confidence of fusion rules.
[0075] Explanation of how the weights of the fusion rules are obtained:
[0076] The data is taken as input, including a set of integrated rules, health tags, and semantically enhanced data.
[0077] The weights of the fusion rules are calculated based on health labels and semantically enhanced features, and the final output is a set of fusion rules with weights.
[0078] In a further embodiment, health scores and semantic similarity of the data are considered, and weighted averages are calculated separately. A weight adjustment coefficient is introduced to further adjust the weight allocation of health scores and semantic similarity. For ease of understanding, the specific calculation formula is W = α1·H + α2·S, where W represents the weight of the fusion rule, H represents the weighted average of the health scores of the data involved in the fusion rule, S represents the weighted average of the semantic similarity of the data involved in the fusion rule, and α1 and α2 are weight adjustment coefficients.
[0079] The confidence level of the fusion rule is obtained as follows:
[0080] The set of fusion rules and their weights are used as input;
[0081] The confidence score of the rules is calculated based on the historical performance of the data source and the adaptability of the fusion rules; the final output is the confidence score of the fusion rules.
[0082] For ease of understanding, the contribution of each data source to the rule confidence and the contribution weight of the data sources associated with the fused rule are considered; that is, the fused rule confidence is a weighted average of the confidence and contribution weight of multiple data sources.
[0083] In this embodiment of the invention, it needs to be further explained that the method for identifying conflicting data is as follows:
[0084] Step S301, Data Matching and Semantic Alignment: Use a semantic embedding model to perform semantic representation on multi-source data, perform semantic alignment on data with the same topic, and output semantically aligned data;
[0085] Step 302, Difference Detection: For semantically aligned data, determine whether there are descriptive differences among multi-source data under the same topic, and output difference indicators;
[0086] For example, in traffic data scenarios, different data sources provide different values for traffic flow on the same road segment; the difference index is the ratio of the numerical difference to the total number of fields.
[0087] Step 303: If the difference index exceeds the preset threshold, it is determined to be conflicting data.
[0088] In this embodiment of the invention, it is necessary to further explain that the quality verification of the fusion result is used to evaluate the accuracy and reliability of the fusion result. The fusion result is quantitatively evaluated based on historical data, business rules, and verification scoring indicators. The core of the verification model is to calculate the verification score by analyzing the deviation between the fusion result and the expected target, and thereby determine the quality and optimization direction of the fusion process. This includes the following steps:
[0089] Step S401: Load historical data and business rules as a benchmark. Historical data is used to provide a reference standard for the fusion results, and business rules are used to define the target accuracy and deviation range. Input historical fusion result data and business rules, analyze the statistical characteristics of historical data (such as distribution and fluctuation range), and extract the requirements for result accuracy and deviation (such as tolerance error and time constraints) in the business rules. Output the benchmark standard data and accuracy target for the validation model. The accuracy target defines the tolerance range for scoring deviation and is applied to the threshold comparison of the validation score.
[0090] Step S402: Format the fusion result to be verified to ensure comparability with the benchmark standard data and rules; clean outliers and redundant entries in the current fusion result data, unify the format, ensure that the data dimensions and features are consistent with the benchmark standard data, and output the cleaned fusion result data.
[0091] Step S403: Calculate the verification score: Calculate the verification score based on the cleaned fusion result data to quantify the reliability of the fusion result; output the comprehensive verification score of the fusion result through the following model:
[0092]
[0093] Among them, Acc t The accuracy of record t This represents the fusion result data; Indicates benchmark data; Q k w represents the validation score of the current fusion result. t This represents the weight of the t-th record;
[0094] Step S404: Compare the comprehensive verification score with the preset threshold to determine whether the fusion result meets the requirements. If the comprehensive verification score is higher than or equal to the threshold, the fusion result is considered qualified and the verification is passed. If the comprehensive verification score is lower than the threshold, the fusion result is considered unqualified and the feedback optimization process is triggered.
[0095] In this embodiment of the invention, it should be further explained that the feedback optimization process includes health assessment model parameter optimization, semantic embedding model adjustment, and fusion rule optimization, specifically:
[0096] Optimize the time dynamic correction term of the health assessment model to improve its adaptability to dynamic fluctuations in data; reassess abnormal situations in data health scores (such as the proportion of low-scoring data or missing fields), adjust the dynamic correction term and abnormal compensation coefficient in the scoring formula, and optimize feature weight allocation to improve the reliability of data health scores and ensure accurate assessment of data quality in subsequent fusion processes.
[0097] The scene feature priority rules of the semantic embedding model are adjusted to enhance the semantic association ability of specific scenes. By introducing domain-specific corpora to improve the training data of the semantic embedding model, adjusting the feature priority weights of the embedding vectors, and adding contextual features, the semantic embedding model is adjusted to a form that is more suitable for specific application scenarios, thereby improving the accuracy of semantic similarity calculation.
[0098] Improve the confidence calculation logic of the fusion rules to enhance the accuracy and consistency of conflict data adjudication;
[0099] Redefine the confidence formula for fusion rules (introduce dynamic factors for data latency or fluctuation), dynamically adjust the weights of fusion rules (correct them according to the performance of different scenarios), and expand the rule logic (such as adding time or domain constraints) to improve the ability of fusion rules to adjudicate conflicting data.
[0100] Using the optimized health assessment model, semantic embedding model, and fusion rules, the semantic consistency index is recalculated and compared with a threshold. If it is higher than the threshold, the optimization is confirmed to be effective; otherwise, the optimization scheme is adjusted to form a dynamic closed-loop optimization mechanism.
[0101] In this embodiment of the invention, it needs to be further explained that the application semantic consistency index adjudication fusion result includes:
[0102] If the semantic consistency index is not lower than the set threshold, the current fusion rule is applied to the conflicting data, and the corresponding fusion result is selected as the final output; if the semantic consistency index is lower than the set threshold, it indicates that the current fusion rule is invalid, and the fusion rule needs to be optimized or a new fusion rule needs to be selected.
[0103] In this embodiment of the invention, step S001 includes the following:
[0104] Let i represent the sequential number of the data, m be the number of data features included in each data entry, and j represent the sequential number of the data features; let n be the number of fields included in each data entry, and k represent the sequential number of the fields. The health score is calculated using the following health assessment model.
[0105]
[0106] φ t =δ*|t a -t0|
[0107] Among them, S i For the health score of the i-th data item; v ij Let α be the j-th data feature of the i-th data; j φ represents the feature weight corresponding to the j-th data feature, indicating the degree of influence of the data feature on the score;t This is the time-based dynamic correction term, where δ is the time correction coefficient, and t a t0 is the current time, t0 is the data collection time; ψ is the anomaly compensation coefficient, used to adjust for the impact of anomalies in the scoring; Var(f k ) represents the variance of the k-th field, measuring the field's volatility.
[0108] For highly dynamic scenarios (such as transportation or finance), the time dynamic correction term can be further optimized using the following formula:
[0109]
[0110] Among them, T max This indicates the maximum value within a time range.
[0111] In this embodiment of the invention, it needs to be further explained that the semantic consistency index is obtained in the following way:
[0112] Let r represent the sequential number of the fusion rule, and Ng represent the number of fusion rules; C r Let W be the confidence level of the r-th fusion rule. r To integrate rule weights, the semantic consistency index is calculated using the following formula:
[0113]
[0114] The confidence level of the fusion rule refers to the ratio of the number of valid fusions to the total number of fusions.
[0115] The semantic consistency index is primarily designed for conflicting data, but in the following scenarios, it can be extended to be part of global data fusion:
[0116] Data fusion involves multi-level processing:
[0117] The consistency index is used to adjudicate conflicting data.
[0118] The data after the ruling will be integrated with the non-conflicting data as a whole;
[0119] High-requirement data fusion scenarios: If all data needs to be strictly reviewed for consistency (such as in medical and financial scenarios), the semantic consistency index can be applied globally to verify the applicability of rules to all data.
[0120] In this embodiment of the invention, it needs to be further explained that the semantic similarity is obtained in the following way:
[0121] Let a set of conflicting data be denoted as (D) i D s ), measuring conflict data through semantic similarity (Di D s The semantic similarity of data is used to assist in determining conflicting data. The semantic similarity Sim(D) is calculated using the following formula. i D s ):
[0122]
[0123] Among them, E ′ (D i ) and E′(D s () are data D i and D s semantic embedding vector; ∥E ′ (D i )∥ represents the vector magnitude, used for normalization; μ is the distance weighting coefficient; Dist(D i D s ) represents the feature distance, f ij and f sj Data D i and D s The j-th feature value, where m is the total number of features.
[0124] In summary, this invention calculates a semantic consistency index through a comprehensive evaluation of semantic similarity, rule confidence, and rule weight, and quantitatively adjudicates conflicting data. It includes optimizing health assessment model parameters, adjusting the semantic embedding model, and improving the fusion rule logic. Closed-loop optimization is formed by dynamically adjusting the key parameters and logic of each module. Health scores are dynamically corrected based on data time characteristics, effectively addressing the volatility and latency issues of real-time data and significantly improving the overall quality of data fusion.
[0125] Example 2, the difference between this embodiment and Example 1 is that the health assessment model is further optimized to adjust the health score by incorporating data noise levels, including:
[0126] By analyzing the noise level N of each data point in real time i The noise weight θ is calculated using the following formula. i This is used to dynamically adjust the data health score;
[0127]
[0128] Where Nose_i represents the noise level of the i-th data point;
[0129] Optimized health score calculation formula:
[0130]
[0131] S represents the adjusted health score. i This indicates the initial health score.
[0132] Multi-source data fusion is implemented based on the adjusted health score.
[0133] Example 3, see Figure 2 The present invention provides a cloud computing-based multi-source data fusion system, as shown in the block diagram of the system structure.
[0134] The data acquisition module acquires real-time data from multiple data sources through a cloud computing interface, constructs a multi-dimensional feature vector of the real-time data, generates a health score using a health assessment model, and generates a health label for each data point based on the health score.
[0135] The data processing module is used to perform format conversion and semantic enhancement on data marked as healthy, generating semantically enhanced structured data;
[0136] The rule construction and adjudication module constructs fusion rules and calculates the confidence of the fusion rules based on the semantically enhanced structured data through knowledge graph reasoning; it constructs a semantic consistency index based on the confidence of the fusion rules and semantic similarity, and applies the semantic consistency index to adjudicate the fusion result for conflicting data;
[0137] The verification and optimization module verifies the quality of the fusion results using historical data and business rules. If the fusion results do not meet the preset requirements, a feedback optimization process is triggered.
[0138] The explanation and summary of the data transfer logic between modules is as follows:
[0139] The data acquisition module outputs multidimensional feature vector data of health scores and health labels for the data processing module to process and convert formats.
[0140] The semantically enhanced structured data generated by the data processing module is used by the rule building and adjudication module for rule building and semantic consistency index calculation.
[0141] The fused data and semantic consistency index output by the rule construction and adjudication module are used to verify the quality verification and score calculation of the optimization module; if the verification fails, optimization feedback is provided to adjust the semantic embedding model or fused rule logic of the health assessment model.
[0142] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A multi-source data fusion method based on cloud computing, characterized in that, Includes the following steps: Step S001: Obtain real-time data from multiple data sources through cloud computing interfaces, construct multi-dimensional feature vectors of the real-time data, generate health scores using a health assessment model, and generate health labels for each data point based on the health scores. Let i represent the sequential number of the data, m be the number of data features included in each data entry, and j represent the sequential number of the data features; let n be the number of fields included in each data entry, and k represent the sequential number of the fields. The health score is calculated using the following health assessment model. , ;in, For the first Health score of each data point; For the i-th data One data feature; For the first Each data feature corresponds to a feature weight, which indicates the degree of influence of the data feature on the score; For time-based dynamic correction items, where This is a time correction factor. For the current time, For data collection time; This is the anomaly compensation coefficient, used to adjust for abnormal effects in the scoring; For the first The variance of each field measures the volatility of that field. Step S002: Perform format conversion and semantic enhancement on the data marked as healthy to generate semantically enhanced structured data; Step S003: Based on the semantically enhanced structured data, construct fusion rules through knowledge graph reasoning and calculate the confidence of the fusion rules. The confidence of the fusion rules refers to the ratio of the number of effective fusions to the total number of fusions. Construct a semantic consistency index based on the confidence of the fusion rules and semantic similarity. Apply the semantic consistency index to adjudicate the fusion result for conflicting data. Let r represent the sequence number of the fusion rules and Ng represent the number of fusion rules. Let be the confidence level of the r-th fusion rule. To integrate rule weights, a set of conflicting data is denoted as... Semantic similarity is denoted as The semantic consistency index is calculated using the following formula: ; Step S004: Verify the quality of the fusion results using historical data and business rules. If the fusion results do not meet the preset requirements, trigger the feedback optimization process.
2. The multi-source data fusion method based on cloud computing according to claim 1, characterized in that, The health assessment model includes a time-based dynamic correction term, which refers to dynamically adjusting the health score based on the time characteristics of the data; the health label is used to mark the status and abnormality category of the data.
3. The multi-source data fusion method based on cloud computing according to claim 1, characterized in that, Based on the semantically enhanced structured data, the construction of fusion rules includes the following steps: Knowledge graphs are generated based on semantic relationships between data using knowledge graph construction technology. Input the defined logical relationships, use logical reasoning techniques to infer the logical connections between data from the knowledge graph, and generate fusion rules; Calculation of weight and confidence of fusion rules: Combine health labels, data source confidence and semantic features to calculate the weight and confidence of fusion rules.
4. The multi-source data fusion method based on cloud computing according to claim 1, characterized in that, The quality verification of the fusion results includes the following steps: Step S401: Load historical data and business rules as benchmarks. Historical data is used to provide a reference standard for the fusion results, and business rules are used to define the target accuracy and deviation range. Input historical fusion result data and business rules, analyze the statistical characteristics of historical data, and extract the requirements for result accuracy and deviation in the business rules. Output the benchmark standard data and accuracy target for validating the model. Step S402: Format the fusion results to be verified to ensure comparability with benchmark data and rules; The current fusion result data is cleaned, outliers and redundant entries are removed from the data, the format is standardized, and the data dimensions and features are consistent with the benchmark standard data. The cleaned fusion result data is then output. Step S403: Calculate the verification score: Calculate the verification score based on the cleaned fusion result data to quantify the reliability of the fusion result; output the comprehensive verification score of the fusion result through the following model: , ;in, No. The accuracy of the record This represents the fusion result data; Represents benchmark data; This represents the validation score of the current fusion result. Indicates the first The weight of each record; Step S404: Compare the comprehensive verification score with the preset threshold to determine whether the fusion result meets the requirements. If the comprehensive verification score is higher than or equal to the threshold, the fusion result is considered qualified and the verification is passed. If the comprehensive verification score is lower than the threshold, the fusion result is considered unqualified and the feedback optimization process is triggered.
5. The multi-source data fusion method based on cloud computing according to claim 4, characterized in that, The feedback optimization process includes health assessment model parameter optimization, semantic embedding model adjustment, and fusion rule optimization, specifically: Optimize the time dynamic correction term of the health assessment model to improve its adaptability to dynamic fluctuations in data; Adjust the scene feature priority rules of the semantic embedding model to enhance the semantic association capability of the scene; Improve the confidence calculation logic of the fusion rules to enhance the accuracy and consistency of conflict data adjudication; Using the optimized health assessment model, semantic embedding model, and fusion rules, the semantic consistency index is recalculated and compared with a threshold. If it is higher than the threshold, the optimization is confirmed to be effective; otherwise, the optimization scheme is adjusted to form a dynamic closed-loop optimization mechanism.
6. The multi-source data fusion method based on cloud computing according to claim 1, characterized in that, The application semantic consistency index adjudication fusion result includes: If the semantic consistency index is not lower than the set threshold, the current fusion rule is applied to the conflicting data, and the corresponding fusion result is selected as the final output; if the semantic consistency index is lower than the set threshold, it indicates that the current fusion rule is invalid, and the fusion rule needs to be optimized or a new fusion rule needs to be selected.
7. The multi-source data fusion method based on cloud computing according to claim 6, characterized in that, The semantic similarity is obtained as follows: Let a set of conflicting data be denoted as Measuring conflict data through semantic similarity The semantic similarity is used to assist in determining conflicting data. The semantic similarity is calculated using the following formula. : , ;in, and Data respectively and semantic embedding vector; The vector magnitude is used for normalization. These are distance-weighted coefficients; For feature distance, and Data respectively and The 1 eigenvalue, The total number of features.
8. A cloud computing-based multi-source data fusion system, characterized in that, include: The data acquisition module acquires real-time data from multiple data sources through a cloud computing interface, constructs a multi-dimensional feature vector of the real-time data, and generates a health score using a health assessment model. Based on the health score, a health label is generated for each data point. Let i represent the sequential number of the data, m be the number of data features included in each data point, and j represent the sequential number of the data features. Let n be the number of fields included in each data point, and k represent the sequential number of the fields. The health score is calculated using the following health assessment model. , ,in, For the first Health score of each data point; For the i-th data One data feature; For the first Each data feature corresponds to a feature weight, which indicates the degree of influence of the data feature on the score; For time-based dynamic correction items, where This is a time correction factor. For the current time, For data collection time; This is the anomaly compensation coefficient, used to adjust for abnormal effects in the scoring; For the first The variance of each field measures the volatility of that field. The data processing module is used to perform format conversion and semantic enhancement on data marked as healthy, generating semantically enhanced structured data; The rule construction and adjudication module, based on semantically enhanced structured data, constructs fusion rules through knowledge graph reasoning and calculates the confidence of the fusion rules. The confidence of the fusion rules refers to the ratio of the number of valid fusions to the total number of fusions. A semantic consistency index is constructed based on the confidence of the fusion rules and semantic similarity. This semantic consistency index is applied to adjudicate the fusion results for conflicting data. The module uses 'r' to represent the sequential number of the fusion rules and 'Ng' to represent the number of fusion rules. Let be the confidence level of the r-th fusion rule. To integrate rule weights, a set of conflicting data is denoted as... Semantic similarity is denoted as The semantic consistency index is calculated using the following formula: ; The verification and optimization module verifies the quality of the fusion results using historical data and business rules. If the fusion results do not meet the preset requirements, a feedback optimization process is triggered.
Citation Information
Patent Citations
Perception data semantic standardization and fusion processing cloud platform
CN117668759A
Multi-source twin data fusion tunnel structure health monitoring and early warning method and system
CN119129077A