Multi-source data fusion method and system based on cloud computing
By using health assessment models, semantic enhancement and knowledge graphs to build fusion rules on the cloud computing platform, dynamic adaptation and semantic consistency problems in multi-source data fusion are solved, and efficient and reliable data fusion effect is achieved.
Patent Information
- Application Number
- CN202510068063.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-16
AI Technical Summary
Existing multi-source data fusion methods are difficult to achieve dynamic adaptation, real-time processing and semantic consistency when processing heterogeneous data, resulting in poor fusion effects, especially lacking effective mechanisms in ruling conflict data.
Through a multi-source data fusion method based on cloud computing, health scores and labels are generated using a health assessment model, format conversion and semantic enhancement are performed, fusion rules are constructed using knowledge graphs and semantic consistency index is calculated to scientifically determine conflicting data.
Dynamic multi-source data fusion is realized, the accuracy of conflicting data ruling and the reliability and semantic consistency of the fusion results are improved, and the adaptability and processing efficiency of the system are enhanced.
Smart Images

Figure CN119989267A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multi-source data fusion, and more specifically, to a multi-source data fusion method and system based on cloud computing. Background Art
[0002] With the rapid development of cloud computing technology, multi-source data fusion has been widely used in smart cities, traffic management, medical health and other fields. Multi-source data fusion supports data-driven decision-making by integrating heterogeneous data from different sources, formats and semantics. However, the diversity and dynamics of heterogeneous data place extremely high demands on fusion methods, especially in real-time data processing, semantic consistency maintenance and conflicting data adjudication. Traditional fusion methods often find it difficult to cope with complex and changing scenarios.
[0003] However, in actual use, it still has the disadvantage of poor fusion effect. For example, it lacks a dynamically adaptive data health assessment mechanism and cannot mark and process low-quality data in real time. In the semantic matching of heterogeneous data, traditional methods lack the accuracy to handle missing fields and inconsistent data, making it difficult to ensure the accuracy of fusion. In the face of conflicts between multi-source data, there is a lack of a fusion rule arbitration mechanism based on confidence and semantic similarity, which easily leads to a decrease in the reliability of the fusion results. Summary of the invention
[0004] In order to overcome the above-mentioned defects of the prior art, the present invention provides a multi-source data fusion method and system based on cloud computing, which generates health scores and labels through a health assessment model, constructs fusion rules based on semantic enhancement, and uses a semantic consistency index to adjudicate conflicting data, so as to solve the problems raised in the above-mentioned background technology.
[0005] To achieve the above object, the present invention provides the following technical solution: a multi-source data fusion method based on cloud computing, comprising:
[0006] Step S001: Acquire real-time data from multiple data sources through a cloud computing interface, construct a multi-dimensional feature vector of the real-time data, generate a health score using a health assessment model, and generate a health label for each piece of data according to the health score;
[0007] Step S002: Perform format conversion and semantic enhancement processing on the data marked as healthy to generate semantically enhanced structured data;
[0008] Step S003: Based on the structured data after semantic enhancement, construct fusion rules through knowledge graph reasoning and calculate the confidence of the fusion rules; construct a semantic consistency index based on the confidence of the fusion rules and the semantic similarity, and apply the semantic consistency index to the conflicting data to determine the fusion result;
[0009] Step S004: Verify the quality of the fusion result through historical data and business rules. If the fusion result does not meet the preset requirements, the feedback optimization process is triggered;
[0010] Through the above steps, dynamic multi-source data fusion and full-process optimization can be achieved.
[0011] Preferably, the health assessment model includes a time dynamic correction term, which refers to dynamically adjusting the health score according to the time characteristics of the data; the health label is used to mark the status and abnormal category of the data.
[0012] Preferably, the format conversion in step S002 refers to structural standardization processing through a pattern matching algorithm; semantic enhancement refers to building semantic associations between data based on a semantic embedding model optimized based on scene features, and the semantic embedding model dynamically adjusts the semantic matching strategy according to the abnormal category of the health label, and dynamically completes the missing or inconsistent fields.
[0013] Preferably, constructing fusion rules based on semantically enhanced structured data includes the following steps:
[0014] Using knowledge graph construction technology, a knowledge graph is generated based on the semantic associations between data;
[0015] Input the defined logical relationship, use logical reasoning technology to infer the logical association between data from the knowledge graph, and generate fusion rules;
[0016] Fusion rule weight and confidence calculation: Combine health labels, data source confidence and semantic features to calculate the weight and confidence of the fusion rule.
[0017] Preferably, the conflicting data is identified by:
[0018] Step S301, data matching and semantic alignment: semantically represent multi-source data using a semantic embedding model, semantically align data on the same subject, and output semantically aligned data;
[0019] Step 302, difference detection: for semantically aligned data, determine whether there are description differences in multi-source data under the same topic, and output difference indicators; Step 303, if the difference indicator exceeds a preset threshold, it is determined to be conflicting data.
[0020] Preferably, the quality verification of the fusion result is used to evaluate the accuracy and reliability of the fusion result, and the fusion result is quantitatively evaluated based on historical data, business rules and verification scoring indicators. The core of the verification model is to calculate the verification score Qk by analyzing the deviation between the fusion result and the expected target, and judge the quality and optimization direction of the fusion process accordingly, including the following steps:
[0021] Step S401, load historical data and business rules as benchmarks, historical data is used to provide reference standards for fusion results, and business rules are used to define target accuracy and deviation range; input historical fusion result data and business rules, analyze the statistical characteristics of historical data, and extract the requirements for result accuracy and deviation in business rules; output the benchmark standard data and accuracy target of the verification model; the accuracy target defines the score deviation tolerance range, which is applied to the threshold comparison of the verification score;
[0022] Step S402: format the fusion result to be verified to ensure comparability with the benchmark standard data and rules; clean the current fusion result data, outliers and redundant entries in the data, unify the format, ensure that the data dimensions and features are consistent with the benchmark standard data, and output the cleaned fusion result data;
[0023] Step S403, calculate the verification score: calculate the verification score based on the cleaned fusion result data to quantify the reliability of the fusion result; output the comprehensive verification score of the fusion result through the following model:
[0024]
[0025] Among them, Acc t The accuracy of the tth record, Represents the fusion result data; Indicates benchmark standard data; Q k represents the verification score of the current fusion result, w t Indicates the weight of the tth record;
[0026] Step S404: compare the comprehensive verification score with the preset threshold to determine whether the fusion result meets the requirements. If the comprehensive verification score is higher than or equal to the threshold, the fusion result is considered qualified and the verification passes; if the comprehensive verification score is lower than the threshold, the fusion result is considered unqualified and the feedback optimization process is triggered.
[0027] Preferably, the feedback optimization process includes health assessment model parameter optimization, semantic embedding model adjustment and fusion rule optimization, specifically:
[0028] Optimize the time dynamic correction item of the health assessment model to improve the adaptability to dynamic fluctuations of data; re-evaluate the abnormal conditions of the data health score (such as the proportion of low-scoring data and missing fields), adjust the dynamic correction item and abnormal compensation coefficient in the scoring formula, and optimize the feature weight distribution to improve the reliability of the data health score and ensure accurate assessment of data quality in the subsequent fusion process;
[0029] Adjust the scene feature priority rules of the semantic embedding model to enhance the semantic association ability of specific scenes; improve the training data of the semantic embedding model by introducing domain-specific corpus, adjust the feature priority weights of the embedding vector, and add contextual features, so as to adjust the semantic embedding model to a form that is more suitable for specific application scenarios, thereby improving the accuracy of semantic similarity calculation;
[0030] Improve the confidence calculation logic of fusion rules to improve the accuracy and consistency of conflicting data adjudication;
[0031] Redefine the confidence formula of fusion rules, dynamically adjust the weight of fusion rules, and expand the rule logic to improve the ability of fusion rules to adjudicate conflicting data;
[0032] Use the optimized health assessment model, semantic embedding model and fusion rules to recalculate and verify the semantic consistency index and compare it with the threshold. If it is higher than the threshold, the optimization is confirmed to be effective. Otherwise, continue to adjust the optimization plan to form a dynamic closed-loop optimization mechanism.
[0033] Preferably, the application of the semantic consistency index to determine the fusion result includes:
[0034] If the semantic consistency index is not lower than the set threshold, the current fusion rule is applied to the conflicting data and the corresponding fusion result is selected as the final output; if the semantic consistency index is lower than the set threshold, it indicates that the current fusion rule is invalid and the fusion rule needs to be optimized or reselected.
[0035] Preferably, the step S001 includes the following contents:
[0036] Let i represent the sequential number of the data, let the number of data features included in each piece of data be m, let j represent the sequential number of the data features; let the number of fields included in each piece of data be n, let k represent the sequential number of the fields, and calculate the health score through the following health assessment model;
[0037]
[0038] φ t =δ*|t a -t 0 |
[0039] Among them, S i is the health score of the i-th data; v ij is the jth data feature of the i-th data; α j is the feature weight corresponding to the jth data feature, indicating the influence of the data feature on the score; φ t is the time dynamic correction term, where δ is the time correction coefficient, t a is the current time, t0 is the data collection time; ψ is the abnormal compensation coefficient, which is used to adjust the abnormal impact in the score; Var(f k ) is the variance of the kth field, which measures the field volatility.
[0040] Preferably, the semantic consistency index Index is obtained in the following manner:
[0041] r is used to represent the order number of the fusion rule, and Ng is used to represent the number of fusion rules; C r is the confidence of the rth fusion rule, W r To integrate the rule weights, the semantic consistency index is calculated using the following formula:
[0042]
[0043] The confidence of the fusion rule refers to the ratio of the number of effective fusions to the total number of fusions.
[0044] Preferably, the semantic similarity is obtained in the following manner:
[0045] A set of conflicting data is recorded as (D i ,D s ), and measure conflicting data by semantic similarity (D i ,D s ) to assist in determining the conflicting data. The semantic similarity Sim(D i ,D s ):
[0046]
[0047] Among them, E ′ (D i ) and E′(D s ) are data D i and D s The semantic embedding vector of ′ (D i )∥ is the vector modulus, used for normalization: μ is the distance weighting coefficient; Dist(D i ,D s ) is the characteristic distance, f ij and f sj The data D i and D s is the j-th eigenvalue of , and m is the total number of features.
[0048] To achieve the above object, the present invention provides the following technical solution: a multi-source data fusion system based on cloud computing, comprising:
[0049] The data acquisition module acquires real-time data from multiple data sources through the cloud computing interface, constructs a multi-dimensional feature vector of the real-time data, generates a health score using a health assessment model, and generates a health label for each piece of data based on the health score;
[0050] A data processing module is used to perform format conversion and semantic enhancement processing on the data marked as healthy, and generate semantically enhanced structured data;
[0051] The rule construction and adjudication module builds fusion rules and calculates the confidence of fusion rules through knowledge graph reasoning based on semantically enhanced structured data; builds a semantic consistency index based on the confidence and semantic similarity of the fusion rules, and applies the semantic consistency index to adjudicate the fusion results for conflicting data;
[0052] The verification and optimization module verifies the quality of the fusion results through historical data and business rules. If the fusion results do not meet the preset requirements, the feedback optimization process is triggered.
[0053] Technical effects and advantages of the present invention:
[0054] (1) The present invention provides a multi-source data fusion method based on cloud computing. By introducing a semantic consistency index and utilizing weighted calculations of rule confidence, rule weight, and semantic similarity, the present invention realizes scientific adjudication of conflicting data. The semantic consistency index can dynamically evaluate the effectiveness of fusion rules to ensure that when conflicts exist between multi-source data, the selected or fused data results have higher reliability and semantic consistency. Compared with traditional methods, the present invention not only improves the accuracy of conflicting data adjudication, but also enhances rule adaptability and scenario flexibility, and is suitable for complex data processing requirements in multiple scenarios such as traffic flow prediction and medical data analysis.
[0055] (2) The present invention provides a multi-source data fusion method based on cloud computing. By designing a closed-loop optimization process for health assessment model parameter optimization, semantic embedding model adjustment, and fusion rule improvement, the key parameters of each module can be dynamically adjusted according to the results of verification scoring. When the quality of the fusion result does not meet the requirements, the optimization process can specifically improve the accuracy of data health assessment, the semantic association ability of the semantic embedding model, and the adjudication efficiency of the fusion rules, thereby ensuring the continuous improvement of the quality of the fusion process. The closed-loop mechanism provides adaptive optimization capabilities for multi-source data fusion, greatly improving the system's adaptability to dynamic scenarios and processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 This is a flow chart of the multi-source data fusion method of the present invention.
[0057] Figure 2 This is a structural block diagram of the multi-source data fusion system of the present invention. DETAILED DESCRIPTION
[0058] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.
[0059] At the same time, it should be understood that for the convenience of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship.
[0060] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the present application, its application, or uses.
[0061] Technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and equipment should be considered part of the specification.
[0062] Example 1, see Figure 1 The embodiment of the present invention provides a multi-source data fusion method based on cloud computing, which includes the following steps:
[0063] Step S001, obtaining real-time data from multiple data sources through a cloud computing interface (the real-time data includes each piece of data, transmission rate, data integrity, and timestamp information), constructing a multidimensional feature vector of the real-time data and generating a health score using a health assessment model (the health score is used to reflect the comprehensive quality index of each piece of data), and generating a health label for each piece of data according to the health score;
[0064] Step S002: Perform format conversion and semantic enhancement processing on the data marked as healthy to generate semantically enhanced structured data;
[0065] Step S003: Based on the structured data after semantic enhancement, construct fusion rules through knowledge graph reasoning and calculate the confidence of the fusion rules; construct a semantic consistency index based on the confidence of the fusion rules and the semantic similarity, and apply the semantic consistency index to the conflicting data to determine the fusion result;
[0066] Explanation: Conflicting data refers to data from different data sources that provide inconsistent descriptions of the same event or object; for example, two data sources provide different values for the traffic flow on a certain road section, or provide conflicting descriptions of the weather conditions at a certain point in time;
[0067] Step S004: Verify the quality of the fusion result through historical data and business rules. If the fusion result does not meet the preset requirements, the feedback optimization process is triggered;
[0068] Through the above steps, dynamic multi-source data fusion and full-process optimization can be achieved.
[0069] What needs to be further explained in the embodiments of the present invention is that the health assessment model includes a time dynamic correction term, and the time dynamic correction term refers to dynamically adjusting the health score according to the time characteristics of the data; the health label is used to mark the status of the data (such as healthy data and abnormal data) and the abnormal category (such as transmission delay, field missing).
[0070] What needs to be further explained in the embodiments of the present invention is that the format conversion in step S002 refers to the structural standardization processing through the pattern matching algorithm; semantic enhancement refers to the construction of semantic associations between data based on the semantic embedding model optimized based on scene features, and the semantic embedding model dynamically adjusts the semantic matching strategy according to the abnormal category of the health label, and dynamically completes the missing or inconsistent fields.
[0071] It needs to be further explained in the embodiment of the present invention that, based on the structured data after semantic enhancement, constructing the fusion rules includes the following steps:
[0072] Use knowledge graph construction technology (such as RDF and SPARQL) to generate knowledge graphs based on semantic associations between data;
[0073] Input the defined logical relationship, use logical reasoning technology (such as rule reasoning, OWL reasoner) to infer the logical association between data from the knowledge graph, and generate fusion rules;
[0074] Fusion rule weight and confidence calculation: Combine health labels, data source confidence and semantic features to calculate the weight and confidence of the fusion rule.
[0075] Explanation: How to obtain the weight of the fusion rule:
[0076] Takes fusion rule set, health labels and semantically enhanced data as input;
[0077] Calculate the fusion rule weights based on health labels and semantically enhanced features, and finally output a set of fusion rules with weights;
[0078] In a further embodiment, the health score and semantic similarity of the data are considered, and the weighted averages are calculated respectively, and a weight adjustment coefficient is introduced to further adjust the weight distribution of the health score and semantic similarity; for ease of understanding, the specific calculation formula is W = α 1 ·H+α 2S, where W represents the weight of the fusion rule, H represents the weighted average of the health scores of the data involved in the fusion rule, S represents the weighted average of the semantic similarity of the data involved in the fusion rule, and α 1 and α 2 is the weight adjustment coefficient;
[0079] The confidence of the fusion rule is obtained as follows:
[0080] Takes the fusion rule set and its weights as input;
[0081] Calculate the rule confidence based on the historical performance of the data source and the adaptability of the fusion rule; finally output the fusion rule confidence;
[0082] For ease of understanding, the contribution of each data source to the rule confidence and the contribution weight of the data source associated with the fusion rule are considered; that is, the fusion rule confidence is the weighted average of the confidences and contribution weights of multiple data sources.
[0083] It needs to be further explained in the embodiment of the present invention that the conflicting data is identified as follows:
[0084] Step S301, data matching and semantic alignment: semantically represent multi-source data using a semantic embedding model, semantically align data on the same subject, and output semantically aligned data;
[0085] Step 302, difference detection: for the semantically aligned data, determine whether there are description differences in the multi-source data under the same topic, and output the difference index;
[0086] For example, in a traffic data scenario, different data sources provide different values for the traffic flow of the same road section; the difference index is the ratio of the value difference to the total number of fields;
[0087] Step 303: If the difference index exceeds a preset threshold, it is determined to be conflicting data.
[0088] It needs to be further explained in the embodiments of the present invention that the quality verification of the fusion result is used to evaluate the accuracy and reliability of the fusion result, and the fusion result is quantitatively evaluated based on historical data, business rules and verification scoring indicators. The core of the verification model is to calculate the verification score by analyzing the deviation between the fusion result and the expected target, and judge the quality and optimization direction of the fusion process based on this, including the following steps:
[0089] Step S401, load historical data and business rules as benchmarks, historical data is used to provide reference standards for fusion results, and business rules are used to define target accuracy and deviation range; input historical fusion result data and business rules, analyze the statistical characteristics of historical data (such as distribution, fluctuation range), and extract the requirements for result accuracy and deviation in business rules (such as tolerance error, time constraint); output the benchmark standard data and accuracy target of the verification model; the accuracy target defines the score deviation tolerance range, which is applied to the threshold comparison of the verification score;
[0090] Step S402: format the fusion result to be verified to ensure comparability with the benchmark standard data and rules; clean the current fusion result data, outliers and redundant entries in the data, unify the format, ensure that the data dimensions and features are consistent with the benchmark standard data, and output the cleaned fusion result data;
[0091] Step S403, calculate the verification score: calculate the verification score based on the cleaned fusion result data to quantify the reliability of the fusion result; output the comprehensive verification score of the fusion result through the following model:
[0092]
[0093] Among them, Acc t The accuracy of the tth record, Represents the fusion result data; Indicates benchmark standard data; Q k represents the verification score of the current fusion result, w t Indicates the weight of the tth record;
[0094] Step S404: compare the comprehensive verification score with the preset threshold to determine whether the fusion result meets the requirements. If the comprehensive verification score is higher than or equal to the threshold, the fusion result is considered qualified and the verification passes; if the comprehensive verification score is lower than the threshold, the fusion result is considered unqualified and the feedback optimization process is triggered.
[0095] It needs to be further explained in the embodiment of the present invention that the feedback optimization process includes health assessment model parameter optimization, semantic embedding model adjustment and fusion rule optimization, specifically:
[0096] Optimize the time dynamic correction item of the health assessment model to improve the adaptability to dynamic fluctuations of data; re-evaluate the abnormal conditions of the data health score (such as the proportion of low-scoring data and missing fields), adjust the dynamic correction item and abnormal compensation coefficient in the scoring formula, and optimize the feature weight distribution to improve the reliability of the data health score and ensure accurate assessment of data quality in the subsequent fusion process;
[0097] Adjust the scene feature priority rules of the semantic embedding model to enhance the semantic association ability of specific scenes; improve the training data of the semantic embedding model by introducing domain-specific corpus, adjust the feature priority weights of the embedding vector, and add contextual features, so as to adjust the semantic embedding model to a form that is more suitable for specific application scenarios, thereby improving the accuracy of semantic similarity calculation;
[0098] Improve the confidence calculation logic of fusion rules to improve the accuracy and consistency of conflicting data adjudication;
[0099] Redefine the confidence formula of fusion rules (introduce data delay or fluctuation dynamic factors), dynamically adjust the weight of fusion rules (modify according to the performance of different scenarios), and expand the rule logic (such as adding time or field constraints) to improve the ability of fusion rules to adjudicate conflicting data;
[0100] Use the optimized health assessment model, semantic embedding model and fusion rules to recalculate and verify the semantic consistency index and compare it with the threshold. If it is higher than the threshold, the optimization is confirmed to be effective. Otherwise, continue to adjust the optimization plan to form a dynamic closed-loop optimization mechanism.
[0101] It needs to be further explained in the embodiment of the present invention that the application semantic consistency index determines the fusion result including:
[0102] If the semantic consistency index is not lower than the set threshold, the current fusion rule is applied to the conflicting data and the corresponding fusion result is selected as the final output; if the semantic consistency index is lower than the set threshold, it indicates that the current fusion rule is invalid and the fusion rule needs to be optimized or reselected.
[0103] It needs to be further explained in the embodiment of the present invention that step S001 includes the following contents:
[0104] Let i represent the sequential number of the data, let the number of data features included in each piece of data be m, let j represent the sequential number of the data features; let the number of fields included in each piece of data be n, let k represent the sequential number of the fields, and calculate the health score through the following health assessment model;
[0105]
[0106] φ t =δ*|t a -t 0 |
[0107] Among them, S i is the health score of the i-th data; v ij is the jth data feature of the i-th data; α jis the feature weight corresponding to the jth data feature, indicating the influence of the data feature on the score; φ t is the time dynamic correction term, where δ is the time correction coefficient, t a is the current time, t 0 is the data collection time; ψ is the abnormal compensation coefficient, which is used to adjust the abnormal impact in the score; Var(f k ) is the variance of the kth field, which measures the field volatility.
[0108] For high-dynamic scenarios (such as transportation or finance), the time dynamic correction term is further optimized by the following formula:
[0109]
[0110] Among them, T max Indicates the maximum value of the time range.
[0111] It needs to be further explained in the embodiment of the present invention that the semantic consistency index Index is obtained in the following manner:
[0112] r is used to represent the order number of the fusion rule, and Ng is used to represent the number of fusion rules; C r is the confidence of the rth fusion rule, W r To integrate the rule weights, the semantic consistency index is calculated using the following formula:
[0113]
[0114] The confidence of the fusion rule refers to the ratio of the number of effective fusions to the total number of fusions.
[0115] Explanation: The semantic consistency index is mainly for conflicting data, but it can be extended to become part of global data fusion in the following scenarios:
[0116] Data fusion involves multiple levels of processing:
[0117] Apply consistency index to adjudicate conflicting data;
[0118] The adjudicated data is integrated with the non-conflicting data;
[0119] High-demand data fusion scenarios: If all data need to be strictly reviewed for consistency (such as medical and financial scenarios), the semantic consistency index can be applied globally to verify the applicability of the rules to all data.
[0120] It needs to be further explained in the embodiment of the present invention that the semantic similarity is obtained in the following manner:
[0121] A set of conflicting data is recorded as (Di ,D s ), and measure conflicting data by semantic similarity (D i ,D s ) to assist in determining the conflicting data. The semantic similarity Sim(D i ,D s ):
[0122]
[0123] Among them, E ′ (D i ) and E′(D s ) are data D i and D s The semantic embedding vector of ′ (D i )∥ is the vector modulus, used for normalization: μ is the distance weighting coefficient; Dist(D i ,D s ) is the characteristic distance, f ij and f sj The data D i and D s is the j-th eigenvalue of , and m is the total number of features.
[0124] Summary: The embodiment of the present invention calculates the semantic consistency index through comprehensive evaluation of semantic similarity, rule confidence and rule weight, and makes quantitative judgments on conflicting data; it includes health assessment model parameter optimization, semantic embedding model adjustment and fusion rule logic improvement, and forms a closed-loop optimization by dynamically adjusting the key parameters and logic of each module; it dynamically corrects the health score according to the time characteristics of the data, effectively copes with the volatility and delay problems of real-time data, and effectively improves the comprehensive quality of data fusion.
[0125] Embodiment 2: The difference between the embodiment of the present invention and embodiment 1 is that the health assessment model is further optimized to adjust the health score in combination with the data noise level, including:
[0126] By analyzing the noise level N of each data in real time i , the noise weight θ is calculated by the following formula i , used to dynamically adjust the data health score;
[0127]
[0128] Among them, Nose_i represents the noise level of the i-th data;
[0129] Optimized health score calculation formula:
[0130]
[0131] represents the adjusted health score, S i Indicates the initial health score.
[0132] Implement multi-source data fusion based on adjusted health scores.
[0133] Example 3, see Figure 2 The embodiment of the present invention provides a multi-source data fusion system based on cloud computing, including:
[0134] The data acquisition module acquires real-time data from multiple data sources through the cloud computing interface, constructs a multi-dimensional feature vector of the real-time data, generates a health score using a health assessment model, and generates a health label for each piece of data based on the health score;
[0135] A data processing module is used to perform format conversion and semantic enhancement processing on the data marked as healthy, and generate semantically enhanced structured data;
[0136] The rule construction and adjudication module builds fusion rules and calculates the confidence of fusion rules through knowledge graph reasoning based on semantically enhanced structured data; builds a semantic consistency index based on the confidence and semantic similarity of the fusion rules, and applies the semantic consistency index to adjudicate the fusion results for conflicting data;
[0137] The verification and optimization module verifies the quality of the fusion results through historical data and business rules. If the fusion results do not meet the preset requirements, the feedback optimization process is triggered.
[0138] Explanation: The logic of data transfer between modules can be summarized as follows:
[0139] The data acquisition module outputs the multi-dimensional feature vector data of health scores and health labels for processing and format conversion by the data processing module;
[0140] The semantically enhanced structured data generated by the data processing module is used by the rule construction and adjudication module to construct rules and calculate the semantic consistency index;
[0141] The fusion data and semantic consistency index output by the rule construction and adjudication module are used to verify the quality verification and score calculation of the optimization module; if the verification fails, optimization feedback is performed to adjust the semantic embedding model of the health assessment model or the fusion rule logic.
[0142] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A multi-source data fusion method based on cloud computing, characterized in that: The following steps are involved: Step S001: Acquire real-time data from multiple data sources through a cloud computing interface, construct a multi-dimensional feature vector of the real-time data, generate a health score using a health assessment model, and generate a health label for each piece of data according to the health score; Step S002: Perform format conversion and semantic enhancement processing on the data marked as healthy to generate semantically enhanced structured data; Step S003: Based on the structured data after semantic enhancement, construct fusion rules through knowledge graph reasoning and calculate the confidence of the fusion rules; The semantic consistency index is constructed based on the confidence and semantic similarity of the fusion rules, and the semantic consistency index is used to determine the fusion result for conflicting data. Step S004: Verify the quality of the fusion result through historical data and business rules. If the fusion result does not meet the preset requirements, the feedback optimization process is triggered.
2. The multi-source data fusion method based on cloud computing according to claim 1 is characterized in that: The health assessment model includes a time dynamic correction term, which refers to dynamically adjusting the health score according to the time characteristics of the data; the health label is used to mark the status and abnormal category of the data.
3. The multi-source data fusion method based on cloud computing according to claim 1 is characterized in that: Based on the semantically enhanced structured data, building fusion rules includes the following steps: Using knowledge graph construction technology, a knowledge graph is generated based on the semantic associations between data; Input the defined logical relationship, use logical reasoning technology to infer the logical association between data from the knowledge graph, and generate fusion rules; Fusion rule weight and confidence calculation: Combine health labels, data source confidence and semantic features to calculate the weight and confidence of the fusion rule.
4. The multi-source data fusion method based on cloud computing according to claim 1 is characterized in that: The quality verification of the fusion result comprises the following steps: Step S401: Load historical data and business rules as a benchmark. Historical data is used to provide a reference standard for fusion results, and business rules are used to define target accuracy and deviation range. Input historical fusion result data and business rules, analyze the statistical characteristics of historical data, and extract the requirements for result accuracy and deviation in business rules. Output the benchmark standard data and accuracy target of the verification model. Step S402: formatting the fusion result to be verified to ensure comparability with the benchmark standard data and rules; The current fusion result data is cleaned of outliers and redundant entries, and the format is unified to ensure that the data dimensions and features are consistent with the benchmark standard data, and the cleaned fusion result data is output; Step S403, calculate the verification score: calculate the verification score based on the cleaned fusion result data to quantify the reliability of the fusion result; output the comprehensive verification score of the fusion result through the following model: Among them, Acc t The accuracy of the tth record, Represents the fusion result data; Indicates benchmark standard data; Q k represents the verification score of the current fusion result, w t Indicates the weight of the tth record; Step S404: compare the comprehensive verification score with the preset threshold to determine whether the fusion result meets the requirements. If the comprehensive verification score is higher than or equal to the threshold, the fusion result is considered qualified and the verification passes; if the comprehensive verification score is lower than the threshold, the fusion result is considered unqualified and the feedback optimization process is triggered.
5. The multi-source data fusion method based on cloud computing according to claim 4 is characterized in that: The feedback optimization process includes health assessment model parameter optimization, semantic embedding model adjustment and fusion rule optimization, specifically: Optimize the time dynamic correction term of the health assessment model to improve the adaptability to dynamic fluctuations of data; Adjust the scene feature priority rules of the semantic embedding model to enhance the semantic association capabilities of specific scenes; Improve the confidence calculation logic of fusion rules to improve the accuracy and consistency of conflicting data adjudication; Use the optimized health assessment model, semantic embedding model and fusion rules to recalculate and verify the semantic consistency index and compare it with the threshold. If it is higher than the threshold, the optimization is confirmed to be effective. Otherwise, continue to adjust the optimization plan to form a dynamic closed-loop optimization mechanism.
6. The multi-source data fusion method based on cloud computing according to claim 1 is characterized in that: The application semantic consistency index determines the fusion result including: If the semantic consistency index is not lower than the set threshold, the current fusion rule is applied to the conflicting data and the corresponding fusion result is selected as the final output; if the semantic consistency index is lower than the set threshold, it indicates that the current fusion rule is invalid and the fusion rule needs to be optimized or reselected.
7. The multi-source data fusion method based on cloud computing according to claim 2 is characterized in that: The step S001 includes the following contents: Let i represent the sequential number of the data, let the number of data features included in each piece of data be m, let j represent the sequential number of the data features; let the number of fields included in each piece of data be n, let k represent the sequential number of the fields, and calculate the health score through the following health assessment model; φ t =δ*|t a -t0| Among them, S i is the health score of the i-th data; v ij is the jth data feature of the i-th data; α j is the feature weight corresponding to the jth data feature, indicating the influence of the data feature on the score; φ t is the time dynamic correction term, where δ is the time correction coefficient, t a is the current time, t0 is the data collection time; ψ is the abnormal compensation coefficient, which is used to adjust the abnormal impact in the score; Var(f k ) is the variance of the kth field, which measures the field volatility.
8. The multi-source data fusion method based on cloud computing according to claim 7 is characterized in that: The semantic consistency index Index is obtained as follows: r is used to represent the order number of the fusion rule, and Ng is used to represent the number of fusion rules; C r is the confidence of the rth fusion rule, W r To integrate the rule weights, the semantic consistency index is calculated using the following formula: The confidence of the fusion rule refers to the ratio of the number of effective fusions to the total number of fusions.
9. The multi-source data fusion method based on cloud computing according to claim 8, characterized in that: The semantic similarity is obtained as follows: A set of conflicting data is recorded as (D i ,D s ), and measure conflicting data by semantic similarity (D i ,D s ) to assist in determining the conflicting data. The semantic similarity Sim(D i ,D s ): Among them, E ′ (D i ) and E′(D s ) are data D i and D s The semantic embedding vector of ′ (D i )∥ is the vector modulus, used for normalization: μ is the distance weighting coefficient; Dist(D i ,D s ) is the characteristic distance, f ij and f sj The data D i and D s is the j-th eigenvalue of , and m is the total number of features.
10. The multi-source data fusion system based on cloud computing is characterized by: include: The data acquisition module acquires real-time data from multiple data sources through a cloud computing interface, constructs a multi-dimensional feature vector of the real-time data, generates a health score using a health assessment model, and generates a health label for each piece of data based on the health score; A data processing module is used to perform format conversion and semantic enhancement processing on the data marked as healthy, and generate semantically enhanced structured data; The rule construction and adjudication module builds fusion rules and calculates the confidence of fusion rules through knowledge graph reasoning based on semantically enhanced structured data; The semantic consistency index is constructed based on the confidence and semantic similarity of the fusion rules, and the semantic consistency index is used to determine the fusion result for conflicting data. The verification and optimization module verifies the quality of the fusion results through historical data and business rules. If the fusion results do not meet the preset requirements, the feedback optimization process is triggered.
Citation Information
Patent Citations
Perception data semantic standardization and fusion processing cloud platform
CN117668759A
Multi-source twin data fusion tunnel structure health monitoring and early warning method and system
CN119129077A
Cited By
Rule base dynamic construction method and device based on large language model and medium
CN120179811A
Method, device and medium for dynamically constructing a rule base based on a large language model
CN120179811B
Desulfurization prediction data preprocessing method based on box plot-knowledge experience double drive
CN120337014A
Bill business online financing management system and method based on Internet platform
CN120543285A
Operation management system and method based on financial marketing big data
CN120580071A