A method and system for quality assessment and encryption of multi-source data fusion in smart cities

By performing sensitivity labeling and encryption on multi-source data from smart cities, combined with quality assessment and trust level assessment, a verifiable proof chain is generated, which solves the dual requirements of data security and quality assessment, and realizes secure data integration and accurate assessment.

CN120849404BActive Publication Date: 2026-01-06JIANGSU FENGYUN TECH SERVICE CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511358598.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-23
Publication Date
2026-01-06
Estimated Expiration
2045-09-23

AI Technical Summary

Technical Problem

Existing technologies have concentrated security risks in multi-source data processing in smart cities. The overall data encryption lacks flexibility, cannot dynamically adjust security strategies based on data quality, and the quality assessment methods are not accurate or comprehensive enough to meet the dual requirements of high-quality and high-security data in complex scenarios.

Method used

By acquiring raw data from multiple functional departments, performing sensitivity labeling and encryption, generating a dataset to be merged, acquiring and normalizing quality indicator data, performing weighted calculations and trust level assessments, generating a verifiable proof chain, and achieving secure data fusion and quality assessment.

Benefits of technology

It achieves secure integration of data from different departments, protects sensitive information from leakage, quantifies the quality and credibility of each data point, generates a verifiable chain of proof, and ensures the accuracy and security of the fusion results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120849404B_ABST
    Figure CN120849404B_ABST
Patent Text Reader

Abstract

This disclosure provides a method and system for quality assessment and encryption of multi-source data fusion in smart cities. Applied to the fields of data security and trusted computing technology, the method includes: acquiring raw data from various functional departments, performing sensitivity marking and encryption to generate a dataset to be fused; normalizing the quality indicators of each data point to generate quality vectors and binding them to the data to form a quality-labeled dataset; weighting the quality vectors to obtain fusion weights and generating a weighted quality encrypted dataset; performing trust level assessment and fusion calculation in the encrypted domain, simultaneously generating zero-knowledge proofs, forming a verifiable proof chain, and sending it to the verifier. This solution can effectively integrate data from different departments while protecting sensitive information from leakage. It can also quantify the quality and trustworthiness of each data point, allowing third parties to verify the correctness and security of the fusion results without accessing the raw data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the fields of data security and trusted computing technology, and in particular to a method and system for quality assessment and encryption of multi-source data fusion in smart cities. Background Technology

[0002] During the accelerated development of smart cities, the demand for data in urban governance and public services has surged, making the integration of multi-source data increasingly crucial for improving urban operations, resource allocation, and the scientific nature of decision-making. Smart cities encompass multiple fields, and various functional departments accumulate massive amounts of raw data. While this data contains information about urban operations, it suffers from problems such as scattered storage, inconsistent formats, and significant differences in security and quality.

[0003] Currently, when processing multi-source data for smart cities, a centralized data aggregation approach is often adopted, which collects data from different departments into a central server and then uses simple statistical methods or models based on preset rules for quality assessment. Data encryption is performed as a whole before data fusion.

[0004] However, the centralized processing of existing technologies is prone to creating security risk hotspots. The overall data encryption lacks flexibility, cannot dynamically adjust security strategies based on data quality, and the quality assessment methods are not accurate or comprehensive enough to meet the dual requirements of high-quality and high-security data in the complex scenarios of smart cities. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this disclosure provides a method and system for quality assessment and encryption of multi-source data fusion in smart cities. This disclosure solves the problems of existing technologies, such as centralized processing leading to concentrated security risks, lack of flexibility in overall data encryption, inability to dynamically adjust security strategies based on data quality, and insufficient precision and comprehensiveness in quality assessment methods, making it difficult to meet the dual requirements of high-quality and high-security data in the complex scenarios of smart cities.

[0006] According to a first aspect of this disclosure, a method for quality assessment and encryption of multi-source data fusion in smart cities is provided, comprising: acquiring raw data from multiple functional departments in a smart city, performing sensitivity labeling and encryption processing on the raw data, and generating a dataset to be fused; wherein the dataset to be fused includes a non-sensitive auxiliary dataset and a ciphertext dataset.

[0007] Obtain the quality index data corresponding to each original data, normalize the quality index data to obtain the quality vector corresponding to each original data, and bind the quality vector to the dataset to be fused to obtain the quality labeled dataset.

[0008] The quality vectors in the quality-labeled dataset are weighted to obtain the fusion weight of each original data, and the fusion weight is bound to the encrypted dataset to obtain the weighted quality encrypted dataset.

[0009] Trustworthiness level assessment and ciphertext domain fusion processing are performed based on a weighted quality ciphertext dataset, and zero-knowledge proof generation is performed to obtain a verifiable proof chain, which is then sent to the verifier.

[0010] According to a second aspect of this disclosure, a smart city multi-source data fusion quality assessment and encryption system is provided, for performing the method as described in the first aspect, comprising: a data preparation module for acquiring raw data from multiple functional departments in a smart city, performing sensitivity marking and encryption processing on the raw data, and generating a dataset to be fused; wherein the dataset to be fused includes a non-sensitive auxiliary dataset and a ciphertext dataset.

[0011] The quality labeling module is used to obtain the quality indicator data corresponding to each original data, perform normalization processing on the quality indicator data to obtain the quality vector corresponding to each original data, and bind the quality vector to the dataset to be fused to obtain the quality labeled dataset.

[0012] The weight calculation module is used to perform weighted calculation on the quality vectors in the quality-labeled dataset to obtain the fusion weight of each original data, and bind the fusion weight to the encrypted dataset to obtain the weighted quality encrypted dataset.

[0013] The fusion verification module is used to perform trust level assessment and ciphertext domain fusion processing based on the weighted quality ciphertext dataset, generate zero-knowledge proofs, obtain a verifiable proof chain, and send the verifiable proof chain to the verifier.

[0014] According to a third aspect of this disclosure, an electronic device is provided, comprising: a memory and a processor, wherein the memory stores a computer program, and the processor executes the program to implement the method described above.

[0015] In the smart city multi-source data fusion quality assessment and encryption method and system provided above, the embodiments of this disclosure can effectively integrate data from different departments, protect sensitive information from being leaked, quantify the quality and credibility of each data point, make the overall result more accurate and reliable through weighted fusion, and generate a verifiable proof chain, allowing third parties to confirm the correctness and security of the fusion result without having access to the original data. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 A schematic flowchart of a method for quality assessment and encryption of multi-source data fusion in smart cities according to an embodiment of the present disclosure is shown.

[0018] Figure 2 A schematic flowchart of a method for quality assessment and encryption of multi-source data fusion in smart cities according to an embodiment of the present disclosure is shown.

[0019] Figure 3 A schematic block diagram of a smart city multi-source data fusion quality assessment and encryption system according to an embodiment of the present disclosure is shown.

[0020] Figure 4 A block diagram of an exemplary electronic device according to an embodiment of the present disclosure is shown. Detailed Implementation

[0021] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values ​​of the components and steps set forth in these embodiments do not limit the scope of the present disclosure.

[0022] Those skilled in the art will understand that the terms "first," "second," etc., in the embodiments of this disclosure are only used to distinguish different steps, devices, or modules, and do not represent any specific technical meaning, nor do they indicate a necessary logical order between them. It should also be understood that in the embodiments of this disclosure, "multiple" can refer to two or more, and "at least one" can refer to one, two, or more. It should also be understood that any component, data, or structure mentioned in the embodiments of this disclosure can generally be understood as one or more unless explicitly defined or given a contrary indication in the context. Furthermore, the term "and / or" in this disclosure is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this disclosure generally indicates that the related objects are in an "or" relationship. It should also be understood that the descriptions of the various embodiments in this disclosure emphasize the differences between them; their similarities or commonalities can be referred to mutually, and for the sake of brevity, they will not be elaborated upon one by one.

[0023] Furthermore, it should be understood that, for ease of description, the dimensions of the various parts shown in the accompanying drawings are not drawn to actual scale. The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit this disclosure or its application or use. Techniques, methods, and apparatus known to those skilled in the art will not be discussed in detail, but where appropriate, such techniques, methods, and apparatus should be considered part of the specification. It should be noted that similar reference numerals and letters in the following drawings denote similar items; therefore, once an item is defined in one drawing, it need not be further discussed in subsequent drawings.

[0024] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0025] Figure 1 This is a schematic flowchart illustrating a method for quality assessment and encryption of multi-source data fusion in smart cities, provided as an embodiment of this disclosure. The method may include the following steps:

[0026] S101, acquire raw data from multiple functional departments in the smart city, perform sensitivity labeling and encryption processing on the raw data, and generate a dataset to be merged; wherein, the dataset to be merged includes a non-sensitive auxiliary dataset and a ciphertext dataset.

[0027] A smart city can be a new form of city. Its core is to apply next-generation information technologies such as the Internet of Things, big data, and artificial intelligence to all aspects of city operations, such as road traffic, public services, energy management, and even environmental protection.

[0028] Functional departments can be government agencies or enterprises that play specific roles in the smart city system, responsible for public management or services in their respective fields, such as transportation, environmental protection, and energy departments.

[0029] Raw data can be first-hand data directly generated and recorded by various functional departments in their daily work, without any processing or reprocessing, such as vehicle trajectories from the transportation department, air quality readings from the environmental protection department, or medical records from the medical department.

[0030] The dataset to be merged can be a collection obtained after preprocessing the original data. Specifically, before entering the merging stage, the original data needs to be marked for sensitivity and encrypted. Sensitive parts are encrypted, while non-sensitive parts can be retained in plaintext or lightly anonymized. Finally, these data are organized together to form a dataset that can be securely used for cross-departmental collaboration.

[0031] Non-sensitive ancillary datasets are records that do not involve privacy or sensitive details. They are typically used to support subsequent analysis or model calculations, such as summaries of road traffic, publicly available air quality indicators, and total electricity consumption. Even if the data in non-sensitive ancillary datasets is used publicly, it will not leak personal privacy.

[0032] A encrypted dataset can be a dataset formed by encrypting data fields identified during the sensitivity marking process. It can contain personal identification, specific medical details, and precise location trajectories. This data is processed into ciphertext using homomorphic encryption or attribute-based encryption, so that even if it is accessed during the fusion process, it cannot be directly restored to the original information.

[0033] Data can be collected from various functional departments, such as traffic congestion and traffic flow statistics, environmental pollution monitoring, energy power load, and medical public health data. After collection, an integrity check should be performed to identify any missing fields or inconsistent formats. Timestamps and source identifiers should be compared to ensure the data is credible in terms of time sequence and source. Next, each data point should be labeled with sensitivity data. Fields involving privacy, such as personal identification characteristics, health details, or records with geographic location features, can be identified and marked as sensitive information according to pre-defined rules. Data like average air quality index and regional power load can be classified as non-sensitive information. Sensitive information is then encrypted. Homomorphic encryption can be used to ensure it can still participate in computations in its encrypted state, or attribute-based encryption can be used to restrict decryption permissions. This part is summarized separately to form an encrypted dataset. Non-sensitive information is retained in plaintext or, if necessary, undergoes mild anonymization, such as numerical range formatting or obfuscation, and then compiled into a non-sensitive auxiliary dataset. Finally, these two types of data are structurally bound together so that each record contains both the encrypted content of the sensitive parts and the auxiliary information of the non-sensitive parts. All such records together constitute the complete dataset to be merged.

[0034] Based on the above technical solution, optionally, sensitivity labeling and encryption processing are performed on the original data to generate a dataset to be fused, including:

[0035] The original data is marked with sensitivity information to obtain an encrypted protected dataset and a non-sensitive auxiliary dataset.

[0036] The encrypted protected dataset is encrypted using a preset encryption method to obtain a ciphertext dataset. The non-sensitive auxiliary dataset is then bound to the ciphertext dataset to generate a dataset to be fused.

[0037] In this scheme, the encrypted protected dataset can be a collection of sensitive information selected from the original data, such as personal privacy or other confidential content, which is collected separately to form a special set.

[0038] The preset encryption method can be a pre-defined encryption technique used to protect sensitive data before data processing begins. For example, homomorphic encryption can be used, allowing data to be processed directly even when encrypted, without prior decryption; attribute-based encryption can be used, where only qualified individuals can decrypt and view the data; symmetric or asymmetric encryption can also be selected as needed.

[0039] After acquiring the raw data provided by various functional departments in the smart city, sensitivity analysis can be performed on each data point. This involves identifying which content involves privacy or confidential information, such as personal identification information, health details, or records with geographic location data. These fields are marked as sensitive, while fields that do not involve personal privacy, such as overall traffic flow statistics, air quality index, or total electricity load, are marked as non-sensitive information. The data marked as sensitive is then aggregated to form an encrypted protected dataset, preparing for subsequent encryption processing. The non-sensitive fields are compiled to form a non-sensitive auxiliary dataset, used to assist analysis in fusion computing. After marking, a pre-defined encryption method, such as homomorphic encryption or attribute-based encryption, is applied to the sensitive fields in the encrypted protected dataset. This ensures that the data can still participate in computation in an encrypted state, but cannot be read without authorization, thus obtaining a ciphertext dataset. The non-sensitive auxiliary dataset remains in plaintext, and can be lightly anonymized or obfuscated if necessary to further reduce privacy risks. The encrypted dataset and the non-sensitive auxiliary dataset are structurally bound to each other, so that each data record contains both the encrypted sensitive part and retains the non-sensitive auxiliary information. All the bound data are then aggregated to form the final dataset to be fused.

[0040] This solution effectively integrates data from various departments while ensuring the security of sensitive information and preventing its leakage. This preserves the business value of the data and assists in analysis and decision-making.

[0041] S102, obtain the quality index data corresponding to each original data, perform normalization processing on the quality index data to obtain the quality vector corresponding to each original data, and bind the quality vector to the dataset to be fused to obtain the quality labeled dataset.

[0042] Quality indicators (QIs) can be quantitative parameters used to measure the quality of raw data. They focus on the reliability of the data itself, such as whether the data is complete, whether the field format is consistent, whether the records are updated in a timely manner, how large the difference is between the values ​​and the actual situation, and whether the source organization is reliable. These indicators can be recorded using numbers or scores.

[0043] A quality vector is a multi-dimensional vector that standardizes and transforms quality indicator data to make it easier to process and compare quality indicator data. In other words, it unifies each quality indicator into a value between 0 and 1 and then combines them together.

[0044] A quality-labeled dataset can be an extended dataset formed by binding each piece of data to be fused with its corresponding quality vector. Each piece of data not only contains business information, but also comes with a set of numerical labels that indicate the quality of the data.

[0045] Each piece of raw data may have quality issues during collection or generation, such as data incompleteness, inconsistent field formats, accurate timestamps, deviations from reality, and the past reliability of the source department. To quantify these issues, this information needs to be extracted to form quality indicators for each piece of data, with each indicator representing its quality level using numbers or scores. Next, to facilitate comparison and processing of these indicators, they need to be standardized, or normalized to a uniform range, such as 0 to 1. Normalization can be done by linear scaling based on the indicator's upper and lower limits, or by adjusting for outliers based on statistical distribution. This not only eliminates dimensional differences between indicators but also better handles extreme values, ensuring data consistency. Then, the normalized indicators for each piece of data are arranged sequentially to form a multi-dimensional vector—the quality vector, with each dimension corresponding to an indicator. This vector acts as a concise label, centralizing the previously scattered quality information. Finally, the quality vector of each piece of data is bound to its corresponding data to be merged, forming a quality-labeled dataset. In this way, each piece of data not only contains business information, but also comes with a set of numerical labels describing its quality level.

[0046] S103, perform weighted calculation on the quality vectors in the quality-labeled dataset to obtain the fusion weight of each original data, and bind the fusion weight to the ciphertext dataset to obtain the weighted quality ciphertext dataset.

[0047] Fusion weights assign a numerical value to each piece of original data, reflecting its proportion in the final fusion result. Higher-quality data pieces have greater fusion weights, meaning they have a stronger influence when calculating the fusion result.

[0048] A weighted quality ciphertext dataset can be a set formed by binding the fusion weight of each piece of original data together with its corresponding encrypted information.

[0049] After obtaining the quality-labeled dataset, the quality vector of each data point needs to be analyzed. Each vector contains several indicators, such as data integrity, field format consistency, timestamp timeliness, numerical accuracy, and the reliability of the source department. Based on pre-defined rules, different weights can be assigned to each indicator. For example, integrity and accuracy are usually more critical, so they have higher weights, while format consistency and timeliness have slightly lower weights. Then, the values ​​of each indicator for each data point are multiplied by their corresponding weights, and these are summed to calculate a total score. This score represents the importance of the data point in the fusion calculation, i.e., the fusion weight. To facilitate subsequent processing, these weights can be standardized so that the sum of all data weights equals 1. This makes it easier to allocate proportions in the fusion calculation and to intuitively compare which data is more important. After calculating the fusion weight, it is mapped to the encrypted data, that is, the encrypted sensitive information is bound to its fusion weight to form a complete record. In this way, even if the data is encrypted, it can still participate in the fusion calculation according to its weight, while sensitive information is also protected. By compiling all such entries together, we obtain a weighted quality encrypted dataset. Each data entry contains both the encrypted portion of the business information and a weight that reflects its role in the fusion process.

[0050] S104. Based on the weighted quality ciphertext dataset, perform trust level assessment and ciphertext domain fusion processing, and generate zero-knowledge proof to obtain a verifiable proof chain, and send the verifiable proof chain to the verifier.

[0051] A verifiable proof chain can be a record generated after completing a trust level assessment and ciphertext domain fusion calculation on a weighted quality ciphertext dataset. It not only preserves the final fusion result but also includes proof information that can verify the calculation process. Third parties can use this chain to confirm the correctness of the calculation and the reliability of the data without needing to examine the original sensitive data. It can be understood as a notarized chain, where each step of the fusion operation has a corresponding proof, making the entire result both traceable and verifiable.

[0052] The verifier can be a third-party entity that receives and checks this verifiable chain of proofs, such as a government regulatory agency, a data auditing department, or a trusted node in a smart city data management platform.

[0053] After obtaining the weighted quality-encrypted dataset, the credibility level of each data point's quality vector can be evaluated. The evaluation comprehensively considers the historical reliability of the data source, scores of various quality indicators, and the previously calculated fusion weights, using a weighted formula to calculate a comprehensive score. These weights can be determined based on historical experience or system settings; for example, historical reliability accounts for 40%, quality indicators for 35%, and fusion weights for 25%, or they can be adjusted according to the actual scenario. A higher score indicates more reliable data, while a lower score suggests potential anomalies or biases, requiring weight reduction in subsequent fusion processes. Each data point receives a credibility level, serving as a reference standard for subsequent encrypted domain fusion calculations. After the evaluation, fusion calculations can be performed in encrypted form, meaning that without decrypting sensitive information, relevant information from different data sources is weighted and superimposed according to the fusion weights or subjected to other statistical processing to obtain a unified fusion result. This process can use homomorphic encryption or computable encryption techniques to ensure data security during computation, while also allowing the results to be used for subsequent analysis. After the fusion calculation is complete, the fusion weight, credibility level, and final fusion result of each data point are input into a zero-knowledge proof algorithm to generate a corresponding proof. This allows external verifiers to confirm the correctness of the entire calculation process without seeing the original data. Finally, these proofs and the fusion result are organized chronologically and by each step of the calculation to form a verifiable proof chain. This chain not only records the source, credibility level, and specific calculation process of each data point but also ensures the traceability of the results. Once generated, the verifiable proof chain is sent to the verifiers via wireless communication technology, allowing them to check the accuracy of the fusion result and the compliance of the data processing without accessing any sensitive information.

[0054] In this embodiment, data from different departments can be effectively integrated while protecting sensitive information from being leaked. It can also quantify the quality and credibility of each piece of data, make the overall result more accurate and reliable through weighted fusion, and generate a verifiable proof chain, allowing third parties to confirm the correctness and security of the fusion result without having to access the original data.

[0055] Based on the above technical solution, optionally, after obtaining the verifiable proof chain, the method further includes:

[0056] Based on each original data, unique identity information and hash digest are determined. The unique identity information, hash digest and quality vector corresponding to each original data are bound together to generate quality traceability tags corresponding to each original data.

[0057] The quality traceability labels are bound to the fusion weights corresponding to each original data to obtain a quality traceability mapping table;

[0058] The verifiable proof chain is combined with the quality traceability mapping table to obtain the quality traceability certificate chain;

[0059] Accordingly, sending the verifiable proof chain to the verifier includes:

[0060] The quality traceability certificate chain is sent to the verifier.

[0061] In this scheme, the unique identification information can be a unique identifier generated for each piece of raw data, which can be understood as an ID card issued to each piece of raw data. It can be composed of the data source number, timestamp, or a randomly generated identifier, with the purpose of ensuring that it can be accurately found among a large amount of data.

[0062] A hash digest is a fixed-length string obtained by performing a hash function on the original data; it can be understood as a fingerprint of the data. By using a hash function to transform the original data into a fixed-length string, this string is virtually impossible to duplicate with other data. Therefore, by simply comparing the hash values, it's immediately possible to determine whether the data has been tampered with.

[0063] A quality traceability label can be a composite tag that binds unique identification information, a hash digest, and the quality vector of the data together. It can identify the source of the data, record the data quality, and ensure that the data content has not been tampered with, which is equivalent to attaching a traceable quality label to each piece of data.

[0064] A quality traceability mapping table can be built upon quality traceability labels, by binding the corresponding fusion weights to form a mapping relationship between labels and weights. This not only allows for the traceability of data sources and quality but also clarifies its contribution to the fusion computation.

[0065] A quality traceability credential chain can be a chain-like data structure generated by combining a verifiable proof chain and a quality traceability mapping table. On the one hand, it preserves the verifiability of the entire fusion computation process, and on the other hand, it provides traceability information on data sources, quality, and weights, thus allowing the verifier to fully verify the correctness and transparency of the fusion results without decrypting the original data.

[0066] Unique identification information is typically composed of elements such as the original data's ID, generation time, and sequence number. This ensures that each piece of data has a unique identifier within the entire system, preventing confusion. A hash digest is generated by performing calculations on the original data using a hash function, resulting in a fixed-length hash value. This value acts like a fingerprint; even a single character change will completely alter the hash value. With these two pieces of information, they are bound to the data's corresponding quality vector to generate a quality traceability tag. This tag contains proof of the data's identity and integrity, as well as an objective description of its quality; it's like attaching a label to the data that identifies its origin and measures its quality. Then, the quality traceability tag and the data's fusion weights are further combined to obtain a quality traceability mapping table. The significance of this step lies in linking information about who the data is and its quality with its contribution to the fusion calculation, allowing subsequent verifiers not only to see the data's tag itself but also to understand its contribution to the entire fusion process. Finally, the verifiable proof chain and the quality traceability mapping table are combined to form a quality traceability credential chain. The verifiable proof chain itself records the correctness and credibility of the fusion computation, while the mapping table supplements this with the identity and quality source information of each data point. Combining the two yields a complete credential chain that not only proves the computation was performed correctly but also traces the origin and role of each data point within the fusion process. This quality traceability credential chain is then sent to the verifier. The verifier doesn't need to access the original data; simply examining the credential chain confirms the compliance of the computation process, the verifiability of the quality of each data point, and the contribution allocation of each data source during the fusion process. This ensures both the credibility and transparency of the results.

[0067] This solution not only verifies the correctness of the fusion results, but also clearly traces the source, quality level, and actual contribution of each data point, making the results both credible and transparent, while avoiding the leakage of original sensitive data.

[0068] Based on the above technical solution, optionally, after sending the quality traceability certificate chain to the verifier, the method further includes:

[0069] Obtain the verification feedback dataset sent by the verification party, parse the feedback information of each original data in the verification feedback dataset, extract the number of low-quality data, the mean of fusion deviation, and the distribution of anomaly labels, and generate a performance index vector for each original data based on the number of low-quality data, the mean of fusion deviation, and the distribution of anomaly labels.

[0070] The performance index vectors of each original data are normalized to obtain normalized performance index vectors;

[0071] The normalized performance index vector is calculated according to the preset feedback weight adjustment rules to obtain the weight correction coefficients of each original data.

[0072] In this scheme, the verification feedback dataset can be a set of data compiled and fed back by the verifier after checking the verifiable proof chain or traceability certificate chain. It records their evaluation of the performance of each original data in the fusion process.

[0073] Feedback information can be a specific piece of raw data, indicating whether it has quality issues, whether there are errors during fusion, or whether it has been marked as an anomaly.

[0074] The "low quality count" refers to the number of times a piece of data is judged as substandard in various validation or fusion processes. A high count indicates that its long-term performance is relatively unstable.

[0075] The mean fusion bias is an average of the deviations of a data point across different fusion calculations, reflecting its overall reliability. A small bias indicates that the data is relatively consistent with the overall trend, while a large bias indicates that it is unreliable.

[0076] Anomaly labeling distribution can be the frequency or distribution of a data point being labeled as an anomaly. For example, whether it occasionally malfunctions or frequently malfunctions; this distribution can help determine whether the problem is sporadic or systemic.

[0077] A performance metric vector can be a set of numerical vectors that combine disparate feedback metrics such as the number of low-quality events, the mean of fusion bias, and the distribution of outlier labels. Each dimension represents a performance metric.

[0078] A normalized performance metric vector can be the result of standardizing a performance metric vector, mapping different metrics to a relatively comparable range, such as 0 to 1. This results in a fairer vector and facilitates subsequent calculations.

[0079] The weighting adjustment coefficient is an adjustment factor derived after processing the normalized vector. It determines whether a data point should have a larger or smaller proportion in the next fusion step. A higher value indicates greater reliability.

[0080] The preset feedback weight adjustment rules can be pre-defined calculation logic used to specify how to convert these normalized indicators into correction coefficients. These rules can be weighted averages, scoring rules, or threshold settings to ensure the results are both reasonable and stable.

[0081] The system can receive verification feedback datasets from the verification party, which may contain feedback information for each piece of raw data. Then, it parses each piece of feedback. This parsing process can utilize a data parser or a custom parsing script to extract key information from the feedback, such as recording the number of times each piece of raw data was judged as low-quality. This typically involves counting statistics or log scanning techniques. Simultaneously, it extracts the deviation between the fusion result and the baseline value. This is usually done by first calculating the deviation value and then averaging all deviations, which can be based on a numerical computing library, such as the NumPy mean function. The distribution of anomaly markers can be obtained using frequency distribution statistics, compiling the proportion and distribution of various anomaly markers to obtain the anomaly marker distribution. With these three types of metrics, a performance metric vector can be constructed for each piece of raw data. This vector contains elements such as the number of low-quality occurrences, the average fusion deviation, and the anomaly marker distribution. These metrics are then normalized to allow data of different dimensions to be compared on the same scale. Specifically, min-max normalization or z-score normalization can be used, depending on the data distribution characteristics. After normalization, a set of normalized performance metric vectors is obtained. Finally, based on the preset feedback weight adjustment rules, the normalized vector is calculated to obtain the weight correction coefficients for each original data point. This calculation can be a weighted summation or a combination of formulas. For example, the final correction value can be determined by comprehensively adjusting the proportion of low-quality occurrences in the overall evaluation, the weight of fusion bias on the stability of the results, and the weight of outlier distribution on data reliability. The final result is the weight correction coefficient for each original data point.

[0082] This solution quantifies data from different sources and of different quality using a unified set of indicators. After normalization and weight correction, it can intuitively identify which data is reliable and which is biased, thereby improving the accuracy and stability of overall data processing and model training.

[0083] Based on the above technical solution, optionally, after obtaining the weight correction coefficients of each original data, the method further includes:

[0084] The fusion weights are updated based on the weight correction coefficients, and the updated fusion weights are re-bound to the ciphertext dataset to obtain the updated weighted quality ciphertext dataset.

[0085] Based on the updated weighted quality ciphertext dataset, the trust level is reassessed and the ciphertext domain is fused. Zero-knowledge proof is then generated to obtain a verifiable proof chain, which is then resent to the verifier.

[0086] In this scheme, after obtaining the weight correction coefficients, they can be used to update the original fusion weights. Specifically, a weighted average or dynamic adjustment method can be used to make the new weights closer to actual performance. Then, these updated weights are re-bound to the encrypted dataset one by one. The entire process is completed under homomorphic encryption, ensuring that the data can be weighted even in an encrypted state without revealing plaintext. This results in a completely new weighted quality encrypted dataset. A new trust level assessment is then performed on this new dataset. Typically, secure multi-party computation or homomorphic operations are used to calculate the trustworthiness of each data point, obtaining the updated score. After the assessment, the data from different sources are fused again in the encrypted domain to ensure that the result reflects the corrected weight distribution, while still keeping all plaintext confidential. Finally, a zero-knowledge proof is generated based on the fused result. The zero-knowledge proof produces a complete and verifiable chain of proofs, demonstrating that all steps are true and valid, while preventing the verifier from accessing the original data. The generated proof chain is sent back to the verifier, thus forming a complete closed loop. This ensures the security and transparency of data processing while continuously optimizing subsequent evaluation and verification.

[0087] In this scheme, the updated fusion weights can more accurately reflect the actual contributions of each data source. At the same time, the entire process is completed in encrypted form, which not only ensures the privacy and security of the data, but also allows the verifier to confirm the credibility and transparency of the results, thereby improving the reliability and traceability of data fusion.

[0088] Based on the above technical solution, optionally, after sending the quality traceability certificate chain to the verifier, the method further includes:

[0089] If an abnormal feedback dataset is received from the verification party, the abnormal feedback dataset is parsed to obtain a list of abnormal data sources;

[0090] Based on the list of abnormal data sources, abnormal data is isolated from the dataset to be merged, resulting in a dataset of isolated abnormal data to be merged.

[0091] The quality index data corresponding to each original data in the dataset to be fused from isolated abnormal data is obtained. The quality index data is re-normalized to obtain the quality vector corresponding to each original data. The quality vector is then bound to the dataset to be fused to obtain a new quality labeled dataset.

[0092] The quality vectors in the quality-labeled dataset are re-weighted to obtain the fusion weights of each original data, and the fusion weights are bound to the encrypted dataset to obtain a new weighted quality encrypted dataset.

[0093] Based on the weighted quality ciphertext dataset, the trust level is reassessed and the ciphertext domain is fused, and zero-knowledge proof is generated to obtain a verifiable proof chain, which is then resent to the verifier.

[0094] In this scheme, the anomaly feedback dataset can be a set of feedback generated by the validator for the original data that exhibits serious anomalies or potential risks during the fusion process. It mainly records data sources that deviate extremely from expectations, may affect the reliability of the fusion results, or pose security risks. This type of dataset helps the system quickly identify key anomalies, isolate these data individually, or perform focused corrections, ensuring the accuracy and reliability of the overall fusion computation.

[0095] An anomalous data source list can be a catalog extracted from anomaly feedback datasets, listing the specific data sources—that is, which original data or data sources are considered anomalous. This list indicates which data requires special attention or temporary isolation.

[0096] Outlier data can be raw data marked in the list of outlier data sources, which may affect the overall quality or reliability of the fusion. This data may be missing information, have significant biases, or exhibit abnormal patterns, requiring separate isolation and processing.

[0097] The dataset to be fused, which has isolated outlier data, can be a collection of data formed by separating the identified outlier data from the original dataset to be fused. In this way, the remaining normal data can continue to be used for fusion calculations, while the outlier data is isolated, making it easier to re-evaluate the quality or make corrections.

[0098] After receiving the abnormal feedback data from the verification party, this feedback information can be analyzed to determine which original data had problems in the previous fusion process. The analysis involves identifying key anomaly indicators, such as values ​​with excessively large deviations, records contributing anomalies to the fusion results, or markers indicating potential security risks. These indicators are then mapped to specific data sources, resulting in a list of abnormal data sources that requires special attention. Based on this list, these abnormal data points need to be isolated from the original dataset to be fused, creating a separate dataset with isolated abnormal data. This allows normal data to continue with standard fusion calculations, while the abnormal data is isolated for later quality review and correction.

[0099] For this isolated data, it's necessary to re-acquire quality metrics for each record, such as completeness, consistency of field formats, timeliness, accuracy of values, and historical reliability of the data source. Normalizing these metrics, unifying different numerical standards into a comparable range, generates a quality vector for each data entry. Then, these quality vectors are linked back to their corresponding data entries to re-obtain the quality-labeled dataset. In this way, each data entry not only contains business information but also a clear quality description.

[0100] With the new quality-annotated dataset, the quality vector of each data point needs to be weighted. Based on the importance of each indicator or pre-defined weighting rules, the fusion weight of each data point is calculated; this weight reflects its contribution to the overall fusion. Then, the fusion weight is bound to the ciphertext data, resulting in an updated weighted quality ciphertext dataset. Each encrypted data point carries not only its weight but also its quality vector information. Trustworthiness assessment and ciphertext domain fusion are then performed again on this dataset. The trustworthiness assessment combines the quality vector, fusion weight, and the historical reliability of the data source to assign a trust score to each data point. The ciphertext domain fusion operation performs weighted accumulation or statistical calculations without decryption, and differential privacy noise can be added according to a privacy budget to ensure both data security and usability.

[0101] Finally, the updated weighted quality fusion result, the fusion weight of each data point, and the privacy constraints are input into the zero-knowledge proof algorithm to generate a new verifiable proof chain. This proof chain records each step of the fusion calculation, trust level assessment, and privacy protection measures, allowing the verifier to confirm the correctness of the fusion result and the accuracy of each data point's contribution without accessing the original sensitive data. After generation, this verifiable proof chain is sent to the verifier to check the reasonableness of the result.

[0102] This solution can promptly identify and isolate anomalous data, ensuring the accuracy of subsequent fusion calculations, while maintaining data security and privacy in encrypted form. The updated weighted quality encrypted dataset and verifiable proof chain allow verifiers to confirm the reliability of the results without accessing the original data, and also provide a transparent and credible basis for auditing and decision-making.

[0103] Figure 2 This is a flowchart illustrating the method for quality assessment and encryption of multi-source data fusion in smart cities provided in this embodiment. The method may include the following steps:

[0104] S201, acquire raw data from multiple functional departments in the smart city, perform sensitivity labeling and encryption processing on the raw data, and generate a dataset to be merged; wherein, the dataset to be merged includes a non-sensitive auxiliary dataset and a ciphertext dataset.

[0105] S202, obtain the quality index data corresponding to each original data, perform normalization processing on the quality index data to obtain the quality vector corresponding to each original data, and bind the quality vector to the dataset to be fused to obtain the quality labeled dataset.

[0106] S203, perform weighted calculation on the quality vectors in the quality-labeled dataset to obtain the fusion weight of each original data, and bind the fusion weight to the ciphertext dataset to obtain the weighted quality ciphertext dataset.

[0107] S204. Based on the quality vector of the weighted quality ciphertext dataset, determine the trust level of each original data, determine the privacy budget of each original data according to the trust level and the preset allocation rule, and embed the privacy budget into the weighted quality ciphertext dataset to obtain the fused constrained ciphertext dataset.

[0108] The quality vector of the weighted quality encrypted dataset and the quality vector of the original data use the same set of metrics, but the fusion weights are considered in the numerical values. It can be understood as a weighted quality vector of the original quality vector.

[0109] Trustworthiness level can be determined by assigning a reliability score to each data point based on the quality vector in the weighted quality encrypted dataset and the historical reliability of the data source. A high score indicates that the data is relatively reliable, while a low score suggests that the data may be problematic or biased.

[0110] The preset allocation rules can be principles or algorithms set in advance during the integration and privacy protection process, used to determine how much privacy budget can be allocated to each piece of data.

[0111] A privacy budget can be defined as the privacy resources allowed for each piece of data in fused computation. It is used to control the risk of leakage of sensitive information in encrypted computation or differential privacy processing, ensuring that the data can play its role without exceeding the scope of privacy and security.

[0112] A fused constrained ciphertext dataset can be a dataset formed by embedding the privacy budget of each data point into a weighted quality ciphertext dataset. In this way, the data can not only undergo weighted fusion computation in an encrypted state, but also meet pre-defined privacy constraints.

[0113] After obtaining the weighted quality encrypted dataset, we can examine the bound quality vectors to determine the reliability of each piece of original data. Specifically, we extract several key pieces of information, such as the historical stability and error rate of the data source, the current quality metric scores (including completeness, accuracy, and timeliness of updates), and their corresponding weights during fusion. Because these metrics originally had different ranges and units, they need to be normalized to unify them into a comparable range. Next, we assign a weight coefficient to each metric, which can be derived from expert experience. Then, we multiply the normalized metric by its coefficient and sum them to obtain a comprehensive score, which is considered the reliability level. The higher the comprehensive score, the more reliable the data. If the score is higher than the set upper limit, such as 0.8, it is classified as high reliability; between 0.5 and 0.8, it is considered medium reliability; and below 0.5, it is considered low reliability.

[0114] After establishing the trust levels, the next step is to allocate privacy budgets to data at different levels. This process follows a pre-defined allocation rule: highly trustworthy data is typically assigned a smaller privacy budget parameter to maintain good usability while protecting privacy. Medium-trustworthy data is assigned a moderate budget, while low-trustworthy data is assigned a larger budget to ensure stricter protection in subsequent computations. The privacy budget allocated to each data point is embedded into the weighted quality ciphertext dataset. Thus, each data point, in its ciphertext state, carries not only its original weights and quality information but also, for the first time, its privacy budget, resulting in a fused constrained ciphertext dataset. This dataset not only enables fused computations in an encrypted environment but also ensures that both privacy protection and trustworthiness assessment are effectively implemented.

[0115] S205, based on the fusion weights and privacy budget, perform fusion calculations on each data point of the fusion constraint ciphertext dataset within a preset ciphertext domain to obtain a weighted quality fusion result.

[0116] A pre-defined ciphertext field can be a computational space that allows operations to be performed even when the data is encrypted. Its purpose is to enable various calculations, such as addition, multiplication, or weighted calculations, while keeping the data encrypted. This often relies on homomorphic encryption or other computable encryption techniques. The ciphertext field acts like a secure computational field; external parties cannot see the true content of the data, but the results remain accurate and usable.

[0117] The weighted quality fusion result can be a unified output obtained by combining data from different sources within this encrypted domain according to pre-defined fusion weights and allocated privacy budgets. It reflects the role of each piece of data in the final result while adhering to privacy protection constraints.

[0118] After allocating the fusion weights and privacy budgets, the next step is to actually use this information in the fusion computation. Although each data point already carries a weight and budget, this is merely labeling and preparation; it doesn't mean the computation is complete. Therefore, a formal computation must be performed in the ciphertext domain to implement these constraints in the result. At this point, the data remains encrypted, and the computation environment is within a pre-defined ciphertext domain. This typically relies on homomorphic encryption or secure multi-party computation techniques to maintain the data's encryption while still allowing necessary mathematical operations. Then, these records are processed one by one. Each data point, in addition to the ciphertext of the business content, also carries the corresponding fusion weight and privacy budget. First, the ciphertext is multiplied by its weight, and then accumulated in the ciphertext domain to obtain the preliminary fusion result. This operation is supported by homomorphic encryption and will not compromise the data's encryption state. Then, the privacy budget is incorporated into the computation logic. For example, differential privacy can be used to add controlled noise to the result, making it difficult for anyone to deduce the specific value of any single sensitive data point even if they see the final fusion result. The privacy budget plays a dual role here: controlling the noise intensity and finding a balance between security and usability for the contribution of each data point. After the above steps, a weighted quality fusion result is finally formed, which integrates information from various data sources while ensuring that data credibility and privacy protection are both implemented.

[0119] S206, perform zero-knowledge proof generation on the weighted quality fusion result, fusion weight, and privacy budget to obtain a verifiable proof chain, and send the verifiable proof chain to the verifier.

[0120] After obtaining the weighted quality fusion result, the next step is to perform zero-knowledge proofs to convince external verifiers that the entire computation process is correct. Although the fusion result already contains the fusion weights and privacy budgets for each data point, this information still needs to be considered when generating the zero-knowledge proof. The weighted quality fusion result is like an exam report card; the report card already shows the total score, but the zero-knowledge proof is like a teacher proving that the scores and grading criteria for each subject were calculated correctly. Even seeing only the total score, one can be certain that there was no cheating in each subject. Specifically, the fusion result can be formatted for cryptographic computation, and then the fusion weights and privacy budgets for each data point can be mapped to the fusion result in a corresponding order. Then, a zero-knowledge proof algorithm is used to prove each step of the fusion computation, including weighted accumulation calculations, the addition of differential privacy noise, and the application of privacy constraints. This allows verifiers to confirm the correctness of the computation logic and results without accessing the original sensitive data. The entire process utilizes cryptographic techniques such as homomorphic encryption or zk-SNARKs to ensure that the proof information fully records the contribution of each data point and the privacy constraints followed, while the original data remains confidential. Finally, these proofs, fusion results, and corresponding fusion weights and privacy budgets are compiled into a complete verifiable proof chain, which is like recording the entire computational trajectory, with the credibility level and contribution of each step clearly verifiable. After generation, the verifiable proof chain is sent to the verifier via wireless communication technology. The verifier can check whether the entire fusion process is correct and meets the privacy protection and credibility level requirements without seeing the original data, providing a reliable basis for subsequent security audits or decisions.

[0121] In this embodiment, even if the original data remains encrypted, the verifier can confidently confirm the correctness of the fusion calculation and the contribution of each data point without worrying about privacy leaks. At the same time, the credibility, transparency, and security of the entire process are guaranteed, making data fusion both safe and reliable.

[0122] Figure 3 This is a schematic block diagram of a smart city multi-source data fusion quality assessment and encryption system provided in an embodiment of this disclosure. The system is characterized by comprising:

[0123] The data preparation module 301 is used to acquire raw data from multiple functional departments in the smart city, perform sensitivity marking and encryption processing on the raw data, and generate a dataset to be fused; wherein, the dataset to be fused includes a non-sensitive auxiliary dataset and a ciphertext dataset.

[0124] The quality labeling module 302 is used to obtain the quality indicator data corresponding to each original data, perform normalization processing on the quality indicator data to obtain the quality vector corresponding to each original data, and bind the quality vector to the dataset to be fused to obtain the quality labeling dataset.

[0125] The weight calculation module 303 is used to perform weighted calculation on the quality vectors in the quality labeled dataset to obtain the fusion weight of each original data, and bind the fusion weight with the encrypted dataset to obtain the weighted quality encrypted dataset.

[0126] The fusion verification module 304 is used to perform trust level assessment and ciphertext domain fusion processing based on the weighted quality ciphertext dataset, generate zero-knowledge proofs, obtain a verifiable proof chain, and send the verifiable proof chain to the verifier.

[0127] like Figure 4 As shown, this application embodiment also provides an electronic device 400, including a processor 401, a memory 402, and a program or instructions stored in the memory 402 and executable on the processor 401. When the program or instructions are executed by the processor 401, they implement the various processes of the above-described smart city multi-source data fusion quality assessment and encryption method embodiment and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0128] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0129] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described smart city multi-source data fusion quality assessment and encryption system embodiment, and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0130] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0131] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element. Furthermore, it should be noted that the scope of the methods and systems in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in reverse order, depending on the functions involved.

[0132] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0133] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

[0134] The above description is merely a preferred embodiment and the technical principles employed in this application. This application is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions that can be made by those skilled in the art will not depart from the scope of protection of this application. Therefore, although this application has been described in detail through the above embodiments, this application is not limited to the above embodiments, and may include more other equivalent embodiments without departing from the concept of this application, the scope of which is determined by the scope of the claims.

Claims

1. A smart city multi-source data fusion quality evaluation and encryption method, characterized in that, The method comprises: obtaining original data of multiple functional departments in a smart city, performing sensitivity marking processing and encryption processing based on the original data to generate a to-be-fused data set; wherein the to-be-fused data set comprises a non-sensitive auxiliary data set and a ciphertext data set; obtaining quality index data corresponding to each original data, performing normalization processing on the quality index data to obtain a quality vector corresponding to each original data, and binding the quality vector with the to-be-fused data set to obtain a quality marked data set; performing weighted calculation on the quality vector in the quality marked data set to obtain a fusion weight of each original data, and binding the fusion weight with the ciphertext data set to obtain a weighted quality ciphertext data set; performing trust level evaluation and ciphertext domain fusion processing based on the weighted quality ciphertext data set, and generating a zero-knowledge proof to obtain a verifiable proof chain, and sending the verifiable proof chain to a verifier; wherein the quality vector based on the weighted quality ciphertext data set is used to determine the trust level of each original data, the privacy budget of each original data is determined according to the trust level and a preset allocation rule, the privacy budget is embedded into the weighted quality ciphertext data set to obtain a fusion constraint ciphertext data set; based on the fusion weight and the privacy budget, performing fusion calculation on each data of the fusion constraint ciphertext data set in a preset ciphertext domain to obtain a weighted quality fusion result; performing zero-knowledge proof generation on the weighted quality fusion result, the fusion weight and the privacy budget to obtain a verifiable proof chain, and sending the verifiable proof chain to a verifier; The method further comprises: determining unique identity information and a hash digest based on each original data, binding the unique identity information, the hash digest and the quality vector corresponding to each original data to generate a quality traceability label corresponding to each original data; binding the quality traceability label with the fusion weight corresponding to each original data to obtain a quality traceability mapping table; combining the verifiable proof chain with the quality traceability mapping table to obtain a quality traceability credential chain; sending the quality traceability credential chain to the verifier.

2. The method of claim 1, wherein, Wherein, performing sensitivity marking processing and encryption processing based on the original data to generate a to-be-fused data set, comprising: performing sensitivity marking on the original data to obtain an encryption protection data set and a non-sensitive auxiliary data set; encrypting the encryption protection data set using a preset encryption method to obtain a ciphertext data set, and binding the non-sensitive auxiliary data set with the ciphertext data set to generate the to-be-fused data set.

3. The method of claim 1, wherein, Wherein, after sending the quality traceability credential chain to the verifier, the method further comprises: obtaining a verification feedback data set sent by the verifier, analyzing the feedback information of each original data in the verification feedback data set, extracting a low quality frequency, a fusion deviation mean and an abnormal marking distribution, and generating a performance index vector of each original data based on the low quality frequency, the fusion deviation mean and the abnormal marking distribution; performing normalization processing on the performance index vector of each original data to obtain a normalized performance index vector; According to a preset feedback weight adjustment rule, the normalized performance index vector is calculated to obtain a weight correction coefficient of each original data.

4. The method of claim 3, wherein, Wherein, After obtaining the weight correction coefficient of each original data, the method further comprises: Based on the weight correction coefficient, update the fusion weight, rebind the updated fusion weight with the ciphertext data set to obtain an updated weighted quality ciphertext data set; Based on the updated weighted quality ciphertext data set, re-perform the trusted level evaluation and the ciphertext domain fusion processing, and generate zero-knowledge proof to obtain a verifiable proof chain, and re-send the verifiable proof chain to the verifier.

5. The method of claim 1, wherein, Wherein, After sending the quality traceability certificate chain to the verifier, the method further comprises: If an abnormal feedback data set transmitted by the verifier is received, the abnormal feedback data set is parsed to obtain an abnormal data source list; Based on the abnormal data source list, isolate abnormal data in the to-be-fused data set to obtain a to-be-fused data set with isolated abnormal data; Obtain the quality index data corresponding to each original data in the to-be-fused data set with isolated abnormal data, re-normalize the quality index data to obtain a quality vector corresponding to each original data, and correspondingly bind the quality vector with the to-be-fused data set to re-obtain a quality labeled data set; Re-perform weighted calculation on the quality vector in the quality labeled data set to obtain the fusion weight of each original data, and bind the fusion weight with the ciphertext data set to re-obtain a weighted quality ciphertext data set; Based on the weighted quality ciphertext data set, re-perform the trusted level evaluation and the ciphertext domain fusion processing, and generate zero-knowledge proof to obtain a verifiable proof chain, and re-send the verifiable proof chain to the verifier.

6. A smart city multi-source data fusion quality evaluation and encryption system, characterized in that, The system comprises: A data preparation module for obtaining original data of multiple functional departments in a smart city, performing sensitivity marking and encryption processing based on the original data, and generating a to-be-fused data set; wherein the to-be-fused data set includes a non-sensitive auxiliary data set and a ciphertext data set; A quality labeling module for obtaining quality index data corresponding to each original data, performing normalization processing on the quality index data to obtain a quality vector corresponding to each original data, and correspondingly binding the quality vector with the to-be-fused data set to obtain a quality labeled data set; A weight calculation module for performing weighted calculation on the quality vector in the quality labeled data set to obtain the fusion weight of each original data, and binding the fusion weight with the ciphertext data set to obtain a weighted quality ciphertext data set; A fusion verification module for performing trusted level evaluation and ciphertext domain fusion processing based on the weighted quality ciphertext data set, and generating zero-knowledge proof to obtain a verifiable proof chain, and sending the verifiable proof chain to the verifier; wherein, based on the quality vector of the weighted quality ciphertext data set, the trusted level of each original data is determined, the privacy budget of each original data is determined according to the trusted level and a preset allocation rule, the privacy budget is embedded into the weighted quality ciphertext data set to obtain a fusion constraint ciphertext data set; Based on the fusion weight and the privacy budget, each data of the fusion constraint ciphertext data set is fused and calculated in a preset ciphertext domain to obtain a weighted quality fusion result; The weighted quality fusion result, the fusion weight and the privacy budget are subjected to zero-knowledge proof generation to obtain a verifiable proof chain, and the verifiable proof chain is sent to a verifier; The system is further configured to: Based on each original data, unique identity information and a hash digest are determined, and the unique identity information, the hash digest and a quality vector corresponding to each original data are bound to generate a quality traceability label corresponding to each original data; The quality traceability label is bound to a fusion weight corresponding to each original data to obtain a quality traceability mapping table; The verifiable proof chain and the quality traceability mapping table are combined to obtain a quality traceability credential chain; The quality traceability credential chain is sent to the verifier.

7. An electronic device, comprising: A processor, a memory and a program or instructions stored on the memory and executable on the processor are included, and the program or instructions are executed by the processor to implement the steps of the smart city multi-source data fusion quality evaluation and encryption method according to any one of claims 1-5.

8. A readable storage medium, characterized by, A program or instructions are stored on the readable storage medium, and the program or instructions are executed by the processor to implement the steps of the smart city multi-source data fusion quality evaluation and encryption method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Tea quality data safety management monitoring method and system

    CN120181876A

  • Data pricing method based on multi-dimensional value evaluation, terminal and storage medium

    CN120218976A

  • Encrypted data quality declaration verification method based on AI and privacy computing technology

    CN120372644A

  • Data quality treatment method and system, electronic equipment and storage medium

    CN120469901A