Data authority evaluation method in multi-source social governance data fusion scene
Through large-scale model technology, the authoritativeness of multi-source social governance data is evaluated, combined with the degree of metadata and domain correlation, the problem of low efficiency in data authoritative evaluation in the existing technology is solved, and efficient data fusion and cross-departmental sharing are achieved.
Patent Information
- Application Number
- CN202510355196.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-18
AI Technical Summary
In the existing multi-source social governance data fusion scenario, the data authoritative evaluation method is inefficient, and it is impossible to accurately evaluate the authoritativeness of data, and it cannot be effectively shared and utilized across departments, resulting in poor data fusion effect.
Using big model technology, by extracting metadata and inputting fine-tuning big model, we evaluate the matching scores between the data and the social governance business field, the professional authority and the standardization of the expression form of the data source department, etc., and combine the correlation degree of the data field and the data source department level to conduct authoritative quantitative evaluation.
It improves the accuracy and efficiency of data authoritative evaluation, ensures the authority and reliability of data fusion in multi-source social governance, and enhances the ability to share and utilize data across departments.
Smart Images

Figure CN120339017A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field related to data processing, and more specifically, relates to a method for evaluating data authority in the scenario of multi-source social governance data fusion. Background Art
[0002] At present, the situation of preventing and controlling social governance risks in China is still severe. Due to reasons such as inconsistent data standards, incompatibility between systems, and inconsistent data authority departments, numerous data islands have been formed, restricting information sharing and collaborative work among different departments. In order to achieve cross-departmental data sharing and efficient utilization, it is necessary to carry out data collaboration and fusion among multiple departments of social governance.
[0003] However, the sources of social governance data are complex and diverse, the quality is uneven, and the authority departments are inconsistent. There may be problems such as large format differences, content duplication, and data conflicts in data from different sources. Around social governance elements such as people, places, and events, it is necessary to evaluate and screen the authority of data in the process of multi-source social governance data fusion. In the field of social governance, data authority is of indispensable importance. Data authority is related to the credibility, reliability of social governance data, and the effectiveness of cross-departmental information sharing. High-authority data can provide a data basis for grass-roots governance and high-level decision-making, and help optimize resource allocation and improve the overall social governance ability. Therefore, it is of great significance to evaluate the data authority during the process of multi-source social governance data fusion, screen out high-authority data from the fields to be fused and conduct fusion.
[0004] The existing methods for evaluating the authority of multi-source social governance data mainly use rules set manually for data extraction, manually defined hierarchical classification rules for data sources, and evaluate whether the data meets the corresponding level or type according to the classification level to evaluate the data authority. This type of method has low efficiency and cannot evaluate the data authority outside the manually defined scope; or conduct data quality evaluation from aspects such as data integrity and data timeliness, without considering the business fields to which social governance data belongs, and the feasibility of quantifying data authority in terms of the business association degree between data and data source departments, resulting in poor evaluation effects of data authority during the fusion process and unable to guarantee data authority well. Therefore, in the scenario of multi-source social governance data fusion process, how to combine the actual business scenarios in the field of social governance to more accurately evaluate the authority of multi-source social governance data and ensure the data authority during the process of multi-source social governance data fusion is an urgent problem to be solved. Summary of the Invention
[0005] In view of the above defects or improvement requirements of the prior art, the present invention provides a method for evaluating data authority in the scenario of multi-source social governance data fusion, aiming to propose a method for evaluating data authority to evaluate data authority in the scenario of multi-source social governance data fusion and ensure the authority of multi-source social governance fusion data.
[0006] To achieve the above object, according to one aspect of the present invention, there is provided a method for evaluating data authority in the scenario of multi-source social governance data fusion, including:
[0007] Extract the field information to which the data to be evaluated belongs, the data table information to which the field belongs, and the database information where the data table is located as the metadata of the data to be evaluated; wherein, the field information includes the field name, field type, and field description, and the data table information includes the name and description of the data table; the database information includes the name, description, collection method, and source department or institution information of the database;
[0008] Input the metadata into the fine-tuned large model to obtain the matching scores between the data to be evaluated and all preset social governance business fields, the professional authority scores of the source department or institution of the data to be evaluated, and the weights of the evaluation scores of the standardization degree of the data representation form, the data field association degree, and the source department level of the data to be evaluated;
[0009] Based on the matching scores between the data to be evaluated and all preset social governance business fields, determine the evaluation scores of the data field association degree between the data to be evaluated and the social governance business field to which the target fusion data belongs through proportional calculation; determine the evaluation scores of the source department level based on the professional authority scores; obtain the evaluation scores of the standardization degree of the data representation form by comparing and verifying the name, type, and data format of the fields to which the data to be evaluated belongs with industry specifications; based on the weights of the evaluation scores, perform weighted summation on the evaluation scores to obtain the quantitative evaluation scores of the authority of the data to be evaluated.
[0010] Further, adopt prompt engineering optimization, input social governance domain knowledge into the pre-trained large model, and fine-tune the large model. The implementation method is as follows:
[0011] Obtain the domain knowledge of the target social governance domain, mark the training data and input it into the large model in the form of prompts; each piece of training data includes: the field information of a piece of social governance data, the data table information to which the field belongs and the database information where the data table is located, the social governance business domain to which the piece of social governance data belongs, the business scope and rights list of the source department to which the piece of social governance data belongs, the matching score of the piece of social governance data with all preset social governance business domains, the professional authority score of the source department or institution of the piece of social governance data, and the weights of the evaluation scores of the standardization degree of the presentation form, the degree of data domain association, and the level of the data source department of the piece of social governance data;
[0012] The field information includes the field name, field type, and field description. The data table information includes the name and description of the data table. The database information includes the name, description, collection method, and source department or institution information of the database.
[0013] Furthermore, the determination method of the evaluation score of the standardization degree of the presentation form is as follows:
[0014] By comparing the semantic similarity between the name of the field to which the data to be evaluated belongs and the industry data specification, determine the field name matching degree score; by comparing the field type of the field to which the data to be evaluated belongs with the industry data specification, obtain the field type consistency score; use an automated tool to check whether the field to which the data to be evaluated belongs conforms to the expected format, and obtain the data format consistency score; add the field name matching degree score, the field type consistency score, and the data format consistency score, and the added result is used as the evaluation score of the standardization degree of the presentation form.
[0015] Furthermore, the evaluation score of the degree of data domain association is expressed as:
[0016]
[0017] Among them, S is the evaluation score of the degree of data domain association, m is the matching score of the data to be evaluated and the target fusion data in the social governance business domain, and M max is the maximum value among the matching scores of the data to be evaluated and all social governance business domains; the matching scores of the data to be evaluated and all social governance business domains are all output by the large model.
[0018] Furthermore, the evaluation score of the level of the data source department is expressed as:
[0019] S = w a A + w p P
[0020] Among them, S is the evaluation score of the level of the data source department; A is the evaluation score of the administrative level of the data source department or institution to be evaluated, which is given a preset score according to the administrative level of the source department to which the data to be evaluated belongs; P is the professional authority score; w a and w p are the weights corresponding to A and P respectively, and satisfy w a + w p = 1.
[0021] Furthermore, the weights obtained from the fine-tuned large model also include the weight of the evaluation score of the credibility of data collection;
[0022] The method for determining the evaluation score of the credibility of data collection is as follows: determine the evaluation score of the transparency of the data collection method to be evaluated, the reliability evaluation score of the collection tool, and the traceability evaluation score of the collection process, and perform a weighted sum of each evaluation score to obtain the evaluation score of the credibility of data collection; each evaluation score is determined through a preset table.
[0023] According to another aspect of the present invention, a multi-source social governance data fusion method is provided, including:
[0024] Through semantic matching, match the fields of the data tables in each data source with the fields of the target fusion data, determine the fields with the same semantics as the target fusion data fields from each data source, and obtain a one-to-many semantic matching relationship between each target fusion data field and the source data fields; divide the source data fields having a matching relationship with the target fusion data field into a heterogeneous same-semantic field set, where each heterogeneous same-semantic field set contains the metadata of each field;
[0025] Regarding each data of each field in the heterogeneous same-semantic field set as a data to be evaluated respectively, for the field to be filled in the data record of each target fusion data, use the data authority evaluation method described above to obtain the authoritative quantitative evaluation score of each data of this field; splice the data with the highest score in this field with the unique identifier of the social governance element entity of the target fusion data to generate a data record guaranteed by authoritative quantitative evaluation.
[0026] According to another aspect of the present invention, an electronic device is provided, including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the above-mentioned method steps and / or the above-mentioned method steps are implemented.
[0027] According to another aspect of the present invention, there is provided a computer-readable storage medium, which includes a stored computer program. When the computer program is run by a processor, it controls the device where the storage medium is located to execute the method steps described above and / or the method steps described above.
[0028] According to another aspect of the present invention, there is provided a computer program product, including a computer program or instruction. When the computer program or instruction is executed by a processor, it implements the method steps described above and / or the method steps described above.
[0029] Generally speaking, compared with the prior art by the above technical solution conceived by the present invention, the technical solution provided by the present invention mainly has the following beneficial effects:
[0030] 1. The present invention proposes an index for quantitatively evaluating the authority of social governance data. The evaluation dimension includes the standardization degree of the presentation form, the degree of association with data fields, and the level of the data source department. Further, the large model technology is introduced. The matching scores of the data to be evaluated and all preset social governance business fields, the professional authority scores of the data source department or institution of the data to be evaluated, and the weights of the evaluation scores of the standardization degree of the presentation form, the degree of association with data fields, and the level of the data source department of the data to be evaluated are obtained by the large model. Based on the matching scores of the data to be evaluated and all preset social governance business fields, the evaluation score of the degree of association with data fields of the data to be evaluated and the social governance business field to which the target fusion data belongs is determined by proportional calculation. The professional authority score is used as the evaluation score of the level of the data source department. By comparing and verifying the name, type, and data format of the fields to which the data to be evaluated belongs with the industry specifications, the evaluation score of the standardization degree of the presentation form is obtained. Based on the weights of the evaluation scores, the evaluation scores are weighted and summed to obtain the quantitative evaluation score of the authority of the data to be evaluated. This method improves the evaluation process and quantitative method of the authority of social instruction data, and improves the effect of evaluating the authority of social governance data.
[0031] 2. The quantitative evaluation index for evaluating the authority of social governance data proposed by the present invention also includes the evaluation of the credibility of data collection. The credibility of data collection is determined based on the transparency of the data collection method, the reliability of the collection tool, and the traceability of the collection process, which improves the reliability of the fusion. Description of the Drawings
[0032] Figure 1 It is a flowchart of a method for evaluating the authority of data in a multi-source social governance data fusion scenario provided by an embodiment of the present invention;
[0033] Figure 2A general framework diagram for data authority evaluation in the scenario of multi-source social governance data fusion provided by the embodiments of the present invention;
[0034] Figure 3 A schematic flowchart of a data fusion method in the scenario of multi-source social governance data fusion provided by the embodiments of the present invention;
[0035] Figure 4 A schematic flowchart of constructing a set of heterogeneous and synonymous fields in the scenario of multi-source social governance data fusion provided by the embodiments of the present invention. Detailed implementation manners
[0036] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0037] Embodiment 1
[0038] A data authority evaluation method in the scenario of multi-source social governance data fusion, as Figure 1 shown, includes:
[0039] Extract the field information to which the data to be evaluated belongs, the data table information to which the field belongs, and the database information where the data table is located as the metadata of the data to be evaluated; wherein, the field information includes the field name, field type, and field description, and the data table information includes the name and description of the data table; the database information includes the name, description, collection method, and source department or institution information of the database;
[0040] Input the metadata into the fine-tuned large model to obtain the matching scores of the data to be evaluated with all preset social governance business fields, the professional authority scores of the source department or institution of the data to be evaluated, and the weights of the evaluation scores of the standardization degree of the data presentation form, the data field association degree evaluation score, and the data source department level evaluation score;
[0041] Based on the matching scores of the data to be evaluated with all preset social governance business fields, determine the data field association degree evaluation score of the data to be evaluated with the social governance business field to which the target fusion data belongs through proportional calculation; take the professional authority score as the data source department level evaluation score; obtain the evaluation score of the standardization degree of the data presentation form by comparing and verifying the name, type, and data format of the fields to which the data to be evaluated belongs with the industry specifications; based on the weights of the various evaluation scores, perform weighted summation on the various evaluation scores to obtain the authoritative quantitative evaluation score of the data to be evaluated.
[0042] In the scenario of multi-source social governance data fusion, there is an urgent need to solve a technical problem: for multi-source social governance data centered around social governance elements such as people, land, and events, propose a data authority evaluation method and a guarantee method to conduct data authority evaluation in the scenario of multi-source social governance data fusion and guarantee the authority of multi-source social governance fusion data. Based on the importance of data authority social governance data in the data processing process, this embodiment proposes a data authority evaluation method applicable to social governance data, providing method support for the data authority evaluation level in the multi-source social governance data fusion scenario. Based on the defects of existing evaluation methods, this embodiment proposes indicators for quantitatively evaluating the authority of social governance data, where the evaluation dimensions include the standardization degree of the presentation form, the degree of association with the data field, and the level of the data source department; further introduce large model technology to obtain the matching scores between the data to be evaluated and all preset social governance business fields, the professional authority scores of the data source departments or institutions of the data to be evaluated, and the weights of the evaluation scores of the standardization degree of the presentation form, the degree of association with the data field, and the level of the data source department of the data to be evaluated from the large model; based on the matching scores between the data to be evaluated and all preset social governance business fields, determine the evaluation score of the degree of association with the data field of the data to be evaluated and the social governance business field to which the target fusion data belongs through proportional calculation; use the professional authority score as the evaluation score of the level of the data source department; obtain the evaluation score of the standardization degree of the presentation form by comparing and verifying the name, type, and data format of the fields to which the data to be evaluated belongs with industry specifications; based on the weights of each evaluation score, sum the weighted evaluation scores to obtain the quantitative evaluation score of the authority of the data to be evaluated. This method improves the social governance data authority evaluation process and quantification method, and improves the effect of social governance data authority evaluation.
[0043] Among them, extract the metadata of the data to be evaluated, specifically as follows: extract and analyze the field information: specifically, accurately extract the field name and field description; check whether the physical type of the field (such as string, numerical value, date) is a regular type or a legal type; sort out the structure of the data table, record key information such as the name of the data table, the total number of fields, and the description of the table; sort out the primary and foreign key constraint relationships between this data table and other tables in the database where this table is located, clarify the data flow and dependency structure; analyze the database where the data table is located to clarify the data collection method (including information about the collection tool) of the database (pre-manually marked, preset).
[0044] The standardization degree of the representation form is determined by the comparative analysis of the consistency of the data format, field names, and attributes between the data to be evaluated and the industry norms or standards; the degree of data domain association is determined by analyzing the matching degree between the data attributes and the social governance business fields to which the target integrated data belongs, and clarifying its actual applicability in the target scenario; the level of the data source department is evaluated based on the level or power scope of the data source department in the social governance administrative system to assess its impact on the data authority.
[0045] The calculation formula for the quantitative evaluation score of the authority of the data to be evaluated is: where n is the number of evaluation dimensions, Score is the quantitative evaluation score of authority, w i is the dimension weight of the i-th dimension, f i is the evaluation score of the i-th dimension, and
[0046] It should be noted that for the dimension evaluation results output by the large model, manual correction can be carried out in combination with the opinions of authoritative domain experts.
[0047] Preferably, prompt engineering optimization is adopted to input social governance domain knowledge into the pre-trained large model for fine-tuning of the large model. The implementation method is as follows:
[0048] Obtain the domain knowledge of the target social governance domain, construct training samples and input them into the large model in the form of prompts; among them, each training sample includes: the field information of a piece of social governance data, the data table information to which the field belongs, the database information where the data table is located, the social governance business field to which the piece of social governance data belongs, and the business scope and rights list of the source department to which the piece of social governance data belongs;
[0049] The field information includes the field name, field type, and field description (i.e., field semantic description), the data table information includes the name and description of the data table, and the database information includes the name, description, collection method, and source department or institution information of the database.
[0050] Based on the metadata of a piece of social governance data, the large model calculates and outputs the matching score between the data to be evaluated and all preset social governance business fields, the professional authority score of the data source department or institution of the data to be evaluated, and the weights of the evaluation scores of the standardization degree of the representation form, the degree of data domain association, and the level of the data source department of the data to be evaluated, based on the social governance business field to which the piece of social governance data belongs and the business scope and rights list of the source department to which the piece of social governance data belongs.
[0051] In specific implementation, prompt engineering optimization is utilized to input domain knowledge to fine-tune the large model. After that, the fine-tuned large model is used for analysis and evaluation. Further, domain expert review can be introduced for correction and supplementation to improve the evaluation results of the content that the large model may miss or misjudge. The weights of the matching scores between the data to be evaluated and all preset social governance business domains, the professional authority scores of the departments or institutions where the data to be evaluated comes from, the evaluation scores of the standardization degree of the presentation form of the data to be evaluated, the evaluation scores of the data domain association degree, and the evaluation scores of the levels of the departments where the data comes from are output to ensure the output of the final authoritative evaluation results of social governance data.
[0052] The main steps for fine-tuning the large model and using the fine-tuned large model for analysis and evaluation may include:
[0053] S1. Plan the selection and coverage areas of data sources, obtain social governance domain knowledge corresponding to the data sources and areas, organize and analyze the social governance domain knowledge, and input it into the large model;
[0054] As a further preference of the present invention, the large language model is ChatGPT4, Qwen2.5 or LLaMA3;
[0055] S2. Label the training data to make the model suitable for the task requirements of data authority evaluation: The process of labeling the training data includes: extracting, organizing and recording the field names, field types, field descriptions, information of the data tables to which they belong, database information, and the departments or institutions of the sources from each data source database, as well as the administrative levels, business scopes and rights lists of the departments or institutions of the sources; cleaning and converting the non-standard or inconsistent format data, removing the missing, duplicate and abnormal data, and unifying the data format; balancing the data from each source in the training data to cover all social governance business domains and ensuring the distinguishability of its matching degree with all preset domains; based on the expert evaluation scores, scoring and weight labeling should be carried out from four dimensions: the credibility of data collection, the standardization degree of the presentation form, the data domain association degree, and the level of the departments where the data comes from.
[0056] S3. Design prompts: The designed prompts need to cover the tasks of dimensional quantitative evaluation and calculation of comprehensive evaluation score weights, including: continuously improving the model evaluation ability and adaptability through prompt optimization methods such as dynamic weight adjustment, multi-round dialogue design, few-shot learning combination and task modularization.
[0057] Specifically, as an optional implementation, for the roles and tasks played by the large model, the following prompt words can be designed: "You are an AI model named 'Data Governance Evaluation Expert', focusing on the evaluation of the authority of social governance data. Your tasks are: gradually analyze the source of the input data and its corresponding business scenarios in the field of social governance, and analyze the relevance between the data to be evaluated and the given business scenarios in the field of social governance; extract relevant information dimension by dimension according to the authority evaluation criteria, and obtain the evaluation scores of each dimension through logical reasoning, and comprehensively calculate the data authority evaluation score; based on the evaluation score results, put forward optimization suggestions (if applicable). Input requirements: I will provide you with data descriptions (including data sources, field descriptions, table uses) and business backgrounds (specific application scenarios of the data), as well as evaluation criteria (dimension evaluation scores and weights related to data authority). Your output needs to include the following structure through step-by-step reasoning: clearly analyze the authority and reliability basis of the data source, gradually analyze the fit between the data and the business scenarios, quantify the degree of domain association of the data and the evaluation scores of the data source department level dimension (in the form of a table) and the appropriate dimension weights for the four dimensions, and finally output the results and conclusion analysis through the multi-dimensional authority calculation formula, and at the same time give possible optimization suggestions to enhance the governance value of social governance data".
[0058] S4. Input the labeled training data into the large model through the prompt words, and perform few-shot fine-tuning on the large model to enable the large model to have the ability to generate matching scores, evaluation scores, and weights. The fine-tuning strategy focuses on optimizing the authority evaluation ability of the model, using multi-task learning and iterative optimization. In the data construction stage, ensure the reasonable distribution of samples with different authority scores (as an optional implementation, select 10,000 initial labeled data and expand it to 100,000 enhanced data); at the same time, freeze the parameters of the basic layer and only fine-tune the high-level weights to retain the language understanding ability. During the actual training process, improve the applicability of the model in various scenarios by dynamically adjusting the sample ratio and adding new training data. In addition, combine the authority evaluation with the domain adaptation task, and achieve efficient fine-tuning by sharing model parameters to ensure that the model can achieve excellent performance in different tasks.
[0059] S5. Based on the optimized large model, use the large model to infer and calculate the weights of each dimension, and combine the data authority quantification evaluation formula to calculate the data authority score to obtain the evaluation result of the data authority.
[0060] As a preferred implementation, the semantic similarity between the name of the field to which the data to be evaluated belongs and the industry data specification is compared to determine the field name matching score; the field type of the field to which the data to be evaluated belongs is compared with the industry data specification to obtain the field type consistency score; an automated tool is used to verify whether the field to which the data to be evaluated belongs conforms to the expected format to obtain the data format consistency score; the field name matching score, the field type consistency score, and the data format consistency score are added together, and the sum is used as the evaluation score for the standardization degree of the presentation form. Among them, the industry data specification includes the national standard "Data Quality Management Specification" (GB / T 35089-2018) and / or the local standard "Data Quality Management Specification" (GB / T 35089-2018).
[0061] As a preferred implementation, the evaluation score for the degree of association in the data field is expressed as:
[0062]
[0063] Among them, S is the evaluation score for the degree of association in the data field, m is the matching score between the data to be evaluated and the social governance business field to which the target fusion data belongs, and M max is the maximum value among the matching scores between the data to be evaluated and all social governance business fields; the matching scores between the data to be evaluated and all social governance business fields are all output by the large model.
[0064] As a preferred implementation, the evaluation score for the level of the data source department can be directly the above-mentioned professional authority score, or it can be:
[0065] S = w a A + w p P
[0066] Among them, S is the evaluation score for the level of the data source department; A is the evaluation score for the administrative level of the data source department or institution to be evaluated, which is given a preset score according to the administrative level of the source department to which the data to be evaluated belongs; P is the professional authority score; w a and w p are the weights corresponding to A and P respectively, and satisfy w a + w p = 1.
[0067] The evaluation score for the administrative level of the data source department or institution to be evaluated is given a preset score according to the administrative level (none, national, provincial, municipal, etc.) of the source department to which the data to be evaluated belongs, and the authority decreases with the level.
[0068] As a preferred implementation, the weights obtained by the fine-tuned large model also include the weight of the evaluation score for the credibility of data collection;
[0069] The method for determining the evaluation score of the credibility of data collection is as follows: determine the evaluation score of the transparency of the data collection method to be evaluated, the evaluation score of the reliability of the collection tool, and the evaluation score of the traceability of the collection process, and perform weighted summation on each evaluation score to obtain the evaluation score of the credibility of data collection; each evaluation score is determined through a preset table.
[0070] This embodiment proposes an index for evaluating the quantitative evaluation of the authority of social governance data, and its evaluation dimension also includes the evaluation of the credibility of data collection. The credibility of data collection is determined based on the transparency of the data collection method, the reliability of the collection tool, and the traceability of the collection process.
[0071] The evaluation of the credibility of data collection takes transparency, reliability, and traceability as the core, assigns different weights, and according to the formula S = w t T + w r R + w c C to calculate the overall credibility evaluation score, where S is the evaluation score of the credibility of data collection, T is the evaluation score of the transparency of the data collection method, R is the evaluation score of the reliability of the collection tool, C is the evaluation score of the traceability of the collection process, w t , w r , w c are the weights corresponding to T, R, and C, and satisfy w t + w r + w c = 1.
[0072] As a preferred implementation, the evaluation score of the transparency of the collection method is calculated based on the following four items in terms of transparency: whether there is a system document to support the manual collection method, whether there is a public description document, whether the responsibilities of the third-party data source department or institution have been disclosed on the government public platform. If all four items are met, the transparency evaluation score is 1.0; if three items are met, it is 0.8; if two items are met, it is 0.6; if one item is met, it is 0.3; if none of the items are met, the evaluation score is 0; in terms of reliability, the evaluation score of the reliability of the collection tool is calculated based on the following five items: (1) whether the accuracy and stability of the collection tool meet the business requirements; (2) whether the manual collection personnel have qualifications; (3) whether the collection time, personnel, and location information are recorded; (4) combined with the creation and modification time of the data table, determine whether the data is continuously updated; (5) whether there is a primary-foreign key dependency relationship between the data table and other tables in the database to form a data closed-loop; where each item met is assigned 0.2, and if all five items are met, the score is 1.0, and if none of them are met, the score is 0; the evaluation score of the traceability of the collection process is calculated based on the following four items: whether the time, location, collection responsible entity, and original record in the collection process are recorded. If the above four items are met, the basic evaluation score is 0.8, if three items are met, it is 0.6, if two items are met, it is 0.4, if one item is met, it is 0.2, and if none of the items are met, the score is 0.
[0073] The standardization degree of the representation form takes the field semantic consistency, field type consistency and data format consistency as the core and assigns different weights; among them, the field semantic consistency means inputting the field name into a pre-trained language model (as a preferred implementation, using the BERT model) and obtaining its vector representation, and calculating the cosine similarity with the standard field name vector in the industry data specification to obtain the matching score; the field type consistency is to compare the standard field type in the industry data specification to check whether the type of the data to be evaluated matches; the data format consistency is based on the regular expression rules for writing the data format in the industry specification or data quality standard, and uses an automated tool to detect whether the field conforms to the specification format. The representation form standardization evaluation score is calculated according to the following formula:
[0074] S = w e E + w t T + w f F
[0075] where E is the semantic consistency evaluation score, T is the type consistency evaluation score, and F is the format consistency evaluation score; w e , w t , w f are the weight coefficients of each dimension, satisfying: w e + w t + w f = 1. Among them, the semantic consistency score evaluation score, type consistency evaluation score and traceability evaluation score can adjust the weights corresponding to the evaluation scores manually according to the data to be evaluated.
[0076] As a preferred implementation, the value range of the field semantic consistency matching score is [-1, 1]; if the field type to be evaluated matches the standard type, the field type consistency is set to 1 point, otherwise it is set to 0 point; use an automated tool to detect whether the field conforms to the specification format. If it conforms, the data format consistency is set to 1 point, otherwise it is set to 0 point; the weights of the three evaluation scores are set to the same value, and the sum of the weights is equal to 1.
[0077] The overall framework of this embodiment can be seen in Figure 2 .
[0078] Embodiment 2
[0079] A multi-source social governance data fusion method, including:
[0080] Through semantic matching, the fields of the data tables in each data source are semantically matched with the fields of the target fusion data, and the fields with the same semantics as the target fusion data fields are determined from each data source, obtaining a one-to-many semantic matching relationship between each target fusion data field and the source data fields; the fields of each data source that have a matching relationship with the target fusion data field are divided into a set of fields with the same semantics from different sources, where each set of fields with the same semantics from different sources contains the metadata of each field.
[0081] Regarding each data of each field in the set of fields with the same semantics from different sources as a data to be evaluated respectively, and adopting the data authority evaluation method described in Embodiment 1 above, obtaining the authoritative quantitative evaluation scores of each data of this field; splicing the data with the highest score in this field with the target fusion data.
[0082] This embodiment provides a fusion method with data authority guarantee in the multi-source social governance data fusion scenario. The purpose is to improve the existing multi-source data fusion process based on the data characteristics in the social governance field, and conduct social governance data authority evaluation on the fields from different data sources with the same semantics under the same entity. Finally, fields are screened and fused to generate a highly credible and highly authoritative fusion data set. As Figure 3 shown, it can be divided into the following steps:
[0083] S1. According to the needs of social governance business, determine the categories of social governance elements in the target fusion data (such as individuals, locations, events, organizations), select the data sources from which data needs to be extracted, and determine the data table structure in the target fusion data;
[0084] Among them, the needs of social governance business refer to the social governance business such as grass-roots data statistics, government macro decision-making, discovery of social risk groups, and prevention of potential risk events by social governance departments based on social governance elements. Social governance elements include people, places, events, things, and organizations.
[0085] As an alternative implementation, for the community personnel statistics service, first determine that the social governance element is an individual, select multiple relevant department data sources related to the social governance element of this person for data fusion, and determine that the target fusion data includes an individual basic information table, an individual family information table, an individual work information table, etc.; for the key place supervision service, determine that the social governance element is a location, select relevant department data sources related to the social governance element of this location for data fusion, and determine that the target fusion data includes a location basic information table, a place business registration information table, a place historical group case event table, etc.; for the contradiction and dispute service, determine that the social governance element is an event, select relevant department data sources related to the social governance element of this event; determine that the target fusion data includes multiple event information tables; according to the needs of the social governance service, design the data table structure in the target fusion data as required, including the data table name, the included fields, and the primary key field, and the primary key field is the unique identifier of the social governance element entity (as an alternative implementation, the unique identifier of a person can be the legally obtained ID number).
[0086] S2. Extract the metadata of the data tables in each data source, including: table name, table description, table primary key, foreign key relationship constraints, field name, field type, and field description; by analyzing the inclusion relationship between the data tables and fields in each data source, as well as the primary key and foreign key relationship constraints, find the field where the unique identifier of this social governance element is located in each data source and all fields associated with this social governance element; and after extracting, merging, and removing duplicates for all values in the field where the unique identifier of this social governance element is located in each data source, store them in the data table of the target fusion data.
[0087] S3. Through semantic matching and manual assistance methods, semantically match the fields of the data tables in each data source with the fields of the target fusion data, find the fields with the same semantics as the target fusion data in each data source, and output the semantic matching relationship for the next step of social governance data authority evaluation and multi-source data fusion.
[0088] As an alternative implementation, use the text-embedding-v3 (general text vector based on large language model) model to obtain the embedding vector; convert the data type of the field into text, and pass the table name, table description, field type, field name, and field description to which the field belongs into the general text vector model, and the model will convert the original text into a high-quality embedding vector; prepare candidate field labels, and define the embedding vectors of each field in the target fusion data as type labels; create a zero-shot classifier, calculate the similarity between the embedding vectors of each source field and the candidate field labels, and select the most matching field label to obtain the semantic matching result.
[0089] Preferably, the methods of manual assistance include: fine-tuning the parameters of the embedding model, replacing the pre-trained classification model, manually correcting the results of semantic matching, manually correcting the wrongly matched ones and manually supplementing the missed matches.
[0090] S4. Using the one-to-many semantic matching relationship between the target fusion data fields and each source data field in S3, for each target fusion data field, divide the source data fields having a matching relationship with it into a set of heterologous and synonymous fields, where each set contains the field metadata of semantically matching fields extracted from each data source.
[0091] S5. For a data record in the data table of the fusion target data, the primary key field of which is the unique identifier of the social governance element stored in the above steps, use the "set of heterologous and synonymous field data" in S4 and this unique identifier to extract the field values of the synonymous fields in each data source, and use the method for quantitatively evaluating the authority of social governance data proposed by the present invention to quantitatively evaluate the authority of the data, so as to obtain the authority evaluation score Score of the field; select the field value of the field with the optimal authority and splice it with the primary key field to generate a data record of a target fusion data. By performing the above operations on all records in the data table of the target fusion data, a new fusion data table can be obtained, and finally the target fusion data of this type of social governance element can be obtained, and the problem of making a choice when there is a conflict between the synonymous fields of multi-source data in the data fusion process is solved, and the authority of the data in the process of multi-source social governance data fusion is guaranteed.
[0092] Illustrated with a specific example, such as Figure 4 shown:
[0093] Taking the generation of the fusion data of personal basic information as an example, first determine that the social governance element required for the target fusion data is "personnel", select the data sources as data source A, data source B and data source C, and determine the table structure of the data table in the target fusion data. Some of the relevant fields that should be included in the table are: ID number, household registration location, place of residence, marital status and health status, where the ID number is the unique identifier of the personnel;
[0094] Extract the data table metadata in each data source, analyze the fields included in the data tables in each data source, find the field where the unique identifier of this type of social governance element is located, extract its field value and perform merging and deduplication processing, and store it in the personal basic information table in the target fusion data as its primary key field;
[0095] As an optional implementation, the household register population table is stored in data source A, and the fields related to this embodiment include ID card number, household registration location, place of residence, and marital status; the grid personnel table is stored in data source B, and the fields related to this embodiment include ID card number, place of residence, and marital status; the regional personnel table is stored in data source C, and the fields related to this embodiment include ID card number, place of residence, health status, and marital status.
[0096] Specifically, in the metadata of the household register personnel table of data source A, the table name is "Household Register Population Information", the table description is "Basic household register information with household registration in the jurisdiction", and the fields include: the field name is idcard_number, the field type is varchar(30), and the field description is ID card number; the field name is domicile_address, the field type is varchar(180), and the field description is household registration location; the field name is residence_full_address, the field type is varchar(180), and the field description is the administrative region name of the current long-term place of residence; the field name is rtial_rkfs, the field type is int, and the field description is marital status;
[0097] Specifically, in the metadata of the grid personnel table of data source B, the table name is "Grid Personnel Information", the table description is "Basic information of personnel with place of residence registered in the community grid", and the fields include: the field name is cardid, the field type is varchar(64), and the field description is certificate number; the field name is residence_address, the field type is varchar(255), and the field description is place of residence; the field name is marital_status, the field type is varchar(2), and the field description is marital status.
[0098] Specifically, in the metadata of the regional personnel table of data source C, the table name is "Regional Personnel Information", the table description is "Basic information of personnel with place of residence registered in the jurisdiction of this department", and the field definitions include: the field name is idcard, the field type is varchar(45), and the field description is ID card number; the field name is residence_place, the field type is varchar(30), and the field description is residential address; the field name is health_condition, the field type is varchar(2), and the field description is current health status; the field name is marital_condition, the field type is int, and the field description is current marital status.
[0099] Through semantic matching and manual assistance methods, perform one-to-many semantic matching between the fields required by the target fusion data and the fields in the data tables of each data source; taking the residence field as an example, the above method matches and generates a set of fields with the same semantics from different sources, which includes the field residence_full_address in data source A, the residence field residence_address in data source B, and the residence field residence_place in data source C.
[0100] Sequentially retrieve a data record from the personal basic information in the target fusion data, where the primary key field of the data record is the ID number, and the previously stored ID number data is stored, while the household registration location, residence, marital status, and health status are empty; for each empty field to be filled in this record, use the ID number to find the data record of this person from the household registration personnel information table in data source A, the grid personnel information table in data source B, and the regional personnel table in data source C.
[0101] For each field to be filled in the data record, use the "set of fields with the same semantics from different sources" in the above process to find the corresponding field and obtain the field value from the data record of this person in each data source, and respectively perform data authority evaluation on the field values of each data source; obtain the authority score Score1 of the field residence_full_address in data source A, the authority score Score2 of the field residence_address in data source B, and the authority score Score3 of the field residence_place in data source C; compare the three authority scores and select the field value with the highest authority score to fill in this field of this record. Finally, generate a data record guaranteed by authoritative quantitative evaluation, and for each generated data record, finally generate the target fusion data of the "personnel" social governance elements guaranteed by the authoritative evaluation method.
[0102] The semantic matching and data authority evaluation methods for the household registration location, marital status, and health status are similar and will not be elaborated here. It should be noted that the field information in the above data sources is obtained legally.
[0103] The method of this embodiment improves the existing multi-source social governance data fusion. By combining semantic matching and manual assistance methods to construct metadata, identify and extract a set of fields with the same semantics from different sources, and perform field-level reconstruction and fusion on social governance elements and their associated field sets based on the results of social governance data authority evaluation scores. It not only realizes the effective integration of multi-source social governance data, but also realizes the data authority of multi-source social governance data, enhances the authority and credibility of social governance fusion data, and provides strong data support for downstream tasks in the field of social governance.
[0104] The related technical solutions are the same as above and will not be elaborated here.
[0105] Embodiment III
[0106] This application also relates to an electronic device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.
[0107] The electronic device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The so-called processor can be a Central Processing Unit (CPU), or other general-purpose processors, Digital Signal Processors (DSPs), Application-Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The memory can be used to store computer programs and / or modules. By running or executing the computer programs and / or modules stored in the memory, and calling the data stored in the memory, various functions of the electronic device can be realized.
[0108] The related technical solutions are the same as above and will not be elaborated here.
[0109] Embodiment IV
[0110] This application also relates to a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above method are implemented.
[0111] Specifically, the memory can include high-speed random access memory, and can also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0112] The related technical solutions are the same as above and will not be elaborated here.
[0113] Embodiment V
[0114] An embodiment of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps of the method in the above embodiment of the present application.
[0115] The related technical solutions are the same as above and will not be elaborated here.
[0116] It is easy for those skilled in the art to understand that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for evaluating data authority in the scenario of multi-source social governance data fusion, characterized in that Including: Extracting the field information to which the data to be evaluated belongs, the data table information to which the field belongs, and the database information where the data table is located as the metadata of the data to be evaluated; among them, the field information includes the field name, field type, and field description, and the data table information includes the name and description of the data table; the database information includes the name, description, collection method, and source department or institution information of the database; Inputting the metadata into the fine-tuned large model to obtain the matching scores of the data to be evaluated with all preset social governance business fields, the professional authority score of the source department or institution of the data to be evaluated, and the weights of the evaluation scores of the standardization degree of the presentation form of the data to be evaluated, the degree of data field association evaluation score, and the source department level evaluation score of the data to be evaluated; Based on the matching scores of the data to be evaluated with all preset social governance business fields, determining the degree of data field association evaluation score of the data to be evaluated with the social governance business field to which the target fusion data belongs through proportional calculation; determining the source department level evaluation score based on the professional authority score; obtaining the evaluation score of the standardization degree of the presentation form by comparing and verifying the name, type, and data format of the fields to which the data to be evaluated belongs with the industry data specifications; based on the weights of the evaluation scores, weighted summing the evaluation scores to obtain the authoritative quantitative evaluation score of the data to be evaluated.
2. The data authority evaluation method according to claim 1, characterized in that, Adopting prompt engineering optimization, inputting social governance domain knowledge into the pre-trained large model, and fine-tuning the large model. The implementation method is as follows: Obtaining the domain knowledge of the target social governance domain, marking the training data and inputting it into the large model in the form of prompts; among them, each piece of training data includes: the field information of a piece of social governance data, the data table information to which the field belongs, and the database information where the data table is located, the social governance business field to which the piece of social governance data belongs, the business scope and rights list of the source department to which the piece of social governance data belongs, the matching scores of the piece of social governance data with all preset social governance business fields, the professional authority score of the source department or institution of the piece of social governance data, and the weights of the evaluation scores of the standardization degree of the presentation form of the piece of social governance data, the degree of data field association evaluation score, and the source department level evaluation score of the piece of social governance data; The field information includes the field name, field type, and field description, the data table information includes the name and description of the data table, and the database information includes the name, description, collection method, and source department or institution information of the database.
3. The data authority evaluation method according to claim 1, wherein The determination method of the evaluation score of the standardization degree of the presentation form is: By comparing the semantic similarity between the name of the field to which the data to be evaluated belongs and the industry data specification, the field name matching degree score is determined; by comparing the field types of the field to which the data to be evaluated belongs and the industry data specification, the field type consistency score is obtained; an automated tool is used to check whether the field to which the data to be evaluated belongs conforms to the expected format, and the data format consistency score is obtained; the field name matching degree score, the field type consistency score, and the data format consistency score are added together, and the added result is used as the evaluation score of the standardization degree of the presentation form.
4. The data authority evaluation method according to claim 1, wherein The evaluation score of the data domain association degree is expressed as: Among them, S is the evaluation score of the data domain correlation degree, m is the matching score of the data to be evaluated and the social governance business domain to which the target fusion data belongs, and M max is the maximum value among the matching scores of the data to be evaluated and all social governance business domains; the matching scores of the data to be evaluated and all social governance business domains are output by the large model.
5. The data authority evaluation method according to claim 1, wherein The evaluation score of the level of the data source department is expressed as: S = w a A + w p P Among them, S is the evaluation score of the level of the data source department; A is the evaluation score of the administrative level of the data source department or institution to be evaluated, which is given a preset score according to the administrative level of the source department to which the data to be evaluated belongs; P is the professional authority score; w a and w p are the weights corresponding to A and P respectively, and satisfy w a +w p = 1.
6. The data authority evaluation method according to any one of claims 1 to 5, characterized in that, Among the weights obtained by the fine-tuned large model, there is also the weight of the evaluation score of the credibility of data collection; The method for determining the evaluation score of the credibility of data collection is: determining the transparency evaluation score of the data collection method to be evaluated, the reliability evaluation score of the collection tool, and the traceability evaluation score of the collection process, and weighted summing the evaluation scores to obtain the evaluation score of the credibility of data collection; each evaluation score is determined through a preset table.
7. A multi-source social governance data fusion method, characterized in that, Including: Through semantic matching, the fields in the data tables of each data source are semantically matched with the fields of the target fusion data, and the fields with the same semantics as the fields of the target fusion data are determined from each data source, obtaining a one-to-many semantic matching relationship between each field of the target fusion data and the fields of the source data; the fields of each data source with a matching relationship with the field of the target fusion data are divided into a set of heterologous fields with the same semantics, where each set of heterologous fields with the same semantics contains the metadata of each field. Each data of each field in the set of heterologous fields with the same semantics is used as a data to be evaluated, and the data authority evaluation method described in any one of claims 1 to 6 is adopted to obtain the authoritative quantitative evaluation score of each data of the field; the data with the highest score in the field is spliced and fused with the target fusion data.
8. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the method steps described in any one of claims 1 to 6 and / or the method steps described in claim 7.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program, wherein when the computer program is run by a processor, it controls the device where the storage medium is located to execute the method steps described in any one of claims 1 to 6 and / or the method steps described in claim 7.
10. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instruction is executed by a processor, it implements the method steps described in any one of claims 1 to 6 and / or the method steps described in claim 7.
Citation Information
Cited By
Unified social credit code data quality control method based on traceability technology
CN120975806A
Contract extraction text quality evaluation method and device
CN121561346A