Data quality evaluation method based on government affair field

By obtaining business scenario portraits and dataset images of government data, the business stage division and quality feature vector quantification are solved, and the problems of low accuracy and insufficient generalization ability in government data evaluation are achieved, and more accurate and comprehensive data quality evaluation is achieved.

CN120387105AActive Publication Date: 2025-07-29JIANGXI YUNQING TECH CO LTD

Patent Information

Application Number
CN202510877734.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-07-29
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

In the prior art, the quality evaluation of government data relies on preset rules checks and single-dimensional evaluation, resulting in low accuracy of evaluation results and insufficient generalization capabilities, making it difficult to adapt to the blind spots and one-sidedness of evaluation in complex business logic and cross-departmental data flow.

Method used

By acquiring business scene portraits and data set images, dividing the data set images based on business scene portraits, determining business segmented data, and obtaining quality feature vectors based on the evaluation dimension system to generate comprehensive data quality evaluation results.

Benefits of technology

It has achieved improvement in government data processing capabilities, enhanced the accuracy and comprehensiveness of evaluation results, and can clarify the core direction of data quality evaluation at each stage and quantify the evaluation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387105A_ABST
    Figure CN120387105A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of government affair data evaluation, and particularly relates to a data quality evaluation method based on the government affair field, and the method comprises the steps: obtaining a business scene portrait and a data set image, and providing basic information for subsequent business stage division and quality evaluation; performing business stage division on the data set image based on the business scene portrait, determining business segment data corresponding to the data set image, and realizing structured classification of government affair data processing stages; obtaining a corresponding evaluation dimension system according to the service segment data, determining a quality feature vector corresponding to the service segment data according to the evaluation dimension system, determining a core direction of data quality evaluation of each stage, and quantifying an evaluation result; and generating a data quality comprehensive evaluation result based on the quality feature vector corresponding to the service segment data, forming an integral evaluation conclusion comprehensively reflecting the government affair data quality, enhancing the government affair data processing capability, and improving the accuracy of the evaluation result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of government data evaluation, and particularly relates to a data quality evaluation method based on the government field. Background Art

[0002] The processing of government data is the full-process management of government data from collection, cleaning, storage to analysis, sharing, and security protection. Data aggregation is achieved through multi-source data collection, data quality is improved through cleaning and standardization processing, secure storage is completed using databases and data warehouses, data value is mined with statistical analysis and machine learning, departmental barriers are broken to achieve data sharing to promote government collaboration, and at the same time, technologies such as encryption and access control are used to ensure data security, ultimately serving government decision-making, policy formulation, and public service optimization, and promoting the construction of digital government and the modernization of governance capabilities.

[0003] In the prior art, the quality evaluation of government data relies on preset rule verification and single-dimensional evaluation. Format rules and evaluation dimensions need to be set manually, parameter adjustment is required for adapting to new business scenarios, and evaluation blind spots and one-sidedness are prone to occur in complex business logics or cross-departmental data flows. There is a lack of multi-dimensional dynamic evaluation tools, resulting in low accuracy of evaluation results and insufficient generalization ability. Summary of the Invention

[0004] The embodiments of this application provide a data quality evaluation method based on the government field, which can solve the problems of lack of multi-dimensional dynamic evaluation tools in the process of government data evaluation, resulting in low accuracy of evaluation results and insufficient generalization ability.

[0005] In a first aspect, the embodiments of this application provide a data quality evaluation method based on the government field, including: Obtain a business scenario portrait and a dataset image; wherein, the business scenario portrait is used to reflect the current business type of government data processing, and the dataset image is used to reflect the data presentation form and its association relationship of government data in the process of government processing; Based on the business scenario portrait, divide the dataset image into business stages to determine the business segmented data corresponding to the dataset image; the business segmented data is used to reflect stages such as data application of government data; Obtain a corresponding evaluation dimension system according to the business segmented data, and determine a quality feature vector corresponding to the business segmented data according to the evaluation dimension system; wherein, the evaluation dimension system is used to reflect the core evaluation direction of data quality, and the quality feature vector is used to quantify the actual performance of data under each evaluation dimension; Generate a comprehensive evaluation result of data quality based on the quality feature vector corresponding to the business segmented data.

[0006] In the embodiments of the present application, the above technical solutions have at least the following technical effects: The data quality assessment method based on the government affairs field provided by the embodiments of the present application provides basic information for subsequent business stage division and quality assessment by obtaining the business scenario portrait and the dataset image. The business stage of the dataset image is divided based on the business scenario portrait, and the business segment data corresponding to the dataset image is determined, realizing the structured classification of the government affairs data processing stage. The corresponding evaluation dimension system is obtained according to the business segment data, and the quality feature vector corresponding to the business segment data is determined according to the evaluation dimension system, clarifying the core direction of data quality assessment in each stage and quantifying the assessment results. Based on the quality feature vector corresponding to the business segment data, a comprehensive data quality evaluation result is generated, forming an overall evaluation conclusion that comprehensively reflects the quality of government affairs data, enhancing the government affairs data processing ability and improving the accuracy of the assessment results.

[0007] In a second aspect, an embodiment of the present application provides a data quality assessment system based on the government affairs field, including: An acquisition unit, configured to acquire a business scenario portrait and a dataset image; wherein, the business scenario portrait is used to reflect the business type of the current processing of government affairs data, and the dataset image is used to reflect the data presentation form and its association relationship of government affairs data in the government affairs processing process; A segmentation unit, configured to divide the business stage of the dataset image based on the business scenario portrait, and determine the business segment data corresponding to the dataset image; the business segment data is used to reflect stages such as data application of government affairs data; A quality unit, configured to obtain a corresponding evaluation dimension system according to the business segment data, and determine a quality feature vector corresponding to the business segment data according to the evaluation dimension system; wherein, the evaluation dimension system is used to reflect the core evaluation direction of data quality, and the quality feature vector is used to quantify the actual performance of data in each evaluation dimension; A result unit, configured to generate a comprehensive data quality evaluation result based on the quality feature vector corresponding to the business segment data.

[0008] In a third aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the processor executes the computer program, the method described in any one of the above aspects is implemented.

[0009] In a fourth aspect, an embodiment of the present application provides a computer program product, which when running on an electronic device, causes the electronic device to execute the method described in any one of the above aspects.

[0010] It can be understood that the beneficial effects of the above second to fourth aspects can be referred to the relevant descriptions in the above aspects, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0012] Figure 1 is a schematic flowchart of a data quality assessment method based on the government affairs field provided by an embodiment of the present application; Figure 2 is a schematic operation diagram of a data quality assessment method based on the government affairs field provided by an embodiment of the present application; Figure 3 is a schematic structural diagram of a data quality assessment system based on the government affairs field provided by an embodiment of the present application; Figure 4 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0013] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0014] It should be understood that when used in the specification and the appended claims of the present application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0015] It should also be understood that the term "and / or" used in the specification and the appended claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0016] As used in the specification of this application and the appended claims, the term "if" may be construed, depending on the context, as "when" or "once" or "in response to determining" or "in response to detecting". Similarly, the phrases "if determined" or "if the described condition or event is detected" may be construed, depending on the context, to mean "once determined" or "in response to determining" or "once the described condition or event is detected" or "in response to detecting the described condition or event".

[0017] In addition, in the description of the specification of this application and the appended claims, the terms "first", "second", "third", etc. are only used for differentiating descriptions and shall not be construed as indicating or implying relative importance.

[0018] The reference to "one embodiment" or "some embodiments" or the like described in the specification of this application means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized in another way.

[0019] Currently, the publicly available technologies similar in the government affairs field mainly include data verification technology based on simple rules and single-dimensional data quality assessment technology. Data verification based on simple rules mainly verifies government affairs data by presetting some basic data format rules, such as whether the data types match, whether the required fields are empty, etc. For example, in the population information registration data, the age field is required to be numeric, and if a character appears, it is determined as a data quality problem. This technology can only detect some of the most basic and superficial data errors and is difficult to detect deeper quality problems such as data integrity, consistency, and relevance. The single-dimensional data quality assessment technology evaluates government affairs data in a certain dimension, such as only focusing on the accuracy of the data. By comparing with an authoritative data source, the accuracy of the data is judged. For example, in tax data, the declared tax amount of an enterprise is compared with the amount recorded in the actual tax system to evaluate the accuracy. This technology ignores other important dimensions of data quality, such as data timeliness, accessibility, etc., and cannot comprehensively reflect the quality status of government affairs data.

[0020] To solve the above problems, an embodiment of the present application provides a data quality assessment method based on the government affairs field. In this method, by obtaining the business scenario portrait and the dataset image, it provides basic information for subsequent business stage division and quality assessment. Based on the business scenario portrait, the dataset image is divided into business stages to determine the business segment data corresponding to the dataset image, realizing the structured classification of the government affairs data processing stage. According to the business segment data, the corresponding evaluation dimension system is obtained, and according to the evaluation dimension system, the quality feature vector corresponding to the business segment data is determined, clarifying the core direction of data quality assessment in each stage and quantifying the assessment results. Based on the quality feature vector corresponding to the business segment data, a comprehensive evaluation result of data quality is generated, forming an overall evaluation conclusion that comprehensively reflects the quality of government affairs data, enhancing the government affairs data processing ability and improving the accuracy of the assessment results.

[0021] The data quality assessment method based on the government affairs field provided by the embodiment of the present application can be applied to an electronic device. At this time, the electronic device is the execution subject of the data quality assessment method based on the government affairs field provided by the embodiment of the present application. The embodiment of the present application does not impose any restrictions on the specific type of the electronic device.

[0022] For example, the electronic device can be an ultra-mobile personal computer (UMPC), a netbook, a desktop computer, a computer, a laptop computer, a communication device, a computing device, a satellite wireless device, etc.

[0023] To better understand the data quality assessment method based on the government affairs field provided by the embodiment of the present application, the following provides an exemplary introduction to the specific implementation process of the data quality assessment method based on the government affairs field provided by the embodiment of the present application.

[0024] Figure 1 Fig. shows a schematic flowchart of the data quality assessment method based on the government affairs field provided by the embodiment of the present application. Figure 2 Fig. shows a schematic operation diagram of the data quality assessment method based on the government affairs field. The data quality assessment method based on the government affairs field includes: S100, obtaining a business scenario portrait and a dataset image; wherein, the business scenario portrait is used to reflect the business type of the current processing of government affairs data, and the dataset image is used to reflect the data presentation form and its association relationship of government affairs data in the government affairs processing process.

[0025] It can be understood that the business scenario portrait refers to an abstract description of the current government data processing business type refined from government business-related data, such as information covering business processes, involved departments, policy bases, etc.; the dataset image is a visual expression of the data form (such as structured, semi-structured) presented by government data during the processing process and its association relationships (such as field dependencies, inter-table associations). In terms of the acquisition method, relevant data can be automatically collected by docking with the government system database, API interfaces, etc. to construct portraits and images; key business information can also be supplemented through manual entry. Use data collection tools (such as Flume) to achieve real-time data collection, and complete data cleaning and transformation through ETL (Extract-Transform-Load) tools (such as Kettle). By obtaining the business scenario portrait and dataset image, it is possible to provide a comprehensive and accurate data basis for subsequent business stage division and quality assessment, help the system understand the business background and data characteristics of government data processing, and improve the pertinence and accuracy of assessment.

[0026] S200, based on the business scenario portrait, divide the dataset image into business stages to determine the business segment data corresponding to the dataset image; the business segment data is used to reflect the stages such as data application of government data.

[0027] It can be understood that the business stage division is to divide the data processing process covered by the dataset image into different business stages, such as the data collection stage, cross-departmental sharing stage, application service stage, etc., according to information such as business processes and policy requirements included in the business scenario portrait. The business type tags (such as "administrative approval", "livelihood subsidy distribution") in the business scenario portrait can be used to match the preset business segment rule library to find the corresponding segmentation strategy; classify the data in the dataset image according to the strategy, and divide the data in the same business processing stage into one business segment data. For example, for the "administrative approval" business, the submitted application data can be divided into "data in the acceptance stage", and the review process data can be divided into "data in the review stage". During the division process, a rule engine (such as Drools) can be used to achieve automatic matching and execution of segmentation rules. Through business stage division, the complex government data processing process can be structured, making the subsequent data quality assessment for different stages more accurate and systematic, and facilitating the discovery of unique data quality problems in each stage.

[0028] In a possible implementation manner, S200, based on the business scenario portrait, divide the dataset image into business stages to determine the business segment data corresponding to the dataset image, including: S210, match the business segment label of the dataset image according to the business scenario portrait.

[0029] It can be understood that the business segmentation label is a keyword or classification identifier used to identify the business stage to which the data in the dataset image belongs, such as "basic business", "cross-domain sharing", "application service", etc. The matching algorithm compares the core business information (such as business process nodes and the scope of applicable policies and regulations) in the business scenario portrait with the pre-set business segmentation label library to calculate the similarity score. For example, if the "data entry system" link is clearly included in the business scenario portrait, the "basic business" label can be matched; if "multi-department data collaborative analysis" is included, the "cross-domain sharing" label is matched. A text similarity calculation algorithm (such as cosine similarity) can be used to quantify the matching degree. When the similarity exceeds a threshold (such as 0.8), it is determined that the matching is successful. By matching the business segmentation label, the business attributes of the data in the dataset image can be quickly located, providing a clear classification basis for subsequent business stage division, reducing manual intervention, and improving the efficiency and accuracy of data processing stage division.

[0030] S220, perform business stage division on the dataset image based on the business segmentation label to determine the business segmentation data corresponding to the dataset image.

[0031] It can be understood that based on the matched business segmentation label, the data with the same label in the dataset image is aggregated together to form different business segmentation data. A data classification algorithm (such as a clustering algorithm) can be used to classify the data records according to the business segmentation label. For example, the data records with the "basic business" label are integrated into "basic business data", including information such as time, collection personnel, and original data format; the data with the "cross-domain sharing" label is divided into "cross-domain sharing data", covering shared interfaces, transmission logs, recipient feedback, etc. During the classification process, a big data processing framework (such as Spark) can be used to achieve fast classification and aggregation of massive data. Through business stage division, the conversion from label matching to actual data segmentation can be completed, enabling the originally mixed government affairs data to be grouped orderly according to business processing logic, laying a foundation for targeted data quality assessment, and facilitating differential management and optimization of data in different stages.

[0032] S300, obtain the corresponding evaluation dimension system according to the business segmentation data, and determine the quality feature vector corresponding to the business segmentation data according to the evaluation dimension system; among them, the evaluation dimension system is used to reflect the core evaluation direction of data quality, and the quality feature vector is used to quantify the actual performance of the data under each evaluation dimension.

[0033] It can be understood that the evaluation dimension system is a set of a series of data quality evaluation directions set for different business segmented data. For example, for the basic business stage, the evaluation dimensions may include integrity and accuracy; for the cross-domain sharing stage, they include consistency, transmission stability, etc. The quality feature vector is a vector representation obtained by quantifying the actual performance of the business segmented data under each evaluation dimension. When obtaining the evaluation dimension system, according to the labels of the business segmented data (such as "basic business data", "cross-domain sharing data", "application service data"), the corresponding dimension set is retrieved from the pre-constructed evaluation dimension rule library. When determining the quality feature vector, corresponding calculation rules can be designed for each evaluation dimension. For example, in the integrity dimension during the data collection stage, the integrity is quantified by calculating the proportion of the number of missing required fields to the total number of fields; in the transmission stability dimension during the data sharing stage, the number of interface call failures per unit time is counted. The specific quantitative calculation is performed using a data calculation framework (such as Pandas), and the quantitative results of each dimension are combined into a quality feature vector, converting the abstract quality evaluation direction into a measurable indicator, realizing the precise characterization of the quality of business segmented data, and facilitating subsequent comparison and analysis of data quality.

[0034] In a possible implementation manner, S300, obtaining the corresponding evaluation dimension system according to the business segmented data, and determining the quality feature vector corresponding to the business segmented data according to the evaluation dimension system, includes: S310, when the business segmented data is basic business data, obtaining the first evaluation dimension system, and determining the first quality feature vector corresponding to the basic business data according to the first evaluation dimension system; wherein, the first evaluation dimension system is used to reflect the missing situation of the basic business data.

[0035] It can be understood that basic business data usually refers to the data in the initial stage of the government affairs data processing process, providing basic support for subsequent services, such as the initially filled personal information, enterprise registration basic materials, etc. The first evaluation dimension system is a set of quality evaluation dimensions specifically set for basic business data, mainly focusing on the missing situation of data, such as the field missing rate, the integrity degree of key information, etc. When obtaining the first evaluation dimension system, according to the classification identifier of the business segmented data ("basic business data"), the corresponding dimension list can be retrieved from the evaluation dimension configuration library, and the first quality feature vector is determined according to the rules of the first evaluation dimension system, timely discovering potential problems in the data entry link, and providing a reliable data basis for subsequent data processing.

[0036] Optionally, in step S310, determining the first quality feature vector corresponding to the basic business data according to the first evaluation dimension system includes: S311. Extract the core field set of the basic business data according to the first evaluation dimension system, calculate the missing value ratio of the core field set within a preset time window, and generate an integrity feature vector.

[0037] It can be understood that the core field set refers to the set of fields in the basic business data that play a key role in subsequent business processing and decision-making. Its selection is based on the evaluation requirements for data integrity in the first evaluation dimension system. For example, in enterprise registration data, fields such as the unified social credit code, enterprise name, and registration address constitute the core field set. When calculating the missing value ratio, a data query tool (such as an SQL statement) can be used to filter the records of the core field set from the basic business data; count the number of missing values of each field within a preset time window (such as the past week), and divide it by the total number of records of the field to obtain the missing ratio of each field; finally, combine these missing ratios into a vector, that is, the integrity feature vector. For example, if the missing rate of the "unified social credit code" in enterprise registration data is 5% and the missing rate of the "enterprise name" is 3%, the integrity feature vector may be expressed as [0.05, 0.03]. By generating the integrity feature vector, the integrity degree of the basic business data at the key information level can be intuitively reflected, helping to quickly locate data missing problems, providing a quantitative basis for data completion and quality improvement, and ensuring the smooth progress of subsequent operations based on complete data.

[0038] S312. Extract the state migration trajectory of the data items in the basic business data according to the first evaluation dimension system, and calculate the average migration duration of each data item from the initial state to the target state, and generate a timeliness feature vector.

[0039] It can be understood that the state migration trajectory of the data item refers to the series of state change processes experienced by each data item in the basic business data from creation (initial state) to meeting specific business requirements (target state) in the business processing flow, such as the state conversion of data from "pending review" to "review passed" or "review not passed". The timeliness evaluation requirements in the first evaluation dimension system determine the state nodes and calculation logics to be concerned about. The state migration trajectory of each data item can be extracted through the log records or status field information of the database, using a data tracking algorithm (such as state sequence analysis based on timestamps); then calculate the time consumed for each data item to convert from the initial state to the target state, accumulate the time consumed by all data items and divide it by the total number of data items to obtain the average migration duration; combine the average migration durations of different types of data items (such as personal information type, business application type) into a timeliness feature vector. For example, the average duration of personal information review is 2 hours, and the average duration of business application review is 4 hours, then the timeliness feature vector is [2, 4]. By generating the timeliness feature vector, the time efficiency of the basic business data in the processing flow can be quantitatively evaluated, bottlenecks in the data processing link can be discovered in a timely manner, and data support can be provided for optimizing the business process and improving data timeliness.

[0040] S313. Feature - fuse the timeliness feature vector and the integrity feature vector to determine the first quality feature vector corresponding to the basic business data.

[0041] It can be understood that feature fusion is to integrate the timeliness feature vector and the integrity feature vector that describe the quality of basic business data from different perspectives, forming a vector representation that more comprehensively reflects data quality. The two vectors can be standardized to scale their values to the same range (such as ), eliminating the influence of dimensional differences on the fusion result. The normalization formula ( ) can be used to achieve this; according to the business characteristics and importance of the basic business data, different weights are assigned to the timeliness feature vector and the integrity feature vector. For example, for emergency data with high real - time requirements, the timeliness feature vector is given a weight of 60% and the integrity feature vector is given a weight of 40%; for basic archive data, the weights of the two are each 50%. The two vectors are fused by weighted summation to obtain the first quality feature vector. For example, if the standardized timeliness feature vector is , and the integrity feature vector is , after fusing with a 50% weight, the first quality feature vector is . Through feature fusion, the performance of basic business data in two key dimensions of time efficiency and information integrity can be comprehensively considered, obtaining a more comprehensive and accurate quality assessment result, providing a more reliable basis for data quality management.

[0042] S320. When the business - segmented data is cross - domain collaborative data, obtain the second evaluation dimension system, and determine the second quality feature vector corresponding to the cross - domain collaborative data according to the second evaluation dimension system; wherein, the second evaluation dimension system is used to reflect the association situation of the cross - domain collaborative data.

[0043] It can be understood that cross - domain collaborative data refers to data involved in the interaction and sharing among multiple departments, institutions, or systems in government affairs operations, such as the enterprise operation data shared between the tax department and the market supervision department. The second evaluation dimension system is a set of evaluation directions set for cross - domain collaborative data, focusing on the association situation of the data, including the consistency of data during cross - domain transmission, the stability of interface docking, etc. When obtaining the second evaluation dimension system, the corresponding dimension list can be retrieved from the evaluation dimension configuration library based on the "cross - domain collaborative data" label of the business - segmented data to determine the second quality feature vector, accurately measuring the quality level of cross - domain collaborative data during the sharing process, timely discovering problems in cross - departmental data interaction, and ensuring the efficient development of government affairs collaborative operations.

[0044] Optionally, in step S320, determining the second quality feature vector corresponding to the cross - domain collaborative data according to the second evaluation dimension system includes: S321. Determine the associated deviation vector corresponding to the cross-domain collaborative data according to the second evaluation dimension system.

[0045] It can be understood that the associated deviation vector is a vector representation used to quantify the degree of deviation of cross-domain collaborative data from the expected rules in the association relationship. The rules regarding data association in the second evaluation dimension system (such as field mapping rules, data consistency verification rules) are the basis for determining the associated deviation vector. The cross-domain association rules in the cross-domain collaborative data can be extracted according to the second evaluation dimension system, and a constraint condition set for data association evaluation can be constructed. For example, it is stipulated that the "taxpayer identification number" in the tax system should correspond one-to-one with the "unified social credit code" in the market supervision system. Then, extract the actual association trajectory of the cross-domain data items, that is, the actual correspondence relationship during the cross-domain transmission and use of the data, and generate an association feature sequence. Compare and analyze the association feature sequence with the constraint condition set. By calculating the deviation degree at each association point between the actual association trajectory and the theoretical association rules (such as the proportion of the number of unmatching fields, the number of association errors, etc.), combine these deviation degree values into an associated deviation vector. For example, if there are two unmatching associated fields and the total number of associated fields is 10, the deviation degree is 20%, and the associated deviation vector may be [0.2] (if there is only one key association dimension) or [0.2, 0, 0] (if there are multiple association dimensions and there is no deviation in other dimensions). By determining the associated deviation vector, the quality problems of cross-domain collaborative data in the association relationship can be intuitively reflected, helping to quickly locate key defects such as data inconsistency and mapping errors, and providing a quantitative reference for the optimization of cross-departmental data collaboration.

[0046] Exemplarily, S321. Determine the associated deviation vector corresponding to the cross-domain collaborative data according to the second evaluation dimension system, including: S3211. Extract the cross-domain association rules in the cross-domain collaborative data according to the second evaluation dimension system, and construct a constraint condition set for data association evaluation.

[0047] It can be understood that cross-domain association rules refer to the standards such as field mapping and business logic correspondence that need to be followed when data is exchanged between different government departments or systems. For example, the "ID number" in the public security system must be exactly the same as the "citizen identification number" in the civil affairs system, and there is a preset association logic between the "enterprise tax number" in the tax system and the "uniform social credit code" in the market supervision system. The second evaluation dimension system includes evaluation criteria for cross-domain data association, such as field matching rules and business process linkage rules. These rules can be extracted from the second evaluation dimension system by parsing business specification documents or invoking predefined rule engines and transformed into a set of computable constraint conditions, such as "the matching rate between field A and field B must be ≥ 95%" and "the status change of data item X must trigger the synchronous update of data item Y". After constructing the set of constraint conditions, it can be formally expressed using structured query language (SQL) or rule description language (such as Drools) to provide a benchmark for subsequent association evaluation. The effect of this step is to clarify the "ideal model" of cross-domain data association, provide a reference standard for quantitatively evaluating the association quality of actual data, and ensure that cross-departmental data collaboration has rules to follow.

[0048] S3212, extract the actual association trajectory of cross-domain data items in cross-domain collaborative data according to the second evaluation dimension system, and generate an association feature sequence.

[0049] It can be understood that the actual association trajectory of cross-domain data items refers to the actual correspondence relationship and status change path of data items between different systems during the cross-departmental transmission and processing of data. For example, when enterprise registration data is synchronized from the market supervision system to the tax system, the change process of the value consistency of fields such as "enterprise name" and "registered address". The second evaluation dimension system defines the association dimensions to be monitored (such as field-level association and business process-level association). The interaction records of cross-domain data items can be captured in real time through data interface log collection, ETL process tracking, etc., and arranged in chronological order to form an association feature sequence. For example, for the "enterprise cancellation application" business, the association feature sequence may include nodes such as "the market supervision system submits a cancellation application → the tax system receives the application and verifies the tax payment status → the market supervision system completes the cancellation according to the tax feedback", and each node records the association status of the data item (such as "field matching successful", "field difference requires manual verification"). After generating the association feature sequence, sequence analysis tools (such as the pandas library in Python) can be used for structured storage and preprocessing to provide a data basis for comparative analysis. The effect of this step is to restore the "real scenario" of cross-domain data association, and provide clues for discovering abnormal points in data association (such as field mapping errors and process interruptions) by recording actual interaction details.

[0050] S3213. Compare and analyze the associated feature sequence with the set of constraint conditions, calculate the deviation degree between the actual associated trajectory and the theoretical association rule, and generate an association deviation vector.

[0051] It can be understood that the deviation degree is a quantitative indicator for measuring the degree of inconsistency between the actual associated trajectory and the theoretical association rule. By comparing and analyzing the rules in the set of constraint conditions with the actual data in the associated feature sequence, the deviation values of each dimension are calculated. The dynamic time warping (DTW) algorithm or the edit distance algorithm can be used to match the associated feature sequence with the standard association process in the set of constraint conditions, identify the nodes that do not conform to the rules (such as field matching failures and missing process steps), and count indicators such as the number of deviations and the duration of deviations. For example, if the constraint condition requires that "the cross-department data field matching rate should be ≥ 95%", and the number of field matching failures in the actual associated feature sequence accounts for 8% of the total interaction times, then the deviation degree is 8%. Combine the deviation degrees of each dimension (such as the field matching deviation degree, the process time sequence deviation degree, and the business logic deviation degree) into an association deviation vector, such as [0.08, 0.03, 0] (assuming that the process time sequence deviation degree is 3% and there is no deviation in business logic). Generating an association deviation vector is to transform the abstract association quality problem into a measurable value. Through the association deviation vector, the weak links in cross-domain data association can be intuitively reflected, such as the business nodes with frequent field mapping errors or the department interfaces with poor process connection, providing a quantitative basis for accurately positioning data collaboration problems and helping to optimize the cross-domain data governance process.

[0052] S322. Obtain the response time distribution of cross-domain collaborative data during cross-domain transmission, and generate a transmission timeliness vector.

[0053] It can be understood that when cross-domain collaborative data is transmitted across departments and systems, the time interval from when the sender initiates a request to when the receiver completes data reception and can process it normally is the transmission response time. The transmission timeliness vector is a vector formed by quantifying the response time characteristics of different data transmission tasks. When obtaining the response time distribution, a monitoring program can be deployed at the data transmission interface (such as using a network packet capture tool or API call logging) to collect the start time and end time of each data transmission in real time, and calculate the response duration; then, statistical analysis is performed on all transmission response durations within a certain period (such as one day or one week) to obtain the distribution of response times, such as indicators like the average response time, the maximum response time, and the standard deviation of the response time; these indicators are combined into a transmission timeliness vector in a preset order, for example, [average response time, maximum response time, standard deviation]. For example, if the average response time is 500ms, the maximum response time is 2000ms, and the standard deviation is 100ms, then the transmission timeliness vector is [500, 2000, 100]. By generating the transmission timeliness vector, the time efficiency and stability of cross-domain collaborative data during transmission can be clearly presented, facilitating the discovery of problems such as excessively high transmission delays and large fluctuations, and providing data support for optimizing the data transmission link and improving cross-domain collaborative efficiency.

[0054] S323, perform feature fusion on the association deviation vector and the transmission timeliness vector to determine the second quality feature vector corresponding to the cross-domain collaborative data.

[0055] It can be understood that feature fusion is to integrate the association deviation vector reflecting the association accuracy of cross-domain collaborative data and the transmission timeliness vector reflecting data transmission timeliness to form a vector for comprehensively evaluating data quality. The two vectors can be normalized to unify their values to the same measurement scale (such as the [0, 1] interval) to eliminate the dimension difference, which can be achieved using the min-max normalization method; according to the actual requirements and data characteristics of cross-domain collaborative services, weights are assigned to the association deviation vector and the transmission timeliness vector. For example, in an emergency data collaboration scenario with extremely high real-time requirements, a weight of 70% is assigned to the transmission timeliness vector and 30% to the association deviation vector; for regular business data sharing, the weights of the two can each account for 50%; finally, the two vectors are fused by weighted summation to obtain the second quality feature vector. For example, the normalized association deviation vector is [0.3], the transmission timeliness vector is [0.8], and after fusion with a 50% weight, the second quality feature vector is [(0.3×0.5 + 0.8×0.5)] = [0.55]. Through feature fusion, the quality performance of cross-domain collaborative data in two key dimensions of association relationship and transmission efficiency can be comprehensively considered, and a more representative quality evaluation result can be obtained, providing a scientific basis for cross-domain data management and optimization.

[0056] In S330, when the business segmented data is application service data, obtain the third evaluation dimension system, and determine the third quality feature vector corresponding to the application service data according to the third evaluation dimension system; wherein, the third evaluation dimension system is used to reflect the compliance status of the application service data.

[0057] It can be understood that application service data refers to data used to support government service applications (such as online services, data analysis and decision-making), such as data of materials submitted by the public for handling affairs and statistical analysis data for government decision-making reference. The third evaluation dimension system is a set of evaluation directions specifically set for application service data, mainly focusing on the compliance status of the data, including the compliance of data usage permissions, the compliance of sensitive information processing, etc. When obtaining the third evaluation dimension system, the corresponding dimension list can be retrieved from the evaluation dimension configuration library according to the "application service data" label of the business segmented data. When determining the third quality feature vector, the access permission change logs of sensitive data items in the application service data can be extracted according to the third evaluation dimension system, and the change frequency of the permission change feature sequence can be calculated; then, anomaly detection can be performed on the access track of the data; finally, the relevant quantitative results can be fused. Through the access log analysis tool and anomaly detection algorithm, the quantitative evaluation of the compliance of application service data can be realized, effectively monitoring the compliance of application service data during use, preventing data security risks, and ensuring the legal and safe operation of government service applications.

[0058] Optionally, in step S330, determining the third quality feature vector corresponding to the application service data according to the third evaluation dimension system includes: S331, extract the access permission change logs of sensitive data items in the application service data according to the third evaluation dimension system, generate a permission change feature sequence, and calculate the change frequency of the permission change feature sequence within a preset time window to generate a change frequency vector.

[0059] It can be understood that sensitive data items refer to the data in application service data that involve sensitive information such as personal privacy, trade secrets, or national security, such as ID numbers, bank account information, etc. The access permission change log records the changes in the access permissions of these sensitive data items at different time points, including operations such as permission granting, modification, and revocation. The permission change feature sequence is a sequence formed by arranging these change operations in chronological order to describe the process of permission changes. The compliance requirements for permission management in the third evaluation dimension system are the basis for extracting and analyzing the log. The access permission change log of sensitive data items in application service data can be extracted from the permission management database using a data query tool (such as SQL); sorted into a permission change feature sequence according to the time stamp order; calculate the total number of permission changes in the preset time window (such as one month) for the permission change feature sequence, and divide it by the length of the time window (such as the number of days) to obtain the average daily change frequency. Combine the change frequencies of different types of sensitive data items (such as personal identity information, financial data) into a change frequency vector. For example, if the monthly change frequency of the sensitive data item of personal identity information is 10 times and the monthly change frequency of the sensitive data item of financial data is 5 times, then the change frequency vector is [10 / 30, 5 / 30]. By generating the change frequency vector, it is possible to quantitatively evaluate the changes in the access permissions of sensitive data, timely detect abnormal and frequent permission change behaviors, prevent the risk of unauthorized data access, and ensure the compliance of the use of application service data.

[0060] S332, perform anomaly detection on the access track of application service data to generate an anomaly feature vector.

[0061] It can be understood that the access track of application service data refers to the operation records during the data access process, including information such as access time, access user, access operation type (such as query, modification, deletion), etc. Anomaly detection is to identify operation behaviors that do not conform to the normal access pattern by analyzing the access track. The requirements for data access security in the third evaluation dimension system provide a standard for anomaly detection. The access track data of application service data can be collected using a log analysis tool. Algorithms such as Isolation Forest, One-Class SVM, or rule-based methods (such as setting a high-frequency access threshold for the same user in a short period of time) can be used to analyze the access track. For each detection dimension (such as abnormal access time, abnormal access frequency, abnormal operation type), calculate the anomaly degree score, and combine these scores into an anomaly feature vector. For example, if the anomaly score for detected abnormal access time is 0.8, the anomaly score for abnormal access frequency is 0.3, and the anomaly score for abnormal operation type is 0.1, then the anomaly feature vector is [0.8, 0.3, 0.1]. By generating the anomaly feature vector, it is possible to quickly locate suspicious behaviors during the access process of application service data, timely warn of the risk of data leakage or malicious operations, and ensure data security.

[0062] S333: Perform feature fusion on the change frequency vector and the abnormal feature vector to determine a third quality feature vector corresponding to the application service data.

[0063] It can be understood that feature fusion combines the change frequency vector, which reflects changes in sensitive data access rights, and the anomaly feature vector, which reflects the security status of data access, to form a vector that comprehensively assesses the compliance quality of application service data. During implementation, both vectors are first normalized to the range [0, 1] to eliminate dimensionality effects. Z-score normalization can be used. Weights can be assigned based on the security importance of the application service scenario. For services involving highly sensitive information (such as medical data query services), the anomaly feature vector is weighted 70% and the change frequency vector 30%. For general public service data, each is weighted 50%. Finally, the two vectors are fused through a weighted sum to obtain a third quality feature vector. For example, if the normalized change frequency vector is [0.4] and the anomaly feature vector is [0.6], after fusion with a 50% weighting, the third quality feature vector is [(0.4 × 0.5 + 0.6 × 0.5)] = [0.5]. Through feature fusion, we can comprehensively consider the quality performance of application service data in terms of permission management and access security, provide an accurate basis for data compliance assessment, and help government services operate safely and stably.

[0064] S400: Generate a comprehensive data quality evaluation result based on the quality feature vector corresponding to the business segment data.

[0065] It can be understood that the quality characteristic vector is a quantitative representation of the quality performance of each business segment data under a specific evaluation dimension. Generating a comprehensive data quality evaluation result is to integrate and analyze the quality characteristics of these segments to obtain an assessment of the overall data quality. The first quality characteristic vector, second quality characteristic vector, and third quality characteristic vector corresponding to basic business data, cross-domain collaborative data, and application service data can be summarized; weights are assigned to each quality characteristic vector based on the importance of different business segments in the overall government data processing and the data application scenarios; through weighted calculations or grading judgment rules, the information of each vector is integrated and mapped to a preset quality grade interval (such as excellent, good, qualified, and unqualified) to obtain a comprehensive data quality evaluation result, thereby achieving a comprehensive and systematic assessment of the overall quality of government data, providing an intuitive reference basis for data governance and optimized decision-making, and helping managers quickly grasp the overall picture of data quality and make targeted improvements.

[0066] In one possible implementation, S400 generates a comprehensive data quality evaluation result based on the quality feature vector corresponding to the business segment data, including: S410 : Generate a comprehensive data quality evaluation result based on the first quality feature vector, the second quality feature vector, and the third quality feature vector.

[0067] It can be understood that the first quality feature vector, the second quality feature vector, and the third quality feature vector respectively quantify and reflect the quality status of data in different dimensions from three key business segments: basic business, cross-domain collaboration, and application services. The process of generating the comprehensive evaluation result is a process of systematically integrating these segmented quality features. The three quality feature vectors can be normalized to be on the same measurement scale; according to the logical relationship of government affairs business processes and data application scenarios, determine the weight allocation strategy for each vector. For example, in a business scenario with data sharing as the core, assign a higher weight to the second quality feature vector; then fuse and calculate the vector information through weighted summation or a complex comprehensive evaluation model (such as a neural network, decision tree ensemble model); divide the data quality level according to the calculation result to generate the comprehensive evaluation result, which can break the boundaries of business segments and achieve a global and comprehensive evaluation of government affairs data quality, providing a more scientific and comprehensive basis for data quality optimization and business decision-making.

[0068] Optionally, in S410, based on the first quality feature vector, the second quality feature vector, and the third quality feature vector, generate a comprehensive data quality evaluation result, including: S411, determine the first quality feature value based on the first quality feature vector, determine the second quality feature value based on the second quality feature vector, and determine the third quality feature value based on the third quality feature vector.

[0069] It can be understood that the quality feature vector is a vector set composed of feature values in multiple dimensions. Determining the quality feature value means extracting key quantitative indicators from the vector for subsequent comprehensive evaluation. For the first quality feature vector, according to the core dimensions of basic business data quality assessment (such as integrity, timeliness), select the corresponding feature values. For example, extract the integrity feature value and the timeliness feature value and perform weighted calculation to obtain the first quality feature value; for the second quality feature vector, select key feature values from dimensions such as the correlation deviation and transmission timeliness of cross-domain collaboration data, and calculate and integrate to obtain the second quality feature value; for the third quality feature vector, extract feature values from dimensions such as the permission change frequency and access anomaly degree of application service data and calculate to obtain the third quality feature value. The feature values of each dimension in the vector can be processed through a preset calculation formula or algorithm model, such as weighted average, extracting key values after principal component analysis dimensionality reduction, etc. By determining each quality feature value, the complex vector information is simplified into a representative single value, which is convenient for subsequent comparison, ranking, and comprehensive evaluation of data quality, improving the evaluation efficiency and accuracy.

[0070] S412, perform basic quality filtering on the first quality feature value, the second quality feature value, and the third quality feature value respectively, and generate a comprehensive data quality evaluation result according to the basic quality filtering result.

[0071] It can be understood that basic quality filtering is to initially screen the quality characteristic values of each business segment by setting key quality thresholds, and quickly identify problems that seriously affect data quality. For the first quality characteristic value, key indicator thresholds such as the integrity and timeliness of basic business data can be set. For example, the integrity threshold is 90%. If the integrity reflected in the first quality characteristic value is lower than this threshold, it is determined that the filtering fails. Similarly, for the second quality characteristic value, the correlation accuracy rate and transmission timeliness thresholds of cross-domain collaborative data are set, and for the third quality characteristic value, the compliance and pass rate threshold of application service data is set. If any quality characteristic value fails the corresponding threshold test, a lower-level comprehensive evaluation result of data quality is directly generated. If all pass, it enters the next stage of comprehensive calculation. The "one-vote veto" mechanism is adopted to preferentially exclude data with serious quality defects, quickly locate the problem business segments, reduce unnecessary complex calculations at the same time, improve the efficiency of data quality assessment, and ensure that the comprehensive evaluation result can preferentially reflect the key quality problems of the data.

[0072] Exemplarily, S412, performing basic quality filtering on the first quality characteristic value, the second quality characteristic value, and the third quality characteristic value respectively, and generating a comprehensive evaluation result of data quality according to the basic quality filtering result, includes: S4121, when the first quality characteristic value, the second quality characteristic value, and the third quality characteristic value do not all pass the basic quality filtering, generating a first comprehensive evaluation result of data quality and determining the first comprehensive evaluation result of data quality as the comprehensive evaluation result of data quality.

[0073] It can be understood that if any one or more of the first, second, and third quality characteristic values do not reach the preset basic quality thresholds (such as insufficient integrity of basic business data, too high error rate of cross-domain collaborative data association), it indicates that there are serious quality problems in the government affairs data in the corresponding business segment. A lower-level first comprehensive evaluation result of data quality (such as "unqualified" or "poor") can be directly generated and determined as the final comprehensive evaluation result of data quality. The three quality characteristic values can be compared one by one through conditional judgment statements (such as if-else statements). Once it is found that a certain characteristic value is lower than the corresponding threshold, the generation logic of the low-level evaluation is immediately triggered, skipping the subsequent complex weighted calculation steps, which can quickly identify and mark the data with serious defects in the initial stage of data quality assessment, provide a clear rectification direction for data governance personnel, avoid affecting the normal development of the overall business due to local data quality problems, and improve the execution efficiency of the evaluation process at the same time.

[0074] S4122. When the first quality characteristic value, the second quality characteristic value, and the third quality characteristic value all pass the basic quality filter, generate a second comprehensive data quality evaluation result and determine the second comprehensive data quality evaluation result as the comprehensive data quality evaluation result.

[0075] It can be understood that when the quality characteristic values corresponding to the basic business data, cross-domain collaborative data, and application service data all meet the preset basic quality threshold requirements, it indicates that there are no serious quality problems in each key business segment of the data. It can enter the refined evaluation stage, generate a second comprehensive data quality evaluation result through more complex calculations, and use it as the final assessment. When executing, first determine the weight distribution of each quality characteristic value according to the business scenario (such as emergency command, statistical analysis, daily government affairs service). For example, in the emergency command scenario, a higher weight is given to the third quality characteristic value of the application service data; then use the weighted summation formula (Comprehensive evaluation result = First quality characteristic value × Weight 1 + Second quality characteristic value × Weight 2 + Third quality characteristic value × Weight 3) to calculate the comprehensive score; finally, map this score to the preset quality grade interval (such as 85 - 100 points is "excellent", 70 - 84 points is "good") to generate a second comprehensive data quality evaluation result. On the premise of ensuring the qualified basic quality of the data, further comprehensively consider the contribution of the data quality of each business segment to the whole, obtain a more scientific and accurate comprehensive evaluation of government affair data quality, and provide a more valuable reference basis for data optimization and business decision-making.

[0076] Exemplarily, S4122. When the first quality characteristic value, the second quality characteristic value, and the third quality characteristic value all pass the basic quality filter, generate a second comprehensive data quality evaluation result and determine the second comprehensive data quality evaluation result as the comprehensive data quality evaluation result, including: S41221. When the first quality characteristic value, the second quality characteristic value, and the third quality characteristic value all pass the basic quality filter, determine the quality assessment weight according to the business segment data.

[0077] It can be understood that the quality assessment weights are a set of parameters used to measure the dynamic contribution degree of the quality characteristic values of basic business data, cross-domain collaborative data, and application service data to the comprehensive evaluation of the overall data quality. The attributes of the business segmented data (such as business type, data volume, real-time requirement, etc.) determine the importance differences of the quality characteristics of each segment. The business scenario to which the current government affairs data belongs (such as emergency command, statistical analysis, people's livelihood service, etc.) can be identified, and the initial weight ratio can be retrieved from the preset weight template library (for example, the weight ratio of application service data in the emergency command scenario accounts for 50%); and by counting the scale of the business segmented data (such as the number of records, data volume), the actual data volume ratios of the basic business data, cross-domain collaborative data, and application service data are calculated through the normalization algorithm to adjust the initial weights (for example, when the data volume ratio of a certain segment exceeds 60%, the weight of its quality characteristic value is increased by 10%-20%) to form the final quality assessment weights. The effect of this step is to make the weight allocation dynamically adapt to the business requirements and data characteristics, highlight the quality impact of key business segments, ensure that the comprehensive evaluation results are more in line with the actual requirements of government affairs data processing, and provide a guidance for precise data governance.

[0078] S41222, weight the first quality characteristic value, the second quality characteristic value, and the third quality characteristic value according to the quality assessment weights, generate the second comprehensive data quality evaluation result, and determine the second comprehensive data quality evaluation result as the comprehensive data quality evaluation result.

[0079] It can be understood that weighted calculation is to quantitatively integrate the quality characteristic values of each business segment according to their corresponding weights to obtain a comprehensive evaluation result that comprehensively reflects the quality of government affairs data. The weighted summation formula can be used: Comprehensive evaluation score = First quality characteristic value × First weight + Second quality characteristic value × Second weight + Third quality characteristic value × Third weight to calculate the three quality characteristic values. For example, if the first quality characteristic value is 80, the corresponding weight is 30%; the second quality characteristic value is 75, the corresponding weight is 20%; the third quality characteristic value is 90, the corresponding weight is 50%, then the comprehensive evaluation score is 80×0.3 + 75×0.2 + 90×0.5 = 84. The comprehensive evaluation score can be mapped to the preset quality grade interval (such as 0-60 points is "unqualified", 61-79 points is "qualified", 80-89 points is "good", 90-100 points is "excellent") to determine the corresponding quality grade, generate the second comprehensive data quality evaluation result, and use it as the final evaluation of the quality of government affairs data. Through weighted calculation, the contribution differences of the data quality of different business segments to the whole are fully considered, making the evaluation result more scientific and reasonable, being able to accurately reflect the actual quality level of government affairs data, and providing a reliable basis for data quality improvement and business decision-making.

[0080] Corresponding to the data quality assessment method in the above-mentioned embodiment in the government affairs field, an embodiment of the present application further provides a data quality assessment system in the government affairs field. Each unit of this system can implement each step of the data quality assessment method in the government affairs field. Figure 3 The structural block diagram of the data quality assessment system in the government affairs field provided by the embodiment of the present application is shown. For the convenience of description, only the parts related to the embodiment of the present application are shown.

[0081] Referring to Figure 3 , the data quality assessment system in the government affairs field includes: An acquisition unit, configured to acquire a business scenario portrait and a dataset image; wherein, the business scenario portrait is used to reflect the business type currently processed by government affairs data, and the dataset image is used to reflect the data presentation form and its association relationship of government affairs data in the government affairs processing process; A segmentation unit, configured to perform business stage division on the dataset image based on the business scenario portrait, and determine the business segmented data corresponding to the dataset image; the business segmented data is used to reflect stages such as data application where government affairs data is located; A quality unit, configured to obtain a corresponding evaluation dimension system according to the business segmented data, and determine a quality feature vector corresponding to the business segmented data according to the evaluation dimension system; wherein, the evaluation dimension system is used to reflect the core evaluation direction of data quality, and the quality feature vector is used to quantify the actual performance of data under each evaluation dimension; A result unit, configured to generate a comprehensive data quality evaluation result based on the quality feature vector corresponding to the business segmented data.

[0082] It should be noted that, regarding the information interaction, execution process, etc. between the above-mentioned systems / units, since they are based on the same concept as the method embodiment of the present application, for their specific functions and the technical effects brought, reference can be specifically made to the method embodiment part, and details are not described here again.

[0083] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules as needed, that is, the internal structure of the system is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit module exists physically alone, or two or more unit modules are integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiment and will not be elaborated herein.

[0084] The embodiment of this application also provides an electronic device, Figure 4 which is a schematic structural diagram of the electronic device provided in an embodiment of this application. As Figure 4 shown, the electronic device 6 in this embodiment includes: at least one processor 60 ( Figure 4 only one is shown herein), at least one memory 61 ( Figure 4 only one is shown herein), and a computer program 62 stored in the at least one memory 61 and executable on the at least one processor 60. When the processor 60 executes the computer program 62, the electronic device 6 implements the steps in any of the above-mentioned data quality assessment method embodiments based on the government affairs field, or the functions of each unit in the above system embodiments are implemented on the electronic device 6.

[0085] Exemplarily, the computer program 62 can be divided into one or more units. The one or more units are stored in the memory 61 and executed by the processor 60 to complete this application. The one or more units can be a series of computer program instruction segments capable of completing specific functions, and these instruction segments are used to describe the execution process of the computer program 62 in the data quality assessment 6 based on the government affairs field.

[0086] The electronic device 6 can be a computing device or a terminal device such as a desktop computer, a notebook, a palm computer, and a cloud server. The electronic device may include, but is not limited to, a processor 60 and a memory 61. Those skilled in the art can understand that Figure 4 merely an example of the electronic device 6, which does not constitute a limitation on the electronic device 6. It may include more or fewer components than shown in the figure, or combine some components, or different components. For example, it may also include input / output devices, network access devices, buses, etc.

[0087] The processor 60 may be a Central Processing Unit (CPU), and the processor 60 may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0088] In some embodiments, the memory 61 may be an internal storage unit of the electronic device 6, such as the hard disk or memory of the electronic device 6. In some other embodiments, the memory 61 may also be an external storage device of the electronic device 6, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the electronic device 6. Further, the memory 61 may also include both the internal storage unit and the external storage device of the electronic device 6. The memory 61 is used to store an operating system, application programs, a BootLoader, data, and other programs, such as the program code of the computer program. The memory 61 may also be used to temporarily store data that has been output or is to be output.

[0089] An embodiment of the present application also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in any of the above method embodiments are implemented.

[0090] An embodiment of the present application provides a computer program product, and when the computer program product runs on an electronic device, the electronic device implements the steps in any of the above method embodiments.

[0091] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of this application, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to an electronic device, a recording medium, a computer memory, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), an electrical carrier signal, a telecommunications signal, and a software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunications signal.

[0092] In the above embodiments, the descriptions of the various embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0093] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0094] In the embodiments provided in this application, it should be understood that the disclosed data quality assessment system / electronic device and method based on the government affairs field can be implemented in other ways. For example, the data quality assessment system / electronic device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.

[0095] The unit described as the separation component may or may not be physically separated. The component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0096] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A data quality assessment method based on the government affairs field, characterized in that Including: Obtain a business scenario portrait and a dataset image; wherein, the business scenario portrait is used to reflect the business types currently processed by government affairs data, and the dataset image is used to reflect the data presentation form and its association relationship of government affairs data during the government affairs processing process; Based on the business scenario portrait, perform business stage division on the dataset image to determine the business segmented data corresponding to the dataset image; the business segmented data is used to reflect the data application stage where the government affairs data is located; Obtain a corresponding evaluation dimension system according to the business segmented data, and determine a quality feature vector corresponding to the business segmented data according to the evaluation dimension system; wherein, the evaluation dimension system is used to reflect the core evaluation direction of data quality, and the quality feature vector is used to quantify the actual performance of the data under each evaluation dimension; Generate a comprehensive data quality evaluation result based on the quality feature vector corresponding to the business segmented data.

2. The data quality assessment method based on the government affairs field according to claim 1, wherein The performing business stage division on the dataset image based on the business scenario portrait to determine the business segmented data corresponding to the dataset image includes: Match the business segmented label of the dataset image according to the business scenario portrait; Based on the business segmented label, perform business stage division on the dataset image to determine the business segmented data corresponding to the dataset image.

3. The data quality assessment method based on the government affairs field according to claim 1, wherein The obtaining a corresponding evaluation dimension system according to the business segmented data and determining a quality feature vector corresponding to the business segmented data according to the evaluation dimension system includes: When the business segmented data is basic business data, obtain a first evaluation dimension system, and determine a first quality feature vector corresponding to the basic business data according to the first evaluation dimension system; wherein, the first evaluation dimension system is used to reflect the missing situation of the basic business data; When the business segmented data is cross-domain collaborative data, obtain a second evaluation dimension system, and determine a second quality feature vector corresponding to the cross-domain collaborative data according to the second evaluation dimension system; wherein, the second evaluation dimension system is used to reflect the association situation of the cross-domain collaborative data; When the business segmented data is application service data, obtain a third evaluation dimension system, and determine a third quality feature vector corresponding to the application service data according to the third evaluation dimension system; wherein, the third evaluation dimension system is used to reflect the compliance status of the application service data; The generating a comprehensive data quality evaluation result based on the quality feature vector corresponding to the business segmented data includes: Generate a comprehensive data quality evaluation result based on the first quality feature vector, the second quality feature vector, and the third quality feature vector.

4. The data quality assessment method based on the government affairs field according to claim 3, characterized in that, The determining a first quality feature vector corresponding to the basic business data according to the first evaluation dimension system includes: Extract the core field set of the basic business data according to the first evaluation dimension system, and calculate the missing value ratio of the core field set within a preset time window to generate an integrity feature vector; Extract the state transition trajectories of data items in the basic service data according to the first evaluation dimension system, calculate the average migration duration of each data item from the initial state to the target state, and generate a timeliness feature vector. Perform feature fusion on the timeliness feature vector and the integrity feature vector to determine the first quality feature vector corresponding to the basic service data.

5. The data quality assessment method based on the government affairs field according to claim 3, wherein Determine the second quality feature vector corresponding to the cross-domain collaborative data according to the second evaluation dimension system, including: Determine the association deviation vector corresponding to the cross-domain collaborative data according to the second evaluation dimension system. Obtain the response time distribution of the cross-domain collaborative data during cross-domain transmission, and generate a transmission timeliness vector. Perform feature fusion on the association deviation vector and the transmission timeliness vector to determine the second quality feature vector corresponding to the cross-domain collaborative data.

6. The data quality assessment method based on the government affairs field according to claim 5, wherein The determining the association deviation vector corresponding to the cross-domain collaborative data according to the second evaluation dimension system includes: Extract the cross-domain association rules in the cross-domain collaborative data according to the second evaluation dimension system, and construct a constraint condition set for data association evaluation. Extract the actual association trajectories of cross-domain data items in the cross-domain collaborative data according to the second evaluation dimension system, and generate an association feature sequence. Perform a comparative analysis on the association feature sequence and the constraint condition set, calculate the deviation degree of the actual association trajectory from the theoretical association rule, and generate an association deviation vector.

7. The data quality assessment method based on the government affairs field according to claim 3, wherein, Determine the third quality feature vector corresponding to the application service data according to the third evaluation dimension system, including: Extract the access permission change logs of sensitive data items in the application service data according to the third evaluation dimension system, generate a permission change feature sequence, and calculate the change frequency of the permission change feature sequence within a preset time window to generate a change frequency vector. Perform anomaly detection on the access trajectory of the application service data to generate an anomaly feature vector. Perform feature fusion on the change frequency vector and the anomaly feature vector to determine the third quality feature vector corresponding to the application service data.

8. The data quality assessment method based on the government affairs field according to claim 3, wherein Generate a comprehensive data quality evaluation result based on the first quality feature vector, the second quality feature vector, and the third quality feature vector, including: Determine the first quality feature value based on the first quality feature vector, the second quality feature value based on the second quality feature vector, and the third quality feature value based on the third quality feature vector. Perform basic quality filtering on the first quality feature value, the second quality feature value, and the third quality feature value respectively, and generate a comprehensive data quality evaluation result according to the basic quality filtering result.

9. The data quality assessment method based on the government affairs field according to claim 8, wherein The performing basic quality filtering on the first quality feature value, the second quality feature value, and the third quality feature value respectively, and generating a comprehensive data quality evaluation result according to the basic quality filtering result includes: When the first quality feature value, the second quality feature value, and the third quality feature value do not all pass the basic quality filtering, generate a first comprehensive data quality evaluation result and determine the first comprehensive data quality evaluation result as the comprehensive data quality evaluation result. When the first quality characteristic value, the second quality characteristic value, and the third quality characteristic value all pass through the basic quality filter, a second comprehensive evaluation result of data quality is generated and the second comprehensive evaluation result of data quality is determined as the comprehensive evaluation result of data quality.

10. The data quality assessment method based on the government affairs field according to claim 9, wherein, The step of, when the first quality characteristic value, the second quality characteristic value, and the third quality characteristic value all pass through the basic quality filter, generating a second comprehensive evaluation result of data quality and determining the second comprehensive evaluation result of data quality as the comprehensive evaluation result of data quality, includes: When the first quality characteristic value, the second quality characteristic value, and the third quality characteristic value all pass through the basic quality filter, determining a quality evaluation weight according to the business segmented data; Weighting the first quality characteristic value, the second quality characteristic value, and the third quality characteristic value according to the quality evaluation weight, generating a second comprehensive evaluation result of data quality and determining the second comprehensive evaluation result of data quality as the comprehensive evaluation result of data quality.

Citation Information

Patent Citations

  • Application data evaluation method and device based on intelligent decision, computer equipment and storage medium

    CN110310079A

  • Government affair big data application maturity evaluation method and system

    CN111832945A

  • Analytic hierarchy process-based power data quality evaluation model construction method

    CN112348695A

  • Comprehensive evaluation system for e-government affairs

    CN112926898A

  • Performance assessment supervision method and system based on working state of telephone operator

    CN116777252A

Cited By

  • Self-adaptive multi-dimensional evaluation high-quality scientific and technological information screening method and system

    CN121009233A