Data quality assessment method based on government affairs

By obtaining business scenario portraits and data set images, dividing the government data processing stages, obtaining the evaluation dimension system and quality feature vectors, and generating comprehensive data quality evaluation results, the problems of low accuracy and insufficient generalization capabilities in government data evaluation are solved, and more accurate and comprehensive evaluation is achieved.

CN120387105BActive Publication Date: 2025-08-29JIANGXI YUNQING TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510877734.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-08-29
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

In the prior art, the quality evaluation of government data relies on preset rules checks and single-dimensional evaluation, resulting in low accuracy of evaluation results and insufficient generalization capabilities, and inability to adapt to the blind spots and one-sidedness of evaluation in complex business logic and cross-departmental data flow.

Method used

By acquiring business scene portraits and data set images, dividing the data set images based on business scene portraits, determining the business segment data corresponding to the data set images, obtaining the evaluation dimension system and determining the quality feature vector, and generating comprehensive data quality evaluation results.

Benefits of technology

The structured classification of the government data processing stage has been realized, the accuracy and comprehensiveness of the evaluation results have been enhanced, the government data processing capabilities have been improved, and multi-dimensional dynamic evaluation tools have been provided, which has enhanced the accuracy and adaptability of the evaluation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387105B_ABST
    Figure CN120387105B_ABST
Patent Text Reader

Abstract

The present application is applicable to the field of government data evaluation technology, and in particular relates to a data quality evaluation method based on the government field, the method comprising: obtaining business scenario portraits and data set images to provide basic information for subsequent business stage division and quality evaluation; dividing the data set images into business stages based on the business scenario portraits, determining the business segmentation data corresponding to the data set images, and realizing structured classification of the government data processing stages; obtaining the corresponding evaluation dimension system according to the business segmentation data, and determining the quality feature vectors corresponding to the business segmentation data according to the evaluation dimension system, clarifying the core direction of data quality evaluation at each stage and quantifying the evaluation results; generating a comprehensive data quality evaluation result based on the quality feature vectors corresponding to the business segmentation data, forming a holistic evaluation conclusion that comprehensively reflects the quality of government data, enhancing the government data processing capabilities, and improving the accuracy of the evaluation results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of government data evaluation technology, and in particular relates to a data quality evaluation method based on the government field. Background Art

[0002] Government data processing is the full-process management of government data from collection, cleaning, storage to analysis, sharing, and security protection. It achieves data aggregation through multi-source data collection, improves data quality through cleaning and standardization, uses databases and data warehouses for secure storage, and uses statistical analysis and machine learning to mine data value. It breaks down departmental barriers to achieve data sharing and promote government collaboration. At the same time, encryption, access control and other technologies are used to ensure data security, ultimately serving government decision-making, policy formulation and public service optimization, and promoting the construction of digital government and the modernization of governance capabilities.

[0003] In existing technologies, the quality assessment of government data relies on preset rule verification and single-dimensional evaluation, which requires manual setting of format rules and evaluation dimensions. Parameters need to be adjusted to adapt to new business scenarios. In complex business logic or cross-departmental data flow, evaluation blind spots and one-sidedness are prone to occur. The lack of multi-dimensional dynamic evaluation tools leads to low accuracy of evaluation results and insufficient generalization capabilities. Summary of the Invention

[0004] The embodiment of the present application provides a data quality assessment method based on the government affairs field, which can solve the problem of lack of multi-dimensional dynamic assessment tools in the government affairs data assessment process, resulting in low accuracy of assessment results and insufficient generalization ability.

[0005] In a first aspect, an embodiment of the present application provides a data quality assessment method based on the government affairs field, including:

[0006] Obtaining a business scenario portrait and a data set image; wherein the business scenario portrait is used to reflect the business type currently processed by the government data, and the data set image is used to reflect the data presentation form and its correlation relationship during the government data processing process;

[0007] Divide the data set image into business stages based on the business scenario portrait, and determine the business segmentation data corresponding to the data set image; the business segmentation data is used to reflect the data application stage of the government data;

[0008] Obtaining a corresponding evaluation dimension system based on the business segment data, and determining a quality feature vector corresponding to the business segment data based on the evaluation dimension system; wherein the evaluation dimension system is used to reflect the core evaluation direction of data quality, and the quality feature vector is used to quantify the actual performance of the data under each evaluation dimension;

[0009] Based on the quality feature vector corresponding to the business segment data, a comprehensive data quality evaluation result is generated.

[0010] The above technical solutions in the embodiments of the present application have at least the following technical effects:

[0011] The data quality assessment method based on the government affairs field provided in the embodiment of the present application provides basic information for subsequent business stage division and quality assessment by obtaining business scenario portraits and data set images. The data set images are divided into business stages based on the business scenario portraits, and the business segmentation data corresponding to the data set images are determined to achieve structured classification of the government affairs data processing stages. According to the business segmentation data, the corresponding evaluation dimension system is obtained, and the quality feature vectors corresponding to the business segmentation data are determined according to the evaluation dimension system, the core direction of the data quality assessment at each stage is clarified and the assessment results are quantified. Based on the quality feature vectors corresponding to the business segmentation data, a comprehensive data quality evaluation result is generated to form a holistic assessment conclusion that comprehensively reflects the quality of government affairs data, enhance the government affairs data processing capabilities, and improve the accuracy of the assessment results.

[0012] In a second aspect, an embodiment of the present application provides a data quality assessment system based on the government affairs field, including:

[0013] An acquisition unit, configured to acquire a business scenario portrait and a data set image; wherein the business scenario portrait is used to reflect the business type currently processed by the government data, and the data set image is used to reflect the data presentation form and its association relationship during the government data processing process;

[0014] A segmentation unit, configured to divide the dataset image into business stages based on the business scenario portrait, and determine business segmentation data corresponding to the dataset image; the business segmentation data is used to reflect the data application stage of the government data;

[0015] A quality unit, configured to obtain a corresponding evaluation dimension system based on the business segment data, and determine a quality feature vector corresponding to the business segment data based on the evaluation dimension system; wherein the evaluation dimension system is used to reflect the core evaluation direction of data quality, and the quality feature vector is used to quantify the actual performance of the data under each evaluation dimension;

[0016] A result unit is used to generate a comprehensive data quality evaluation result based on the quality feature vector corresponding to the business segment data.

[0017] In a third aspect, an embodiment of the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described in any one of the above aspects when executing the computer program.

[0018] In a fourth aspect, an embodiment of the present application provides a computer program product, which, when executed on an electronic device, enables the electronic device to execute the method according to any one of the above aspects.

[0019] It can be understood that the beneficial effects of the second to fourth aspects mentioned above can be found in the relevant descriptions of the above aspects and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0021] Figure 1 This is a flowchart of a data quality assessment method based on the government affairs field provided by an embodiment of the present application;

[0022] Figure 2 This is a schematic diagram of the operation of a data quality assessment method based on the government affairs field provided by an embodiment of the present application;

[0023] Figure 3 This is a schematic diagram of the structure of a data quality assessment system based on the government affairs field provided by an embodiment of the present application;

[0024] Figure 4 It is a structural diagram of an electronic device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0025] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.

[0026] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.

[0027] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0028] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if the described condition or event is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of the described condition or event" or "in response to detecting the described condition or event," depending on the context.

[0029] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0030] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.

[0031] Currently, similar publicly available technologies in the government sector include simple rule-based data validation and single-dimensional data quality assessment. Simple rule-based data validation primarily verifies government data by pre-setting basic data formatting rules, such as ensuring data type compatibility and ensuring mandatory fields are left blank. For example, in population registration data, the age field must be numeric; the presence of characters is considered a data quality issue. This technology only detects the most basic, superficial data errors and is unaware of deeper quality issues such as data integrity, consistency, and relevance. Single-dimensional data quality assessment techniques assess a single dimension of government data, such as accuracy. Data accuracy is determined by comparing it with authoritative data sources. For example, in tax data, accuracy is assessed by comparing the tax amounts declared by enterprises with the actual amounts recorded in the tax system. This technique ignores other important dimensions of data quality, such as data timeliness and accessibility, and fails to fully reflect the quality of government data.

[0032] In order to solve the above problems, an embodiment of the present application provides a data quality assessment method based on the government affairs field. In this method, by obtaining business scenario portraits and data set images, basic information is provided for subsequent business stage division and quality assessment. Based on the business scenario portraits, the data set images are divided into business stages, and the business segmentation data corresponding to the data set images are determined to achieve structured classification of the government affairs data processing stages. According to the business segmentation data, the corresponding evaluation dimension system is obtained, and the quality feature vector corresponding to the business segmentation data is determined according to the evaluation dimension system, the core direction of the data quality assessment at each stage is clarified and the assessment results are quantified. Based on the quality feature vector corresponding to the business segmentation data, a comprehensive data quality evaluation result is generated to form a holistic assessment conclusion that comprehensively reflects the quality of government affairs data, enhance the government affairs data processing capabilities, and improve the accuracy of the assessment results.

[0033] The data quality assessment method based on the government affairs field provided in the embodiment of the present application can be applied to electronic devices. In this case, the electronic device is the executor of the data quality assessment method based on the government affairs field provided in the embodiment of the present application. The embodiment of the present application does not impose any restrictions on the specific type of electronic device.

[0034] For example, the electronic device may be an ultra-mobile personal computer (UMPC), a netbook, a desktop computer, a computer, a laptop computer, a communication device, a computing device, a satellite wireless device, etc.

[0035] In order to better understand the data quality assessment method based on the government affairs field provided in the embodiment of the present application, the specific implementation process of the data quality assessment method based on the government affairs field provided in the embodiment of the present application is exemplarily introduced below.

[0036] Figure 1 The following is a schematic flow chart of a data quality assessment method based on the government affairs field provided by an embodiment of the present application. Figure 2 The following is a schematic diagram showing the operation of the data quality assessment method based on the government affairs field provided by an embodiment of the present application. The data quality assessment method based on the government affairs field includes:

[0037] S100, obtaining a business scenario portrait and a data set image; wherein, the business scenario portrait is used to reflect the business type currently processed by the government data, and the data set image is used to reflect the data presentation form and its correlation relationship of the government data during the government processing process.

[0038] A business scenario portrait is an abstract description of the current government data processing business type, extracted from government business data. It covers, for example, business processes, involved departments, and policy basis. A dataset image is a visual representation of the data format (e.g., structured or semi-structured) and its relationships (e.g., field dependencies and inter-table associations) during government data processing. Data collection can be automated through integration with government system databases and APIs to construct the portrait and image. Alternatively, data can be supplemented by manually entering key business information. Data collection tools such as Flume are used for real-time data collection, while ETL (Extract-Transform-Load) tools such as Kettle are used for data cleansing and transformation. Obtaining business scenario portraits and dataset images provides a comprehensive and accurate data foundation for subsequent business phase delineation and quality assessment. This helps the system understand the business context and data characteristics of government data processing, improving the relevance and accuracy of assessments.

[0039] S200, divide the data set image into business stages based on the business scenario portrait, and determine the business segmentation data corresponding to the data set image; the business segmentation data is used to reflect the data application and other stages of the government data.

[0040] Business stage segmentation, as understood, divides the data processing process covered by the dataset image into different business stages, such as data collection, cross-departmental sharing, and application services, based on information such as business processes and policy requirements contained in the business scenario portrait. Based on the business type labels in the business scenario portrait (e.g., "administrative approval" or "civilian subsidy issuance"), the pre-set business segmentation rule library can be matched to the corresponding segmentation strategy. The data in the dataset image is categorized according to the strategy, and data in the same business processing stage is grouped into a single business segment. For example, for the "administrative approval" business, application submission data can be grouped into "acceptance stage data," and review process data into "review stage data." During this segmentation process, a rule engine (such as Drools) can be used to automatically match and execute segmentation rules. Business stage segmentation can structure the complex government data processing process, making subsequent data quality assessments for each stage more accurate and systematic, and facilitating the identification of data quality issues specific to each stage.

[0041] In one possible implementation, S200, dividing the dataset image into business stages based on the business scenario portrait and determining the business segment data corresponding to the dataset image, includes:

[0042] S210, matching business segmentation labels of dataset images according to business scenario portraits.

[0043] It can be understood that business segmentation labels are keywords or classification identifiers used to identify the business stage to which the data in the dataset image belongs, such as "basic business," "cross-domain sharing," and "application services." The matching algorithm compares the core business information in the business scenario image (such as business process nodes and the scope of application of policies and regulations) with a pre-defined library of business segmentation labels and calculates a similarity score. For example, if the business scenario image clearly includes the "data entry system" link, the "basic business" label can be matched; if it includes "multi-department data collaborative analysis," the "cross-domain sharing" label can be matched. Text similarity calculation algorithms (such as cosine similarity) can be used to quantify the degree of match. A match is considered successful when the similarity exceeds a threshold (e.g., 0.8). By matching business segmentation labels, the business attributes of the data in the dataset image can be quickly located, providing a clear classification basis for subsequent business stage division, reducing manual intervention, and improving the efficiency and accuracy of data processing stage division.

[0044] S220 , dividing the data set image into business stages based on the business segmentation labels, and determining business segmentation data corresponding to the data set image.

[0045] It can be understood that based on the matched business segmentation labels, data with the same label in the dataset image is aggregated together to form different business segmentation data. Data classification algorithms (such as clustering algorithms) can be used to classify data records according to business segmentation labels. For example, data records with the "basic business" label are integrated into "basic business data", which includes information such as time, collection personnel, and original data format; data with the "cross-domain sharing" label is divided into "cross-domain sharing data", which covers shared interfaces, transmission logs, receiver feedback, and other content. During the classification process, big data processing frameworks (such as Spark) can be used to achieve rapid classification and aggregation of massive data. Through business stage division, the conversion from label matching to actual data segmentation can be completed, so that the originally mixed government data can be orderly grouped according to business processing logic, laying the foundation for targeted data quality assessment and facilitating differentiated management and optimization of data at different stages.

[0046] S300, obtain the corresponding evaluation dimension system based on the business segment data, and determine the quality feature vector corresponding to the business segment data based on the evaluation dimension system; wherein the evaluation dimension system is used to reflect the core evaluation direction of data quality, and the quality feature vector is used to quantify the actual performance of the data under each evaluation dimension.

[0047] It can be understood that the evaluation dimension system is a set of data quality assessment dimensions defined for different business segment data. For example, for the basic business phase, evaluation dimensions might include completeness and accuracy; for the cross-domain sharing phase, these might include consistency and transmission stability. A quality feature vector is a vector representation of the actual performance of the business segment data under each evaluation dimension. To obtain the evaluation dimension system, the corresponding dimension set is retrieved from a pre-built evaluation dimension rule library based on the business segment data's label (e.g., "basic business data," "cross-domain sharing data," or "application service data"). To determine the quality feature vector, specific calculation rules can be designed for each evaluation dimension. For example, in the data collection phase, the completeness dimension is quantified by calculating the ratio of missing required fields to the total number of fields; in the data sharing phase, the transmission stability dimension is calculated by counting the number of failed interface calls per unit time. Using a data computing framework (such as Pandas), specific quantitative calculations are performed, combining the quantitative results of each dimension into a quality feature vector. This transforms abstract quality assessment dimensions into measurable indicators, accurately characterizing the quality of business segment data and facilitating subsequent data quality comparison and analysis.

[0048] In one possible implementation, S300, obtaining a corresponding evaluation dimension system based on the business segment data, and determining a quality feature vector corresponding to the business segment data based on the evaluation dimension system, includes:

[0049] S310, when the business segmentation data is basic business data, obtain a first evaluation dimension system, and determine a first quality feature vector corresponding to the basic business data according to the first evaluation dimension system; wherein the first evaluation dimension system is used to reflect the missing status of the basic business data.

[0050] It can be understood that basic business data usually refers to data that is at the initial stage of the government data processing process and provides basic support for subsequent business, such as initially reported personal information, basic enterprise registration information, etc. The first evaluation dimension system is a set of quality assessment dimensions specifically set for basic business data, mainly focusing on data missing conditions, such as field missing rate, key information completeness, etc. When obtaining the first evaluation dimension system, the corresponding dimension list can be retrieved from the evaluation dimension configuration library based on the classification identifier of the business segment data ("basic business data"), and the first quality feature vector can be determined according to the rules of the first evaluation dimension system, so as to promptly discover potential problems in the data entry link and provide a reliable data foundation for subsequent data processing.

[0051] Optionally, in step S310, determining a first quality feature vector corresponding to the basic business data according to the first evaluation dimension system includes:

[0052] S311, extracting the core field set of the basic business data according to the first evaluation dimension system, and calculating the missing value ratio of the core field set within a preset time window to generate a completeness feature vector.

[0053] It can be understood that the core field set refers to the set of fields in the underlying business data that are critical to subsequent business processing and decision-making. This set is selected based on the data integrity assessment requirements of the first evaluation dimension. For example, in enterprise registration data, fields such as the unified social credit code, enterprise name, and registered address constitute the core field set. To calculate the missing value ratio, data query tools (such as SQL statements) can be used to filter records in the core field set from the underlying business data. The number of missing values ​​for each field within a preset time window (e.g., the past week) is counted and divided by the total number of records for that field to obtain the missing value ratio for each field. Finally, these missing value ratios are combined into a vector, the completeness feature vector. For example, if the missing value rate for "unified social credit code" in enterprise registration data is 5% and the missing value rate for "enterprise name" is 3%, the completeness feature vector might be [0.05, 0.03]. Generating a completeness feature vector can intuitively reflect the completeness of the underlying business data at the critical information level, helping to quickly identify missing data issues, providing a quantitative basis for data completion and quality improvement, and ensuring that subsequent business operations are carried out smoothly based on complete data.

[0054] S312: extract the state transition trajectory of the data items in the basic business data according to the first evaluation dimension system, calculate the average transition time of each data item from the initial state to the target state, and generate a time-sensitive feature vector.

[0055] It can be understood that the state transition trajectory of a data item refers to the series of state changes that each data item in the basic business data undergoes during the business processing process, from creation (initial state) to meeting specific business requirements (target state), such as the transition of data from "pending review" to "approved" or "review failed." The timeliness assessment requirements in the first evaluation dimension system determine the state nodes to focus on and the calculation logic. The state transition trajectory of each data item can be extracted using database log records or status field information, using data tracking algorithms (such as timestamp-based state sequence analysis). The time it takes for each data item to transition from the initial state to the target state is then calculated. The time taken for all data items is accumulated and divided by the total number of data items to obtain the average transition duration. The average transition durations for different types of data items (e.g., personal information and business application) are combined to form a timeliness feature vector. For example, if the average review duration for personal information is 2 hours and the average review duration for business applications is 4 hours, then the timeliness feature vector is [2, 4]. By generating time-efficiency feature vectors, we can quantitatively evaluate the time efficiency of basic business data in the processing process, promptly identify bottlenecks in the data processing link, and provide data support for optimizing business processes and improving data timeliness.

[0056] S313: Fusing the timeliness feature vector and the integrity feature vector to determine a first quality feature vector corresponding to the basic business data.

[0057] It can be understood that feature fusion is to integrate the timeliness feature vector and integrity feature vector that describe the quality of basic business data from different perspectives to form a vector representation that more comprehensively reflects the data quality. The two vectors can be normalized and their values ​​can be scaled to the same range (such as ), to eliminate the influence of dimension difference on the fusion result, the normalization formula ( ) is implemented; according to the business characteristics and importance of the basic business data, different weights are assigned to the timeliness feature vector and the integrity feature vector. For example, for emergency data with high real-time requirements, the timeliness feature vector is given a weight of 60% and the integrity feature vector is given a weight of 40%; for basic archive data, the two weights are 50% each. The two vectors are fused by weighted summation to obtain the first quality feature vector. For example, if the standardized timeliness feature vector is , the integrity feature vector is , after fusion with a weight of 50%, the first quality feature vector is Through feature fusion, we can comprehensively consider the performance of basic business data in two key dimensions: time efficiency and information completeness, obtain more comprehensive and accurate quality assessment results, and provide a more reliable basis for data quality management.

[0058] S320, when the business segmentation data is cross-domain collaborative data, obtain a second evaluation dimension system, and determine a second quality feature vector corresponding to the cross-domain collaborative data based on the second evaluation dimension system; wherein the second evaluation dimension system is used to reflect the correlation status of the cross-domain collaborative data.

[0059] It can be understood that cross-domain collaborative data refers to data that is interactively shared between multiple departments, institutions or systems in government affairs, such as business operation data shared by tax departments and market supervision departments. The second evaluation dimension system is a set of evaluation directions set for cross-domain collaborative data, focusing on the correlation of data, including the consistency of data in cross-domain transmission, the stability of interface docking, etc. When obtaining the second evaluation dimension system, you can retrieve the corresponding dimension list from the evaluation dimension configuration library based on the "cross-domain collaborative data" label of the business segment data, determine the second quality characteristic vector, and accurately measure the quality level of cross-domain collaborative data in the sharing process, promptly discover problems in cross-departmental data interaction, and ensure the efficient development of government collaborative business.

[0060] Optionally, in step S320, determining a second quality feature vector corresponding to the cross-domain collaborative data according to the second evaluation dimension system includes:

[0061] S321: Determine the correlation deviation vector corresponding to the cross-domain collaborative data according to the second evaluation dimension system.

[0062] It can be understood that the association deviation vector is a vector representation used to quantify the degree of deviation of cross-domain collaborative data from the expected rules in terms of association relationships. The rules for data association in the second evaluation dimension system (such as field mapping rules and data consistency verification rules) are the basis for determining the association deviation vector. According to the second evaluation dimension system, cross-domain association rules in cross-domain collaborative data can be extracted to construct a set of constraints for data association assessment. For example, it is stipulated that the "taxpayer identification number" in the tax system must correspond one-to-one with the "unified social credit code" in the market supervision system; then the actual association trajectory of the cross-domain data items is extracted, that is, the actual correspondence between the data during cross-domain transmission and use, to generate an association feature sequence; the association feature sequence is compared and analyzed with the constraint set, and by calculating the deviation between the actual association trajectory and the theoretical association rule at each association point (such as the proportion of mismatched fields, the number of association errors, etc.), these deviation values ​​are combined into an association deviation vector. For example, if two associated fields do not match, and the total number of associated fields is 10, the deviation is 20%, and the associated deviation vector might be [0.2] (if there is only one key associated dimension) or [0.2, 0, 0] (if there are multiple associated dimensions and no deviation in the other dimensions). By determining the associated deviation vector, we can intuitively reflect the quality issues of cross-domain collaborative data in terms of association relationships, help quickly identify key defects such as data inconsistencies and mapping errors, and provide a quantitative reference for optimizing cross-departmental data collaboration.

[0063] Exemplarily, S321, determining the correlation deviation vector corresponding to the cross-domain collaborative data according to the second evaluation dimension system, includes:

[0064] S3211: Extract cross-domain association rules from cross-domain collaborative data based on the second evaluation dimension system, and construct a constraint condition set for data association evaluation.

[0065] Cross-domain association rules refer to standards for field mapping and business logic correspondences that must be followed when data is exchanged between different government departments or systems. For example, the "ID Number" in the public security system must be completely consistent with the "Citizen Identity Number" in the civil affairs system, and the "Enterprise Taxpayer Number" in the tax system must have pre-defined association logic with the "Unified Social Credit Code" in the market supervision system. The second evaluation dimension system includes assessment criteria for cross-domain data association, such as field matching rules and business process linkage rules. These rules can be extracted from the second evaluation dimension system by parsing business specification documents or invoking a predefined rule engine and converted into a computable set of constraints. For example, "The match rate between field A and field B must be ≥95%," "A change in the status of data item X must trigger a synchronous update of data item Y," etc. Once the constraint set is constructed, it can be formalized using Structured Query Language (SQL) or a rule description language (such as Drools) to provide a benchmark for subsequent association evaluation. This step effectively defines an "ideal model" for cross-domain data association, providing a reference standard for quantitatively evaluating the association quality of actual data and ensuring that cross-departmental data collaboration is governed by rules.

[0066] S3212: extracting actual association trajectories of cross-domain data items in the cross-domain collaborative data according to the second evaluation dimension system, and generating an association feature sequence.

[0067] It can be understood that the actual correlation trajectory of cross-domain data items refers to the actual correspondence and status change path of data items between different systems during cross-departmental transmission and processing. For example, when synchronizing enterprise registration data from the market supervision system to the tax system, the consistency of values ​​in fields such as "Enterprise Name" and "Registered Address" changes. The second evaluation dimension system defines the correlation dimensions to be monitored (such as field-level correlation and business process-level correlation). Through methods such as data interface log collection and ETL process tracking, cross-domain data item interaction records can be captured in real time and arranged chronologically to form a correlation feature sequence. For example, for the "Enterprise Deregistration Application" business, the correlation feature sequence may include nodes such as "Deregistration application submitted to the market supervision system → Tax system receives the application and verifies the tax status → Market supervision system completes deregistration based on the tax feedback." Each node records the correlation status of the data item (such as "Field match successful" or "Field discrepancies require manual verification"). After generating the correlation feature sequence, sequence analysis tools (such as Python's pandas library) can be used for structured storage and preprocessing, providing a data foundation for comparative analysis. The effect of this step is to restore the "real scenario" of cross-domain data association. By recording the actual interaction details, it provides clues for discovering anomalies in data association (such as field mapping errors and process interruptions).

[0068] S3213, compare and analyze the association feature sequence with the constraint condition set, calculate the deviation between the actual association trajectory and the theoretical association rule, and generate an association deviation vector.

[0069] Deviation is a quantitative indicator that measures the degree of inconsistency between actual association traces and theoretical association rules. By comparing and analyzing the rules in the constraint set with the actual data in the association feature sequence, the deviation values ​​for each dimension are calculated. Dynamic time warping (DTW) or edit distance algorithms can be used to match the association feature sequence with the standard association process in the constraint set. Non-compliant nodes (such as field matching failures or missing process steps) are identified, and metrics such as the number of deviations and the duration of the deviations are calculated. For example, if a constraint requires "cross-departmental data field matching rate ≥ 95%," and the number of field matching failures in the actual association feature sequence accounts for 8% of the total number of interactions, the deviation is 8%. The deviations for each dimension (such as field matching deviation, process timing deviation, and business logic deviation) are combined into an association deviation vector, for example, [0.08, 0.03, 0] (assuming a 3% process timing deviation and no business logic deviation). Generating an association deviation vector is to convert abstract association quality issues into measurable numerical values. The association deviation vector can intuitively reflect the weak links in cross-domain data associations, such as business nodes with frequent field mapping errors or department interfaces with poor process connections. This provides a quantitative basis for accurately locating data collaboration problems and helps optimize cross-domain data governance processes.

[0070] S322: Obtain the response time distribution of cross-domain collaborative data during cross-domain transmission and generate a transmission time efficiency vector.

[0071] It can be understood that when cross-domain collaborative data is transmitted across departments and systems, the time interval from the sender initiating a request to the receiver completing data reception and enabling normal processing is the transmission response time. The transmission time efficiency vector is a vector formed by quantifying the response time characteristics of different data transmission tasks. To obtain the response time distribution, a monitoring program can be deployed at the data transmission interface (such as using a network packet capture tool or API call logging) to collect the start and end times of each data transmission in real time and calculate the response time. Statistical analysis is then performed on all transmission response times over a period of time (such as a day or a week) to obtain the response time distribution, such as average response time, maximum response time, and standard deviation of response time. These indicators are then combined in a preset order to form a transmission time efficiency vector, such as [average response time, maximum response time, standard deviation]. For example, if the average response time is 500ms, the maximum response time is 2000ms, and the standard deviation is 100ms, the transmission time efficiency vector is [500, 2000, 100]. By generating a transmission time efficiency vector, the time efficiency and stability of cross-domain collaborative data during transmission can be clearly presented, making it easier to discover problems such as excessive transmission delay and excessive fluctuation, and providing data support for optimizing data transmission links and improving cross-domain collaborative efficiency.

[0072] S323: Perform feature fusion on the correlation deviation vector and the transmission time efficiency vector to determine a second quality feature vector corresponding to the cross-domain collaborative data.

[0073] It can be understood that feature fusion combines the correlation deviation vector, which reflects the accuracy of cross-domain collaborative data correlation, and the transmission timeliness vector, which reflects the timeliness of data transmission, to form a vector that comprehensively assesses data quality. Both vectors can be normalized to the same scale (e.g., [0, 1]) to eliminate dimensional differences. This can be achieved using the min-max normalization method. Weights are assigned to the correlation deviation vector and the transmission timeliness vector based on the actual needs of cross-domain collaborative services and data characteristics. For example, for emergency data collaboration scenarios with extremely high real-time requirements, the transmission timeliness vector can be given a weight of 70% and the correlation deviation vector a weight of 30%. For routine business data sharing, each can be weighted 50%. Finally, the two vectors are fused through a weighted summation to obtain a second quality feature vector. For example, if the normalized correlation deviation vector is [0.3] and the transmission timeliness vector is [0.8], after fusion with a 50% weighting, the second quality feature vector is [(0.3 × 0.5 + 0.8 × 0.5)] = [0.55]. Through feature fusion, we can comprehensively consider the quality performance of cross-domain collaborative data in two key dimensions: correlation and transmission efficiency, obtain more representative quality assessment results, and provide a scientific basis for cross-domain data management and optimization.

[0074] S330, when the business segment data is application service data, obtain a third evaluation dimension system, and determine a third quality feature vector corresponding to the application service data based on the third evaluation dimension system; wherein the third evaluation dimension system is used to reflect the compliance status of the application service data.

[0075] Application service data refers to data used to support government service applications (such as online services and data analysis and decision-making). Examples include data submitted by the public for public services and statistical analysis data used for government decision-making. The third evaluation dimension system is a set of assessment areas specifically designed for application service data, focusing on the data's compliance status, including compliance with data usage permissions and sensitive information handling. To obtain the third evaluation dimension system, the corresponding dimension list can be retrieved from the evaluation dimension configuration library based on the "Application Service Data" tag in the business segment data. To determine the third quality feature vector, the third evaluation dimension system can be used to extract access permission change logs for sensitive data items within the application service data and analyze the frequency of permission changes. Anomaly detection can then be performed on the data access traces, and the relevant quantitative results can be integrated. Using access log analysis tools and anomaly detection algorithms, a quantitative assessment of the compliance of application service data can be achieved, effectively monitoring the compliance of application service data during use, preventing data security risks, and ensuring the legal and secure operation of government service applications.

[0076] Optionally, in step S330, determining a third quality feature vector corresponding to the application service data according to the third evaluation dimension system includes:

[0077] S331, extracting access permission change logs of sensitive data items in application service data according to the third evaluation dimension system, generating a permission change feature sequence and calculating the change frequency of the permission change feature sequence within a preset time window, and generating a change frequency vector.

[0078] Sensitive data items refer to data within application service data that contains sensitive information such as personal privacy, trade secrets, or national security, such as identification card numbers and bank account information. Access rights change logs record changes in access rights to these sensitive data items at different points in time, including operations such as granting, modifying, and revoking permissions. A permission change feature sequence is a chronological sequence of these change operations, used to describe the process of permission changes. The compliance requirements for permission management within the third evaluation dimension system serve as the basis for extracting and analyzing logs. Data query tools (such as SQL) can be used to extract access rights change logs for sensitive data items within application service data from the permission management database. These logs are organized into a permission change feature sequence in timestamp order. The total number of permission changes within a preset time window (e.g., one month) is calculated and divided by the length of the time window (e.g., number of days) to obtain the average daily change frequency. The change frequencies of different types of sensitive data items (such as personal identity information and financial data) are combined into a change frequency vector. For example, if the monthly change frequency of sensitive personal identity information items is 10 times, and the monthly change frequency of sensitive financial data items is 5 times, the change frequency vector is [10 / 30,5 / 30]. By generating a change frequency vector, we can quantitatively assess changes in sensitive data access permissions, promptly identify abnormally frequent permission changes, prevent the risk of unauthorized data access, and ensure the compliance of application service data usage.

[0079] S332: Perform anomaly detection on the access trace of the application service data to generate an anomaly feature vector.

[0080] As can be understood, the access trajectory of application service data refers to the operational record of data access, including information such as access time, accessing user, and access operation type (e.g., query, modify, delete). Anomaly detection analyzes access trajectories to identify operational behaviors that do not conform to normal access patterns. The requirements for data access security in the third evaluation dimension provide a standard for anomaly detection. Log analysis tools can be used to collect access trajectory data for application service data. Access trajectories can be analyzed using algorithms such as the Isolation Forest algorithm, One-Class Support Vector Machine (SVM), or rule-based methods (e.g., setting a threshold for frequent access by the same user within a short period of time). For each detection dimension (e.g., access time anomalies, access frequency anomalies, and operation type anomalies), an anomaly score is calculated and combined to form an anomaly feature vector. For example, if the access time anomaly score is 0.8, the access frequency anomaly score is 0.3, and the operation type anomaly score is 0.1, then the anomaly feature vector is [0.8, 0.3, 0.1]. By generating an anomaly feature vector, suspicious behavior in the application service data access process can be quickly identified, providing timely warnings of data leakage or malicious operation risks, and ensuring data security.

[0081] S333: Perform feature fusion on the change frequency vector and the abnormal feature vector to determine a third quality feature vector corresponding to the application service data.

[0082] It can be understood that feature fusion combines the change frequency vector, which reflects changes in sensitive data access rights, and the anomaly feature vector, which reflects the security status of data access, to form a vector that comprehensively assesses the compliance quality of application service data. During implementation, both vectors are first normalized to the range [0, 1] to eliminate dimensionality effects. Z-score normalization can be used. Weights can be assigned based on the security importance of the application service scenario. For services involving highly sensitive information (such as medical data query services), the anomaly feature vector is weighted 70% and the change frequency vector 30%. For general public service data, each is weighted 50%. Finally, the two vectors are fused through a weighted sum to obtain a third quality feature vector. For example, if the normalized change frequency vector is [0.4] and the anomaly feature vector is [0.6], after fusion with a 50% weighting, the third quality feature vector is [(0.4 × 0.5 + 0.6 × 0.5)] = [0.5]. Through feature fusion, we can comprehensively consider the quality performance of application service data in terms of permission management and access security, provide an accurate basis for data compliance assessment, and help government services operate safely and stably.

[0083] S400: Generate a comprehensive data quality evaluation result based on the quality feature vector corresponding to the business segment data.

[0084] It can be understood that the quality characteristic vector is a quantitative representation of the quality performance of each business segment data under a specific evaluation dimension. Generating a comprehensive data quality evaluation result is to integrate and analyze the quality characteristics of these segments to obtain an assessment of the overall data quality. The first quality characteristic vector, second quality characteristic vector, and third quality characteristic vector corresponding to basic business data, cross-domain collaborative data, and application service data can be summarized; weights are assigned to each quality characteristic vector based on the importance of different business segments in the overall government data processing and the data application scenarios; through weighted calculations or grading judgment rules, the information of each vector is integrated and mapped to a preset quality grade interval (such as excellent, good, qualified, and unqualified) to obtain a comprehensive data quality evaluation result, thereby achieving a comprehensive and systematic assessment of the overall quality of government data, providing an intuitive reference basis for data governance and optimized decision-making, and helping managers quickly grasp the overall picture of data quality and make targeted improvements.

[0085] In one possible implementation, S400 generates a comprehensive data quality evaluation result based on the quality feature vector corresponding to the business segment data, including:

[0086] S410 : Generate a comprehensive data quality evaluation result based on the first quality feature vector, the second quality feature vector, and the third quality feature vector.

[0087] It can be understood that the first, second, and third quality characteristic vectors quantitatively reflect the quality of data in different dimensions, respectively, from the three key business segments of basic business, cross-domain collaboration, and application services. The process of generating comprehensive evaluation results is the process of systematically integrating the quality characteristics of these segments. The three quality characteristic vectors can be normalized to the same measurement scale. The weight distribution strategy for each vector is determined based on the logical relationship of the government business process and the data application scenario. For example, in business scenarios centered on data sharing, the second quality characteristic vector is given a higher weight. The vector information is then fused and calculated through weighted summation or complex comprehensive evaluation models (such as neural networks and decision tree integration models). The data quality level is divided according to the calculation results, and a comprehensive evaluation result is generated. This can break the boundaries of business segments and achieve a global and comprehensive assessment of government data quality, providing a more scientific and comprehensive basis for data quality optimization and business decision-making.

[0088] Optionally, S410, generating a comprehensive data quality evaluation result based on the first quality feature vector, the second quality feature vector, and the third quality feature vector, includes:

[0089] S411 : Determine a first quality characteristic value based on the first quality characteristic vector, determine a second quality characteristic value based on the second quality characteristic vector, and determine a third quality characteristic value based on the third quality characteristic vector.

[0090] It can be understood that a quality feature vector is a vector set composed of eigenvalues ​​from multiple dimensions. Determining the quality feature value involves extracting key quantitative indicators from the vector for subsequent comprehensive evaluation. For the first quality feature vector, corresponding eigenvalues ​​can be selected based on core dimensions of basic business data quality assessment (such as completeness and timeliness). For example, the integrity and timeliness eigenvalues ​​can be extracted and weighted to obtain the first quality feature value. For the second quality feature vector, key eigenvalues ​​are selected from dimensions such as correlation deviation and transmission timeliness of cross-domain collaborative data, and then integrated and calculated to obtain the second quality feature value. The third quality feature vector is calculated by extracting eigenvalues ​​from dimensions such as permission change frequency and access anomaly level of application service data. The eigenvalues ​​of each dimension in the vector can be processed using a preset calculation formula or algorithmic model, such as weighted averaging or principal component analysis followed by dimensionality reduction to extract key values. By determining each quality feature value, complex vector information is simplified into a single, representative value, facilitating subsequent data quality comparison, ranking, and comprehensive assessment, improving evaluation efficiency and accuracy.

[0091] S412: Perform basic quality filtering on the first quality feature value, the second quality feature value, and the third quality feature value respectively, and generate a comprehensive data quality evaluation result according to the basic quality filtering results.

[0092] It can be understood that basic quality filtering is to conduct a preliminary screening of the quality characteristic values ​​of each business segment by setting key quality thresholds, and quickly identify problems that seriously affect data quality. For the first quality characteristic value, key indicator thresholds such as the integrity and timeliness of basic business data can be set. For example, the integrity threshold is 90%. If the integrity reflected in the first quality characteristic value is lower than this threshold, it is determined that it has not passed the filtering; similarly, the correlation accuracy and transmission timeliness thresholds of cross-domain collaborative data are set for the second quality characteristic value, and the compliance rate threshold of application service data is set for the third quality characteristic value. If any quality characteristic value fails the corresponding threshold test, a lower-level data quality comprehensive evaluation result is directly generated; if all pass, it enters the next stage of comprehensive calculation, using a "one-vote veto" mechanism to prioritize the exclusion of data with serious quality defects, quickly locate the problem business segment, and at the same time reduce unnecessary complex calculations, improve the efficiency of data quality assessment, and ensure that the comprehensive evaluation results can prioritize the key quality issues of the data.

[0093] Exemplarily, in S412, basic quality filtering is performed on the first quality feature value, the second quality feature value, and the third quality feature value, and a comprehensive data quality evaluation result is generated according to the basic quality filtering result, including:

[0094] S4121 : When the first quality feature value, the second quality feature value, and the third quality feature value do not all pass the basic quality filter, generate a first data quality comprehensive evaluation result and determine the first data quality comprehensive evaluation result as the data quality comprehensive evaluation result.

[0095] It can be understood that if any one or more of the first, second, and third quality characteristic values ​​fail to reach the preset basic quality threshold (such as insufficient integrity of basic business data, or too high an error rate in cross-domain collaborative data association), it indicates that there are serious quality problems with government data in the corresponding business segment. A lower-level comprehensive evaluation result of the first data quality (such as "unqualified" or "poor") can be directly generated and determined as the final comprehensive evaluation result of data quality. The three quality characteristic values ​​can be compared one by one through conditional judgment statements (such as if-else statements). Once a characteristic value is found to be lower than the corresponding threshold, the low-level evaluation generation logic is immediately triggered, skipping the subsequent complex weighted calculation steps. This allows for rapid identification and marking of data with serious defects in the early stages of data quality assessment, providing data governance personnel with a clear direction for rectification, and avoiding the impact of local data quality issues on the normal development of the overall business, while improving the execution efficiency of the assessment process.

[0096] S4122: When the first quality feature value, the second quality feature value, and the third quality feature value all pass the basic quality filter, a second data quality comprehensive evaluation result is generated and the second data quality comprehensive evaluation result is determined as the data quality comprehensive evaluation result.

[0097] It can be understood that when the quality characteristic values ​​corresponding to basic business data, cross-domain collaboration data, and application service data all meet the preset basic quality threshold requirements, it indicates that the data does not have serious quality issues in each key business segment. The refined assessment phase can then be entered, where a more complex calculation is used to generate a second comprehensive data quality evaluation result, which serves as the final assessment. During execution, the weights of each quality characteristic value are first determined based on the business scenario (such as emergency command, statistical analysis, and daily government services). For example, in the emergency command scenario, the third quality characteristic value of application service data is given a higher weight. A weighted summation formula (comprehensive evaluation result = first quality characteristic value × weight 1 + second quality characteristic value × weight 2 + third quality characteristic value × weight 3) is then used to calculate a comprehensive score. Finally, this score is mapped to a preset quality grade range (e.g., 85-100 is "excellent" and 70-84 is "good") to generate the second comprehensive data quality evaluation result. While ensuring that the basic data quality meets the requirements, the contribution of the data quality of each business segment to the overall data quality is further comprehensively considered to produce a more scientific and accurate comprehensive assessment of government data quality, providing a more valuable reference for data optimization and business decision-making.

[0098] Exemplarily, at S4122, when the first quality feature value, the second quality feature value, and the third quality feature value all pass the basic quality filter, generating a second data quality comprehensive evaluation result and determining the second data quality comprehensive evaluation result as the data quality comprehensive evaluation result includes:

[0099] S41221: When the first quality characteristic value, the second quality characteristic value, and the third quality characteristic value all pass the basic quality filter, determine a quality assessment weight according to the service segment data.

[0100] It can be understood that quality assessment weights are a set of parameters used to measure the dynamic contribution of the quality characteristics of basic business data, cross-domain collaboration data, and application service data to the overall comprehensive data quality evaluation. The attributes of business segment data (such as business type, data volume, and real-time requirements) determine the varying importance of the quality characteristics of each segment. The business scenario to which the current government data belongs (such as emergency command, statistical analysis, and public services) can be identified, and initial weight allocations can be retrieved from a pre-set weight template library (for example, a 50% weight for application service data in the emergency command scenario). By analyzing the data size of each business segment (such as the number of records and data volume), a normalization algorithm is used to calculate the actual data volume share of basic business data, cross-domain collaboration data, and application service data. These initial weights are then adjusted (for example, if a segment's data volume exceeds 60%, its quality characteristic weight is increased by 10%-20%) to form the final quality assessment weights. This step dynamically adapts weight allocation to business needs and data characteristics, highlighting the quality impact of key business segments, ensuring that comprehensive evaluation results better meet the actual needs of government data processing and providing guidance for precise data governance.

[0101] S41222, weighting the first quality characteristic value, the second quality characteristic value, and the third quality characteristic value according to the quality assessment weight, generating a second data quality comprehensive evaluation result and determining the second data quality comprehensive evaluation result as the data quality comprehensive evaluation result.

[0102] It can be understood that weighted calculation involves quantifying and integrating the quality characteristic values ​​of each business segment according to their corresponding weights to produce a comprehensive evaluation result that comprehensively reflects the quality of government data. The three quality characteristic values ​​can be calculated using the weighted summation formula: Comprehensive evaluation score = first quality characteristic value × first weight + second quality characteristic value × second weight + third quality characteristic value × third weight. For example, if the first quality characteristic value is 80, corresponding to a weight of 30%; the second quality characteristic value is 75, corresponding to a weight of 20%; and the third quality characteristic value is 90, corresponding to a weight of 50%, the comprehensive evaluation score is 80 × 0.3 + 75 × 0.2 + 90 × 0.5 = 84. The comprehensive evaluation score can be mapped to a pre-set quality grade range (e.g., 0-60 is "unqualified," 61-79 is "qualified," 80-89 is "good," and 90-100 is "excellent") to determine the corresponding quality grade and generate a second comprehensive data quality evaluation result, which serves as the final government data quality assessment. Through weighted calculation, the differences in the contribution of data quality of different business segments to the overall situation are fully taken into account, making the evaluation results more scientific and reasonable, and able to accurately reflect the actual quality level of government data, providing a reliable basis for data quality improvement and business decision-making.

[0103] Corresponding to the data quality assessment method based on the government affairs field in the above embodiment, the embodiment of the present application also provides a data quality assessment system based on the government affairs field, and each unit of the system can implement each step of the data quality assessment method based on the government affairs field. Figure 3 A structural block diagram of a data quality assessment system based on the government affairs field provided in an embodiment of the present application is shown. For the sake of convenience, only the parts related to the embodiment of the present application are shown.

[0104] Reference Figure 3 , the data quality assessment system based on the government affairs field includes:

[0105] An acquisition unit, configured to acquire a business scenario portrait and a data set image; wherein the business scenario portrait is used to reflect the business type currently processed by the government data, and the data set image is used to reflect the data presentation form and its association relationship during the government data processing process;

[0106] A segmentation unit, configured to divide the dataset image into business stages based on the business scenario portrait, and determine business segmentation data corresponding to the dataset image; the business segmentation data is used to reflect the data application stage of the government data;

[0107] A quality unit, configured to obtain a corresponding evaluation dimension system based on the business segment data, and determine a quality feature vector corresponding to the business segment data based on the evaluation dimension system; wherein the evaluation dimension system is used to reflect the core evaluation direction of data quality, and the quality feature vector is used to quantify the actual performance of the data under each evaluation dimension;

[0108] A result unit is used to generate a comprehensive data quality evaluation result based on the quality feature vector corresponding to the business segment data.

[0109] It should be noted that the information interaction, execution process, etc. between the above-mentioned systems / units are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment section and will not be repeated here.

[0110] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit module can exist physically alone, or two or more unit modules can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.

[0111] The embodiment of the present application also provides an electronic device, Figure 4 This is a schematic diagram of the structure of an electronic device provided in one embodiment of the present application. Figure 4 As shown, the electronic device 6 of this embodiment includes: at least one processor 60 ( Figure 4 Only one is shown), at least one memory 61 ( Figure 4 Only one is shown in the figure) and a computer program 62 stored in the at least one memory 61 and executable on the at least one processor 60. When the processor 60 executes the computer program 62, the electronic device 6 implements the steps of any of the above-mentioned data quality assessment method embodiments based on the government affairs field, or implements the functions of the units in the above-mentioned system embodiments.

[0112] For example, the computer program 62 may be divided into one or more units, which are stored in the memory 61 and executed by the processor 60 to implement the present application. The one or more units may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program 62 in the government-based data quality assessment 6.

[0113] The electronic device 6 can be a computing device or terminal device such as a desktop computer, a notebook, a PDA, or a cloud server. The electronic device may include, but is not limited to, a processor 60 and a memory 61. Those skilled in the art will understand that Figure 4 It is only an example of the electronic device 6 and does not constitute a limitation on the electronic device 6. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, it may also include input and output devices, network access devices, buses, etc.

[0114] The processor 60 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0115] In some embodiments, the memory 61 may be an internal storage unit of the electronic device 6, such as a hard drive or memory of the electronic device 6. In other embodiments, the memory 61 may also be an external storage device of the electronic device 6, such as a plug-in hard drive, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. equipped on the electronic device 6. Furthermore, the memory 61 may include both an internal storage unit of the electronic device 6 and an external storage device. The memory 61 is used to store an operating system, application programs, a boot loader, data, and other programs, such as the program code of the computer program. The memory 61 may also be used to temporarily store data that has been output or is about to be output.

[0116] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above method embodiments are implemented.

[0117] An embodiment of the present application provides a computer program product. When the computer program product is run on an electronic device, the electronic device implements the steps of any of the above method embodiments.

[0118] If the integrated unit is implemented as a software functional unit and sold or used as a standalone product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the process steps in the above-mentioned method embodiments by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When executed by a processor, the computer program can implement the steps of each of the above-mentioned method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to an electronic device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signals, telecommunication signals, and software distribution media. Examples include USB flash drives, removable hard drives, magnetic disks, or optical disks. In some jurisdictions, based on legislation and patent practice, computer-readable media cannot be electric carrier signals or telecommunication signals.

[0119] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0120] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0121] In the embodiments provided in the present application, it should be understood that the disclosed data quality assessment system / electronic device and method based on the government affairs field can be implemented in other ways. For example, the data quality assessment system / electronic device embodiment based on the government affairs field described above is only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0122] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0123] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A data quality assessment method based on the government affairs field, characterized in that: include: Obtaining a business scenario portrait and a data set image; wherein the business scenario portrait is used to reflect the business type currently processed by the government data, and the data set image is used to reflect the data presentation form and its correlation relationship during the government data processing process; Dividing the data set image into business stages based on the business scenario portrait, and determining the business segmentation data corresponding to the data set image; the business segmentation data is used to reflect the data application stage of the government data; Obtaining a corresponding evaluation dimension system based on the business segment data, and determining a quality feature vector corresponding to the business segment data based on the evaluation dimension system; wherein the evaluation dimension system is used to reflect the core evaluation direction of data quality, and the quality feature vector is used to quantify the actual performance of the data under each evaluation dimension; generating a comprehensive data quality evaluation result based on the quality feature vector corresponding to the business segment data; The step of acquiring a corresponding evaluation dimension system according to the business segment data and determining a quality feature vector corresponding to the business segment data according to the evaluation dimension system includes: When the business segment data is basic business data, a first evaluation dimension system is obtained, and a first quality feature vector corresponding to the basic business data is determined according to the first evaluation dimension system; wherein the first evaluation dimension system is used to reflect the missing status of the basic business data; When the business segment data is cross-domain collaborative data, a second evaluation dimension system is obtained, and a second quality feature vector corresponding to the cross-domain collaborative data is determined according to the second evaluation dimension system; wherein the second evaluation dimension system is used to reflect the correlation of the cross-domain collaborative data; When the business segment data is application service data, obtaining a third evaluation dimension system, and determining a third quality feature vector corresponding to the application service data based on the third evaluation dimension system; wherein the third evaluation dimension system is used to reflect the compliance status of the application service data; Determining a first quality feature vector corresponding to the basic business data according to the first evaluation dimension system includes: Extracting a core field set of the basic business data according to the first evaluation dimension system, and calculating a missing value ratio of the core field set within a preset time window to generate a completeness feature vector; Extracting state transition trajectories of data items in the basic business data according to the first evaluation dimension system, and calculating the average transition time of each data item from an initial state to a target state to generate a timeliness feature vector; Performing feature fusion on the timeliness feature vector and the integrity feature vector to determine a first quality feature vector corresponding to the basic business data; Determining a second quality feature vector corresponding to the cross-domain collaborative data according to the second evaluation dimension system includes: Determining a correlation deviation vector corresponding to the cross-domain collaborative data according to the second evaluation dimension system; Obtaining a response time distribution of the cross-domain collaborative data during cross-domain transmission, and generating a transmission time efficiency vector; Performing feature fusion on the correlation deviation vector and the transmission time efficiency vector to determine a second quality feature vector corresponding to the cross-domain collaborative data; The determining, according to the second evaluation dimension system, the correlation deviation vector corresponding to the cross-domain collaborative data includes: Extracting cross-domain association rules from the cross-domain collaborative data according to the second evaluation dimension system, and constructing a constraint condition set for data association evaluation; Extracting actual association trajectories of cross-domain data items in the cross-domain collaborative data according to the second evaluation dimension system, and generating an association feature sequence; The association feature sequence is compared and analyzed with the constraint condition set, the deviation between the actual association trajectory and the theoretical association rule is calculated, and an association deviation vector is generated.

2. The data quality assessment method based on the government affairs field according to claim 1 is characterized in that: The dividing the data set image into business stages based on the business scenario portrait and determining the business segment data corresponding to the data set image includes: Matching business segmentation labels of the dataset images according to the business scenario portrait; The data set image is divided into business stages based on the business segmentation label, and business segmentation data corresponding to the data set image is determined.

3. The data quality assessment method based on the government affairs field as claimed in claim 1 is characterized in that: Generating a comprehensive data quality evaluation result based on the quality feature vector corresponding to the business segment data includes: A comprehensive data quality evaluation result is generated based on the first quality feature vector, the second quality feature vector, and the third quality feature vector.

4. The data quality assessment method based on the government affairs field according to claim 1 is characterized in that: Determining a third quality feature vector corresponding to the application service data according to the third evaluation dimension system includes: Extracting access permission change logs for sensitive data items in the application service data based on the third evaluation dimension system, generating an access permission change feature sequence, and calculating the change frequency of the access permission change feature sequence within a preset time window to generate a change frequency vector; Performing anomaly detection on the access trace of the application service data to generate an anomaly feature vector; The change frequency vector and the abnormal feature vector are subjected to feature fusion to determine a third quality feature vector corresponding to the application service data.

5. The data quality assessment method based on the government affairs field as claimed in claim 3 is characterized in that: Generating a comprehensive data quality evaluation result based on the first quality feature vector, the second quality feature vector, and the third quality feature vector includes: determining a first quality characteristic value based on the first quality characteristic vector, determining a second quality characteristic value based on the second quality characteristic vector, and determining a third quality characteristic value based on the third quality characteristic vector; The first quality characteristic value, the second quality characteristic value, and the third quality characteristic value are respectively subjected to basic quality filtering, and a comprehensive data quality evaluation result is generated according to the basic quality filtering results.

6. The data quality assessment method based on the government affairs field according to claim 5 is characterized in that: The performing basic quality filtering on the first quality feature value, the second quality feature value, and the third quality feature value respectively, and generating a comprehensive data quality evaluation result according to the basic quality filtering results, includes: When the first quality feature value, the second quality feature value, and the third quality feature value do not all pass the basic quality filter, generating a first comprehensive data quality evaluation result and determining the first comprehensive data quality evaluation result as the comprehensive data quality evaluation result; When the first quality feature value, the second quality feature value, and the third quality feature value all pass the basic quality filter, a second data quality comprehensive evaluation result is generated and determined as the data quality comprehensive evaluation result.

7. The data quality assessment method based on the government affairs field according to claim 6 is characterized in that: When the first quality feature value, the second quality feature value, and the third quality feature value all pass the basic quality filter, generating a second comprehensive data quality evaluation result and determining the second comprehensive data quality evaluation result as the comprehensive data quality evaluation result, includes: When the first quality feature value, the second quality feature value, and the third quality feature value all pass the basic quality filter, determining a quality assessment weight according to the service segment data; The first quality characteristic value, the second quality characteristic value and the third quality characteristic value are weighted according to the quality assessment weight to generate a second comprehensive data quality evaluation result and the second comprehensive data quality evaluation result is determined as the comprehensive data quality evaluation result.

Citation Information

Patent Citations

  • Application data evaluation method and device based on intelligent decision, computer equipment and storage medium

    CN110310079A

  • Government affair big data application maturity evaluation method and system

    CN111832945A