Data quality management method for judicial field
Through data quality management methods, the problems of data silos and lack of standards in the judicial system have been solved, cross-departmental data consistency verification and full-process quality management have been achieved, and the efficiency and security of data interoperability have been improved.
Patent Information
- Application Number
- CN202510737958.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-09-26
AI Technical Summary
The existing judicial system is plagued by data silos, lack of standards, and insufficient quality control, resulting in low efficiency and poor security in cross-departmental data interaction, and a lack of intelligent rule configuration and dynamic quality assessment capabilities.
Adopting data quality management methods, based on multi-dimensional quality detection rules and metadata models, through data warehouse technology and cross-departmental metadata verification, we can achieve full-process data quality detection and evaluation, generate detection reports, and support data quality supervision and traceability.
It improves the efficiency of data intercommunication, reduces the computing pressure of the central system, realizes the consistency verification of cross-departmental data and full-process quality management, and supports the improvement of data quality and business collaboration.
Smart Images

Figure CN120705137A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer software and relates to a data quality management method oriented to the judicial field. Background Art
[0002] The world is steadily deepening its research and application of key data fusion technologies in the judicial field within a digitally connected environment. While significant progress has been made in the informatization and intelligentization of the judicial field, issues such as data silos and disconnected business processes persist in the coordination of legal, prosecutorial, and judicial operations. Breakthroughs are urgently needed in big data governance and services, data interaction management and control, and virtual data space and distributed data fusion.
[0003] Metadata: It is data about data. It refers to structured data extracted from information resources to describe their characteristics and content. It is used to organize, describe, retrieve, save, manage, identify, evaluate, track resource changes, manage large amounts of data, and achieve effective search and management of information resources.
[0004] Data governance is an organization's systematic management of data assets. It aims to ensure data quality, security, and compliance through the development of policies, standards, processes, and technical tools, and to promote efficient data utilization throughout its lifecycle. Essentially, it transforms dispersed data resources into trusted strategic assets through the allocation of authority and process control, supporting enterprise decision-making and business innovation.
[0005] Data quality refers to the degree to which data meets requirements for accuracy, completeness, consistency, and reliability in specific usage scenarios. High-quality data supports effective analysis and decision-making, while low-quality data may lead to erroneous conclusions or business risks.
[0006] Data accuracy: Whether the data accurately and accurately reflects real-world entities or events. Inaccurate data may lead to incorrect analytical results (such as incorrect user profiles and biased financial forecasts).
[0007] Data integrity: Check whether the data contains missing or empty values (NULL values). Incomplete data will reduce the reliability of analysis and may even make it impossible to complete specific tasks (such as user portrait analysis).
[0008] Data consistency: Whether data remains consistent across different systems, time or logical relationships. Inconsistent data can lead to difficulties in cross-departmental collaboration or contradictory statistical results.
[0009] Data timeliness: Is data available when needed and updated frequently enough to meet business needs? Outdated data can reduce the timeliness of decision-making and lead to missed opportunities.
[0010] Data uniqueness: Check whether there are duplicate records. Duplicate data can lead to inflated statistical results (such as user numbers and sales).
[0011] Data reliability: Is the data source credible and the collection process standardized? Unreliable data can lead to systemic risks (such as misjudgment of industrial equipment).
[0012] Data interpretability: Does the data have clear metadata (such as units, definitions, and coding rules)? Difficult-to-understand data increases analysis costs or risks of misuse.
[0013] Data compliance: Whether the data complies with laws and regulations (such as GDPR, CCPA) or industry standards. Compliance issues may result in legal penalties or reputational damage.
[0014] At present, my country's judicial and procuratorial departments have simultaneously carried out the construction of digital network big data platforms. Although preliminary customized development and application have been carried out at the level of cross-departmental collaboration, the above-mentioned judicial and procuratorial big data platforms are all self-contained systems, and are still insufficient in terms of systematicity, intelligence, and accuracy. Therefore, the three-party data interaction between the judicial and procuratorial departments still uses the traditional paper-based material submission method, resulting in low data interaction efficiency, difficulty in mutual recognition of materials and documents, and poor data security protection. It is urgent to make breakthroughs in big data governance and services, data interaction management and control, virtual data space and distributed data integration, break down data silos, and connect the online data interaction business processes of the judicial and procuratorial departments.
[0015] Huayu Software provides the Ruiyuan Judicial Big Data Platform, which offers modules such as case management and document generation. However, it focuses on the internal data governance of a single judicial institution, and cross-departmental collaboration only supports basic file transfer, lacking semantic alignment and dynamic quality assessment capabilities.
[0016] The collaborative case-handling system provided by Taiji Computer is primarily designed for the coordinated flow of public security, procuratorial, and judicial cases, supporting the cross-departmental distribution of electronic files. It primarily focuses on business process integration, and data quality testing only covers format compliance (such as PDF format verification) without in-depth analysis of content consistency and relevance.
[0017] iFLYTEK provides an intelligent verification system for judicial documents. Its business scope mainly focuses on NLP-based document content error correction and supports prompts for the integrity of the evidence chain. However, it focuses on text content rather than data structured governance and cannot solve the problem of inconsistent data standards across systems.
[0018] The existing solutions currently have the following technical defects:
[0019] Insufficient cross-domain adaptability: Rigid data models cannot cope with the differences in legal and prosecutorial business, and the field mapping error rate is generally higher than 20%;
[0020] Lack of quality assessment dimensions: 87% of the tools only cover basic dimensions such as completeness and timeliness, and lack core judicial indicators such as the relevance of the chain of evidence;
[0021] Lack of intelligent evolution mechanism: Rule base updates rely on manual intervention. Statistics from a provincial justice department show that 65% of rules become invalid three years after the system was put into use.
[0022] Security coordination defects: The centralized architecture leads to unclear ownership of data assets. Cross-domain tracing requires coordination of log permissions of multiple departments, and the average processing cycle exceeds 72 hours.
[0023] With the development of judicial informatization, the demand for business collaboration between courts, prosecutors and other departments is growing. However, the existing judicial system has the following problems:
[0024] 1. Data silos: The business systems of courts, procuratorates, and the Ministry of Justice operate independently, data cannot be interoperable, and some processes rely on the circulation of paper documents, which is inefficient.
[0025] 2. Lack of standards: There is a lack of unified metadata standards for cross-departmental data, leading to problems such as semantic ambiguity and inconsistent formats.
[0026] 3. Insufficient quality control: There is a lack of automated tools to perform consistency verification and quality monitoring on the entire data process, making it difficult to detect data anomalies or trace the root cause of the problem.
[0027] In existing technologies, traditional data verification tools are mostly designed for a single system and cannot adapt to the complex needs of cross-departmental and multi-business scenarios in the judicial field. They also lack intelligent rule configuration and dynamic quality assessment capabilities. Summary of the Invention
[0028] In view of the problems existing in the prior art, the purpose of the present invention is to provide a data quality management method for the judicial field. The present invention performs data alignment and consistency mapping from multiple dimensions such as data semantics, content, and measurement. The present invention takes data resources as the evaluation object, is based on data quality evaluation standards and management specifications, and takes quality detection rules as the detection basis. From the aspects of data integrity, data accuracy, data uniqueness, data timeliness, data consistency, data rationality, and data relevance, the present invention performs quality detection on the entire process of data collection, governance, and application, generates detection result data, and analyzes and processes the result data of data quality detection on this basis to form a data quality assessment report. By discovering data quality problems, the data quality can be effectively supervised, controlled, and traced, thereby forming a closed-loop process of data quality problem discovery, monitoring and tracking, and result analysis, providing support for promoting the improvement of data quality.
[0029] The technical solution of the present invention is:
[0030] A data quality management method for the judicial field, comprising the following steps:
[0031] 1) Generate corresponding quality inspection rules based on the set quality inspection business rule template configuration. The quality inspection rules are bound to the quality dimension and used to evaluate the quality dimension score and quality index; the quality inspection rules are bound to the rule object and used to mark the type of the inspected object; the quality inspection business rules are configured according to the defined quality dimension, judicial business requirements and court standards;
[0032] 2) setting a plurality of quality inspection indicators for the data source to be inspected according to the quality inspection business rules and the set quality inspection method and quality inspection template, generating a quality inspection task; and setting a corresponding threshold value for each of the quality inspection indicators;
[0033] 3) Determine the data range from the data source to be tested according to the data screening conditions in the quality inspection task, and then obtain the indicator data of each quality inspection indicator from the data range; calculate the indicator value of the corresponding quality inspection indicator according to the calculation formula of each quality inspection indicator in the quality inspection task and compare it with the corresponding threshold value to generate a data quality assessment report for the data source to be tested.
[0034] Furthermore, the quality dimensions include completeness, accuracy, uniqueness, timeliness, consistency, rationality and relevance.
[0035] Furthermore, the quality inspection business rules define the business rules and descriptions of data quality inspection based on data quality inspection requirements from a business perspective.
[0036] Furthermore, the rule object is a column or an entire table in the data source to be detected.
[0037] A server, characterized in that it includes a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing the above method.
[0038] A computer-readable storage medium stores a computer program thereon, wherein the computer program implements the above method when executed by a processor.
[0039] The advantages of the present invention are as follows:
[0040] 1. To address the problem of data silos (independent systems of the law enforcement and prosecutorial departments, and inability to communicate with each other), the data processing of the present invention is based on the lake-warehouse integrated architecture and data warehouse technical methods, providing tools for data warehouse planning and definition, and data processing tools to help quickly build a data warehouse system and conduct data hierarchical governance, thereby improving data development and governance efficiency; data quality provides data quality management tools to support the detection and evaluation of data quality, facilitate the discovery of data quality problems, and provide decision-making support for improving data quality.
[0041] Based on databases, data sets, data files, data APIs and other channels, we provide a variety of data services such as data sharing services, data recommendation services, data theme services, cross-network data collaboration services, etc. to meet different types of data service needs and support data empowerment business.
[0042] 2. In response to the problem of missing standards (semantic ambiguity caused by inconsistent metadata), the data consistency quality assessment of the present invention takes collaborative data resources as the assessment object, is based on data assessment standards and management specifications, and takes detection rules as the detection basis. From the aspects of data integrity, data accuracy, data uniqueness, data timeliness, data consistency, data rationality, data relevance, etc., the entire process of data collection, governance, and application is quality tested, and the test result data is generated. On this basis, the result data of the data quality test is analyzed and processed to form a data quality assessment report. By discovering data quality problems, effective supervision, control and traceability of data quality are achieved, thereby forming a closed-loop process of data quality problem discovery, monitoring and tracking, and result analysis, providing support for promoting the improvement of data quality.
[0043] Based on a collaborative metadata model for structured and unstructured data, we research verification methods for cross-departmental fusion metadata, supporting automated cross-departmental consistency verification of collaborative metadata with semantic, content, and metric ambiguities, thereby generating case information data that can be shared across departments. By studying metadata models for no fewer than six types of structured and unstructured data, including basic case information, case files, evidentiary materials, parties involved / finances, judicial documents, and audio and video materials, we develop multidimensional metadata fusion rules, collaborative file and audio and video fusion standards, and collaborative case information fusion standards, and determine the mapping relationships between multi-departmental data. We also research data identification methods for cross-departmental collaborative data, enabling the precise positioning of key information within multidimensional data models.
[0044] Different from the traditional "point-to-point" data interaction method, based on the industry standard results of the project, the "total-to-total" data interaction method of the three parties of the Law Enforcement, Procuratorate and Justice Department is realized. Without changing the respective standards of the three departments of the Law Enforcement, Procuratorate and Justice Department, unified data mapping and consistency verification are carried out in the consistency verification system. In order to reduce the data calculation pressure of the central system, the data is first mapped to the industry standard in the node systems of the three departments. The central system is only responsible for the storage of standard conversion rules and data quality verification, which greatly reduces the calculation pressure of the central system.
[0045] 3. In response to the full-dimensional data quality management issues, the present invention can conduct quality verification and management of data throughout the entire data governance process from seven aspects: data integrity, data accuracy, data uniqueness, data timeliness, data consistency, data rationality, and data relevance, through 10 different quality verification methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 This is a flow chart of the data quality management method.
[0047] Figure 2 This is the flow chart of quality inspection tasks. DETAILED DESCRIPTION
[0048] The present invention will be described in further detail below with reference to the accompanying drawings. The examples given are only used to explain the present invention and are not used to limit the scope of the present invention.
[0049] The data quality management method process of the present invention is as follows Figure 1 As shown in Table 3 and Table 4, specific examples are shown. Calculation formulas need to be formulated in combination with the business.
[0050] 1. Data quality inspection rule management implements the management of quality dimensions and quality inspection business rules. Quality inspection business rules support the maintenance of quality inspection rules from a business perspective. Quality inspection rules need to be bound to quality dimensions to calculate quality dimension scores and quality indexes in evaluation plans. Quality inspection rules need to be bound to rule objects to mark the types of quality inspection objects that can be implemented by the quality inspection rules.
[0051] First, configure and manage data quality detection rules. Overall, data quality is divided into seven quality dimensions: completeness, accuracy, uniqueness, timeliness, consistency, rationality, and relevance. Quality dimensions provide a set of vocabulary to define data quality requirements. These dimension definitions can be used to evaluate initial data quality and the effectiveness of continuous improvement. Therefore, they need to be defined in the system in advance for the configuration of subsequent data quality assessment plans and the presentation of assessment report results.
[0052] Next, configure quality inspection business rules based on actual judicial business needs and court standards. Quality inspection business rules are based on data quality testing requirements from a business perspective and define the business rules and descriptions for data quality testing. Therefore, they must be defined in the system in advance and associated with the data quality dimensions associated with the business rules. This allows for the configuration and editing of business rule names for subsequent data quality testing tasks. The quality inspection business rule template in Table 1 can be used to generate the corresponding quality testing rules.
[0053] Table 1 shows the quality inspection rules and templates
[0054]
[0055]
[0056] Different quality inspection methods require different configuration considerations for quality inspection indicators, so they are explained in detail.
[0057] Null value detection
[0058] Null value detection is to detect the number of rows in which the specified column is empty.
[0059] ●Select database, table, column (single)
[0060] ●The problem data is a detail with a specified column empty
[0061] Regular expression detection
[0062] Detects the number of rows that match a regular expression and checks whether the format of the value in the specified column meets the requirements.
[0063] ●Select database, table, column (single).
[0064] ●Quality inspection templates can be selected from built-in templates or customized. When customized is selected, enter a regular expression.
[0065] ●The problem data is detailed in accordance with the regular expression.
[0066] Value Detection
[0067] Detect the number of rows that meet the value detection rules, and check whether the value of the specified column meets the requirements
[0068] ●Select database, table, column (single).
[0069] ●Quality inspection templates can choose built-in templates and custom templates. When choosing custom templates, select rules and add up to two groups.
[0070] You can switch and click [And / Or] to switch the logical relationship between the two sets of conditions. Note that when the input value is a string, you should enter '2022-01-05'.
[0071] ●Problem data is the details that meet the value detection rules.
[0072] Field length detection
[0073] Detect the number of rows that meet the field length rules and whether the length of the specified column meets expectations.
[0074] ●Select database, table, column (single).
[0075] ●Quality inspection templates can choose built-in templates and custom templates. When choosing custom templates, select rules and add up to two groups.
[0076] You can switch and click [And / Or] to switch the logical relationship between two sets of conditions.
[0077] ●The problem data is a detail that complies with the field length rules.
[0078] Enumeration value detection
[0079] Detects the number of rows that match the enumeration value range.
[0080] ●Select database, table, column (single).
[0081] ●Quality inspection templates can be selected from built-in templates or customized. When choosing customized, fill in the enumeration value, and separate multiple values with English commas.
[0082] ●The problem data is a detail that conforms to the enumeration value range.
[0083] Cross-column null value detection
[0084] Detects the number of rows where both specified columns are empty.
[0085] ●Select database, data table, and columns (2).
[0086] ●The problem data is the details that both columns are empty.
[0087] Custom SQL detection
[0088] The detection results are obtained based on the input SQL.
[0089] ●View databases, data tables, and columns.
[0090] Θ is calculated based on the input SQL, and the full amount / incremental amount cannot be calculated.
[0091] ΘThe problem data is the details of the input calculation SQL, and the field cannot be specified.
[0092] Uniqueness Detection
[0093] Detect the number of rows with unique fields and whether there are duplicates in the specified column.
[0094] ΘSelect database, table, column (single).
[0095] ●The problem data is the details of the duplicate specified columns.
[0096] Table row count detection
[0097] Check the number of table rows.
[0098] ●Select the database and data table.
[0099] No problem data.
[0100] 2. Data quality detection is done by configuring detection tasks, such as Figure 2 As shown, it includes manual execution and scheduling of inspection tasks, generates inspection results, and supports result viewing and detailed viewing of problem data. When configuring quality inspection tasks, quality inspection indicators are generated according to quality inspection business rules, quality inspection methods, quality inspection templates, and data sources to be inspected. At the same time, thresholds are set for quality inspection indicators for display and statistics of inspection results to determine whether the quality inspection indicator results are qualified.
[0101] A quality inspection task contains multiple quality inspection indicators. Quality inspection indicators are constructed based on business rules and quality inspection methods. The data information to be inspected is selected, and the calculation logic and other information are configured. The quality inspection indicators are then calculated through task execution.
[0102] Quality inspection indicator name: unique.
[0103] Business rules: Select existing business rules as the basis to build quality inspection indicators.
[0104] Quality inspection method: Select the appropriate quality inspection method based on business rules. Currently, 9 quality inspection methods are supported.
[0105] Data filtering conditions: Enter a where statement (excluding where) to define the data range. It is not required.
[0106] Full / Incremental: If you select Full, the full data will be tested each time. If you select Incremental, the full data will be tested for the first time, and the incremental data will be tested subsequently.
[0107] Maximum and minimum calculation fields: Calculate the maximum and minimum values of a specified column. We recommend selecting numeric and time type data columns to identify the range of the data to be tested. For example, when selecting a time field, you can indicate that the range of the data to be tested is from 2023-01-01 to 2023-10-01 (this is an example).
[0108] Quality inspection template: used to guide users to directly configure simple tasks; some quality inspection methods can have built-in common templates, which users can use directly, or choose custom mode. If there are common templates that are not in the system, you can contact the data center team to add them.
[0109] Detection SQL: Based on the selected content, a detection SQL will be generated to determine whether the calculation logic meets expectations.
[0110] Calculation formula: It is calculated based on the SQL test and is the final calculation result of the quality inspection indicator. It provides four types: actual value, actual value / total number of rows, 1-actual value / total number of rows, and total number of rows-actual value.
[0111] Actual value: detects the calculation result of SQL;
[0112] Total number of rows: The total number of rows of selected data (including data filtering conditions, the same below).
[0113] Quality Inspection Indicator Threshold: This function determines whether the final calculated results of the quality inspection indicators meet expectations and displays them in the quality inspection indicator results list. The calculated results and thresholds are used to determine if the equation is true, indicating a pass; otherwise, a failure. For example, if the calculated null value rate for the party's ID card number is 20%, and the threshold is set to <10%, then 20% < 10% is not true, so the indicator result is a failure.
[0114] Whether to save problem data: For the problem data of quality inspection, you can choose whether to save it. Select yes and select the data columns to be saved. By default, 50 problem data are saved.
[0115] The specific configuration of quality inspection tasks is shown in Table 2 below:
[0116] Table 2 shows the specific configuration of quality inspection tasks.
[0117]
[0118]
[0119]
[0120] 3. Data quality assessment generates an assessment report through assessment plan configuration. It also supports visual query and comparative analysis of assessment reports, as well as downloading of report results and viewing of detailed data. The assessment plan configuration is based on the quality inspection indicators in the quality inspection task, and the quality inspection indicators under the same quality dimension are configured with formulas and weights to obtain the quality dimension score. Based on the quality dimension score, the quality index formula and weight are configured to obtain the quality index.
[0121] The quality assessment plan is based on the quality inspection and testing tasks and the results of quality inspection indicators. First, the calculation formula and weight configuration are performed for the quality inspection indicators under the quality dimension to obtain the score calculation formula for each quality dimension. Secondly, the weight is configured for each quality dimension to obtain the calculation formula for the quality assessment index, form an assessment plan, and generate the index and quality dimension score results.
[0122] In [Quality Assessment - Quality Assessment Plan], click New and save.
[0123] Select the quality inspection task and quality inspection indicators. Multiple selections are allowed.
[0124] Quality dimension score formula and weight configuration: Each quality dimension may have multiple quality inspection indicators. Calculation weights need to be configured to ensure that the sum of weights under the same quality dimension is equal to 1. Since the quality inspection indicator results may be the number of rows or percentages, it is not possible to uniformly calculate the quality dimension score. In this case, you can enter four arithmetic operations to perform a secondary conversion calculation on the quality inspection indicators.
[0125] Index weight configuration: Based on the quality dimension score results calculated in the above configuration, the index is calculated by configuring the quality dimension score weights. The sum of the weights must be equal to 1.
[0126] 4. Data quality monitoring is the statistical analysis and visual display of quality inspection business rules, quality inspection indicators, quality inspection results, and quality assessment results.
[0127] This technology enables list display and query of maintained information such as quality dimension codes, quality dimension names, and descriptions for data quality detection quality dimensions. The saved quality dimensions are displayed in a list, sorted in descending order by creation time, with support for paging and fuzzy queries based on codes and quality dimensions. This technology enables list display and query of maintained information such as codes, business rule names, descriptions, and quality dimensions for data quality detection business rules. The saved business rules are displayed in a list, sorted in descending order by creation time, with support for paging and fuzzy queries based on codes, business rule names, quality dimensions, and rule objects.
[0128] Based on the quality inspection task, a quality inspection report can be generated, realizing the generation and visualization of the report of the evaluation plan. Click View Report to open the View Report page, select the inspection result (select according to the time when the inspection result was generated), click View Report to display the visualization results, including index, quality dimension score, quality inspection indicator results, and support download. At the same time, the data quality assessment process (business rule statistics, quality inspection indicator statistics, quality inspection result statistics) can be analyzed and visualized, statistically analyzing the business rules, quality inspection indicators, quality inspection results, and quality assessment report, and visually displaying them.
[0129] Table 3 is an example of quality inspection indicators
[0130]
[0131]
[0132] Table 4 is the second example of quality inspection indicators
[0133]
[0134]
[0135] While specific embodiments of the present invention have been disclosed for illustrative purposes, intended to facilitate understanding and implementation of the present invention, those skilled in the art will appreciate that various substitutions, variations, and modifications are possible without departing from the spirit and scope of the present invention and the appended claims. Therefore, the present invention should not be limited to the disclosure of the preferred embodiments, and the scope of protection claimed in the present invention shall be determined by the scope of the claims.
Claims
1. A data quality management method for the judicial field, comprising the following steps: 1) Generate corresponding quality inspection rules based on the set quality inspection business rule template configuration. The quality inspection rules are bound to the quality dimension and used to evaluate the quality dimension score and quality index; the quality inspection rules are bound to the rule object and used to mark the type of the inspected object; the quality inspection business rules are configured according to the defined quality dimension, judicial business requirements and court standards; 2) setting multiple quality inspection indicators for the data source to be inspected according to the quality inspection business rules and the set quality inspection method and quality inspection template, and generating a quality inspection task; and setting a corresponding threshold value for each of the quality inspection indicators; 3) Determine the data range from the data source to be tested according to the data screening conditions in the quality inspection task, and then obtain the indicator data of each quality inspection indicator from the data range; calculate the indicator value of the corresponding quality inspection indicator according to the calculation formula of each quality inspection indicator in the quality inspection task and compare it with the corresponding threshold value to generate a data quality assessment report for the data source to be tested.
2. The method according to claim 1, characterized in that The quality dimensions include completeness, accuracy, uniqueness, timeliness, consistency, rationality and relevance.
3. The method according to claim 1, characterized in that The quality inspection business rules define the business rules and descriptions of data quality inspection based on the data quality inspection requirements from a business perspective.
4. The method according to claim 1, wherein The rule object is a column or an entire table in the data source to be detected.
5. A server, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program comprises instructions for executing the method according to any one of claims 1 to 4.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.