Government affair data resource directory intelligent management method, system, equipment and medium
By constructing a dedicated large-scale model and real-time monitoring mechanism for government data, the problems of low merging efficiency and poor cleaning quality in the management of government data resource catalogs have been solved, achieving automated catalog management and efficient data utilization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-11
- Publication Date
- 2026-04-03
AI Technical Summary
The existing management of government data resource catalogs relies on manual operation, which results in low catalog merging efficiency, poor data cleaning quality, and high maintenance costs. In particular, it is difficult to achieve efficient and accurate catalog merging and data cleaning in cross-departmental heterogeneous data scenarios.
We construct a dedicated large model for government data, optimize the general large model through the training set, and realize field standardization, data item matching, redundant data identification and deletion, error data correction and missing field completion. Combined with multi-level composite indexes and real-time monitoring mechanisms, we automatically process the integration, cleaning and maintenance of government data resource catalogs.
It significantly improves the management efficiency and quality of the government data resource catalog, ensures the timeliness and consistency of data, provides efficient data retrieval and utilization capabilities, and reduces manual maintenance costs.
Smart Images

Figure CN121786032A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of data processing technology, and more specifically relates to a method, system, device and medium for intelligent management of government data resource catalogs. Background Technology
[0002] With the continuous advancement of digital government construction, government departments at all levels have accumulated massive amounts of government data. This data is scattered across different business systems and databases, forming multiple independent government data resource catalogs. Currently, the management of these government data resource catalogs mainly relies on manual operation, which presents the following problems: 1. Low efficiency of directory merging: The formats of government data resource directories of different departments are not uniform and the field definitions are different. Manual identification and matching of data items in the directories are required one by one before merging. For directories containing tens of thousands of data items, it often takes several weeks or even months, which seriously affects the progress of government data integration.
[0003] 2. Poor data cleaning quality: During manual cleaning, subjective factors can easily influence the process, resulting in incomplete identification of redundant data (such as duplicate organization names and identical data indicators) and erroneous data (such as incorrectly formatted dates and logically contradictory values). This leads to a large number of problematic data remaining in the cleaned catalog, affecting the subsequent use of government data.
[0004] 3. High maintenance costs: Government data is constantly being updated. New data is constantly being added and old data needs to be adjusted. Manual updates, merging, and cleaning of the catalog are required continuously. In the long run, a lot of manpower and time costs are required, and it is difficult to guarantee the timeliness and consistency of management.
[0005] In existing technologies, there are also some government data catalog management solutions based on traditional algorithms, such as rule matching algorithms and similarity calculation algorithms. However, traditional algorithms rely on manually preset rules, which have poor adaptability to complex government data scenarios, such as heterogeneous data across departments and ambiguous field descriptions. They cannot achieve efficient and accurate catalog merging and data cleaning, and are difficult to meet the current intelligent and efficient needs of government data management. Summary of the Invention
[0006] To address the above issues, the present invention aims to provide an intelligent management method, system, device, and medium for government data resource catalogs. By constructing a large-scale model that enhances knowledge in the government domain, the invention automates the entire process of government data resource catalog management, from intelligent integration and quality cleaning to dynamic maintenance. This significantly improves data governance efficiency and quality, and provides reliable technical support for the efficient management and in-depth utilization of government data.
[0007] To achieve the above objectives, the present invention employs the following technical solution: In a first aspect, embodiments of this application provide a method for intelligent management of government data resource catalogs, including: A training set is constructed based on samples from the government data resource catalog, a government data field dictionary, and historical manual records. The training set is then used to train a general large model to generate a large model specifically for government data. Collect government data resource catalogs from multiple sources, and use a dedicated big data model for government data to standardize fields and match data items in the government data resource catalogs to generate a merged catalog; The merged catalog is input into the dedicated big data model for government data. Redundant data is identified and deleted, erroneous data is identified and corrected, and missing fields are filled in through the dedicated big data model for government data, thereby generating a standardized government data resource catalog. The standardized government data resource catalog is stored and indexed; the data update operations of each government data source are continuously monitored; when a data change is detected, the updated government data resource catalog is automatically collected, and field standardization, data item matching, redundant data identification and deletion, erroneous data identification and correction, and missing field completion are triggered to update the stored standardized government data resource catalog.
[0008] In an optional implementation, the step of constructing a training set based on government data resource catalog samples, government data field dictionaries, and historical manual records, and using the training set to train a general large model to generate a government data-specific large model, includes: We collected samples of government data resource catalogs, government data field dictionaries, and historical manual merging and cleaning records from government departments at all levels. We then performed format unification and noise reduction on the collected data to obtain the training dataset. Using a general large model as the base model, the training dataset is input into the base model, and the base model is trained using few-shot learning and incremental training methods to optimize the base model's ability to understand semantics in the field of government data, field matching ability, and data error identification ability. A validation dataset is constructed, comprising a preprocessed catalog of government data resources according to predetermined merging and cleaning rules. The validation dataset is input into the base model, and the model output is compared with the preprocessed catalog of government data resources to calculate the catalog matching accuracy and the data cleaning accuracy. If either the catalog matching accuracy or the data cleaning accuracy fails to reach its corresponding preset threshold, a supplementary training dataset is added, and the model parameters are adjusted for retraining until both the catalog matching accuracy and the data cleaning accuracy reach their respective preset thresholds. Finally, a large-scale model specifically designed for government data is output.
[0009] In an optional implementation, the collection of multi-source government data resource catalogs involves standardizing fields and matching data items using a dedicated government data model to generate a merged catalog, including: The catalog of government data resources to be merged is collected from the business systems or databases of various government departments through API calls or file uploads. The catalog of government data resources to be merged is converted into standard JSON format, and the field names and field types are standardized according to the preset government data field standards using the dedicated large model of government data. The semantic and attribute similarity between data items in different government data resource catalogs are calculated using the government data-specific big data model, and matching is performed using knowledge of the government data domain. For government data resource catalogs with hierarchical relationships, the hierarchical structure is automatically identified and the data items in the catalogs are matched according to the rules of supplementing the upper-level catalogs with the lower-level catalogs. For directory data items that match successfully, a strategy of retrieving all fields and removing duplicate values is used to merge the directory data items; for directory data items that do not match successfully, they are treated as new data items. After the directory data items are merged, a merge log is generated that records the source of the directory data items and the merging method, and the merged directory is output.
[0010] In an optional implementation, the merged catalog is input into a dedicated government data model. The model then performs redundant data identification and deletion, erroneous data identification and correction, and missing field completion to generate a standardized government data resource catalog, including: The merged catalog is input into the government data exclusive big model. The government data exclusive big model identifies completely duplicate and partially duplicate catalog data items from the data content dimension and redundant fields from the data attribute dimension. Based on the preset rules of retaining the latest data and core fields, redundant fields are automatically deleted to generate the first process catalog. The first process catalog is used to identify errors, including formatting errors, logical errors, and numerical errors, through a dedicated big data model for government data. For formatting errors, automatic format correction is performed. For logical and numerical errors, correction is performed by querying the government data standard database. If correction is not possible, an error message is generated. After the correction is completed, the second process catalog is generated. The model, dedicated to government data, identifies missing fields in the second-stage catalog. Based on existing data and knowledge of the government data domain, it automatically completes the missing field content, marks all completed missing field content, and outputs a standardized government data resource catalog and cleaning log.
[0011] In an optional implementation, storing the standardized government data resource catalog and establishing an index includes: The standardized government data resource catalog is persistently stored in a distributed NoSQL database in document form, and a globally unique catalog identifier is assigned to each catalog data item in the standardized government data resource catalog. Parse the directory data items, extract the names and data types of the core business fields, and build a key-value pair index as the field index; Based on the source department metadata recorded in the catalog data items, construct an index that maps department codes to corresponding catalog identifiers, which serves as the government department index; Based on the preset government data classification standards, the business themes and classification tags of the catalog data items are extracted, and an index mapping the business classification to the catalog identifier is constructed as a data type index. A multi-level composite index is constructed based on field indexes, government department indexes, and data type indexes; The directory identifiers are associated with and stored in a multi-level composite index to form a benchmark directory library that can be retrieved by field, government department, or data type.
[0012] In an optional implementation, the continuous monitoring of data update operations from various government data sources, upon detecting data changes, automatically collects the updated government data resource catalog and triggers field standardization, data item matching, redundant data identification and deletion, erroneous data identification and correction, and missing field completion processing to update the stored standardized government data resource catalog, including: By monitoring the data update events of the business systems or databases of various government departments through the monitoring agent program, the updated government data resource catalog can be captured. The updated government data resource catalog is input into the government data-specific big data model, and the following steps are executed in sequence: field standardization, data item matching, redundant data identification and deletion, erroneous data identification and correction, and missing field completion, to generate the target catalog. The directory is written to the distributed NoSQL database, and the corresponding multi-level composite index is updated synchronously to complete the incremental update of the baseline directory database. The integrity verification and consistency check of the government data resource catalog in the benchmark catalog are periodically performed through a scheduled task service. By calling the data service interface through the front-end application, the government data resource catalog and associated merged and cleaned logs are obtained. The government data resource catalog is then organized and rendered in a tree topology structure, and a corresponding table view is generated.
[0013] In one optional implementation, the general large model adopts any one of the following: GPT series model, ERNIE series model, ChatGLM series model, LLaMA series model, and Tongyi Qianwen series model.
[0014] Secondly, embodiments of this application also provide an intelligent management system for government data resource catalogs, including: The model building module is used to build a training set based on samples from the government data resource catalog, a government data field dictionary, and historical manual records. The training set is then used to train a general large model to generate a large model specific to government data. The catalog merging module is used to collect government data resource catalogs from multiple sources. It performs field standardization and data item matching on the government data resource catalogs using a dedicated large model for government data, and generates a merged catalog. The data cleaning module is used to input the merged catalog into the government data-specific big data model. The government data-specific big data model is used to identify and delete redundant data, identify and correct erroneous data, and complete missing fields to generate a standardized government data resource catalog. The directory management module is used to store the standardized government data resource directory and establish an index; continuously monitor the data update operations of each government data source; when a data change is detected, automatically collect the updated government data resource directory and trigger field standardization, data item matching, redundant data identification and deletion, erroneous data identification and correction, and missing field completion processing to update the stored standardized government data resource directory.
[0015] Thirdly, embodiments of this application also provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the intelligent management method for the government data resource catalog as described in any of the above.
[0016] Fourthly, embodiments of this application also provide a storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the intelligent management method for the government data resource catalog as described in any of the above claims.
[0017] As can be seen from the above technical solutions, the present invention has the following advantages: The intelligent management method for government data resource catalogs provided in this application significantly improves the accuracy and efficiency of standardization of multi-source heterogeneous catalog fields and matching of data items through semantic understanding and domain knowledge of large models; it realizes automatic identification and deletion of redundant data, intelligent correction of erroneous data, and automatic completion of missing fields, greatly improving data quality; through real-time monitoring and automated update mechanisms, it ensures the timeliness and consistency of catalog data; and finally, through a visualization platform and index optimization, it provides reliable technical support for the unified management, efficient retrieval, and effective utilization of government data resources, greatly improving the automation level and decision support capabilities of government data governance.
[0018] This application significantly improves the intelligent processing level of government data by constructing a large-scale model specifically for government data. Based on samples from the government data resource catalog, a government data field dictionary, and historical manual records, the model undergoes targeted training, enabling it to possess deep semantic understanding capabilities in the government domain. This allows it to accurately parse multi-source, heterogeneous government data and achieve automatic standardization of field names and data types, fundamentally solving the understanding barriers and inefficiencies faced by traditional methods when processing government data of different formats and standards.
[0019] This application achieves high-quality fusion of cross-departmental government data resource catalogs through a large-scale intelligent matching algorithm. The method is based on a comprehensive calculation of semantic and attribute similarity, combined with knowledge of the government domain for accurate matching, and employs an optimized merging strategy of "taking all fields and removing duplicate values." This ensures the integrity of data items while effectively eliminating data redundancy, generating a unified and complete merged catalog, laying a solid foundation for the in-depth utilization of government data.
[0020] This application establishes a comprehensive automated data cleaning system. It identifies redundant data from both the content and attributes dimensions using a large-scale model, and performs intelligent cleaning based on preset rules. Simultaneously, it performs multi-type error detection and correction for formatting errors, logical errors, and numerical errors, and verifies these errors using a government data standard database. Furthermore, it achieves intelligent completion of missing fields, forming a complete data quality improvement loop, significantly improving the accuracy, completeness, and reliability of government data.
[0021] This application constructs an event-triggered automated update mechanism. A monitoring agent program monitors the data updates of various government departments in real time, automatically capturing changed data and triggering a complete processing flow, including field standardization, data item matching, and data cleaning. This enables real-time synchronous updates of the government data resource catalog, ensuring data timeliness and consistency while significantly reducing the manual costs and workload of data maintenance.
[0022] This application significantly enhances the usability of government data by establishing a multi-level composite index and a visualization platform. Based on field indexes, government department indexes, and data type indexes, it constructs an efficient retrieval system. Combined with a dual display method of tree-structured topology and tabular views, it provides comprehensive technical support for the rapid location, multi-dimensional analysis, and intuitive display of government data resources, effectively promoting the sharing and utilization of government data and enhancing decision support capabilities.
[0023] The large-scale model for government data in this application supports incremental training. As the types of government data increase and business needs change, only new training data needs to be added to optimize the model performance without large-scale adjustments to the overall methodology architecture, which can meet the long-term needs of government data catalog management. Attached Figure Description
[0024] To more clearly illustrate the technical solution of the present invention, the accompanying drawings used in the description will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0025] Figure 1 A flowchart illustrating the intelligent management method for the government data resource catalog provided in this application.
[0026] Figure 2 A schematic diagram of the structure of the intelligent management system for the government data resource catalog provided in this application.
[0027] Figure 3 A schematic diagram of the structure of the electronic device provided in this application. Detailed Implementation
[0028] The various embodiments of this disclosure will be described more fully in the following detailed description of the specific steps of the intelligent management method for government data resource catalogs. This disclosure may have various embodiments, and adjustments and changes may be made therein. However, it should be understood that there is no intention to limit the various embodiments of this disclosure to the specific embodiments disclosed herein, but rather this disclosure should be understood to cover all adjustments, equivalents, and / or alternatives falling within the spirit and scope of the various embodiments of this disclosure.
[0029] In the following, the terms “comprising” or “may include”, which may be used in various embodiments of this disclosure, indicate the presence of the disclosed functions, operations, or elements, and do not limit the addition of one or more functions, operations, or elements. Furthermore, as used in various embodiments of this disclosure, the terms “comprising,” “having,” and their cognates are intended only to indicate a particular feature, number, step, operation, element, component, or combination of the foregoing, and should not be construed as primarily excluding the presence of one or more other features, numbers, steps, operations, elements, components, or combinations of the foregoing, or the possibility of adding one or more combinations of the foregoing.
[0030] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0031] Please see Figure 1 The diagram shown is a flowchart of a method for intelligent management of government data resource catalogs in a specific embodiment. The method includes: S1: Construct a training set based on government data resource catalog samples, government data field dictionary and historical manual records, and use the training set to train a general large model to generate a government data-specific large model.
[0032] In a specific implementation, data preprocessing is performed first. Sample catalogs of government data resources from various levels of government departments are collected, including catalog files in various formats such as Excel, JSON, and XML. Simultaneously, government data field dictionaries are collected, specifically including important dictionary resources such as institution name dictionaries, data indicator dictionaries, and administrative division dictionaries, as well as historical records of manual merging and cleaning. This data undergoes format standardization processing, converting catalog files of different formats into standard JSON format and performing noise reduction, including removing irrelevant characters and cleaning blank fields. For example, after converting an Excel-format catalog collected from a department into JSON format, the remarks column and null value fields are removed, retaining the core data content to form a standardized training dataset.
[0033] Next, model training was conducted. GPT-4 was selected as the base model, and the preprocessed training dataset was input into the model. A few-shot learning technique was employed, initially using a small amount of high-quality labeled data for guided training, enabling the model to quickly grasp the basic characteristics of government data. Then, incremental training was used, inputting more training data in batches to gradually optimize the model's understanding of the semantics of government data. During training, the focus was on improving the model's field matching ability; for example, training the model to recognize that different expressions such as agency codes, unit numbers, and organization codes actually refer to the same field. Simultaneously, the model's data error detection ability was enhanced, enabling it to identify common problems such as incorrect date formats and out-of-bounds numerical values.
[0034] Finally, model validation and optimization were performed. A validation dataset containing 5000 labeled samples was constructed. These samples were preprocessed according to the government data metadata catalog standard, and the correct merging and cleaning results were labeled. The validation dataset was input into the trained model, and the model's catalog matching accuracy and data cleaning accuracy were calculated. Matching accuracy thresholds of 95% and cleaning accuracy thresholds of 98% were set. If either indicator failed to meet the threshold, erroneous samples were analyzed, corresponding training data was supplemented, training parameters such as the learning rate were adjusted, and retraining was performed. After three rounds of iterative optimization, the model performance reached the preset thresholds, and the final government data-specific large-scale model was output. Throughout the model construction process, any one of the following models can be used: GPT series, ERNIE series, ChatGLM series, LLaMA series, and Tongyi Qianwen series. The most suitable base model is selected based on the specific application scenario.
[0035] S2: Collect government data resource catalogs from multiple sources, and perform field standardization and data item matching on the government data resource catalogs through a dedicated government data big model to generate a merged catalog.
[0036] In the specific implementation, the first step is to collect the catalog. This is done through RESTful API calls and web file uploads, collecting the government data resource catalogs to be merged from the business systems and databases of various government departments. During the collection process, detailed metadata such as the source department, creation time, and data format of each catalog is recorded. For example, by calling the open interface of a city's data exchange platform, the latest data resource catalog is obtained, along with a record of the catalog being created by a certain bureau at a specific time, and its original format being XML.
[0037] Next, field standardization is performed. The collected directories in different formats are uniformly converted to standard JSON format and input into the dedicated government data model. The model performs semantic recognition on each field in the directory and standardizes it according to the government data metadata directory standard. For example, if the model identifies that the unit number field in a certain directory has a 92% semantic similarity to the organization code field in the standard, it is standardized to an organization code; simultaneously, the date fields in each directory are uniformly converted to the YYYY-MM-DD standard format.
[0038] At this point, data item matching analysis is conducted. Specifically, a dedicated large-scale model for government data is used to match data items in different standardized directories. Semantic and attribute similarities between data items are calculated using knowledge from the government data domain. For example, the model identifies that the number of urban residents enrolled in basic medical insurance in directory A and the number of urban and rural residents enrolled in medical insurance in directory B have a semantic similarity of 88%, and their data types, units of measurement, and other attributes are consistent, thus determining them as matching data items. For directories with hierarchical relationships, such as provincial directories containing municipal directories, the model automatically identifies the hierarchical structure and matches data items from lower-level directories according to the rule of supplementing data items from higher-level directories.
[0039] Finally, the directory merging operation is performed. For data items that match successfully, a strategy of retrieving all fields and removing duplicate values is used for merging. For example, if data item A contains data name and update time fields, and data item B contains data name and data source fields, the merged data item will contain data name, update time, and data source fields, while removing duplicate data name values. Data items that do not match successfully are considered new data items and are directly added to the merged directory. After the merge is complete, a detailed merge log is generated, recording information such as the source directory, merge time, and merge method for each data item.
[0040] S3: Input the merged catalog into the dedicated government data model. The dedicated government data model is used to identify and delete redundant data, identify and correct erroneous data, and complete missing fields to generate a standardized government data resource catalog.
[0041] In the specific implementation, redundant data identification and deletion are performed first. The merged catalog is input into a dedicated large-scale model for government data. The model identifies redundant data from two dimensions: data content and data attributes. In terms of data content, it identifies completely duplicated data items with all field values identical, and partially duplicated data items with identical core field values, differing only in non-core fields. For example, if two population statistics records are identified that are identical in all fields except for their update time (difference of one day), and the update time difference is within a preset 7-day range, they are determined to be partially duplicated data. In terms of data attributes, redundant fields are identified, such as descriptive fields with multiple identical values. After identification, redundant data is automatically deleted according to the rules of retaining the latest data and deleting duplicate data, and retaining core fields and deleting redundant fields, with deletion logs recorded.
[0042] Next, error data identification and correction are performed. The dedicated large-scale model for government data identifies errors in the deredundant catalog data, including format errors, logical errors, and numerical errors. For format errors, such as a date format of 2025 / 13 / 01, the model automatically corrects it to 2026 / 01 / 01 according to government data format standards; for mobile phone numbers with fewer than 11 digits, they are marked and not processed temporarily. For logical errors, such as data updated daily but the last update time being more than 30 days from the current time, the model queries the business database to confirm the actual situation. For numerical errors, such as negative population figures, the model performs a reasonableness check based on administrative division statistics and corrects them.
[0043] Finally, data completion is performed. The model identifies missing fields in the catalog and intelligently completes them based on existing data information and knowledge of the government data domain. For example, if the data source field for a data item is found to be empty, the model determines the source department based on the organization code of the data item and automatically completes the data source field. For cases where the data description field is not filled in, the model generates concise and accurate data description content based on the data name and relevant business scenarios. All completed data is specially marked for easy manual verification later. After completion, the final standardized government data resource catalog and detailed cleaning logs are output.
[0044] S4: Store the standardized government data resource catalog and establish an index; continuously monitor the data update operations of each government data source; when a data change is detected, automatically collect the updated government data resource catalog and trigger field standardization, data item matching, redundant data identification and deletion, erroneous data identification and correction, and missing field completion processing to update the stored standardized government data resource catalog.
[0045] In a specific implementation, on the one hand, directory storage management is performed. The cleaned government data resource directory is stored in an HBase distributed database, using columnar storage to optimize query performance. A globally unique directory identifier is assigned to each directory data item, using a combination of department code, timestamp, and sequence number. Simultaneously, a multi-level composite index is established, including: parsing the directory data item to extract the names and data types of core business fields, constructing a key-value pair index as a field index; constructing a mapping index between department codes and corresponding directory identifiers based on the source department metadata recorded in the directory data item, serving as a government department index; and extracting the business theme and category tags of the directory data item according to a preset government data classification standard, constructing a mapping index between business categories and directory identifiers, serving as a data type index. Through these indexes, millisecond-level multi-dimensional directory retrieval is achieved.
[0046] On the other hand, a real-time monitoring and update mechanism is established. A government data catalog monitoring system based on a microservice architecture is built. Through monitoring agents deployed in various government departments, the system monitors data updates in business systems and databases in real time. When a data addition, modification, or deletion event is detected, the system automatically captures the updated catalog data, generates a structured message, and pushes it to a Kafka message queue. After the message consumer retrieves the message from the queue, it automatically triggers a complete catalog processing flow, including field standardization, data item matching, catalog merging, and data cleaning of relevant government data resource catalogs using the methods in steps S2 and S3, through a dedicated large-scale model for government data. Simultaneously, the system performs a comprehensive monthly review of the stored catalog to ensure the continuous accuracy of the data.
[0047] Furthermore, this method also enables visualization and interaction. Specifically, a government data resource catalog visualization platform is developed based on the Vue.js and SpringBoot frameworks, providing data services to the front end via a RESTful API. The platform displays the government data resource catalog in a multi-dimensional manner using tree structures and tables, supporting intelligent queries and filtering by department, data type, update time, etc. Staff can view detailed catalog merging and data cleaning logs through the platform, confirming or adjusting the data automatically corrected by the model. The platform also provides a catalog file upload function, supporting manual triggering of the catalog merging and cleaning process, forming a complete human-machine collaborative working mechanism.
[0048] In this embodiment, a large-scale knowledge-enhanced model in the government domain is constructed to achieve intelligent processing of the entire process of data integration, quality cleaning, and dynamic maintenance. Based on few-shot learning and incremental training techniques, the model's semantic understanding and error recognition capabilities for government data are optimized. A multi-source heterogeneous data standardization and intelligent matching mechanism is adopted to ensure the quality of data fusion. A real-time monitoring and automated update system is established to ensure data timeliness. Combined with multi-level composite indexes and a visualization platform, the efficiency of data services is improved. Ultimately, the beneficial effects of significantly improving the efficiency of government data governance, effectively ensuring data quality, and realizing unified management and efficient utilization of data resources are achieved.
[0049] like Figure 2 As shown, the following are embodiments of the intelligent management system for government data resource catalog provided in this disclosure. This system and the intelligent management method for government data resource catalog in the above embodiments belong to the same inventive concept. For details not described in detail in the embodiments of the intelligent management system for government data resource catalog, please refer to the embodiments of the intelligent management method for government data resource catalog described above.
[0050] A smart management system for government data resource catalogs, comprising: The model building module is used to build a training set based on samples from the government data resource catalog, a government data field dictionary, and historical manual records. The training set is then used to train a general large model to generate a large model specifically for government data.
[0051] The directory merging module is used to collect government data resource directories from multiple sources. It performs field standardization and data item matching on the government data resource directories using a dedicated government data big model to generate a merged directory.
[0052] The data cleaning module is used to input the merged catalog into the dedicated government data model. The dedicated government data model is used to identify and delete redundant data, identify and correct erroneous data, and complete missing fields to generate a standardized government data resource catalog.
[0053] The directory management module is used to store the standardized government data resource directory and establish an index; continuously monitor the data update operations of each government data source; when a data change is detected, automatically collect the updated government data resource directory and trigger field standardization, data item matching, redundant data identification and deletion, erroneous data identification and correction, and missing field completion processing to update the stored standardized government data resource directory.
[0054] The intelligent management system for the government data resource catalog provided in this embodiment automates the entire process of government data resource cataloging, from intelligent integration and quality cleaning to dynamic maintenance, by constructing a large-scale model that enhances knowledge in the government domain. This effectively solves the standardization problem of multi-source heterogeneous data and significantly improves the accuracy and efficiency of data processing. By establishing an intelligent matching and quality verification mechanism, it ensures the integrity and consistency of data fusion. With the help of an automated monitoring and update system, it ensures the real-time performance and accuracy of catalog data, ultimately providing reliable technical support for the unified management, efficient sharing, and in-depth utilization of government data resources.
[0055] Figure 3 A schematic diagram of the hardware structure of an electronic device for implementing various embodiments of the present invention.
[0056] The intelligent management method for government data resource catalogs provided in this application can be applied to electronic devices. Those skilled in the art will understand that the electronic device structure involved in the embodiments of this invention does not constitute a limitation on the electronic device. An electronic device may include more or fewer components than illustrated, or combine certain components, or have different component arrangements. In the embodiments of this invention, electronic devices include, but are not limited to, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic devices may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the embodiments of this application described and / or claimed herein.
[0057] Electronic devices may include processors, external memory interfaces, internal memory, universal serial bus (USB) interfaces, charging management modules, power management modules, batteries, wireless communication modules, audio modules, speakers, microphones, sensor modules, buttons, cameras, displays, and SIM card interfaces, etc.
[0058] A processor may include one or more processing units, such as: a central processing unit (CPU), an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). Different processing units may be independent devices or integrated into one or more processors.
[0059] The processor can serve as the nerve center and command center of an electronic device. The controller can generate operation control signals based on the instruction opcode and timing signals to control the fetching and execution of instructions.
[0060] The processor may also include memory for storing instructions and data. In some embodiments, the memory in the processor is a cache memory. This memory can store instructions or data that the processor has just used or that are used repeatedly. If the processor needs to use the instruction or data again, it can retrieve it directly from this memory. This avoids repeated accesses, reduces processor latency, and thus improves system efficiency.
[0061] An external storage interface (ESI) can be used to connect external memory cards, such as microSD cards, to expand the storage capacity of electronic devices. The external memory card communicates with the processor through the ESI to perform data storage functions, such as saving music and video files on the external memory card.
[0062] Internal memory can be used to store computer executable program code, which includes instructions. The processor executes various functional applications and data processing of electronic devices by running the instructions stored in internal memory. Internal memory can include a program storage area and a data storage area. Internal memory can include high-speed random access memory, and can also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.
[0063] Wireless communication functionality in electronic devices can be achieved through antennas, wireless communication modules, modem processors, and baseband processors.
[0064] Wireless communication modules can provide solutions for wireless communication applications in electronic devices, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies.
[0065] Electronic devices can implement audio functions through audio modules, speakers, receivers, microphones, headphone jacks, and application processors.
[0066] Electronic devices can achieve shooting functions through ISPs, cameras, video codecs, GPUs, displays, and application processors.
[0067] Electronic devices can achieve display functions through GPUs, displays, and application processors.
[0068] A GPU is a microprocessor for image processing, connected to the display screen and application processor. GPUs are used to perform mathematical and geometric calculations for graphics rendering. A processor may include one or more GPUs, which execute program instructions to generate or modify display information.
[0069] A display screen is used to display images, videos, etc. A display screen includes a display panel.
[0070] The aforementioned electronic equipment realizes the intelligent management method of the government data resource catalog of this application. By adopting a large model enhanced with knowledge in the government domain, it achieves intelligent processing of the entire process of data integration, quality verification and dynamic maintenance, which has achieved the beneficial effects of significantly improving the efficiency of government data governance, effectively ensuring data quality, and realizing unified management and efficient utilization of data resources.
[0071] The storage medium provided in this application stores a program product capable of implementing an intelligent management method for government data resource catalogs.
[0072] Intelligent management methods for government data resource catalogs include: A training set is constructed based on samples from the government data resource catalog, a government data field dictionary, and historical manual records. The training set is then used to train a general large model to generate a large model specifically for government data. Collect government data resource catalogs from multiple sources, and use a dedicated big data model for government data to standardize fields and match data items in the government data resource catalogs to generate a merged catalog; The merged catalog is input into the dedicated big data model for government data. Redundant data is identified and deleted, erroneous data is identified and corrected, and missing fields are filled in through the dedicated big data model for government data, thereby generating a standardized government data resource catalog. The standardized government data resource catalog is stored and indexed; the data update operations of each government data source are continuously monitored; when a data change is detected, the updated government data resource catalog is automatically collected, and field standardization, data item matching, redundant data identification and deletion, erroneous data identification and correction, and missing field completion are triggered to update the stored standardized government data resource catalog.
[0073] In some possible implementations, the intelligent management method for the government data resource catalog of this disclosure can be implemented as a program product, which includes program code. When the program product is run on a terminal device, the program code is used to cause the terminal device to perform the steps described in the "Exemplary Methods" section above according to various exemplary embodiments of this disclosure.
[0074] The storage medium disclosed herein may be any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof.
[0075] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for intelligent management of government data resource catalogs, characterized in that, include: A training set is constructed based on samples from the government data resource catalog, a government data field dictionary, and historical manual records. The training set is then used to train a general large model to generate a large model specifically for government data. Collect government data resource catalogs from multiple sources, and use a dedicated big data model for government data to standardize fields and match data items in the government data resource catalogs to generate a merged catalog; The merged catalog is input into the dedicated big data model for government data. Redundant data is identified and deleted, erroneous data is identified and corrected, and missing fields are filled in through the dedicated big data model for government data, thereby generating a standardized government data resource catalog. Store the standardized government data resource catalog and create an index; Continuously monitor data update operations from various government data sources. When data changes are detected, automatically collect the updated government data resource catalog and trigger field standardization, data item matching, redundant data identification and deletion, erroneous data identification and correction, and missing field completion processing to update the stored standardized government data resource catalog.
2. The intelligent management method for government data resource catalog according to claim 1, characterized in that, The process involves constructing a training set based on samples from the government data resource catalog, a government data field dictionary, and historical manual records. This training set is then used to train a general-purpose large-scale model to generate a government data-specific large-scale model. This includes: We collected samples of government data resource catalogs, government data field dictionaries, and historical manual merging and cleaning records from government departments at all levels. We then performed format unification and noise reduction on the collected data to obtain the training dataset. Using a general large model as the base model, the training dataset is input into the base model, and the base model is trained using few-shot learning and incremental training methods to optimize the base model's ability to understand semantics in the field of government data, field matching ability, and data error identification ability. A validation dataset is constructed, comprising a preprocessed catalog of government data resources according to predetermined merging and cleaning rules. The validation dataset is input into the base model, and the model output is compared with the preprocessed catalog of government data resources to calculate the catalog matching accuracy and the data cleaning accuracy. If either the catalog matching accuracy or the data cleaning accuracy fails to reach its corresponding preset threshold, a supplementary training dataset is added, and the model parameters are adjusted for retraining until both the catalog matching accuracy and the data cleaning accuracy reach their respective preset thresholds. Finally, a large-scale model specifically designed for government data is output.
3. The intelligent management method for government data resource catalog according to claim 2, characterized in that, The aforementioned collection of multi-source government data resource catalogs involves standardizing fields and matching data items using a dedicated large-scale model for government data, resulting in a merged catalog, including: The catalog of government data resources to be merged is collected from the business systems or databases of various government departments through API calls or file uploads. The catalog of government data resources to be merged is converted into standard JSON format, and the field names and field types are standardized according to the preset government data field standards using the dedicated large model of government data. The semantic and attribute similarity between data items in different government data resource catalogs are calculated using the government data-specific big data model, and matching is performed using knowledge of the government data domain. For government data resource catalogs with hierarchical relationships, the hierarchical structure is automatically identified and the data items in the catalogs are matched according to the rules of supplementing the upper-level catalogs with the lower-level catalogs. For directory data items that match successfully, a strategy of retrieving all fields and removing duplicate values is used to merge the directory data items; for directory data items that do not match successfully, they are treated as new data items. After the directory data items are merged, a merge log is generated that records the source of the directory data items and the merging method, and the merged directory is output.
4. The intelligent management method for government data resource catalog according to claim 3, characterized in that, The merged catalog is then input into a dedicated government data model. This model performs redundant data identification and deletion, erroneous data identification and correction, and missing field completion, generating a standardized government data resource catalog, including: The merged catalog is input into the government data exclusive big model. The government data exclusive big model identifies completely duplicate and partially duplicate catalog data items from the data content dimension and redundant fields from the data attribute dimension. Based on the preset rules of retaining the latest data and core fields, redundant fields are automatically deleted to generate the first process catalog. The first process catalog is used to identify errors, including formatting errors, logical errors, and numerical errors, through a dedicated big data model for government data. For formatting errors, automatic format correction is performed. For logical and numerical errors, correction is performed by querying the government data standard database. If correction is not possible, an error message is generated. After the correction is completed, the second process catalog is generated. The model, dedicated to government data, identifies missing fields in the second-stage catalog. Based on existing data and knowledge of the government data domain, it automatically completes the missing field content, marks all completed missing field content, and outputs a standardized government data resource catalog and cleaning log.
5. The intelligent management method for government data resource catalog according to claim 4, characterized in that, The process of storing the standardized government data resource catalog and establishing an index includes: The standardized government data resource catalog is persistently stored in a distributed NoSQL database in document form, and a globally unique catalog identifier is assigned to each catalog data item in the standardized government data resource catalog. Parse the directory data items, extract the names and data types of the core business fields, and build a key-value pair index as the field index; Based on the source department metadata recorded in the catalog data items, construct an index that maps department codes to corresponding catalog identifiers, which serves as the government department index; Based on the preset government data classification standards, the business themes and classification tags of the catalog data items are extracted, and an index mapping the business classification to the catalog identifier is constructed as a data type index. A multi-level composite index is constructed based on field indexes, government department indexes, and data type indexes; The directory identifiers are associated with and stored in a multi-level composite index to form a benchmark directory library that can be retrieved by field, government department, or data type.
6. The intelligent management method for government data resource catalog according to claim 5, characterized in that, The continuous monitoring of data update operations from various government data sources automatically collects the updated government data resource catalog when data changes are detected, and triggers field standardization, data item matching, redundant data identification and deletion, erroneous data identification and correction, and missing field completion processing to update the stored standardized government data resource catalog, including: By monitoring the data update events of the business systems or databases of various government departments through the monitoring agent program, the updated government data resource catalog can be captured. The updated government data resource catalog is input into the government data-specific big data model, and the following steps are executed in sequence: field standardization, data item matching, redundant data identification and deletion, erroneous data identification and correction, and missing field completion, to generate the target catalog. The directory is written to the distributed NoSQL database, and the corresponding multi-level composite index is updated synchronously to complete the incremental update of the baseline directory database. The integrity verification and consistency check of the government data resource catalog in the benchmark catalog are periodically performed through a scheduled task service. By calling the data service interface through the front-end application, the government data resource catalog and associated merged and cleaned logs are obtained. The government data resource catalog is then organized and rendered in a tree topology structure, and a corresponding table view is generated.
7. The intelligent management method for government data resource catalog according to claim 1, characterized in that, The general large model adopts any one of the following: GPT series model, ERNIE series model, ChatGLM series model, LLaMA series model, and Tongyi Qianwen series model.
8. A smart management system for government data resource catalogs, characterized in that, The system adopts the intelligent management method for the government data resource catalog as described in any one of claims 1 to 7; The system includes: The model building module is used to build a training set based on samples from the government data resource catalog, a government data field dictionary, and historical manual records. The training set is then used to train a general large model to generate a large model specific to government data. The catalog merging module is used to collect government data resource catalogs from multiple sources. It performs field standardization and data item matching on the government data resource catalogs using a dedicated large model for government data, and generates a merged catalog. The data cleaning module is used to input the merged catalog into the government data-specific big data model. The government data-specific big data model is used to identify and delete redundant data, identify and correct erroneous data, and complete missing fields to generate a standardized government data resource catalog. The directory management module is used to store the standardized government data resource directory and establish an index; continuously monitor the data update operations of each government data source; when a data change is detected, automatically collect the updated government data resource directory and trigger field standardization, data item matching, redundant data identification and deletion, erroneous data identification and correction, and missing field completion processing to update the stored standardized government data resource directory.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the intelligent management method for the government data resource catalog as described in any one of claims 1 to 7.
10. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the intelligent management method for the government data resource catalog as described in any one of claims 1 to 7.