Government affair big model-based data government method and system
Through the data governance method of the government big model, the problem of data governance in government scenarios has been solved, intelligent analysis and decision support of cross-departmental data have been realized, and the efficiency and security of government data processing have been improved.
Patent Information
- Application Number
- CN202510782403.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-09-12
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The data governance of existing large models in government scenarios is difficult to adapt to the needs of professional fields, especially when facing unstructured policy texts, multi-source heterogeneous people's livelihood data and privacy information. The lack of effective data integration and security guarantees leads to insufficient efficiency of smart government affairs.
Through the data governance method of the government big model, including government data collection, governance and construction, a multi-level processing mechanism is used for pre-processing, combined with a deep learning framework and semantic understanding engine, intelligent analysis and decision support of cross-departmental data can be achieved.
It reduces the difficulty of data governance in government scenarios, promotes the deep transformation of smart government affairs, improves the efficiency of data analysis and the convenience of decision-making support, and realizes unified governance and security management of data.
Smart Images

Figure CN120632787A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning technology, and specifically to a data governance method and system based on a government affairs big model. Background Art
[0002] With the rapid advancement of artificial intelligence (AI) technology, the application of big models is penetrating deeply into various verticals. However, it's worth noting that current mainstream big models often utilize generalized architectural designs, lacking deep adaptation and targeted training for specialized scenarios. This directly results in their practical effectiveness in specialized fields like government services falling short of expectations. This is particularly true when faced with the unique data characteristics of government scenarios, where unstructured policy texts, multi-source, heterogeneous public welfare data, and highly sensitive privacy information are intertwined, creating unique data governance challenges. This complex data ecosystem not only tests the semantic understanding capabilities of intelligent systems but also places stringent demands on data security mechanisms. Achieving a deep transformation towards intelligent government requires overcoming barriers to cross-departmental data integration, building a specialized algorithmic framework for government knowledge graphs, and establishing an intelligent decision-making support system that aligns with government workflows. This represents both a breakthrough point for technological innovation and a key focus for driving government digital transformation. Summary of the Invention
[0003] The technical task of the present invention is to address the above shortcomings and provide a data governance method and system based on the government big model, which can reduce the difficulty of data governance in government scenarios, promote the deep transformation of smart government affairs and the establishment of an intelligent decision-making support system.
[0004] The technical solution adopted by the present invention to solve its technical problem is:
[0005] A data governance method based on a government affairs big model, the implementation of which includes:
[0006] Government data collection: Obtain basic data from data sources that implement data governance. Basic data includes historical data and current data to ensure efficient data aggregation and consistency, providing a foundation for subsequent data processing and analysis.
[0007] Government data governance: Preprocess the collected basic data and integrate it to obtain data sequences. A multi-level processing mechanism is used to sequentially perform preprocessing operations, including structured transformation, outlier correction, and missing value filling. A standardized training set is generated through feature dimension reconstruction and semantic alignment algorithms.
[0008] Building a large government model: The platform adaptively selects a pre-set deep learning framework based on the characteristics of government scenarios, and combines it with a phased parameter optimization strategy to train and generate a government decision-making model with domain cognition capabilities.
[0009] Intelligent analysis: Deploy a semantic understanding engine to convert unstructured government texts into machine-parseable data vector representations, and output data analysis results through multimodal feature fusion technology.
[0010] This approach acquires multi-source government data and leverages government big data models and natural language processing technologies to empower data processing and analysis. This approach reduces the difficulty of data governance in government scenarios, promotes the deep transformation of smart government, and promotes the establishment of intelligent decision-making support systems. By building a domain-aware framework, it enables intelligent analysis and decision-making support for cross-departmental government data.
[0011] Furthermore, the government data collection, by building a multi-source heterogeneous information access channel, dynamically integrates historical business records, real-time monitoring data, and cross-departmental shared information based on the actual needs of specific government functional areas; the basic data acquisition process is as follows:
[0012] Based on data governance goals and needs, identify the types of data sources that need to be connected for data governance implementation, including internal business systems, external partner systems, social media platforms, and IoT devices, and establish connections with the corresponding data sources;
[0013] Based on business needs and data governance strategies, data extraction rules are defined to determine the data types, data fields, data formats, and data frequencies to be extracted. Priority and sequencing of data extraction are established to ensure priority processing of critical data. Based on the defined rules, basic data, including historical and current data, is extracted from connected data sources. Data types include structured, unstructured, and time series data. For historical data, batch collection is used to obtain the required data stored in the database in one go. For current data, real-time collection is used to continuously obtain data through real-time data streams.
[0014] Convert the collected data into a unified format and standardize the data to eliminate differences between different data sources and ensure data consistency.
[0015] Furthermore, the governance process of government data is as follows:
[0016] The collected basic data are uniformly connected to the data governance module, and the data are preliminarily analyzed to clarify the structure, type and distribution of the data; the basic data dataset is cleaned, including missing value processing, outlier detection and duplicate data processing, and the fields and records where the missing values are located are identified according to the characteristics of the data and business needs, and the missing values are filled or deleted; the data records of the same entity from different data sources are horizontally integrated, and the data records of different time periods in the same data source are integrated to form a continuous time series dataset, and a unique ID number is created for each set of data. Based on the timestamp association attributes between the data, the association relationship between each data and the ID number is established to form a data sequence.
[0017] Furthermore, the implementation of the process of filling missing values or deleting missing values:
[0018] For numeric fields, use the mean, median, or mode of the field to fill missing values; for categorical fields, use the mode to fill missing values; if the data has obvious trends or periodicity, implement interpolation to fill missing values; for records or fields with many missing values that cannot be accurately filled, choose to delete them to avoid a significant impact on subsequent analysis and model training;
[0019] For detected outliers, appropriate processing measures are taken based on the nature of the outliers and business needs, including correcting the outliers, deleting the outliers, and retaining the outliers. Data deduplication algorithms are used to compare the values of each field in the data records to identify duplicate records. After identifying duplicate data, redundant duplicate records are deleted based on business needs and data uniqueness requirements, leaving only one copy of valid data.
[0020] Furthermore, the data set of the basic data is cleaned.
[0021] Scan the data set integrated with basic data, delete duplicate data in the historical government industry data through a hash table to obtain processed data; and perform data cleaning on the processed data.
[0022] Furthermore, the construction of the government affairs big model includes:
[0023] Target model selection: Determine the target model to be trained from the preset model architecture based on the data type of the data to be trained and the business requirements;
[0024] Model training: performing model training on the target model to be trained based on the training data to obtain a trained government affairs model;
[0025] Model evaluation: Evaluate the trained government affairs model based on a preset evaluation model to obtain an evaluation result, and determine whether the evaluation result meets the training requirements of the preset model;
[0026] If the evaluation result meets the preset model training requirements, the trained government affairs big model will be determined as the target government affairs big model; if the evaluation result does not meet the preset model training requirements, the model parameters of the trained government affairs big model will be optimized based on the preset model optimization strategy to obtain the target government affairs big model.
[0027] Furthermore, the intelligent analysis includes:
[0028] Data identification: obtaining the data to be analyzed and identifying the data to be analyzed based on natural language processing technology to obtain model input data;
[0029] Data analysis: input the model input data into the target government affairs big model, and generate data analysis results corresponding to the data to be analyzed based on the data type and business requirements of the model input data.
[0030] The present invention also claims protection for a data governance system based on a government affairs big model, comprising:
[0031] The government data collection module is used to obtain basic data from data sources that implement data governance. The basic data includes historical data and current data;
[0032] The government data governance module is used to pre-process the collected basic data and integrate the data to obtain data sequences;
[0033] The government affairs model building module is used by the platform to adaptively select a pre-set deep learning framework based on the characteristics of government affairs scenarios, and combine it with a phased parameter optimization strategy to train and generate a government affairs decision-making model with domain cognition capabilities;
[0034] The intelligent analysis module deploys a semantic understanding engine to transform unstructured government documents into machine-parseable data vector representations and output data analysis results through multimodal feature fusion technology. This solution builds a domain-aware framework to enable intelligent analysis and decision support for cross-departmental government data.
[0035] The system implements government data governance through the above method.
[0036] The present invention also claims protection for a data governance device based on a government affairs big model, comprising: at least one memory and at least one processor;
[0037] The at least one memory is configured to store a machine-readable program;
[0038] The at least one processor is configured to call the machine-readable program to implement the above method.
[0039] The present invention also claims protection for a computer-readable medium having computer instructions stored thereon, which implement the above method when executed by a processor.
[0040] Compared with the existing technology, the data governance method and system based on the government affairs big model of the present invention have the following beneficial effects:
[0041] 1. Lowering the threshold for data analysis: Users can interact with the system through natural language and quickly obtain the required data insights without having to possess professional data analysis skills.
[0042] 2. Data-driven decision acceleration: Respond to complex and changing data needs in real time, help enterprises make data-based decisions quickly, and reduce the time cost of waiting for assistance from IT departments or data analysis teams.
[0043] 3. Improved the efficiency and convenience of data analysis: The system can quickly understand the user's query requirements and provide accurate analysis results. It also supports data visualization to help users better understand the data.
[0044] 4. Realized unified data governance: Integrating massive data resources to form a unified data view, improving data management efficiency and ease of use. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 This is a diagram illustrating the architecture of a data governance method based on a government affairs big model provided by an embodiment of the present invention;
[0046] Figure 2 It is a flowchart of the data governance method based on the government affairs big model provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0047] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0048] An embodiment of the present invention provides a data governance method based on a government affairs big model. The implementation of the method includes:
[0049] Government data collection: By building multi-source heterogeneous information access channels, we dynamically integrate historical business records, real-time monitoring data, and cross-departmental shared information based on the actual needs of specific government functional areas;
[0050] Government data governance: This uses a multi-level processing mechanism to sequentially perform pre-processing operations such as structured transformation, outlier correction, and missing value filling. It then generates a standardized training set through feature dimension reconstruction and semantic alignment algorithms.
[0051] Building a large government model: The platform adaptively selects a pre-set deep learning framework based on the characteristics of government scenarios, and combines it with a phased parameter optimization strategy to train and generate a government decision-making model with domain cognition capabilities.
[0052] Intelligent analysis: Deploy a semantic understanding engine to convert unstructured government texts into machine-parseable data vector representations, and output data analysis results through multimodal feature fusion technology.
[0053] The specific implementation of this method is as follows:
[0054] S1. Collection of government data.
[0055] Obtain basic data from data sources that implement data governance, including historical data and current data, to ensure efficient data aggregation and consistency, and provide a basis for subsequent data processing and analysis.
[0056] The basic data acquisition process is as follows: according to the goals and needs of data governance, identify the types of data sources that need to be connected to implement data governance, and establish connections with the corresponding data sources; according to business needs and data governance strategies, define data extraction rules, determine the types of data to be extracted, etc., and formulate the priority and order of data extraction to ensure priority processing of key data. According to the defined rules, extract basic data including historical data and current data from the connected data source; data types include structured data types, unstructured data types and time series data types; for historical data, adopt batch collection to obtain the required data stored in the database at one time; for current data, adopt real-time collection to continuously obtain data through real-time data streams.
[0057] S2. Government data governance.
[0058] Preprocess the collected basic data and integrate the data to obtain data sequences, enhance the integrity and consistency of the data, and provide high-quality data support for subsequent data analysis and model training.
[0059] The governance process is as follows: uniformly import the collected basic data into the government data governance module, and conduct a preliminary analysis of the data to clarify the structure, type, distribution and potential problems of the data; scan the data set that integrates the basic data, and delete the duplicate data in the historical government industry data through the hash table to obtain the first processed data; perform data cleaning on the first processed data, including missing value processing, outlier detection and duplicate data processing, and identify the fields and records where the missing values are located according to the characteristics of the data and business needs, and implement processing measures to fill in the missing values or delete the missing values; convert the collected data into a unified format to facilitate subsequent data processing and analysis, and standardize the data to obtain the second processed data; horizontally integrate the second processed data of the same entity from different data sources, and integrate the second processed data of different time periods in the same data source to form a continuous time series data set; create a unique ID number for each set of data, and establish an association relationship between each data and the ID number based on the timestamp association attribute between the data, and integrate to form a data sequence.
[0060] S3. Construction of a large government model. This includes:
[0061] Select a model and determine the target model to be trained from the preset model architecture based on historical data in the data series and relevant business needs.
[0062] Model training: Based on the historical data in the data sequence, the target model to be trained is trained to obtain a large government model.
[0063] Model evaluation: evaluate the trained government affairs model based on the preset evaluation model to obtain the evaluation results, and determine whether the evaluation results meet the preset model training requirements.
[0064] If the evaluation result meets the preset model training requirements, the trained government affairs big model will be determined as the target government affairs big model. If the evaluation result does not meet the preset model training requirements, the model parameters of the trained government affairs big model will be optimized based on the preset model optimization strategy to obtain the target government affairs big model.
[0065] S4. Intelligent analysis. Including:
[0066] Data identification: obtaining the data to be analyzed, and identifying the data to be analyzed based on natural language processing technology to obtain model input data.
[0067] Data analysis: input the model input data into the target government affairs big model, and generate data analysis results corresponding to the data to be analyzed based on the data type and business requirements of the model input data.
[0068] This approach acquires multi-source government data and leverages government big data models and natural language processing technologies to empower data processing and analysis. This approach reduces the difficulty of data governance in government scenarios, promotes the deep transformation of smart government, and promotes the establishment of intelligent decision-making support systems. By building a domain-aware framework, it enables intelligent analysis and decision-making support for cross-departmental government data.
[0069] The embodiment of the present invention further provides a data governance system based on a government affairs big model, including:
[0070] The government data collection module, by building multi-source heterogeneous information access channels, dynamically integrates historical business records, real-time monitoring data and cross-departmental shared information based on the actual needs of specific government functional areas.
[0071] The government data governance module adopts a multi-level processing mechanism to perform pre-processing operations such as structured transformation, outlier correction, and missing value filling in sequence, and then generates a standardized training set through feature dimension reconstruction and semantic alignment algorithm.
[0072] In the government affairs big model construction module, the platform adaptively selects a preset deep learning framework based on the characteristics of government affairs scenarios, combines it with a phased parameter optimization strategy, and trains and generates a government affairs decision-making model with domain cognition capabilities.
[0073] The intelligent analysis module deploys a semantic understanding engine to convert unstructured government texts into machine-parseable data vector representations, and outputs data analysis results through multimodal feature fusion technology.
[0074] The system implements government data governance through the data governance method based on the government big model described in the above embodiment.
[0075] 1. The government data collection module is used to obtain basic data from data sources that implement data governance. Basic data includes historical data and current data to ensure efficient data aggregation and consistency, and provide a basis for subsequent data processing and analysis. According to the goals and needs of data governance, identify the types of data sources that need to be connected to implement data governance, including internal business systems, external partner systems, social media platforms, and IoT devices, and establish connections with corresponding data sources. The ways to establish connections include API interface docking, database connection configuration, and network protocol adaptation. For data sources that provide API interfaces, obtaining data by calling the API interface requires negotiation with the data source to obtain the API key, understand the API usage restrictions and data format, etc., to ensure that the required data can be obtained smoothly through the API interface. For data sources stored in the database, configure the database connection parameters, including database type, server address, port number, user The system uses username and password to establish a connection with the database to extract data from the database, define data extraction rules, determine the data fields to be extracted, data format and update frequency, etc., and set data extraction priorities to ensure priority processing of key data. According to the defined rules, basic data is extracted from the connected data source; for historical data, batch collection is used to obtain the required data stored in the database at one time; for current data, real-time collection is used to continuously obtain data through real-time data streams; the collected data is converted into a unified format to facilitate subsequent data processing and analysis, and the data is standardized to eliminate differences between different data sources and ensure data consistency;
[0076] 2. The government data governance module pre-processes the collected basic data and integrates the data to obtain a data sequence. The collected basic data are uniformly connected to the data governance module, and the data is preliminarily analyzed to clarify the structure, type and distribution of the data; the basic data dataset is cleaned, including missing value processing, outlier detection and duplicate data processing, and the fields and records where the missing values are located are identified according to the characteristics of the data and business needs, and the processing measures of filling in missing values or deleting missing values are implemented. For numerical fields, the mean, median or mode of the field is used to fill the missing values. For categorical fields, the mode is used to fill in. For data with obvious trends or periodicity, interpolation filling measures are implemented. For records or fields with many missing values that cannot be accurately filled, they are deleted to avoid causing problems in subsequent analysis and model training. A significant impact. For detected outliers, appropriate processing measures are taken based on the nature of the outliers and business needs, including correcting, deleting, and retaining outliers. Data deduplication algorithms are used to compare the values of each field in data records to identify duplicate records. After identifying duplicate data, redundant duplicate records are deleted based on business needs and data uniqueness requirements, leaving only one valid copy of the data. Data records of the same entity from different data sources are horizontally integrated, and data records from different time periods in the same data source are integrated to form a continuous time series data set. A unique ID number is created for each set of data, and based on the timestamp association attribute between the data, an association relationship is established between each data set and the ID number to form a data sequence.
[0077] 3. Government affairs big model construction module, including:
[0078] Select a model and determine the target model to be trained from the preset model architecture based on historical data in the data series and relevant business needs;
[0079] Model training: training the target model to be trained based on historical data in the data sequence to obtain a trained government affairs model;
[0080] Model evaluation: evaluate the trained government affairs model based on the preset evaluation model to obtain the evaluation results, and determine whether the evaluation results meet the preset model training requirements;
[0081] If the evaluation result meets the preset model training requirements, the trained government affairs big model will be determined as the target government affairs big model. If the evaluation result does not meet the preset model training requirements, the model parameters of the trained government affairs big model will be optimized based on the preset model optimization strategy to obtain the target government affairs big model.
[0082] Systematic planning is required in the early stages of model building, with the following dimensions to consider: First, data features must be analyzed, and an appropriate algorithm framework must be selected based on different data types (such as time series data, spatial data, and multimodal data); second, application requirements must be matched to ensure that the architecture design effectively corresponds to business scenarios (such as risk prediction, customer segmentation, and image recognition); hardware conditions must be evaluated to design an appropriate network depth and parameter scale within the constraints of video memory capacity and computing power. Model training involves four core steps: first, error metric configuration, where evaluation indicators must be selected based on the nature of the task (such as squared error for regression scenarios and cross-entropy loss for classification tasks); second, iterative strategy formulation, where an efficient convergence mechanism is established by comparing different optimizers (such as Adam and SGD with momentum) and adjusting parameters such as the learning rate; third, generalization capability enhancement, where weight constraints (L1 / L2 norms) or random dropout techniques are used to suppress overfitting; and fourth, training process design, where key parameters such as the number of iterations, sample batch size, and early termination mechanism must be comprehensively determined. The model validation system needs to construct a multi-dimensional evaluation matrix: in addition to basic performance indicators (precision, recall, F-value, and area under the ROC curve), a k-fold cross-validation must be implemented to ensure the reliability of the results, and a confusion matrix must be used to analyze specific misclassification cases. Based on the evaluation feedback, performance enhancement measures can be implemented: using automated parameter adjustment tools (grid method, Bayesian optimizer) to explore the optimal parameter combination; using network pruning technology to remove redundant connections to improve inference efficiency; integrating multi-model prediction results to form a more robust target government model. A real-time monitoring module should also be configured in the operating environment to dynamically track key indicators such as prediction accuracy, response latency, and resource utilization. A regular update mechanism should be established to address data distribution drift and ensure that the model continues to meet business needs.
[0083] 4. Intelligent analysis module, which uses natural language processing technology to determine model input data from the data to be analyzed, and inputs the model input data into the target government affairs model to obtain data analysis results corresponding to the data to be analyzed. It includes:
[0084] Data identification: obtaining the data to be analyzed and identifying the data to be analyzed based on natural language processing technology to obtain model input data. Specific natural language processing technologies include topic modeling and keyword extraction.
[0085] Data analysis: input the model input data into the target government affairs big model, and generate data analysis results corresponding to the data to be analyzed based on the data type and business requirements of the model input data.
[0086] The intelligent analysis module provides an API interface, allowing integration with other systems to achieve data exchange and function expansion. When processing sensitive data, the intelligent analysis module needs to ensure data security and user privacy protection.
[0087] The intelligent analysis module also includes data visualization and reporting capabilities, which help users intuitively understand analysis results. Data and analysis results are presented in the form of charts, graphs, and dashboards. The system can automatically generate reports containing key indicators, trend analysis, and recommendations.
[0088] An embodiment of the present invention further provides a data governance device based on a government affairs big model, comprising: at least one memory and at least one processor;
[0089] The at least one memory is configured to store a machine-readable program;
[0090] The at least one processor is used to call the machine-readable program to implement the data governance method based on the government affairs big model described in the above embodiment.
[0091] An embodiment of the present invention further provides a computer-readable medium having computer instructions stored thereon. When executed by a processor, the computer instructions implement the data governance method based on the government affairs big model described in the above embodiments. Specifically, a system or device equipped with a storage medium can be provided. The storage medium stores software program code that implements the functions of any of the above embodiments, and the computer (or CPU or MPU) of the system or device can read and execute the program code stored in the storage medium.
[0092] In this case, the program code itself read from the storage medium can realize the function of any one of the above-mentioned embodiments, and thus the program code and the storage medium storing the program code constitute part of the present invention.
[0093] Examples of storage media for providing program code include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, the program code can be downloaded from a server computer via a communication network.
[0094] In addition, it should be clear that the functions of any of the above embodiments can be achieved not only by executing the program code read by the computer, but also by enabling the operating system operating on the computer to complete part or all of the actual operations based on the instructions of the program code.
[0095] In addition, it can be understood that the program code read from the storage medium is written into the memory provided in the expansion board inserted into the computer or into the memory provided in the expansion unit connected to the computer, and then based on the instructions of the program code, the CPU installed on the expansion board or expansion unit is enabled to perform part or all of the actual operations, thereby realizing the functions of any of the above embodiments.
[0096] The present invention has been shown and described in detail above through the accompanying drawings and preferred embodiments. However, the present invention is not limited to these disclosed embodiments. Based on the above multiple embodiments, those skilled in the art can know that the code review methods in the above different embodiments can be combined to obtain more embodiments of the present invention, and these embodiments are also within the scope of protection of the present invention.
Claims
1. A data governance method based on a government affairs big model, characterized by: The implementation of this method includes: Government data collection: Obtain basic data from data sources that implement data governance. Basic data includes historical data and current data; Government data governance: Preprocess the collected basic data and integrate it to obtain data sequences. A multi-level processing mechanism is used to sequentially perform preprocessing operations, including structured transformation, outlier correction, and missing value filling. A standardized training set is generated through feature dimension reconstruction and semantic alignment algorithms. Building a large government model: The platform adaptively selects a pre-set deep learning framework based on the characteristics of government scenarios, and combines it with a phased parameter optimization strategy to train and generate a government decision-making model with domain cognition capabilities. Intelligent analysis: Deploy a semantic understanding engine to convert unstructured government texts into machine-parseable data vector representations, and output data analysis results through multimodal feature fusion technology.
2. A data governance method based on a government affairs big model according to claim 1, characterized in that: The government data collection mentioned above builds a multi-source heterogeneous information access channel to dynamically integrate historical business records, real-time monitoring data, and cross-departmental shared information based on the actual needs of specific government functional areas. The basic data acquisition process is as follows: Based on the goals and requirements of data governance, identify the types of data sources that need to be connected to implement data governance and establish connections with the corresponding data sources; Based on business needs and data governance strategies, data extraction rules are defined to determine the data types, data fields, data formats, and data frequencies to be extracted. The data extraction priority and sequence are then established. Based on the defined rules, basic data, including historical and current data, is extracted from connected data sources. Data types include structured, unstructured, and time series data. For historical data, batch collection is used to obtain the required data stored in the database in one go. For current data, real-time collection is used to continuously obtain data through real-time data streams. Convert the collected data into a unified format and standardize the data to eliminate differences between different data sources.
3. The data governance method based on the government affairs big model according to claim 1 is characterized in that: The governance process of government data is as follows: The collected basic data is uniformly connected to the data governance module, and a preliminary analysis is performed on the data to clarify the structure, type, and distribution of the data. The basic data dataset is cleaned, including missing value processing, outlier detection, and duplicate data processing. Based on the characteristics of the data and business needs, the fields and records where missing values are located are identified, and the missing values are filled or deleted. The data records of the same entity from different data sources are horizontally integrated, and the data records of different time periods in the same data source are integrated to form a continuous time series data set. A unique ID number is created for each set of data. Based on the timestamp association attributes between the data, the association relationship between each data and the ID number is established to integrate and form a data sequence.
4. A data governance method based on a government affairs big model according to claim 3, characterized in that: The implementation of the process of filling missing values or deleting missing values: For numeric fields, use the mean, median, or mode of the field to fill missing values; for categorical fields, use the mode to fill missing values; if the data has obvious trends or periodicity, use interpolation to fill missing values; for records or fields with many missing values that cannot be accurately filled, choose to delete; For detected outliers, appropriate processing measures are taken based on the nature of the outliers and business needs, including correcting the outliers, deleting the outliers, and retaining the outliers. Data deduplication algorithms are used to compare the values of each field in the data records to identify duplicate records. After identifying duplicate data, redundant duplicate records are deleted based on business needs and data uniqueness requirements, leaving only one copy of valid data.
5. The data governance method based on the government affairs big model according to claim 3 is characterized in that: The data set of the basic data is cleaned. Scan the data set integrated with basic data, delete duplicate data in the historical government industry data through a hash table to obtain processed data; and perform data cleaning on the processed data.
6. The data governance method based on the government affairs big model according to claim 1 is characterized in that: The construction of the government affairs big model includes: Target model selection: Determine the target model to be trained from the preset model architecture based on the data type of the data to be trained and the business requirements; Model training: performing model training on the target model to be trained based on the training data to obtain a trained government affairs model; Model evaluation: Evaluate the trained government affairs model based on a preset evaluation model to obtain an evaluation result, and determine whether the evaluation result meets the training requirements of the preset model; If the evaluation result meets the preset model training requirements, the trained government affairs big model will be determined as the target government affairs big model; if the evaluation result does not meet the preset model training requirements, the model parameters of the trained government affairs big model will be optimized based on the preset model optimization strategy to obtain the target government affairs big model.
7. A data governance method based on a government affairs big model according to claim 1 or 6, characterized in that: The intelligent analysis includes: Data identification: obtaining the data to be analyzed and identifying the data to be analyzed based on natural language processing technology to obtain model input data; Data analysis: input the model input data into the target government affairs big model, and generate data analysis results corresponding to the data to be analyzed based on the data type and business requirements of the model input data.
8. A data governance system based on a government affairs big model, characterized by: include: The government data collection module is used to obtain basic data from data sources that implement data governance. The basic data includes historical data and current data; The government data governance module is used to pre-process the collected basic data and integrate the data to obtain data sequences; The government affairs model building module is used by the platform to adaptively select a pre-set deep learning framework based on the characteristics of government affairs scenarios, and combine it with a phased parameter optimization strategy to train and generate a government affairs decision-making model with domain cognition capabilities; The intelligent analysis module deploys a semantic understanding engine to transform unstructured government documents into machine-parseable data vector representations and output data analysis results through multimodal feature fusion technology. This solution builds a domain-aware framework to enable intelligent analysis and decision support for cross-departmental government data. The system implements government data governance through any of the methods described in claims 1 to 7.
9. A data governance device based on a government affairs big model, characterized in that: include: at least one memory and at least one processor; The at least one memory is configured to store a machine-readable program; The at least one processor is configured to call the machine-readable program to implement the method according to any one of claims 1 to 7.
10. A computer-readable medium, characterized in that The computer-readable medium stores computer instructions, which, when executed by a processor, implement the method according to any one of claims 1 to 7.
Citation Information
Cited By
Government affair data processing method and device based on large model, equipment and storage medium
CN120873069A
Data aided analysis, prediction, research and judgment system based on index management
CN121258285A