Government affair data sharing and exchanging method
Through standardized data processing, multi-level feature extraction and BERT-based deep learning classification, combined with security level control and efficient data sharing mechanism, the problems of low classification efficiency, irregular exchange and insufficient security in government data sharing are solved, and efficient, secure and intelligent government data sharing and exchange are achieved.
Patent Information
- Application Number
- CN202510232330.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-17
AI Technical Summary
The existing government data sharing methods have problems such as low data classification efficiency, irregular data exchange and insufficient data security.
Through standardized data processing, multi-level feature extraction, deep learning classification based on BERT, security level control and efficient data sharing mechanism, efficient government data classification, standardized exchange and security management are achieved.
It significantly improves the efficiency, security and intelligence level of government data sharing and exchange, meets the high-quality, security and traceability requirements in government data management, and provides strong support for cooperation and decision-making among government departments.
Smart Images

Figure CN120162644A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of government data processing, and particularly to a method for sharing and exchanging government data. Background Art
[0002] With the advancement of digital government construction, the sharing and exchange of government data have become important means to improve government service efficiency and optimize resource allocation. However, government data is usually scattered in different departments, with inconsistent data formats and standards, and complex data content, involving various types of text information (such as policy documents, statistical data, approval records, etc.). The traditional methods for sharing government data have the following problems:
[0003] (1) Low data classification efficiency: There are numerous types of government data, and manual classification is time-consuming and laborious, and prone to errors.
[0004] (2) Irregular data exchange: The data standards between different departments are not unified, resulting in low data exchange efficiency.
[0005] (3) Insufficient data security: Government data involves sensitive information, and the lack of effective classification and permission management mechanisms may lead to data leakage.
[0006] Therefore, there is an urgent need for a sharing and exchange method that can efficiently classify government data, standardize the data exchange process, and ensure data security. Summary of the Invention
[0007] In order to overcome the deficiencies of the prior art, the purpose of the present invention is to provide a method for sharing and exchanging government data, which significantly improves the efficiency, security, and intelligence level of government data sharing and exchange through standardized data processing, multi-level feature extraction, deep learning classification based on BERT, security level control, and an efficient data sharing mechanism. At the same time, it meets the requirements of high quality, security, and traceability in government data management, and provides strong support for the collaboration and decision-making among government departments.
[0008] To achieve the above purpose, the present invention provides the following solution:
[0009] A method for sharing and exchanging government data, comprising:
[0010] Collecting raw data from each government department, and cleaning and formatting the raw data to obtain standardized data;
[0011] Inputting the standardized data into a text context relationship feature extraction layer to extract context relationship feature information;
[0012] Inputting the standardized data into a global feature extraction layer to extract global feature information;
[0013] Fuse the context relationship feature information and the global feature information to obtain fused text features;
[0014] Input the fused text features into a deep learning model based on BERT to obtain classified data;
[0015] Set data access permissions for the classified data according to a preset security level, and perform data sharing and exchange based on the data access permissions.
[0016] Preferably, input the standardized data into a text context relationship feature extraction layer to extract context relationship feature information, including:
[0017] Use a pre-trained language model to extract initial feature information of the text to be classified;
[0018] Input the initial feature information into a forward gated recurrent unit and a backward gated recurrent unit;
[0019] Concatenate the outputs of the forward gated recurrent unit and the backward gated recurrent unit to obtain context relationship feature information.
[0020] Preferably, concatenate the outputs of the forward gated recurrent unit and the backward gated recurrent unit to obtain context relationship feature information, including:
[0021] Adopt the formula Concatenate the outputs of the forward gated recurrent unit and the backward gated recurrent unit to obtain context relationship feature information; where, H t ={H1, H2, … H l} represents the initial feature information input at time t, represents the output of the forward gated recurrent unit at time t, represents the output of the backward gated recurrent unit at time t, GRU represents the gated recurrent unit, and G represents the context relationship feature information.
[0022] Preferably, input the text to be classified into a global feature extraction layer to extract global feature information, including:
[0023] Input the initial feature information of the text to be classified into a convolutional layer and a pooling layer in sequence to obtain global feature information; where, the global feature information extraction formula is:
[0024] c i = f(ω·H + b)
[0025]
[0026] where f is the activation function, ω is the convolution kernel, h is the convolution kernel size, b is the bias, and c i is the i-th extracted feature vector of the convolutional layer, represents the values of the features extracted by 3 different convolution kernels after passing through the max pooling layer, and C represents the extracted global feature information.
[0027] Preferably, fusing the context relationship feature information and the global feature information to obtain fused text features includes:
[0028] concatenating the context relationship feature information and the global feature information to obtain a concatenated feature; where the concatenated feature O is:
[0029] performing max pooling processing on the concatenated feature to obtain a max pooling feature;
[0030] performing average pooling processing on the concatenated feature to obtain an average pooling feature;
[0031] combining the concatenated feature, the max pooling feature and the average pooling feature to obtain the fused text features.
[0032] Preferably, inputting the fused text features into a deep learning model based on BERT to obtain classified data includes:
[0033] inputting the fused text features into the embedding layer of a deep learning model based on BERT to encode the input fused text features and generate context-related embedding representations;
[0034] inputting the embedding representations into the multi-layer Transformer encoder of a deep learning model based on BERT to obtain deep semantic features;
[0035] inputting the deep semantic features into the classification layer of a deep learning model based on BERT to obtain the classified data.
[0036] Preferably, the classification layer consists of a fully connected layer and a Softmax activation function.
[0037] Preferably, the types of the data access rights include three categories; the first category is: all users can access; the second category is: only specific roles can access; the third category is: only authorized users can access.
[0038] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:
[0039] The present invention provides a method for sharing and exchanging government affairs data, including: collecting raw data from each government department, cleaning and formatting the raw data to obtain standardized data; inputting the standardized data into a text context relationship feature extraction layer to extract context relationship feature information; inputting the standardized data into a global feature extraction layer to extract global feature information; fusing the context relationship feature information and the global feature information to obtain fused text features; inputting the fused text features into a deep learning model based on BERT to obtain classified data; setting data access permissions for the classified data according to a preset security level, and performing data sharing and exchange based on the data access permissions. Through standardized data processing, multi-level feature extraction, deep learning classification based on BERT, security level control, and an efficient data sharing mechanism, the present invention significantly improves the efficiency, security, and intelligence level of government affairs data sharing and exchange. At the same time, it meets the requirements of high quality, security, and traceability in government affairs data management, providing strong support for collaboration and decision-making among government departments. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0041] Figure 1 It is a flowchart of the method provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0043] The object of the present invention is to provide a method for sharing and exchanging government affairs data. Through standardized data processing, multi-level feature extraction, deep learning classification based on BERT, security level control, and an efficient data sharing mechanism, the present invention significantly improves the efficiency, security, and intelligence level of government affairs data sharing and exchange. At the same time, it meets the requirements of high quality, security, and traceability in government affairs data management, providing strong support for collaboration and decision-making among government departments.
[0044] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0045] Figure 1 The flowchart of the method provided by the embodiment of the present invention is as Figure 1 shown. The present invention provides a method for sharing and exchanging government affairs data, including:
[0046] Step 100: Collect raw data from each government department, and clean and format the raw data to obtain standardized data;
[0047] Step 200: Input the standardized data into the text context relationship feature extraction layer to extract context relationship feature information;
[0048] Step 300: Input the standardized data into the global feature extraction layer to extract global feature information;
[0049] Step 400: Integrate the context relationship feature information and the global feature information to obtain integrated text features;
[0050] Step 500: Input the integrated text features into the deep learning model based on BERT to obtain classified data;
[0051] Step 600: Set data access permissions for the classified data according to the preset security levels, and perform data sharing and exchange based on the data access permissions.
[0052] Specifically, step 100 of this embodiment includes:
[0053] Step 101: Data collection
[0054] When collecting raw data from each government department, multiple collection methods need to be adopted according to the data sources and formats of different departments. Common collection methods include obtaining real-time data through API interfaces, exporting structured data from databases, and extracting unstructured or semi-structured data from files (such as Excel, CSV, JSON, etc.). During the collection process, it is necessary to ensure the security and integrity of data transmission. For example, data transmission is carried out through encrypted transmission protocols (such as HTTPS or SFTP), and the collected data is verified (such as verifying the hash value or the number of records of the file) to ensure that the data has not been tampered with or lost.
[0055] Step 102: Data cleaning
[0056] The collected raw data may have problems such as redundancy, inconsistency, missing, or errors. Therefore, it is necessary to clean the data. The cleaning process includes the following steps:
[0057] Delete duplicate records to ensure data uniqueness; handle missing values: fill in missing data (such as using the mean, median, or default value) or delete incomplete records; detect outliers in the data through statistical analysis or rules and correct or remove them; check whether the data format meets the expectations (such as date format, numerical range, etc.) and correct the data that does not meet the requirements.
[0058] Through data cleaning, the quality and reliability of the data can be improved, providing high-quality input for subsequent processing.
[0059] Step 103: Data Formatting and Standardization
[0060] The cleaned data needs to be formatted to unify the data structure and format to obtain standardized data. The formatting process includes:
[0061] (1) Field mapping and renaming: Map the data field names of different departments to unified standard field names.
[0062] (2) Data type conversion: Convert the data to a unified data type (such as unifying the date to the "YYYY - MM - DD" format and the numerical value to the floating-point type).
[0063] (3) Data encoding standardization: Standardize the categorical data (such as gender, area code, etc.) to ensure that the data of different departments has consistent encoding rules.
[0064] (4) Data structure unification: Organize the data into a unified structured format (such as table form or JSON format) for easy subsequent processing.
[0065] Through formatting and standardization processing, the present invention can eliminate the differences between the data of different departments and ensure the consistency and availability of the data during sharing and exchange.
[0066] Preferably, input the standardized data into the text context relationship feature extraction layer to extract context relationship feature information, including:
[0067] Use a pre-trained language model to extract the initial feature information of the text to be classified;
[0068] Input the initial feature information into the forward gated recurrent unit and the backward gated recurrent unit;
[0069] Concatenate the outputs of the forward gated recurrent unit and the backward gated recurrent unit to obtain context relationship feature information.
[0070] Specifically, in practical applications, this embodiment may adopt the XLNet model to complete the initial feature extraction step. The XLNet model is a pre-trained language model based on the combination of autoregression and autoencoding. Its core part is the Transformer-XL architecture, which solves the defect of the input length constraint of the context in the BERT model (Bidirectional Encoder Representations from Transformers), enabling the pre-trained model to learn more distant context information.
[0071] Preferably, splicing the outputs of the forward gated recurrent unit and the backward gated recurrent unit to obtain context relationship feature information, including:
[0072] Using the formula Splice the outputs of the forward gated recurrent unit and the backward gated recurrent unit to obtain context relationship feature information; where, H t ={H1, H2, … H l} represents the initial feature information input at time t, represents the output of the forward gated recurrent unit at time t, represents the output of the backward gated recurrent unit at time t, GRU represents the gated recurrent unit, and G represents the context relationship feature information.
[0073] Specifically, by splicing the outputs of the forward gated recurrent unit and the backward gated recurrent unit, the present invention can more fully learn the text context relationship and obtain context information.
[0074] Preferably, input the text to be classified into the global feature extraction layer to extract global feature information, including:
[0075] Input the initial feature information of the text to be classified into the convolutional layer and the pooling layer in sequence to obtain global feature information; where the global feature information extraction formula is:
[0076] c i = f(ω·H + b)
[0077]
[0078] In the formula, f is the activation function, ω is the convolution kernel, h is the convolution kernel size, b is the bias, and c i is the i-th feature vector extracted by the convolutional layer, represents the values of the features extracted by 3 different convolution kernels after passing through the max pooling layer, and C represents the extracted global feature information.
[0079] Preferably, fusing the context relationship feature information and the global feature information to obtain fused text features, including:
[0080] Concatenating the context relationship feature information and the global feature information to obtain a concatenated feature; wherein, the concatenated feature O is:
[0081] Performing max pooling on the concatenated feature to obtain a max pooling feature;
[0082] Performing average pooling on the concatenated feature to obtain an average pooling feature;
[0083] Combining the concatenated feature, the max pooling feature, and the average pooling feature to obtain the fused text features.
[0084] Specifically, in this embodiment, the two are concatenated to obtain a concatenated feature O, and its formula is: O = [G; C], where G represents context relationship feature information and C represents global feature information. The concatenation operation combines the two features along the feature dimension into an overall feature matrix. Then, pooling processing is performed on the concatenated feature O, including max pooling and average pooling operations. The max pooling processing extracts the maximum value of each dimension in the concatenated feature to obtain a max pooling feature; the average pooling processing calculates the average value of each dimension in the concatenated feature to obtain an average pooling feature.
[0085] Subsequently, the concatenated feature, the max pooling feature, and the average pooling feature are combined to obtain the final fused text feature F_final, and its formula is: F_final = [O; MaxPool(O); MeanPool(O)]. This fusion method combines the global information of the concatenated feature, the significant information of the max pooling feature, and the overall trend information of the average pooling feature, and can express text features more comprehensively, providing rich input features for subsequent deep learning models.
[0086] Preferably, inputting the fused text features into a deep learning model based on BERT to obtain classified data, including:
[0087] Inputting the fused text features into the embedding layer of a deep learning model based on BERT to encode the input fused text features and generate context-related embedding representations;
[0088] Inputting the embedding representations into the multi-layer Transformer encoder of a deep learning model based on BERT to obtain deep semantic features;
[0089] Inputting the deep semantic features into the classification layer of a deep learning model based on BERT to obtain the classified data.
[0090] Preferably, the classification layer consists of a fully connected layer and a Softmax activation function.
[0091] Specifically, in this embodiment, the fused text features are input into the embedding layer of the BERT model. The role of the embedding layer is to map the input fused text features into a high-dimensional semantic space and generate context-related embedding representations. Specifically, the embedding layer encodes the input features, captures their semantic information and context relationships, and generates an embedding representation E as the input for subsequent processing.
[0092] Next, the embedding representation E is input into the multi-layer Transformer encoder of the BERT model. The Transformer encoder consists of multiple self-attention mechanisms and feed-forward neural networks, and can perform in-depth semantic modeling on the input embedding representation. Through the processing of multiple layers of Transformer, the model can capture the global dependencies and deep semantic information of the input features, and finally generate deep semantic features T_L, where L represents the last layer of the Transformer encoder.
[0093] Then, the deep semantic features T_L are input into the classification layer of the BERT model. The classification layer consists of a fully connected layer and a Softmax activation function. The role of the fully connected layer is to map the deep semantic features into the classification space and generate the scores for each category; the Softmax activation function then converts these scores into a probability distribution, representing the probability that the input data belongs to each category. Finally, the classification layer outputs the classification result Y_cls, that is, the classified data.
[0094] In this way, the BERT model can make full use of the context information and global features of the fused text features, combined with its powerful semantic modeling ability, to achieve accurate classification of the input data. The classification result can be directly used for subsequent tasks, such as data sharing, permission setting, or other application scenarios.
[0095] Preferably, the types of the data access permissions include three categories; the first category is: all users can access; the second category is: only specific roles can access; the third category is: only authorized users can access.
[0096] Optionally, in this embodiment, security level rules and access permission policies are first defined. According to business requirements, a mapping relationship between preset security levels (such as "public", "internal", "confidential", etc.) and classified data categories is established. For example, data classified as "public" corresponds to the lowest security level and allows all users to access; data classified as "confidential" corresponds to the highest security level and only allows specific authorized users to access. Through this rule mapping, the classified data (Y_cls) is associated with the security level (Security_Level), providing a basis for subsequent permission settings.
[0097] Next, specific access permissions are set according to the security level of the data. The access permission policy needs to combine user roles (such as ordinary users, internal users, administrators, etc.) and security level rules to determine which users or systems can access which data. For example, using an access control list (ACL) or a role-based access control (RBAC) mechanism, the security level is bound to the user role to generate access permissions (Access_Permission). At the same time, access permission meta-information can be attached to each piece of data to ensure that permissions can be dynamically verified during data sharing and exchange.
[0098] Then, the data is encrypted to ensure the security of the data during sharing and exchange. An appropriate encryption algorithm is selected according to the security level of the data. For example, simple encryption or no encryption is used for "public" data, and advanced encryption algorithms (such as AES or RSA) are used for "confidential" data. The encrypted data (Encrypted_Data) is stored or transmitted together with the access permission information to ensure that only users or systems that meet the permission requirements can decrypt and access the data.
[0099] Finally, data sharing and exchange are carried out based on access permissions. The encrypted data is shared with the target user or system through a secure transmission method (such as HTTPS, SFTP, or API interface). At the data receiving end, the user or system needs to perform permission verification and decryption operations to access the data. At the same time, log information about each data access and sharing (such as access time, user identity, data category, etc.) is recorded to achieve the traceability and security supervision of operations. This mechanism ensures the security, compliance, and efficiency of data sharing.
[0100] The beneficial effects of the present invention are as follows:
[0101] (1) Through the cleaning and formatting of the original data, the present invention eliminates redundancy, errors, and inconsistencies in the data, ensuring the standardization and normalization of the data; improving the data quality and laying a solid foundation for subsequent data analysis and sharing.
[0102] (2) By combining the text context relationship feature extraction layer and the global feature extraction layer, the present invention can extract deep features of data from different dimensions: the context relationship feature extraction layer can capture the semantic associations and context information of the data. The global feature extraction layer can extract the overall features and global patterns of the data. The fusion of context features and global features further enriches the feature expression of the data, improving the accuracy and robustness of the classification model.
[0103] (3) The present invention uses a deep learning model based on BERT to classify the fused text features, making full use of the powerful semantic understanding ability of the BERT model in natural language processing. The BERT model can capture the deep semantic information of the data, significantly improving the classification accuracy and generalization ability, and is applicable to complex government data scenarios.
[0104] (4) The present invention sets access permissions for the classified data according to the preset security levels, ensuring the security and compliance of the data: data with different security levels are only accessible to users with specific permissions, preventing the leakage of sensitive data. The refined management of data access permissions meets the security requirements in government data sharing.
[0105] (5) The present invention conducts data sharing and exchange based on data access permissions, ensuring the efficient circulation of data between different departments or systems: the data sharing process follows the security level rules, avoiding unnecessary security risks. The present invention improves the data collaboration efficiency between government departments and promotes cross-departmental business collaboration.
[0106] (6) The access permission settings and log records in the data sharing and exchange process of the present invention ensure the traceability and transparency of operations: each data access and sharing operation can be recorded and audited, facilitating subsequent supervision and problem tracing. The present invention improves the transparency of government data management and enhances the standardization of data use.
[0107] (7) The present invention can adapt to the diversity and complexity of government data, supporting different types of data processing, classification, and sharing requirements. The present invention provides a flexible feature extraction and classification mechanism, applicable to a variety of government scenarios (such as public services, policy analysis, risk assessment, etc.).
[0108] (8) Through efficient data sharing and exchange, government departments can quickly obtain the required data, improving the efficiency of government services; the intelligent processing of data classification and sharing provides accurate data support for government decision-making, facilitating scientific decision-making.
[0109] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other.
[0110] In this article, specific examples are used to illustrate the principles and implementation modes of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation modes and application scopes. To sum up, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A method for sharing and exchanging government data, characterized in that: include: Collecting raw data from various government departments, and cleaning and formatting the raw data to obtain standardized data; Inputting the standardized data into a text context feature extraction layer to extract context feature information; Inputting the standardized data into a global feature extraction layer to extract global feature information; Fusing the contextual feature information and the global feature information to obtain a fused text feature; Inputting the fused text features into a BERT-based deep learning model to obtain classified data; Data access permissions are set for the classified data according to a preset security level, and data sharing and exchange are performed based on the data access permissions.
2. The government data sharing and exchange method according to claim 1 is characterized in that: The standardized data is input into the text context feature extraction layer to extract context feature information, including: Use the pre-trained language model to extract the initial feature information of the text to be classified; Inputting the initial feature information into a forward gated recurrent unit and a reverse gated recurrent unit; The outputs of the forward gated recurrent unit and the reverse gated recurrent unit are concatenated to obtain contextual relationship feature information.
3. The government data sharing and exchange method according to claim 2 is characterized in that: The outputs of the forward gated recurrent unit and the reverse gated recurrent unit are concatenated to obtain contextual feature information, including: Using formula The outputs of the forward gated recurrent unit and the reverse gated recurrent unit are concatenated to obtain contextual relationship feature information; wherein, H t ={H1,H2,…H l } represents the initial feature information input at time t, represents the output of the positive gated recurrent unit at time t, Represents the output of the reverse gated recurrent unit at time t, GRU represents the gated recurrent unit, and G represents the contextual relationship feature information.
4. The government data sharing and exchange method according to claim 3 is characterized in that: Inputting the text to be classified into the global feature extraction layer to extract global feature information includes: The initial feature information of the text to be classified is input into the convolution layer and the pooling layer in sequence to obtain the global feature information; wherein, the global feature information extraction formula is: c i =f(ω·H+b) In the formula, f is the activation function, ω is the convolution kernel, h is the convolution kernel size, b is the bias value, c is i is the feature vector extracted by the i-th convolutional layer, It represents the value of the features extracted by three different convolution kernels after the maximum pooling layer, and C represents the extracted global feature information.
5. The government data sharing and exchange method according to claim 4 is characterized in that: The context feature information and the global feature information are fused to obtain a fused text feature, including: The contextual feature information and the global feature information are spliced to obtain a spliced feature; wherein the spliced feature O is: Performing maximum pooling processing on the spliced features to obtain maximum pooling features; Performing mean pooling processing on the spliced features to obtain mean pooling features; The concatenated feature, the maximum pooling feature and the mean pooling feature are combined to obtain the fused text feature.
6. The government data sharing and exchange method according to claim 1 is characterized in that: The fused text features are input into a BERT-based deep learning model to obtain classified data, including: Inputting the fused text features into an embedding layer of a BERT-based deep learning model to encode the input fused text features and generate a context-related embedding representation; Inputting the embedded representation into a multi-layer Transformer encoder of a BERT-based deep learning model to obtain deep semantic features; The deep semantic features are input into the classification layer of the BERT-based deep learning model to obtain the classified data.
7. The government data sharing and exchange method according to claim 6 is characterized in that: The classification layer consists of a fully connected layer and a Softmax activation function.
8. The government data sharing and exchange method according to claim 1 is characterized in that: The types of data access permissions include three categories: the first category is: accessible to all users; the second category is: accessible only to specific roles; and the third category is: accessible only to authorized users.