Government affair data processing method and device based on large model, equipment and storage medium
By integrating a multimodal input human-computer interaction interface and a large model for government data processing, the problem of discovering cross-departmental and cross-level data resources has been solved, and efficient and secure government data sharing and personalized needs have been achieved.
Patent Information
- Application Number
- CN202511350785.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-22
- Publication Date
- 2025-10-31
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the current sharing of government data, it is difficult to discover and understand data resources across departments, levels, and regions. Traditional search methods are unable to meet the diverse and personalized needs of users, and the variety of forms of the same type of data resources increases the difficulty of screening and evaluation.
Data requests are received through a human-computer interaction interface that integrates multimodal inputs. Context-aware AI agents are used to clarify and refine the requests. A large model is combined for semantic understanding and entity extraction. A data-sharing database is constructed, semantic relevance is determined, and a recommended list of target data resources is generated. Quality analysis and anonymized sample generation are then performed.
It has improved the efficiency of government data processing, enhanced the user experience, and enabled the accurate, efficient processing and secure sharing of government data.
Smart Images

Figure CN120873069A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device, and storage medium for processing government data based on a large model. Background Technology
[0002] Currently, with the rapid development of information technology, data has become a fundamental strategic resource, and its importance is increasingly prominent. To fully unleash the value of data elements, it is crucial to attach great importance to and systematically promote the sharing of government data. Since 2015, a series of top-level design documents have been successively issued, providing clear policy guidance and legal guarantees for government data sharing. These documents have established the basic principles of unified catalog management and dynamic updates for government data, requiring all departments to compile and maintain their own data catalogs in accordance with standards and specifications. The documents also clarify that data catalogs should include key information such as data name, providing unit, data format, sharing attributes, sharing method, and data classification and grading.
[0003] Driven by strong policy support, significant progress has been made in the construction of the government data sharing system. According to data, as of September 2022, a preliminary government data catalog system covering the national, provincial, municipal, and county levels has been established, compiling over 3 million government data catalog entries and more than 20 million information items. Simultaneously, significant progress has been made in the construction of the integrated government data sharing hub, which has connected nearly 6,000 government departments at all levels, released a large amount of data resources, and cumulatively supported over 400 billion data sharing calls. These data fully demonstrate that the "hardware" foundation for government data sharing—namely, the platform and network system—has been basically completed, and the "assets" of data resources are becoming increasingly clear.
[0004] However, compared to the vast amount of data resources and the ever-increasing demand for sharing, the actual effectiveness of current government data sharing still faces many challenges. For data-using departments, especially when facing complex application scenarios involving multiple departments, levels, and regions, accurately and efficiently discovering and understanding the required data resources has become a prominent problem. While the data catalog is extensive, its retrieval and usage methods remain relatively traditional, failing to meet the diverse and personalized needs of users. Furthermore, the same type of data resource may exist in multiple forms such as files, databases, and API interfaces, further increasing the difficulty for users in screening and evaluating it. Therefore, how to leverage cutting-edge technologies, especially the latest achievements in artificial intelligence, to overcome these challenges and enhance the "software" capabilities of government data sharing has become crucial for promoting the in-depth development of digital infrastructure.
[0005] As can be seen from the above, how to improve the efficiency of processing government data in the process of processing government data based on large models is an urgent problem to be solved. Summary of the Invention
[0006] In view of this, the purpose of this invention is to provide a method, apparatus, device, and storage medium for processing government data based on a large model, which can improve the efficiency of processing government data in the process of processing government data based on a large model. The specific solution is as follows:
[0007] Firstly, this application provides a method for processing government data based on a large model, including:
[0008] The system receives initial data requests, including government data to be processed, through a human-computer interaction interface that integrates multimodal input functions. The initial data requests are then transformed into structured data requests to be processed. Finally, an AI agent with context awareness is used to clarify ambiguous areas and refine the information elements of the data requests to be processed, resulting in refined data requests.
[0009] The large model is used to perform semantic understanding, intent recognition and entity extraction operations on the refined data requirements to obtain target data requirements. Then, the large model is used to process the target data requirements to obtain structured query vectors. A data sharing database is constructed based on the preset government data resource catalog and the data features in the preset government data resource catalog.
[0010] The semantic relevance of the query vector to the data resource vector corresponding to each data feature in the data sharing database is determined to obtain a semantic relevance score. The query vectors corresponding to each semantic relevance score are sorted in descending order of score to obtain a target data resource recommendation list.
[0011] The government data resource library is updated based on the target data resource recommendation list. The updated government data resource library is used to perform quality analysis and target anonymization sample generation operations on the target data requirements. The resulting quality report and target anonymization sample corresponding to the government data to be processed are then displayed on the human-computer interaction interface.
[0012] Optionally, the process involves receiving initial data requests, including government data to be processed, through a human-computer interaction interface integrating multimodal input functionality. The initial data requests are then transformed into structured data requests to be processed. A context-aware AI agent is then used to clarify ambiguous areas and refine information elements within the data requests, resulting in refined data requests, including:
[0013] The system receives initial data requests through a human-computer interaction interface that integrates multimodal input functionality, and determines whether the data format corresponding to the initial data requests is a speech format. If the data format corresponding to the initial data requests is a speech format, the system uses a preset speech recognition service to convert the data corresponding to the initial data requests into text format data, thereby obtaining structured data requests to be processed.
[0014] A dialogue management system based on deep reinforcement learning is used to provide intelligent dialogue guidance in the human-computer interaction interface, so that users can perform operations such as supplementing key missing information, eliminating semantic ambiguity, and determining the range of query conditions for data requirements, thereby obtaining refined data requirements. The refined data requirements include data topic classification, time range, geographical range, and data item description information.
[0015] Optionally, the step of using the large model to perform semantic understanding, intent recognition, and entity extraction operations on the refined data requirements to obtain the target data requirements includes:
[0016] The large model is used to perform semantic understanding on the refined data requirements to obtain semantic understanding results, and a preset multi-head attention mechanism is used to determine the user intent type corresponding to the semantic understanding results; the user intent type includes data query type, directory structure browsing type, quality assessment type, and data request type.
[0017] Based on the user intent type, a corresponding entity extraction operation rule is determined, and the entity extraction operation rule is used to perform entity extraction operation on the refined data requirements to obtain entity extraction results. The target data requirements are then determined based on the user intent type and the entity extraction results. The entity extraction results include one or more of the following indicators: the qualification of the data provider, the level of sharing attributes, the data volume, and the update frequency, which correspond to the government data to be processed.
[0018] Optionally, the step of processing the target data requirement using the large model to obtain a structured query vector, and constructing a data sharing database based on a preset government data resource catalog and data features in the preset government data resource catalog, includes:
[0019] Text information extraction is performed on the target data requirements to obtain the original text information. The original text information is then subjected to deep semantic encoding processing using the large model to obtain the encoding processing result. Then, the encoding processing result is processed in a high-dimensional space using a preset vector index optimization technique to obtain a structured query vector.
[0020] Based on the target data requirements, standardized interfaces conforming to preset security standards and encryption protocols corresponding to the target data requirements are determined. Through each standardized interface and the encryption protocol, corresponding data features are determined from the data sources corresponding to each government system based on preset timed tasks and preset event triggering methods. The data features include metadata features and instance data features. The data sources include a preset government data resource catalog and a preset government data resource library.
[0021] A data sharing database is constructed based on the aforementioned data features; wherein, the metadata features include data resource naming conventions, descriptive text information, and sharing classification types; and the instance data features include data size features, field type definition features, data distribution statistics features, and update status identifier features.
[0022] Optionally, the semantic relevance of the query vector to the data resource vectors corresponding to each data feature in the data sharing database is determined to obtain a semantic relevance score. The query vectors corresponding to each semantic relevance score are then sorted in descending order of score to obtain a target data resource recommendation list, including:
[0023] The target fusion algorithm is determined based on the preset cosine similarity algorithm and the preset Euclidean distance algorithm. The target fusion algorithm is then used to determine the semantic relevance between the query vector and the data resource vectors corresponding to each data feature in the data sharing database, and the corresponding semantic relevance score is obtained.
[0024] A preset dynamic threshold filtering mechanism is used to determine the corresponding threshold determination result based on the target data requirements. The query vectors corresponding to the semantic relevance scores that are greater than the threshold determination results are set as target query vectors. An initial data resource recommendation list is determined based on each target query vector.
[0025] The target query vectors in the initial data resource recommendation list are sorted in descending order of semantic relevance scores to obtain the target data resource recommendation list, and the target data resource recommendation list is sent to the human-computer interaction interface; the target data resource recommendation list includes one or more combinations of data resource name, providing unit information, core quality indicators and secure access links.
[0026] Optionally, the step of updating the government data resource library based on the target data resource recommendation list, and using the updated government data resource library to perform quality analysis and target anonymization sample generation operations on the target data requirements, so as to display the obtained quality report and target anonymization sample corresponding to the government data to be processed on the human-computer interaction interface, includes:
[0027] The government data resource library is updated based on the target data resource recommendation list to obtain the updated government data resource library. Then, a machine learning-based data quality detection algorithm is used to perform missing value statistics and outlier detection on the data resources in the target data requirements to obtain the initial quality detection results.
[0028] The initial quality inspection results are compared with the data source that meets the preset high reliability conditions in several dimensions to obtain the comparison results. Then, the internal logical contradictions of the data resources of the target data requirement are checked using the preset logical rule engine and based on the comparison results to obtain the inspection results.
[0029] Based on historical records, timeliness indicators are determined. A preset weighting model is used to determine dimensional scores based on the timeliness indicators and the comparison results. The corresponding dimensional scores are obtained, and the target quality detection result is determined based on the dimensional scores, the initial quality detection results, and the inspection results.
[0030] Based on the target quality detection results, a quality report is determined, and an intelligent desensitization algorithm based on pattern recognition and a preset desensitization engine are used to perform content replacement, scope generalization and secure deletion operations on personal identity information, corporate trade secrets and security sensitive fields in the target data requirements to obtain an initial desensitization sample.
[0031] The distribution characteristics and structural relationships corresponding to the initial de-identified samples are determined using the large model. Then, a target de-identified sample is generated based on the distribution characteristics and structural relationships using generative adversarial network technology. The quality report and the target de-identified sample are then displayed on the human-computer interaction interface. The target de-identified sample does not contain any real sensitive information, and the data distribution is consistent.
[0032] Optionally, after displaying the obtained quality report corresponding to the government data to be processed and the target anonymization sample on the human-computer interaction interface, the method further includes:
[0033] The connection management component corresponding to the data resources in the target de-identification sample is determined, and a secure encrypted connection channel between the connection management component and the data source of the government system is established. Based on the data source characteristics and business requirements of the data source, a resource synchronization strategy and a resource update mechanism are determined, and the real-time consistency between the data sharing database and the data source is determined using the synchronization strategy and the resource update mechanism.
[0034] The connection management component supports unified access adaptation to several heterogeneous data sources and supports data extraction, data format conversion and content cleaning through a preset standardized data interface and the secure encrypted connection channel.
[0035] Secondly, this application provides a government data processing device based on a large model, comprising:
[0036] The data requirement refinement module is used to receive initial data requirements, including government data to be processed, through a human-computer interaction interface with integrated multimodal input function, and transform the initial data requirements into structured data requirements to be processed. Then, an artificial intelligence agent with context awareness is used to clarify the ambiguous areas of the data requirements to be processed and refine the information elements of the requirements to be processed, so as to obtain the refined data requirements.
[0037] The shared database construction module is used to perform semantic understanding, intent recognition and entity extraction operations on the refined data requirements using the large model to obtain target data requirements. Then, the target data requirements are processed using the large model to obtain structured query vectors, and a data sharing database is constructed based on the preset government data resource catalog and the data features in the preset government data resource catalog.
[0038] The data resource recommendation list determination module is used to determine the semantic relevance between the query vector and the data resource vectors corresponding to each data feature in the data sharing database, obtain a semantic relevance score, and sort the query vectors corresponding to each semantic relevance score in descending order of score to obtain the target data resource recommendation list.
[0039] The desensitization sample generation module is used to update the government data resource library based on the target data resource recommendation list, so as to use the updated government data resource library to perform quality analysis and target desensitization sample generation operations on the target data requirements, and display the obtained quality report and target desensitization sample corresponding to the government data to be processed on the human-computer interaction interface.
[0040] Thirdly, this application provides an electronic device, comprising:
[0041] Memory, used to store computer programs;
[0042] A processor is used to execute the computer program to implement the aforementioned large-model-based government data processing method.
[0043] Fourthly, this application provides a computer-readable storage medium for storing a computer program, wherein the computer program, when executed by a processor, implements the aforementioned large-model-based government data processing method.
[0044] As can be seen from the above, before performing government data processing based on a large model, this application needs to receive initial data requirements, including the government data to be processed, through a human-computer interaction interface integrating multimodal input functions. These initial data requirements are then transformed into structured data requirements. Next, an AI agent with context awareness clarifies ambiguous areas and refines information elements within the data requirements, resulting in refined data requirements. The large model then performs semantic understanding, intent recognition, and entity extraction on the refined data requirements to obtain target data requirements. Finally, the large model processes these target data requirements to obtain structured query vectors, which are then used to perform the processing based on preset government data requirements. A data sharing database is constructed by combining data features from a data resource catalog and a pre-defined government data resource database. The semantic relevance of query vectors with the corresponding data resource vectors for each data feature in the data sharing database is determined, resulting in a semantic relevance score. The query vectors corresponding to each semantic relevance score are then sorted in descending order to obtain a target data resource recommendation list. The government data resource database is updated based on this target data resource recommendation list. The updated database is then used to perform quality analysis and target anonymization sample generation for target data requirements. The resulting quality report and target anonymization sample corresponding to the government data to be processed are then displayed on the human-computer interaction interface.
[0045] Therefore, this application first needs to receive initial data requirements, including government data to be processed, through a human-computer interaction interface integrating multimodal input functionality, and transform the initial data requirements into structured data requirements to be processed. Then, an AI agent with context awareness is used to clarify ambiguous areas and refine the information elements of the data requirements to obtain refined data requirements. Second, a large model is used to perform semantic understanding, intent recognition, and entity extraction operations on the refined data requirements to obtain target data requirements. Then, the large model is used to process the target data requirements to obtain structured query vectors, and these vectors are then used based on a preset government data resource catalog and a preset... A data sharing database is constructed using data features from the government data resource repository. Then, the semantic relevance between query vectors and corresponding data resource vectors in the data sharing database is determined, resulting in a semantic relevance score. The query vectors corresponding to each semantic relevance score are then sorted in descending order to obtain a target data resource recommendation list. Finally, the government data resource repository is updated based on this recommendation list. The updated repository is then used to perform quality analysis and target anonymization sample generation for target data requirements. The resulting quality report and anonymized samples corresponding to the government data to be processed are then displayed on the user interface. This approach improves the efficiency of government data processing based on a large model, thereby enhancing the user experience. Attached Figure Description
[0046] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0047] Figure 1 This application discloses a flowchart of a government data processing method based on a large model.
[0048] Figure 2 This is a schematic diagram of a specific topology for government data processing based on a large model, as disclosed in this application.
[0049] Figure 3 This is a schematic diagram of a specific architecture for government data processing based on a large model disclosed in this application;
[0050] Figure 4 This is a schematic diagram of a specific user interaction flow disclosed in this application;
[0051] Figure 5 This is a schematic diagram of the structure of a government data processing device based on a large model disclosed in this application;
[0052] Figure 6 This is a structural diagram of an electronic device disclosed in this application. Detailed Implementation
[0053] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0054] Currently, with the rapid development of information technology, data has become a fundamental strategic resource, and its importance is increasingly prominent. However, compared with the vast total amount of data resources and the ever-growing demand for sharing, the actual effectiveness of current government data sharing still faces many challenges. For data-using departments, especially when facing complex application scenarios involving multiple departments, levels, and regions, accurately and efficiently discovering and understanding the required data resources has become a prominent problem. Although the data catalog is vast, its retrieval and usage methods remain relatively traditional, making it difficult to meet the diverse and personalized needs of users. Furthermore, the same type of data resource may exist in multiple forms such as files, databases, and API interfaces, further increasing the difficulty for users in screening and evaluating it. Therefore, this application provides a government data processing method based on a large model, which can improve the efficiency of processing government data in the process of large model-based government data processing.
[0055] See Figure 1 As shown in the figure, this invention discloses a government data processing method based on a large model, including:
[0056] Step S11: Receive initial data requirements including government data to be processed through a human-computer interaction interface with integrated multimodal input function, and transform the initial data requirements into structured data requirements to be processed. Then, use an AI agent with context awareness to clarify the ambiguous areas of the data requirements to be processed and refine the information elements of the requirements to obtain the refined data requirements.
[0057] In this embodiment, the overall architecture of this application mainly consists of four core modules, and the corresponding topology diagram is shown below. Figure 2 As shown: the interactive interface and question-answering robot module, the data sharing dynamic knowledge base module, the large-scale AI model module (as described in this application embodiment), and the data resource connection and synchronization module, and the architecture diagram is as follows. Figure 3 As shown, the four modules work together to form a complete closed loop from user demand input to intelligent matching and recommendation of data resources.
[0058] Furthermore, the data-sharing dynamic knowledge base module serves as the entry point for direct interaction between the user and the implementation of this application. Its core is a human-computer interaction interface mounted on the government data resource sharing portal, supporting multiple input methods, including text and voice input. When a user submits a data request via voice, this implementation first invokes integrated Automatic Speech Recognition (ASR) technology to convert the speech into text. Subsequently, the text is sent to the backend question-answering robot. It is worth noting that this question-answering robot is not a simple keyword matching program, but an intelligent agent developed based on Artificial Intelligence Agent (Automated Agent) technology. It can understand the user's context, conduct multi-turn dialogues, and gradually clarify and refine the user's data needs, thereby collecting more accurate and complete information about the user's requirements.
[0059] Specifically, the process involves receiving initial data requests, including government data to be processed, through a human-computer interaction interface integrating multimodal input functionality. These initial data requests are then transformed into structured data requests. A context-aware AI agent is then used to clarify ambiguous areas and refine information elements within the data requests, resulting in refined data requests. This process may include: receiving initial data requests through the multimodal input interface and determining whether the data format is speech. If speech is present, a pre-defined speech recognition service is used to convert the data into text, resulting in structured data requests. A deep reinforcement learning-based dialogue management system is then used to guide intelligent dialogue within the interface, enabling users to perform operations such as supplementing missing information, eliminating semantic ambiguity, and defining query criteria, thus obtaining refined data requests. These refined data requests include data subject classification, time range, geographical range, and data item descriptions.
[0060] Step S12: Use the large model to perform semantic understanding, intent recognition and entity extraction operations on the refined data requirements to obtain target data requirements. Then, use the large model to process the target data requirements to obtain structured query vectors, and construct a data sharing database based on the data features in the preset government data resource catalog and the preset government data resource database.
[0061] In this embodiment, the AI (Artificial Intelligence) model module is the core driving force of this embodiment, integrating the powerful capabilities of a Large Language Model (LLM). It is primarily responsible for two key tasks: first, deep semantic understanding and transformation of user needs; and second, intelligent analysis and processing of data resources. Upon receiving user requests from the interactive interface, the large model performs semantic parsing, intent recognition, and slot filling to transform ambiguous natural language requests into structured query conditions. Simultaneously, the AI model module is also responsible for in-depth processing of information obtained from the data resource repository. For example, it performs multi-dimensional analysis of data quality to generate quality rating reports; or, based on user needs, automatically generates anonymized data samples that comply with security standards.
[0062] Subsequently, after a user submits a data request through the interactive interface, this embodiment of the application will initiate an intelligent recognition and semantic conversion process, as follows: First, if the data request corresponds to voice input, this embodiment of the application will call a voice recognition service to convert it into text. Then, the text is fed into a large-scale language model so that the model can perform intent recognition on the text and determine whether the user wants to query specific data, understand the data directory structure, or conduct data quality assessment, etc. Next, the model will perform entity extraction, identifying key information from the text such as data topic (e.g., "population" or "enterprise"), time range (e.g., "2025" or "last five years"), geographical range (e.g., "a certain city" or "a certain district"), and data items (e.g., "operating revenue" or "population size"). Finally, the model combines these extracted entities and the recognized intent into a structured query request, preparing for subsequent retrieval and matching.
[0063] Specifically, the step of using the large model to perform semantic understanding, intent recognition, and entity extraction operations on the refined data requirements to obtain the target data requirements may include: using the large model to perform semantic understanding on the refined data requirements to obtain semantic understanding results, and using a preset multi-head attention mechanism to determine the user intent type corresponding to the semantic understanding results; the user intent type includes data query type, directory structure browsing type, quality assessment type, and data application type; determining the corresponding entity extraction operation rules based on the user intent type, and using the entity extraction operation rules to perform entity extraction operations on the refined data requirements to obtain entity extraction results, and determining the target data requirements based on the user intent type and the entity extraction results; wherein, the entity extraction results include one or more combinations of indicators such as the data provider qualification, sharing attribute level, data volume scale, and update frequency corresponding to the government data to be processed.
[0064] Furthermore, the data sharing dynamic knowledge base module is the "intelligent brain" in this embodiment, responsible for storing and managing all knowledge related to government data sharing. It is not a static database, but a dynamically updated and continuously learning knowledge base. Its core content originates from the joint operation and in-depth analysis of the government data resource catalog and the data resource database. This embodiment uses a large-scale model processor and database connection technologies such as JDBC (Java Database Connectivity) to periodically or in real-time obtain the latest data catalog information (such as data resource name, providing unit, data item, sharing method, etc.) and the core indicators of the data resources themselves (such as data volume, data version, update frequency, etc.) from the data source. After vectorization processing by the large model, this information is constructed into a high-dimensional vector knowledge base, providing strong semantic support for subsequent intelligent retrieval and matching.
[0065] Furthermore, the data resource connection and synchronization module serves as a bridge between the embodiments of this application and the underlying government data resources, responsible for establishing and maintaining secure and stable connections with government data resource catalogs and resource repositories at all levels. It achieves unified access to heterogeneous data sources through standardized interfaces (such as JDBC) and protocols. This module has a built-in flexible synchronization mechanism, which can configure scheduled or real-time data synchronization tasks according to the update frequency of the data source and business needs, ensuring that the information in the dynamic knowledge base remains consistent with the source data. It is worth mentioning that the aforementioned direct and dynamic connection method avoids the information inconsistency problems caused by data copying and delayed synchronization in traditional platforms, providing a solid data foundation for the intelligence and accuracy of the entire embodiments of this application.
[0066] Specifically, the process of using the large model to process the target data requirement to obtain a structured query vector, and constructing a data sharing database based on the data features in the preset government data resource catalog and preset government data resource library, may include: extracting text information from the target data requirement to obtain raw text information, performing deep semantic encoding on the raw text information using the large model to obtain encoding results, and then processing each encoding result in a high-dimensional space using preset vector index optimization technology to obtain a structured query vector; determining standardized interfaces that conform to preset security standards and encryption protocols corresponding to the target data requirement based on the target data requirement, so as to determine corresponding data features from the data sources corresponding to each government system through each standardized interface and the encryption protocol based on preset timed tasks and preset event triggering methods; the data features include metadata features and instance data features; the data sources include the preset government data resource catalog and preset government data resource library; constructing a data sharing database based on each of the data features; wherein, the metadata features include data resource naming conventions, descriptive text information, and sharing classification types; the instance data features include data size features, field type definition features, data distribution statistics features, and update status identifier features.
[0067] Step S13: Determine the semantic relevance between the query vector and the data resource vectors corresponding to each data feature in the data sharing database, obtain a semantic relevance score, and sort the query vectors corresponding to each semantic relevance score in descending order of score to obtain a target data resource recommendation list.
[0068] In this embodiment, after completing the semantic transformation of user needs and constructing a dynamic knowledge base, the embodiment proceeds to the core intelligent retrieval and demand matching stage. First, the structured query request is vectorized to generate a corresponding query vector. Then, the similarity (e.g., cosine similarity) between the query vector and all data resource vectors in the knowledge base is calculated to determine the most relevant data resources to the user's needs. This vector similarity-based retrieval method surpasses traditional keyword matching, understanding the deeper meaning of user needs. Even if the query terms do not perfectly match the description in the data directory, semantically relevant resources can be found. Subsequently, the retrieved results are sorted according to similarity scores, and a recommendation list containing the data resource name, providing organization, and core indicators is returned to the user.
[0069] Specifically, determining the semantic relevance between the query vector and the data resource vectors corresponding to each data feature in the data sharing database to obtain a semantic relevance score, and sorting the query vectors corresponding to each semantic relevance score in descending order to obtain a target data resource recommendation list, may include: determining a target fusion algorithm based on a preset cosine similarity algorithm and a preset Euclidean distance algorithm, using the target fusion algorithm to determine the semantic relevance between the query vector and the data resource vectors corresponding to each data feature in the data sharing database to obtain the corresponding semantic relevance score; and using a preset dynamic threshold filtering mechanism and based on... A threshold determination result is determined for the target data requirement, and the query vectors corresponding to the semantic relevance scores that are greater than the threshold determination result are set as target query vectors. An initial data resource recommendation list is determined based on each target query vector. The target query vectors in the initial data resource recommendation list are sorted in descending order of semantic relevance scores to obtain a target data resource recommendation list, which is then sent to the human-computer interaction interface. The target data resource recommendation list includes one or more combinations of information such as data resource name, providing unit information, core quality indicators, and secure access links.
[0070] Step S14: Update the government data resource library based on the target data resource recommendation list, and use the updated government data resource library to perform quality analysis and target anonymization sample generation operations on the target data requirements, so as to display the obtained quality report and target anonymization sample corresponding to the government data to be processed on the human-computer interaction interface.
[0071] In this embodiment, to help users better assess the availability of data resources, this application embodiment also provides a data quality intelligent analysis and rating function. When this application embodiment identifies that a user is interested in a specific data resource, the data quality intelligent analysis and rating function will be triggered. That is, the large model will analyze the data resource from multiple preset dimensions, such as: completeness (checking whether there are a large number of missing values in the data), accuracy (by comparing with authoritative data sources or performing logical verification), consistency (checking whether there are contradictions within the data or with other related data), and timeliness (assessing the frequency and delay of data updates), etc. After the data analysis is completed, this application embodiment will score each dimension according to preset weights and algorithms, and finally generate a comprehensive data quality rating (such as "excellent" or "poor"). This rating result will be presented to the user in an intuitive way to provide important reference for their decision-making.
[0072] Furthermore, to resolve the conflict between security and convenience in data sharing, this application proposes an innovative data sample sharing mechanism. When a user needs to view the specific content of a data resource to assess its applicability, this application does not directly provide the original data. Instead, it generates a rigorously anonymized sample dataset using large-scale modeling technology. That is, this application can call specialized data anonymization algorithms to replace, generalize, or delete sensitive information (such as personal identification information, corporate trade secrets, etc.) in the original data. Alternatively, leveraging the generative capabilities of large-scale models, based on an understanding of the original data structure, a batch of "synthetic data" with similar statistical characteristics to the original data but completely devoid of any real sensitive information can be generated. These anonymized or synthesized sample data not only help users understand the data format and content but also fundamentally eliminate the risk of sensitive information leakage, achieving a balance between security and convenience. In one specific implementation, a schematic diagram of the interaction process between the user and this application is shown below. Figure 4 As shown.
[0073] Specifically, the step of updating the government data resource library based on the target data resource recommendation list, and using the updated government data resource library to perform quality analysis and target anonymization sample generation operations on the target data requirements, so as to display the obtained quality report and target anonymization sample corresponding to the government data to be processed on the human-computer interaction interface, may include: updating the government data resource library based on the target data resource recommendation list to obtain the updated government data resource library, and using a machine learning-based data quality detection algorithm to perform missing value statistics and outlier detection on the data resources in the target data requirements to obtain an initial quality detection result; comparing the initial quality detection result with a data source that meets preset high reliability conditions in several dimensions to obtain a comparison result, and then using a preset logical rule engine and based on the comparison result to check the internal logical contradictions of the data resources of the target data requirements to obtain a check result; and determining timeliness based on historical records. The system uses a preset weighting model to determine dimensional scores based on the timeliness indicators and the comparison results, obtaining corresponding dimensional scores. Based on these dimensional scores, the initial quality inspection results, and the inspection results, a target quality inspection result is determined. A quality report is generated based on the target quality inspection result. An intelligent de-identification algorithm based on pattern recognition and a preset de-identification engine are used to perform content replacement, scope generalization, and secure deletion operations on personal identity information, corporate trade secrets, and security-sensitive fields in the target data requirements, resulting in an initial de-identification sample. The large model is used to determine the distribution characteristics and structural relationships corresponding to the initial de-identification sample. Then, generative adversarial network technology is used to generate a target de-identification sample based on the distribution characteristics and structural relationships. The quality report and the target de-identification sample are displayed on the human-computer interaction interface. The target de-identification sample does not contain real sensitive information, and the data distribution is consistent.
[0074] Furthermore, after displaying the obtained quality report corresponding to the government data to be processed and the target anonymized sample to the human-computer interaction interface, the process may further include: determining the connection management component corresponding to the data resources in the target anonymized sample, establishing a secure encrypted connection channel between the connection management component and the data source of the government system, and determining a resource synchronization strategy and resource update mechanism based on the data source characteristics and business requirements of the data source, so as to determine the real-time consistency between the data sharing database and the data in the data source using the synchronization strategy and resource update mechanism; wherein, the connection management component supports unified access adaptation to several heterogeneous data sources, and supports data extraction, data format conversion and content cleaning through a preset standardized data interface and the secure encrypted connection channel.
[0075] It is worth mentioning that this application's embodiments choose to analyze two mainstream data sharing models—front-end server switching and Web service interfaces—to deeply explore their technical principles, application scenarios, and inherent limitations. Subsequently, the technical solutions of this application's embodiments are compared with these existing technologies from multiple dimensions and in-depth levels. From aspects such as intelligent retrieval, dynamic knowledge base, intelligent question answering, data quality analysis, secure sharing, and the architecture of this application's embodiments, the application comprehensively demonstrates how this invention solves the pain points of existing technologies and brings significant beneficial effects. Finally, specific embodiments will further illustrate the feasibility and value of this invention in practical applications.
[0076] In one specific implementation, this application analyzes the implementation methods of existing government data sharing platforms. With the continuous advancement of "digital" construction, breaking down "information silos" between departments and achieving efficient sharing and exchange of government data has become crucial for improving governance capabilities and public service levels. To address this challenge, various regions and relevant departments have invested significant resources in building various types of government data sharing platforms. It is worth noting that these platforms primarily rely on two traditional implementation methods: front-end machine exchange and Web service interfaces. While these two methods made significant contributions to the interconnection of government data in specific historical periods and application scenarios, their inherent limitations have become increasingly prominent with the explosive growth of data volume and the increasing complexity of application demands, becoming a bottleneck restricting the further release of the value of government data.
[0077] Front-end server exchange is a relatively traditional data sharing model. Its core idea is to establish a physical or virtual "transfer station," or front-end server, between the data provider and the data requester. This front-end server serves as a window for information sharing and a relay station for data exchange between departments, and is an important component of the data sharing and exchange platform. Its basic workflow is as follows: The data provider first uses a bridging tool to extract shared data from its department's business database and push it to the front-end server deployed within its department. Subsequently, the data sharing and exchange platform uses middleware to convert the data format in the front-end server into a format that the data requester can read and push it to the requester's front-end server. Finally, the requester imports the data into its local business system. In other words, this implementation method is very suitable for large-volume batch data interaction scenarios. When the amount of data to be exchanged is huge, such as during historical data migration or large-scale centralized data cleaning and integration, the front-end server exchange method can provide a stable and reliable data transmission channel. Secondly, this method supports data storage on-premises, facilitating subsequent offline analysis and processing. Furthermore, the front-end server mode demonstrates good adaptability when a single data collection needs to serve multiple data requesters, or when mixed exchange of data and files is required. Especially when data exchange is required across network domains (such as between government intranets and extranets), the front-end server can act as a secure isolation zone, ensuring network security to a certain extent.
[0078] In another specific implementation, Web service interfaces are another widely used data sharing method. They encapsulate data into standardized service interfaces, enabling application integration and data exchange between heterogeneous embodiments of this application. Unlike front-end machines that operate directly at the data layer, Web service interfaces interact at the application layer. Data providers define and develop public data service interfaces based on sharing needs, encapsulating the content and protocols of data exchange. Data requesters then obtain the required data in the form of services by calling these interfaces. This approach is typically based on open technology standards such as SOAP (Simple Object Access Protocol) and XML (Extensible Markup Language), featuring platform independence, loose coupling, and self-containment, enabling interoperability between different machines and applications. It is worth noting that the main advantages of Web service interfaces lie in their flexibility and real-time performance. They are particularly suitable for scenarios with relatively small data volumes but high real-time requirements for data transmission, such as cross-departmental data queries and information verification applications. By making data available as a service, the fast, efficient, and diverse data service collection needs across departments, domains, and heterogeneous sources can be easily met. This approach embeds data deep within an independent and closed system, providing services externally through standardized interfaces, thus reducing direct dependencies between systems.
[0079] In this embodiment, the present application constructs a highly accurate, timely, low-cost, and secure intelligent inter-agency data response system for government affairs. RAG technology, by combining information retrieval with text generation, effectively solves the "illusion" problem that may occur with large models and significantly reduces the cost of model training and updates.
[0080] In the first specific implementation, the specific implementation process is as follows: First, a knowledge base is constructed. The system first slices a large number of unstructured or semi-structured documents, such as government data catalogs, policies and regulations, service guides, and data dictionaries. Then, it uses an embedding model to convert these text fragments into high-dimensional vectors and stores them in a vector database, thus constructing a localized, dynamically updated government knowledge base. Second, data requests or business inquiries are submitted to the question-answering robot via natural language (text or voice). After receiving the user's question in this embodiment, the question is also converted into a vector using an embedding model. Then, a similarity search is performed in the vector database to quickly find the knowledge fragment most relevant to the question. Then, this embodiment uses the retrieved knowledge fragment as context and inputs it along with the user's question into a locally deployed Large Language Model (LLM). The LLM generates accurate, reliable, and verifiable answers based on this precise and up-to-date knowledge. Finally, this embodiment returns the generated answer to the user and can simultaneously display the knowledge sources on which the answer is based, enhancing the credibility and transparency of the answer.
[0081] In a second specific implementation, this application embodiment uses LLM data augmentation for government affairs retrieval, aiming to solve the problems of "inaccurate and incomplete searches" caused by insufficient training data and semantic understanding bias in government affairs data retrieval. By utilizing the powerful generative capabilities of Large Language Models (LLM), the limited labeled data is augmented, thereby training a more accurate and discriminative semantic retrieval model. The specific implementation process is as follows: First, addressing the pain point of scarce labeled data in the government affairs field, this application's embodiment adopts structured prompt engineering and few-shot examples to guide LLM to generate a large number of positive and negative sample data pairs that are highly relevant to government affairs retrieval tasks. For example, given a real question, "How to transfer social security?", LLM can generate semantically similar positive samples (such as "How to transfer social security relationship to another place?", "What to do about social security when working across provinces?") and irrelevant negative samples (such as "How to apply for a housing provident fund loan?"). Subsequently, the model is fine-tuned, that is, using these high-quality, large-scale triple data (query, positive sample, negative sample) generated by LLM, on a pre-trained semantic model, such as ROBERT (Robustly Optimized BERT Pretraining Approach), a supervised contrastive learning framework, such as SimCSE (Simple Contrastive Sentence), is adopted. Embeddings (i.e., simple contrastive sentence vector embeddings) are used for fine-tuning. Simultaneously, to reduce the interference of potential "spurious negative samples" in the generated data on model training, this embodiment also designs a dynamic weighted masking mechanism and a bias-free contrastive loss function to further improve the model's robustness. It is worth mentioning that the fine-tuned model can map the descriptive text of government issues and data resources into a more uniform and discriminative semantic space. When a user inputs a query, this embodiment can quickly and accurately retrieve the semantically most matching result from massive data resources.
[0082] In the third specific implementation, this application embodiment combines proactive data services with an AI Agent. The goal of this embodiment is to promote a paradigm shift in government services from "passive response" to "proactive service." By introducing AI Agent technology, this application embodiment can proactively perceive user needs, predict user intentions, and provide personalized data recommendations and services. The specific implementation process is as follows: First, user profile construction is performed. That is, this application embodiment analyzes users' historical query records, browsing behavior, downloaded data types, and other information, and uses machine learning algorithms to build a dynamic user profile for each user, tagging users' interest areas, business needs, etc. Second, intent prediction and proactive push are performed: When new data resources are added to the database, or existing data resources are updated, the AI Agent will proactively determine which users may be interested in the data based on the user profile. For example, when new "tax incentive policies for high-tech enterprises" data is released, this application embodiment can proactively push it to users who frequently query related data such as "enterprise registration" and "tax processing." Subsequently, multi-round dialogue and task execution are performed: The AI Agent has the ability to conduct multi-round dialogues and invoke tools. Users can issue complex commands using natural language, such as "Find all policy documents, company directories, and patent data related to the 'artificial intelligence industry,' and generate an analysis report." Finally, the AI Agent will break down this complex task into multiple sub-tasks, autonomously invoking tools for data retrieval, data analysis, and report generation to ultimately complete the user-specified task.
[0083] As can be seen from the above, the embodiments of this application first need to receive initial data requirements, including government data to be processed, through a human-computer interaction interface integrating multimodal input functions, and transform the initial data requirements into structured data requirements to be processed. Then, an artificial intelligence agent with context awareness is used to clarify ambiguous areas and refine the information elements of the data requirements to be processed, resulting in refined data requirements. Secondly, a large model is used to perform semantic understanding, intent recognition, and entity extraction operations on the refined data requirements to obtain target data requirements. Then, the large model is used to process the target data requirements to obtain structured query vectors, and based on a preset government data resource catalog and... A data-sharing database is constructed based on data features from a pre-defined government data resource repository. Then, the semantic relevance between query vectors and corresponding data resource vectors in the data-sharing database is determined, resulting in a semantic relevance score. The query vectors corresponding to each semantic relevance score are then sorted in descending order to obtain a target data resource recommendation list. Finally, the government data resource repository is updated based on this recommendation list. The updated repository is then used to perform quality analysis and target anonymization sample generation for target data requirements. The resulting quality report and anonymized samples corresponding to the government data to be processed are then displayed on the user interface. This approach improves the efficiency of government data processing in large-scale model-based government data processing, thereby enhancing the user experience.
[0084] Accordingly, see Figure 5 As shown, this application also provides a government data processing device based on a large model, including:
[0085] The data requirement refinement module 11 is used to receive initial data requirements including government data to be processed through a human-computer interaction interface with integrated multimodal input function, and transform the initial data requirements into structured data requirements to be processed. Then, it uses an artificial intelligence agent with context awareness to clarify the ambiguous areas of the data requirements to be processed and refine the information elements of the requirements to be processed, so as to obtain the refined data requirements.
[0086] The shared database construction module 12 is used to perform semantic understanding, intent recognition and entity extraction operations on the refined data requirements using the large model to obtain target data requirements. Then, the target data requirements are processed using the large model to obtain structured query vectors, and a data sharing database is constructed based on the preset government data resource catalog and the data features in the preset government data resource catalog.
[0087] The data resource recommendation list determination module 13 is used to determine the semantic relevance between the query vector and the data resource vectors corresponding to each data feature in the data sharing database, obtain a semantic relevance score, and sort the query vectors corresponding to each semantic relevance score in descending order of score to obtain a target data resource recommendation list.
[0088] The desensitization sample generation module 14 is used to update the government data resource library based on the target data resource recommendation list, so as to use the updated government data resource library to perform quality analysis and target desensitization sample generation operations on the target data requirements, and to display the obtained quality report and target desensitization sample corresponding to the government data to be processed on the human-computer interaction interface.
[0089] In some specific embodiments, the data requirement refinement processing module 11 may specifically include:
[0090] The data format determination unit is used to receive initial data requirements through a human-computer interaction interface with integrated multimodal input function, and determine whether the data format corresponding to the initial data requirements is a voice format. If the data format corresponding to the initial data requirements is a voice format, the data corresponding to the initial data requirements is converted into text format data using a preset voice recognition service to obtain structured data requirements to be processed.
[0091] The data requirement refinement subunit is used to conduct intelligent dialogue guidance on the human-computer interaction interface using a dialogue management system based on deep reinforcement learning, so that users can perform operations such as supplementing key missing information, eliminating semantic ambiguity, and determining the range of query conditions for data requirements, thereby obtaining refined data requirements; the refined data requirements include data topic classification, time range, geographical range, and data item description information.
[0092] In some specific embodiments, the shared database construction module 12 may specifically include:
[0093] The user intent type determination unit is used to perform semantic understanding on the refined data requirements using the large model, obtain semantic understanding results, and use a preset multi-head attention mechanism to determine the user intent type corresponding to the semantic understanding results; the user intent type includes data query type, directory structure browsing type, quality assessment type, and data request type.
[0094] The entity extraction result determination unit is used to determine the corresponding entity extraction operation rules based on the user intent type, and to perform entity extraction operations on the refined data requirements using the entity extraction operation rules to obtain entity extraction results, and to determine the target data requirements based on the user intent type and the entity extraction results; wherein, the entity extraction results include one or more combinations of indicators such as the data provider qualification, sharing attribute level, data volume scale and update frequency corresponding to the government data to be processed.
[0095] In some specific embodiments, the shared database construction module 12 may specifically include:
[0096] The query vector generation unit is used to extract text information from the target data requirements to obtain the original text information, and to perform deep semantic encoding on the original text information using the large model to obtain the encoding processing result. Then, the preset vector index optimization technology is used to process each of the encoding processing results in a high-dimensional space to obtain a structured query vector.
[0097] The data feature determination unit is used to determine standardized interfaces that conform to preset security standards and encryption protocols corresponding to the target data requirements based on the target data requirements. Then, through each standardized interface and the encryption protocol, and based on preset timed tasks and preset event triggering methods, it determines corresponding data features from the data sources corresponding to each government system. The data features include metadata features and instance data features. The data sources include a preset government data resource catalog and a preset government data resource library.
[0098] The shared database construction subunit is used to construct a data sharing database based on the aforementioned data characteristics; wherein, the metadata characteristics include data resource naming conventions, descriptive text information, and sharing classification types; the instance data characteristics include data size characteristics, field type definition characteristics, data distribution statistics characteristics, and update status identifier characteristics.
[0099] In some specific embodiments, the data resource recommendation list determination module 13 may specifically include:
[0100] The semantic relevance score determination unit is used to determine the target fusion algorithm based on a preset cosine similarity algorithm and a preset Euclidean distance algorithm, so as to use the target fusion algorithm to determine the semantic relevance between the query vector and the data resource vectors corresponding to each data feature in the data sharing database, and obtain the corresponding semantic relevance score.
[0101] The initial data resource recommendation list determination unit is used to determine the corresponding threshold determination result based on the target data requirements using a preset dynamic threshold filtering mechanism, and to set the query vectors corresponding to the semantic relevance scores that are greater than the threshold determination results as target query vectors, and to determine the initial data resource recommendation list based on each target query vector;
[0102] The resource recommendation list distribution unit is used to sort the target query vectors in the initial data resource recommendation list in descending order of semantic relevance score to obtain the target data resource recommendation list, and distribute the target data resource recommendation list to the human-computer interaction interface; the target data resource recommendation list includes one or more combinations of data resource name, providing unit information, core quality indicators and secure access links.
[0103] In some specific embodiments, the desensitized sample generation module 14 may specifically include:
[0104] An outlier detection unit is used to update the government data resource library based on the target data resource recommendation list, obtain the updated government data resource library, and use a machine learning-based data quality detection algorithm to perform missing value statistics and outlier detection on the data resources in the target data requirements to obtain the initial quality detection results.
[0105] The inspection result determination unit is used to compare the initial quality inspection result with the data source that meets the preset high reliability conditions in several dimensions to obtain the comparison result, and then use the preset logical rule engine to check the internal logical contradictions of the data resources of the target data requirement based on the comparison result to obtain the inspection result.
[0106] The quality inspection result determination unit is used to determine the timeliness index based on historical records, and to determine the dimension score by using a preset weight model and based on the timeliness index and each of the comparison results, so as to obtain the corresponding dimension score, and to determine the target quality inspection result based on the dimension score, the initial quality inspection result and the inspection result.
[0107] The initial desensitization sample determination unit is used to determine a quality report based on the target quality detection results, and to use a pattern recognition-based intelligent desensitization algorithm and a preset desensitization engine to perform content replacement, scope generalization and secure deletion operations on personal identity information, enterprise trade secrets and security sensitive fields in the target data requirements to obtain an initial desensitization sample;
[0108] The target desensitization sample determination unit is used to determine the distribution characteristics and structural relationships corresponding to the initial desensitization sample using the large model, and then generate target desensitization samples using generative adversarial network technology based on the distribution characteristics and structural relationships, so as to display the quality report and the target desensitization samples to the human-computer interaction interface; the target desensitization samples do not contain real sensitive information and the data distribution is consistent.
[0109] In some specific embodiments, the government data processing device based on a large model may further include:
[0110] The connection management component determination unit is used to determine the connection management component corresponding to the data resources in the target de-identification sample, so as to establish a secure and encrypted connection channel with the data source of the government system using the connection management component, and determine the resource synchronization strategy and resource update mechanism based on the data source characteristics and business requirements, so as to determine the real-time consistency of the data sharing database and the data in the data source using the synchronization strategy and resource update mechanism; wherein, the connection management component supports unified access adaptation to several heterogeneous data sources, and supports data extraction, data format conversion and content cleaning through a preset standardized data interface and the secure and encrypted connection channel.
[0111] Furthermore, embodiments of this application also disclose an electronic device, Figure 6 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application. Specifically, the electronic device 20 may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the large-model-based government data processing method disclosed in any of the foregoing embodiments. Furthermore, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0112] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.
[0113] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon can include operating system 221, computer program 222, etc., and the storage method can be temporary storage or permanent storage.
[0114] The operating system 221 is used to manage and control the various hardware devices on the electronic device 20 and the computer program 222, which may be Windows Server, Netware, Unix, Linux, etc. In addition to including computer programs capable of performing the large-model-based government data processing method executed by the electronic device 20 as disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs capable of performing other specific tasks.
[0115] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned disclosed method for processing government data based on a large model. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.
[0116] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.
[0117] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0118] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0119] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0120] The technical solutions provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A method for processing government data based on a large model, characterized in that, include: The system receives initial data requests, including government data to be processed, through a human-computer interaction interface that integrates multimodal input functions. The initial data requests are then transformed into structured data requests to be processed. Finally, an AI agent with context awareness is used to clarify ambiguous areas and refine the information elements of the data requests to be processed, resulting in refined data requests. The large model is used to perform semantic understanding, intent recognition and entity extraction operations on the refined data requirements to obtain target data requirements. Then, the large model is used to process the target data requirements to obtain structured query vectors. A data sharing database is constructed based on the preset government data resource catalog and the data features in the preset government data resource catalog. The semantic relevance of the query vector to the data resource vector corresponding to each data feature in the data sharing database is determined to obtain a semantic relevance score. The query vectors corresponding to each semantic relevance score are sorted in descending order of score to obtain a target data resource recommendation list. The government data resource library is updated based on the target data resource recommendation list. The updated government data resource library is used to perform quality analysis and target anonymization sample generation operations on the target data requirements. The resulting quality report and target anonymization sample corresponding to the government data to be processed are then displayed on the human-computer interaction interface.
2. The government data processing method based on a large model according to claim 1, characterized in that, The system receives initial data requests, including government data to be processed, through a human-computer interaction interface integrating multimodal input functionality. It then transforms these initial data requests into structured data requests for processing. Next, an AI agent with context-aware capabilities clarifies ambiguous areas and refines information elements within the data requests, resulting in refined data requests, including: The system receives initial data requests through a human-computer interaction interface that integrates multimodal input functionality, and determines whether the data format corresponding to the initial data requests is a speech format. If the data format corresponding to the initial data requests is a speech format, the system uses a preset speech recognition service to convert the data corresponding to the initial data requests into text format data, thereby obtaining structured data requests to be processed. A dialogue management system based on deep reinforcement learning is used to provide intelligent dialogue guidance in the human-computer interaction interface, so that users can perform operations such as supplementing key missing information, eliminating semantic ambiguity, and determining the range of query conditions for data requirements, thereby obtaining refined data requirements. The refined data requirements include data topic classification, time range, geographical range, and data item description information.
3. The government data processing method based on a large model according to claim 1, characterized in that, The process of using the large model to perform semantic understanding, intent recognition, and entity extraction on the refined data requirements to obtain the target data requirements includes: The large model is used to perform semantic understanding on the refined data requirements to obtain semantic understanding results, and a preset multi-head attention mechanism is used to determine the user intent type corresponding to the semantic understanding results; the user intent type includes data query type, directory structure browsing type, quality assessment type, and data request type. Based on the user intent type, a corresponding entity extraction operation rule is determined, and the entity extraction operation rule is used to perform entity extraction operation on the refined data requirements to obtain entity extraction results. The target data requirements are then determined based on the user intent type and the entity extraction results. The entity extraction results include one or more of the following indicators: the qualification of the data provider, the level of sharing attributes, the data volume, and the update frequency, which correspond to the government data to be processed.
4. The government data processing method based on a large model according to claim 1, characterized in that, The process of using the large model to process the target data requirements to obtain a structured query vector, and constructing a data sharing database based on a preset government data resource catalog and data features in the preset government data resource catalog, includes: Text information extraction is performed on the target data requirements to obtain the original text information. The original text information is then subjected to deep semantic encoding processing using the large model to obtain the encoding processing result. Then, the encoding processing result is processed in a high-dimensional space using a preset vector index optimization technique to obtain a structured query vector. Based on the target data requirements, standardized interfaces conforming to preset security standards and encryption protocols corresponding to the target data requirements are determined. Through each standardized interface and the encryption protocol, corresponding data features are determined from the data sources corresponding to each government system based on preset timed tasks and preset event triggering methods. The data features include metadata features and instance data features. The data sources include a preset government data resource catalog and a preset government data resource library. A data sharing database is constructed based on the aforementioned data features; wherein, the metadata features include data resource naming conventions, descriptive text information, and sharing classification types; and the instance data features include data size features, field type definition features, data distribution statistics features, and update status identifier features.
5. The government data processing method based on a large model according to claim 1, characterized in that, The semantic relevance of the query vector to the data resource vectors corresponding to each data feature in the data sharing database is determined to obtain a semantic relevance score. The query vectors corresponding to each semantic relevance score are then sorted in descending order of score to obtain a target data resource recommendation list, including: The target fusion algorithm is determined based on the preset cosine similarity algorithm and the preset Euclidean distance algorithm. The target fusion algorithm is then used to determine the semantic relevance between the query vector and the data resource vectors corresponding to each data feature in the data sharing database, and the corresponding semantic relevance score is obtained. A preset dynamic threshold filtering mechanism is used to determine the corresponding threshold determination result based on the target data requirements. The query vectors corresponding to the semantic relevance scores that are greater than the threshold determination results are set as target query vectors. An initial data resource recommendation list is determined based on each target query vector. The target query vectors in the initial data resource recommendation list are sorted in descending order of semantic relevance scores to obtain the target data resource recommendation list, and the target data resource recommendation list is sent to the human-computer interaction interface; the target data resource recommendation list includes one or more combinations of data resource name, providing unit information, core quality indicators and secure access links.
6. The government data processing method based on a large model according to claim 1, characterized in that, The step of updating the government data resource library based on the target data resource recommendation list, and using the updated government data resource library to perform quality analysis and target anonymization sample generation operations on the target data requirements, and displaying the resulting quality report and target anonymization sample corresponding to the government data to be processed on the human-computer interaction interface, includes: The government data resource library is updated based on the target data resource recommendation list to obtain the updated government data resource library. Then, a machine learning-based data quality detection algorithm is used to perform missing value statistics and outlier detection on the data resources in the target data requirements to obtain the initial quality detection results. The initial quality inspection results are compared with the data source that meets the preset high reliability conditions in several dimensions to obtain the comparison results. Then, the internal logical contradictions of the data resources of the target data requirement are checked using the preset logical rule engine and based on the comparison results to obtain the inspection results. Based on historical records, timeliness indicators are determined. A preset weighting model is used to determine dimensional scores based on the timeliness indicators and the comparison results. The corresponding dimensional scores are obtained, and the target quality detection result is determined based on the dimensional scores, the initial quality detection results, and the inspection results. Based on the target quality detection results, a quality report is determined, and an intelligent desensitization algorithm based on pattern recognition and a preset desensitization engine are used to perform content replacement, scope generalization and secure deletion operations on personal identity information, corporate trade secrets and security sensitive fields in the target data requirements to obtain an initial desensitization sample. The distribution characteristics and structural relationships corresponding to the initial de-identified samples are determined using the large model. Then, a target de-identified sample is generated based on the distribution characteristics and structural relationships using generative adversarial network technology. The quality report and the target de-identified sample are then displayed on the human-computer interaction interface. The target de-identified sample does not contain any real sensitive information, and the data distribution is consistent.
7. The government data processing method based on a large model according to claim 1, characterized in that, After displaying the obtained quality report corresponding to the government data to be processed and the target anonymization sample on the human-computer interaction interface, the process further includes: The connection management component corresponding to the data resources in the target de-identification sample is determined, and a secure encrypted connection channel between the connection management component and the data source of the government system is established. Based on the data source characteristics and business requirements of the data source, a resource synchronization strategy and a resource update mechanism are determined, and the real-time consistency between the data sharing database and the data source is determined using the synchronization strategy and the resource update mechanism. The connection management component supports unified access adaptation to several heterogeneous data sources and supports data extraction, data format conversion and content cleaning through a preset standardized data interface and the secure encrypted connection channel.
8. A government data processing device based on a large model, characterized in that, include: The data requirement refinement module is used to receive initial data requirements, including government data to be processed, through a human-computer interaction interface with integrated multimodal input function, and transform the initial data requirements into structured data requirements to be processed. Then, an artificial intelligence agent with context awareness is used to clarify the ambiguous areas of the data requirements to be processed and refine the information elements of the requirements to be processed, so as to obtain the refined data requirements. The shared database construction module is used to perform semantic understanding, intent recognition and entity extraction operations on the refined data requirements using the large model to obtain target data requirements. Then, the target data requirements are processed using the large model to obtain structured query vectors, and a data sharing database is constructed based on the preset government data resource catalog and the data features in the preset government data resource catalog. The data resource recommendation list determination module is used to determine the semantic relevance between the query vector and the data resource vectors corresponding to each data feature in the data sharing database, obtain a semantic relevance score, and sort the query vectors corresponding to each semantic relevance score in descending order of score to obtain the target data resource recommendation list. The desensitization sample generation module is used to update the government data resource library based on the target data resource recommendation list, so as to use the updated government data resource library to perform quality analysis and target desensitization sample generation operations on the target data requirements, and display the obtained quality report and target desensitization sample corresponding to the government data to be processed on the human-computer interaction interface.
9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the large-model-based government data processing method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, Used to store computer programs, wherein the computer programs, when executed by a processor, implement the government data processing method based on a large model as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Government affair intelligent big data center system architecture
CN112015962A
Government affair industry intelligent information retrieval and pushing system and method based on large model
CN120336504A
Intelligent government affair number asking method, device and equipment based on large model and medium
CN120596509A
Government affair big model-based data government method and system
CN120632787A
Content generation method based on multimedia content, device and medium
US20250053590A1
Cited By
Government affair data processing system and method
CN121073402A