Method and system for unified semantic representation and sharing of cross-domain community data

CN122838366APending Publication Date: 2026-09-29CETC BIGDATA RES INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610873316.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-17
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0004]本发明提供一种跨域社区数据统一语义表征与共享方法及系统,以解决大规模、多源、异构条件下社区跨部门数据安全共享、语义互通、高效融合的行业共性难题

Benefits of technology

[0021]本发明提供的跨域社区数据统一语义表征与共享方法及系统,采用双层数据模型建模、统一语义表征、分布式知识图谱构建的方式,以“标准统一、语义互通、知识关联”为目标,先对社区多部门大规模异构数据进行分布式分层建模,形成统一数据规范;再通过语义识别、对齐、融合构建统一语义空间,消除跨域歧义;随后基于语义结果构建分布式跨域知识图谱,实现多源知识深度关联,进而可以根据跨域知识图谱对外提供数据共享服务。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122838366A_ABST
    Figure CN122838366A_ABST
Patent Text Reader

Abstract

The application discloses a cross-domain community data unified semantic representation and sharing method and system, and the method comprises the following steps: accessing community department data sources to obtain multi-source heterogeneous data; constructing a double-layer data model, including a community core data model and department field data models, and establishing a two-way mapping rule library between the department field data models and the community core data model; according to a pre-established community governance dictionary and the double-layer data model, mapping the multi-source heterogeneous data to a unified semantic space to obtain a standardized semantic library; constructing a distributed cross-domain knowledge graph based on the standardized semantic library, and storing multiple subgraphs of the distributed cross-domain knowledge graph in multiple storage nodes of a distributed graph database; and providing data sharing services to the outside by using the distributed graph database. According to the application, multi-department data sharing services can be conveniently realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data governance, specifically to a method and system for unified semantic representation and sharing of cross-domain community data. Background Technology

[0002] As the core unit of urban grassroots governance, communities bear multiple functions, including population management, government services, security monitoring, livelihood security, property services, environmental sanitation, and emergency response. These involve various departments and entities, such as civil affairs, health, housing and urban-rural development, urban management, sub-district offices, community residents' committees, property management companies, and third-party service providers. With the advancement of smart community construction, the information systems independently built by various departments have accumulated petabytes (PB) of massive heterogeneous data, including: tens of millions of structured household registration data from the household registration management department, millions of semi-structured resident health records from the health department, hundreds of millions of unstructured text data entries related to resident payments and repair requests from the property management department, and tens of millions of patrol records from community grids.

[0003] The core challenges facing large-scale cross-domain data governance in communities today are "data silos," "semantic islands," and "imbalance in secure sharing," with the problem of large-scale cross-domain data sharing being particularly prominent. Specifically, this manifests in three ways: First, inconsistent data standards. Different departments use independent data elements, classification rules, and coding systems. The same entity (such as residents, houses, or events) has completely different names, formats, and field definitions in different departmental systems, leading to a surge in error rates during large-scale data mapping. Second, diverse data formats, including structured database data, semi-structured Excel / XML data, and unstructured text / image / video data, resulting in high computational consumption and extremely low efficiency when merging large-scale heterogeneous data. Third, semantic... The problems are: First, there are serious ambiguities, with the same term having different meanings and different terms referring to the same object. Large-scale data semantic alignment requires a lot of manual intervention and cannot meet the needs of real-time sharing. Second, large-scale data cross-domain secure sharing is difficult. Traditional sharing models either adopt centralized aggregation (which is prone to privacy leaks and single points of failure) or complete isolation (which cannot achieve data interoperability). Fine-grained security control and efficient sharing cannot be taken into account at the same time, and sensitive data (resident health, household registration, property information) has a very high risk of leakage in large-scale sharing. Third, there is no standardized process for cross-domain exchange, and there is a lack of unified data models and exchange protocols. Large-scale data interoperability has poor compatibility and low reuse rate, which cannot support grassroots governance collaborative decision-making. Summary of the Invention

[0004] This invention provides a method and system for unified semantic representation and sharing of cross-domain community data, in order to solve the common industry problem of secure sharing, semantic interoperability and efficient integration of cross-departmental data in communities under large-scale, multi-source and heterogeneous conditions.

[0005] Therefore, the present invention provides the following technical solution: On one hand, embodiments of the present invention provide a method for unified semantic representation and sharing of cross-domain community data, the method comprising: By accessing data sources from various departments within the community, we obtained multi-source heterogeneous data. A two-layer data model is constructed, which includes a community core data model and departmental domain data models, and a bidirectional mapping rule base is established between the departmental domain data models and the community core data model. Based on the pre-established community governance dictionary and the two-layer data model, the multi-source heterogeneous data is mapped to a unified semantic space to obtain a standardized semantic library; the community governance dictionary contains community-specific vocabulary. A distributed cross-domain knowledge graph is constructed based on the standardized semantic library, and multiple subgraphs of the distributed cross-domain knowledge graph are stored in multiple storage nodes of the distributed graph database respectively. The distributed graph database is used to provide data sharing services to the outside world. The data sharing services include any one or more of the following: query services, shared access, and visualization.

[0006] Optionally, constructing the community's core data model includes: Identify the core entities that are universally applicable across the entire community; The attribute information and association rules of the core entities are determined to obtain the core data model of the community.

[0007] Optionally, the core entities include: residents, houses, grids, events, organizations, and devices; the resident entity is uniquely identified by a hash value; the attribute information includes: standard attributes, data types, encoding rules, and uniqueness constraints.

[0008] Optionally, constructing data models for each department includes: expanding the attributes of the core community data model according to the business characteristics of each department to obtain data models for each department.

[0009] Optionally, the bidirectional mapping rule base includes: an entity mapping rule base, an attribute mapping rule base, and a relationship mapping rule base.

[0010] Optionally, the step of mapping the multi-source heterogeneous data to a unified semantic space based on a pre-established community governance dictionary and the two-layer data model to obtain a standardized semantic library includes: Community instance information is extracted from the multi-source heterogeneous data according to the community governance dictionary. The community instance information includes: entity, attribute and relationship information. The similarity of community instance information from different departments is calculated based on the two-layer data model. Based on the calculation results, the multi-source heterogeneous data is mapped to a unified semantic space to obtain a standardized semantic library.

[0011] Optionally, constructing a distributed cross-domain knowledge graph based on the standardized semantic library includes: Knowledge triples are extracted from the standardized semantic library, and the knowledge triples are represented in a structured form. The knowledge triples are fused to ensure their accuracy and generate target triples. Construct a distributed cross-domain knowledge graph based on the target triples.

[0012] Optionally, the method further includes: A knowledge graph semantic hierarchy is established in the distributed graph database; the semantic hierarchy includes: a core semantic layer, a business semantic layer, and a sensitive semantic layer; Participating in federated learning according to the semantic hierarchy includes: using local data nodes of each department as participants, and using the semantic representation results of the corresponding level in the cross-domain knowledge graph as the basic data for federated learning; wherein, core semantic layer data serves as the common foundation for federated computation of each department, business semantic layer data serves as the core basis for exclusive computation of each department, and sensitive semantic layer data is only computed on local nodes and does not participate in global parameter sharing; and / or configuring access permissions according to the semantic hierarchy and de-identifying the accessed data.

[0013] Optionally, the method further includes: incrementally updating and online upgrading the two-layer data model, the community governance dictionary, and the cross-domain knowledge graph.

[0014] On the other hand, embodiments of the present invention also provide a cross-domain community data unified semantic representation and sharing system, the system comprising: The data access module is used to access data sources from various departments in the community to obtain multi-source heterogeneous data. The model and rule base construction module is used to construct a two-layer data model, which includes a community core data model and departmental domain data models, and to establish a bidirectional mapping rule base between the departmental domain data models and the community core data model. The standardization module is used to map the multi-source heterogeneous data to a unified semantic space based on the pre-established community governance dictionary and the two-layer data model, thereby obtaining a standardized semantic library; the community governance dictionary contains community-specific vocabulary. The knowledge graph generation module is used to construct a distributed cross-domain knowledge graph based on the standardized semantic library, and to store multiple subgraphs of the distributed cross-domain knowledge graph in multiple storage nodes of the distributed graph database respectively. The service providing module is used to provide data sharing services to the outside world using the distributed graph database. The data sharing services include any one or more of the following: query services, shared calls, and visualization displays.

[0015] Optionally, the standardization module includes: The information extraction unit is used to extract community instance information from the multi-source heterogeneous data according to the community governance dictionary. The community instance information includes: entity, attribute and relationship information. The mapping unit is used to calculate the similarity of community instance information from different departments based on the two-layer data model, and to map the multi-source heterogeneous data to a unified semantic space based on the calculation results to obtain a standardized semantic library.

[0016] Optionally, the knowledge graph generation module includes: The triple extraction unit is used to extract knowledge triples from the standardized semantic library, wherein the knowledge triples adopt a structured representation. The fusion processing unit is used to perform fusion processing on the knowledge triples to ensure the accuracy of the knowledge triples and generate target triples. The knowledge graph building unit is used to construct a distributed cross-domain knowledge graph based on the target triples.

[0017] Optionally, the system further includes: The operation and maintenance management module is used to provide operation and maintenance management operations, which include any one or more of the following: user management, permission configuration, resource monitoring, log viewing, and alarm notification.

[0018] Optionally, the distributed graph database is equipped with a knowledge graph semantic hierarchy; the semantic hierarchy includes: a core semantic layer, a business semantic layer, and a sensitive semantic layer; the system also includes: a federated learning management module and / or an access control module; The federated learning management module is used to participate in federated learning with local data nodes of each department as participants and the semantic representation results of the corresponding level in the cross-domain knowledge graph as the basic data for federated learning. Among them, the core semantic layer data serves as the common foundation for federated computing of each department, the business semantic layer data serves as the core basis for exclusive computing of each department, and the sensitive semantic layer data is only calculated on the local node and does not participate in global parameter sharing. The access control module is used to configure access permissions according to the semantic hierarchy and to perform anonymization processing on the accessed data.

[0019] Optionally, the system further includes a system management module for incrementally updating and online upgrading the two-layer data model, the community governance dictionary, and the cross-domain knowledge graph.

[0020] On the other hand, the present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when run by a processor, executes the steps of the cross-domain community data unified semantic representation and sharing method.

[0021] The present invention provides a method and system for unified semantic representation and sharing of cross-domain community data. It adopts a two-layer data model modeling, unified semantic representation, and distributed knowledge graph construction approach. With the goal of "standard unification, semantic interoperability, and knowledge association", it first performs distributed layered modeling on large-scale heterogeneous data from multiple departments in the community to form a unified data standard. Then, it constructs a unified semantic space through semantic recognition, alignment, and fusion to eliminate cross-domain ambiguity. Subsequently, it constructs a distributed cross-domain knowledge graph based on the semantic results to achieve deep association of multi-source knowledge. Then, it can provide data sharing services to external parties based on the cross-domain knowledge graph.

[0022] Furthermore, privacy computing technology can be combined to build a fine-grained secure sharing channel, enabling secure and efficient sharing of large-scale cross-domain data that is "usable but invisible, with controllable and traceable permissions". Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the embodiments will be briefly described below. Obviously, the drawings described below are merely some embodiments of the present invention, and those skilled in the art can obtain other drawings based on these drawings without any creative effort.

[0024] Figure 1 This is a flowchart of a method for unified semantic representation and sharing of cross-domain community data provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of a cross-domain community data unified semantic representation and sharing system provided in an embodiment of the present invention; Figure 3 This is another structural diagram of the cross-domain community data unified semantic representation and sharing system provided in the embodiments of the present invention. Detailed Implementation

[0025] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0026] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0027] To address the pain points in the current cross-domain data governance and sharing among multiple departments in communities, this invention provides a method and system for unified semantic representation and sharing of cross-domain community data, enabling secure and efficient sharing of large-scale cross-domain data that is "usable but invisible, with controllable and traceable permissions".

[0028] The present invention addresses one or more of the following problems: (1) To solve the problem of low efficiency and high error rate in cross-domain fusion caused by inconsistent standards, different formats and prominent semantic ambiguity of large-scale heterogeneous data (structured, semi-structured and unstructured) from multiple departments in the community.

[0029] (2) To address risks such as privacy leaks, unauthorized access, and loose permissions during large-scale cross-domain data sharing, and to achieve a balance between security compliance and efficient sharing.

[0030] (3) Solve the engineering implementation problem that existing technologies cannot adapt to the rapid semantic alignment, incremental update and high concurrency call of PB-level community data, and improve the system's practicality and scalability.

[0031] (4) Solve the problems of disconnect between knowledge graph and business data model, inability to deeply associate cross-domain knowledge, and weak intelligent reasoning ability, and improve the ability to reuse data and support governance decisions.

[0032] like Figure 1 The diagram shown is a flowchart of a cross-domain community data unified semantic representation and sharing method provided by an embodiment of the present invention, which includes the following steps: Step 101: Connect to data sources from various departments in the community to obtain multi-source heterogeneous data.

[0033] The multi-source heterogeneous data may be, but is not limited to, structured databases, semi-structured files, unstructured text, or images, or videos, etc., and the embodiments of the present invention do not limit this.

[0034] In this embodiment of the invention, data access methods can include API (Application Programming Interface) integration, direct database connection, file import, message queues, real-time acquisition, and other methods. Specifically, distributed parallel acquisition technology can be used to acquire multi-source heterogeneous data from various sectors of society.

[0035] Furthermore, to ensure data quality, preliminary preprocessing can be performed on the acquired petabyte-level massive data, including missing value imputation, outlier filtering, duplicate data deletion, and format standardization conversion, in order to shield the underlying data format differences, reduce the subsequent fusion computing power consumption, and ensure the stability and efficiency of the data.

[0036] Step 102: Construct a two-layer data model, which includes a community core data model and departmental domain data models, and establish a bidirectional mapping rule base between the departmental domain data models and the community core data model.

[0037] The process of constructing the community core data model is as follows: First, identify the core entities that are common to the entire community; then, determine the attribute information and association rules of the core entities to obtain the community core data model.

[0038] Specifically, core entities can be defined based on prior knowledge such as general knowledge of community governance, industry standards, and laws and regulations. For example, in a non-limiting embodiment, six categories of core entities are identified: residents, houses, grids, events, organizations, and equipment; among them, the resident entity is uniquely identified by a hash value to avoid exposing sensitive information such as ID card numbers in plaintext.

[0039] The attribute information of the core entity may include, but is not limited to: standard attributes, data types, encoding rules, uniqueness constraints, etc.

[0040] The process of constructing data models for each department is as follows: Based on the business characteristics of each department, the core community data model is extended with attributes to obtain data models for each department.

[0041] For example, based on the specific business characteristics of different departments such as civil affairs, health, property management, urban management, and street offices, dedicated attributes are extended on the core model to obtain corresponding departmental domain data models. This ensures that the core structure of the data models across different departments remains consistent, guaranteeing compatibility.

[0042] A bidirectional mapping rule base is established between departmental domain data models and the community core data model. Its core function is to build a rule base for converting entities, attributes, and relationships between these two models based on ontology mapping technology. This mechanism enables, on the one hand, the automatic and rapid conversion of heterogeneous data from various departments into the standard core model format, achieving cross-domain integration; on the other hand, it allows the unified query results of the community core data model to be mapped back to the domain formats familiar to each department, facilitating direct use by the relevant department's business systems. This achieves automatic cross-departmental data conversion and standardization, resolving issues of inconsistent data standards and mapping difficulties.

[0043] In this embodiment of the invention, the bidirectional mapping rule base may include: an entity mapping rule base, an attribute mapping rule base, and a relationship mapping rule base.

[0044] Specifically, the bidirectional mapping rule base can be established through a combination of fully automated mapping and manual review, which can significantly reduce labor costs, ensure the accuracy and compatibility of large-scale data transformation, and truly achieve cross-departmental data "semantic interoperability and rapid transformation".

[0045] Step 103: Based on the pre-established community governance dictionary and the two-layer data model, the multi-source heterogeneous data is mapped to a unified semantic space to obtain a standardized semantic library.

[0046] The community governance dictionary contains community-specific terms, covering terms such as empty-nest elderly, low-income households, people in extreme poverty, grid areas, micro-governance, and incident reporting.

[0047] Specifically, graph neural networks can be used to achieve cross-departmental semantic alignment, eliminating ambiguity and conflict. By mapping all objects to a unified semantic space, a consistent semantic representation is formed across the entire community, providing a unified semantic foundation for subsequent knowledge fusion and sharing.

[0048] For example, in a non-limiting embodiment, the baseline model BERT (Bidirectional Encoder Representation based on Transformer)-BiLSTM (Bidirectional Long Short-Term Memory Network)-CRF (Conditional Random Field) is first used to initially extract community instance information from free text or semi-structured fields of data from various domains. This community instance information includes entities (e.g., "Uncle Zhang"), attributes (e.g., "low-income households"), and relationships (e.g., "residing in"). This process relies on thousands of community-specific terms (e.g., "empty-nest elderly" and "micro-governance") provided by the governance dictionary to improve recognition accuracy and avoid domain ambiguity. Specifically, BERT (bottom layer / embedding layer) converts the input text (i.e., multi-source heterogeneous data) into high-quality contextual word vectors; BiLSTM (intermediate layer / encoding layer) scans the vectors generated by BERT from left to right and from right to left to capture the dependencies of the entire sentence; and CRF (output layer / decoding layer) predicts the globally optimal label.

[0049] In some embodiments, the above functions can also be achieved by using BERT + linear layers + CRF network, or by using a pre-trained large model (such as BERT, RoBERTa) plus a classification head. This embodiment of the invention does not limit the scope of the invention.

[0050] Then, a GNN (Graph Neural Network) is used to receive the above preliminary extraction results, and the similarity of community instance information from different departments (such as different expressions like "Uncle Zhang" and "Zhang San") is calculated using the standard entity types, attribute definitions and relation frameworks given by the two-layer model. This completes entity alignment, attribute normalization and relation conflict resolution, thereby mapping the multi-source heterogeneous data to a unified semantic space and generating standardized semantic representations.

[0051] In the above mapping process, the community governance dictionary ensures the "recall rate" and "precision rate" of the extraction, the two-layer data model provides a "target coordinate system" for the alignment of different data, and GNN is responsible for mapping the "dialect" coordinates of each department to this coordinate system, ultimately generating a unified semantic representation and saving it to a standardized semantic library.

[0052] Step 104: Construct a distributed cross-domain knowledge graph based on the standardized semantic library, and store multiple subgraphs of the distributed cross-domain knowledge graph in multiple storage nodes of the distributed graph database.

[0053] Knowledge triples are extracted from the standardized semantic database. These knowledge triples are represented in a structured format, such as subject-verb-object, entity-relationship-entity or attribute, or entity-attribute-value. They contain three types of information: a head entity (e.g., a specific object like a resident, house, or grid), a relation (e.g., an entity's attribute relation "possesses age," a behavioral relation "reports events," or a spatial relation "located in a grid"), and a tail entity or attribute value (e.g., another entity or a specific numerical value / text). For example, "Resident (Zhang San) - Belongs to - Grid (Third Grid)" or "Resident (Li Si) - Health Status - Chronic Disease".

[0054] The knowledge triples characterize both the inherent attributes and relationships of static basic data (such as population and housing) and the behavioral links of dynamic business data (such as event handling and patrol records). Covering static basic data such as population, housing, and grids, as well as dynamic business data such as events, services, and patrols, it provides a structured foundation for cross-domain knowledge fusion and reasoning.

[0055] Because multi-source heterogeneous data, after undergoing unified semantic representation, generates a large number of triples, these triples inevitably contain duplicate references (e.g., multiple names for the same entity), attribute conflicts (e.g., inconsistent ages for the same resident), low-quality information, or errors. Therefore, in some embodiments, these extracted knowledge triples can be further fused. This fusion process includes removing erroneous and redundant knowledge triples using preset quality verification rules to ensure the accuracy of the knowledge triples and generate target triples. A distributed cross-domain knowledge graph is then constructed based on the target triples. Multiple subgraphs of the distributed cross-domain knowledge graph are stored on different nodes and linked through the cross-domain relationships in the target triples.

[0056] The above-mentioned fusion processes can be handled in various ways, such as one or more of the following.

[0057] Coreference resolution refers to determining whether different expressions of the head or tail entities in different knowledge triples (such as "Zhang San", "Zhang San (ID card number ending in 055X)", "owner Zhang") refer to the same real object by using entity alignment results and similarity calculations (such as string similarity, semantic vector cosine similarity). If they do, they are merged into a unified identifier, and the entity references in all related triples are updated.

[0058] Conflict detection refers to the comprehensive evaluation of knowledge triples (such as <Resident A-Age-35> and <Resident A-Age-36>) that share the same entity-attribute-value pair, based on factors such as data source authority, update timestamp, and confidence level. For example, if public safety system data has higher priority than property management system data, the higher priority or most recent timestamp value is used, and conflict records are marked for manual review.

[0059] Quality assessment: For example, dimensions such as completeness (whether a primary key or key attribute is missing), accuracy (whether the value conforms to the data type and range, such as age 0-120), and consistency (whether it contradicts existing authoritative facts) can be set to calculate a quality score for each knowledge triple. Knowledge triples with quality scores below the set threshold are automatically removed or sent to the waiting queue for review.

[0060] Merging: refers to merging valid knowledge triples that have been resolved, tested, and evaluated into a distributed knowledge graph with a unified identifier. For multiple versions of the same fact, the final knowledge is generated by "majority voting" or "selecting the optimal rule", while retaining the source information for subsequent auditing and updating.

[0061] It should be noted that, in this embodiment of the invention, when the triple entity information contains personal privacy information, consent or authorization must be obtained when obtaining information involving personal privacy.

[0062] Step 105: Utilize the distributed graph database to provide data sharing services to external parties.

[0063] The data sharing services include, but are not limited to, any one or more of the following: query services, shared access, and visualization.

[0064] The embodiments of the present invention can support petabyte-level knowledge graph storage and high-concurrency querying, and support real-time updates of incremental data and periodic updates of full data.

[0065] Efficient semantic-based indexes can be built in distributed graph databases to support rapid querying and logical reasoning of knowledge graphs. Accordingly, for user query requests, the distributed graph database can be used to mine implicit cross-domain relationships based on rule engines and graph reasoning algorithms, such as "low-income residents should have priority in receiving community assistance," "elderly people living alone need to be regularly visited," and "important relevant personnel correspond to key houses," providing reasoning capabilities for intelligent governance.

[0066] The rule engine and graph reasoning algorithm described above form a complementary intelligent reasoning mechanism. The rule engine loads predefined business logic rules (e.g., "low-income residents AND no assistance records → generate assistance event") from the knowledge graph through forward or backward chain matching, triggering specific actions to achieve interpretable and easily adjustable deterministic reasoning. The graph reasoning algorithm (e.g., path sorting, graph neural networks) automatically explores multi-hop connections between entities in the knowledge graph, discovering hidden statistical patterns in the data. For example, through multi-step paths such as "resident-residence-house-owner-key related personnel," it uncovers the potential association between "key related personnel and key houses." The combination of these two approaches enables precise intervention using expert knowledge and automatic discovery of unknown patterns from large-scale graph structures, jointly supporting intelligent community governance decision-making.

[0067] Furthermore, to ensure secure data sharing, a semantic hierarchy of knowledge graphs can be established within the distributed graph database. For example, in a non-limiting embodiment, based on a community governance business scenario, the entities, attributes, and relationships in the knowledge graph are divided into three layers: a core semantic layer, a business semantic layer, and a sensitive semantic layer. The core semantic layer contains basic entities and relationships common to all departments, the business semantic layer contains entities and relationships related to the specific business of each department, and the sensitive semantic layer contains entities and attributes involving resident privacy and core departmental data, providing semantic hierarchy support for subsequent secure cross-domain data sharing.

[0068] Accordingly, access permissions can be configured according to the semantic hierarchy, and the accessed data can be anonymized.

[0069] For example, based on the semantic hierarchy, fine-grained access permissions can be configured according to the community governance business scenarios and the sensitivity levels of each semantic level. For resident privacy and core business data of departments involved in the sensitive semantic layers in the knowledge graph, precise semantic desensitization processing can be carried out in combination with the entity association relationships in the knowledge graph. During the desensitization process, the integrity of the association between the core semantic layer and the business semantic layer is preserved, ensuring that the desensitized data can still be used for subsequent governance decisions and federated computing.

[0070] Furthermore, based on the knowledge structure of cross-domain knowledge graphs, standardized cross-domain data exchange interfaces and services can be built. Cross-domain exchange is only performed on compliant data in the core semantic layer and business semantic layer of the knowledge graph. Relying on the efficient indexing of the knowledge graph, data can be quickly matched and called, improving the efficiency of cross-domain exchange of petabyte-scale large-scale data and solving the technical problems of high latency and low efficiency in cross-domain exchange of large-scale data in existing technologies.

[0071] Furthermore, federated learning can be conducted based on the aforementioned semantic hierarchy. Specifically, local data nodes of each department participate, and the semantic representation results of the corresponding levels in the cross-domain knowledge graph serve as the foundational data for federated learning. Core semantic layer data serves as the common foundation for federated computation across departments, business semantic layer data serves as the core basis for departmental-specific computation, and sensitive semantic layer data is computed only on local nodes and does not participate in global parameter sharing. This achieves the technical effect of local computation of data across departments and global sharing of model parameters while avoiding the privacy leakage risks associated with cross-domain transmission of sensitive semantic data. Furthermore, homomorphic encryption algorithms can be employed to encrypt the transmission of model parameters related to the core semantic layer and business semantic layer shared by each department, as well as intermediate computation results, throughout the federated learning process. Combined with the division of sensitive semantic layers in the knowledge graph, local encryption isolation is implemented for computational processes related to sensitive semantic layers, thereby ensuring data privacy and computational security and solving the security protection technical issues in cross-domain data sharing.

[0072] Furthermore, a full-process operation log recording mechanism can be established to record every step of the operation, including knowledge graph data calls, federated computing, and cross-domain exchange, in detail. This log can be linked to the semantic level and entity information corresponding to the knowledge graph, enabling traceability of operations and responsibilities. Ultimately, this will achieve secure, compliant, and efficient cross-domain data sharing based on knowledge graphs, fully leveraging the core value of cross-domain knowledge graphs and solving the technical problems of low efficiency, poor security, inability to adapt to large-scale data processing, and ineffective application of knowledge graphs after construction in existing technologies.

[0073] In some embodiments, the two-layer data model, the community governance dictionary, and the cross-domain knowledge graph can also be incrementally updated and upgraded online.

[0074] For example, new data collected and preprocessed in real time can be semantically represented and then integrated into the current cross-domain knowledge graph in real time, thus ensuring the timeliness of the knowledge graph. At the same time, based on preset rules and graph reasoning algorithms, the implicit entity relationships in the knowledge graph are mined to form a complete knowledge system that can directly support community governance decisions, providing core knowledge support for subsequent cross-domain data sharing.

[0075] In some embodiments, operational metrics such as data sharing efficiency, semantic alignment accuracy, system response speed, and concurrent load can be monitored in real time. For example, operational logs and performance counters can be collected at key nodes of the system (data access layer, semantic representation layer, knowledge sharing layer, security management module, etc.). For data sharing efficiency, the number of successfully processed cross-domain query or exchange requests (TPS), average response latency, and throughput can be statistically analyzed and collected using time-series databases such as Prometheus. For semantic alignment accuracy, a sampling verification mechanism can be used to periodically randomly sample a certain proportion of entity alignment results from a unified semantic library, compare them with a manually annotated benchmark set, calculate precision, recall, and F1 score, and record cases that fail the verification for analysis. For system response speed, for standardized interfaces, the timestamps of each request from receipt to return (including network, processing, and queuing time) can be recorded, and the P99 and P95 quantiles can be calculated. For concurrent load, the utilization rates of system CPU, memory, network I / O interfaces, and database connection pools can be monitored, and the request queuing length of each microservice can be detected in conjunction with a service mesh (such as Istio). All metrics are visualized on the operations and maintenance dashboard, and threshold alarms are set to trigger automatic scaling up or down or notify operations and maintenance personnel, ensuring long-term stable operation of the system.

[0076] Based on the monitoring results of various operational indicators, we determine the community's business expansion needs and dynamically expand the data model, governance dictionary, and knowledge graph. We optimize distributed computing power scheduling to improve the speed of large-scale data processing. We continuously iterate semantic algorithms and security models to ensure the long-term stability, efficiency, and availability of the system, meeting the ever-growing data scale and business needs.

[0077] Accordingly, embodiments of the present invention also provide a unified semantic representation and sharing system for cross-domain community data, such as... Figure 2 The diagram shown is a structural schematic of the system.

[0078] In this embodiment, the cross-domain data unified semantic representation and sharing system 200 includes the following modules: Data access module 201 is used to access data sources from various departments in the community to obtain multi-source heterogeneous data; The model and rule base construction module 202 is used to construct a two-layer data model, which includes a community core data model and departmental domain data models, and to establish a bidirectional mapping rule base between the departmental domain data models and the community core data model. The standardization module 203 is used to map the multi-source heterogeneous data to a unified semantic space based on the pre-established community governance dictionary and the two-layer data model to obtain a standardized semantic library; the community governance dictionary contains community-specific vocabulary. The knowledge graph generation module 204 is used to construct a distributed cross-domain knowledge graph based on the standardized semantic library, and store multiple subgraphs of the distributed cross-domain knowledge graph in multiple storage nodes of the distributed graph database respectively. Service providing module 205 is used to provide data sharing services to the outside world using the distributed graph database. The data sharing services include any one or more of the following: query service, shared call, and visualization display.

[0079] Specifically, it allows for a comprehensive analysis of all data sources across the community, clarifying the type, structure, field meanings, business rules, and update frequency of each data category. The data access module 201 can connect to various data sources through a distributed acquisition framework, enabling the access of petabyte-scale full data. Furthermore, a data preprocessing module can be set up within the cross-domain data unified semantic representation and sharing system 200 to perform unified preprocessing of the data, such as removing invalid, duplicate, and erroneous data, and standardizing formats and encodings, resulting in a high-quality dataset to be integrated. Employing distributed parallel processing can effectively improve efficiency, ensuring that large-scale data can be preprocessed within a specified timeframe.

[0080] In this embodiment of the invention, the model and rule base construction module 202 can determine a core entity common to the entire community, and then determine the attribute information and association rules of the core entity to obtain a community core data model. Then, according to the business characteristics of each department, the community core data model is extended with attributes to obtain departmental domain data models.

[0081] A non-limiting structure of the standardized module 203 may include: an information extraction unit and a mapping unit. Wherein: The information extraction unit is used to extract community instance information from the multi-source heterogeneous data according to the community governance dictionary. The community instance information includes: entity, attribute and relationship information. The mapping unit is used to calculate the similarity of community instance information from different departments based on the two-layer data model, and to map the multi-source heterogeneous data to a unified semantic space based on the calculation results to obtain a standardized semantic library.

[0082] A non-restricted structure of the knowledge graph generation module 204 may include: a triple extraction unit, a fusion processing unit, and a knowledge graph building unit. Wherein: The triple extraction unit is used to extract knowledge triples from the standardized semantic library, wherein the knowledge triples adopt a structured representation. The fusion processing unit is used to perform knowledge purification processing on the knowledge triples to ensure the accuracy of the knowledge triples and generate target triples; the fusion processing may specifically include one or more processing methods, such as coreference resolution, conflict detection, quality assessment and merging, etc. The knowledge graph building unit is used to construct a distributed cross-domain knowledge graph based on the target triples.

[0083] In some embodiments, the cross-domain data unified semantic representation and sharing system 200 may further include: an operation and maintenance management module (not shown), used to provide operation and maintenance management operations, the operation and maintenance management operations including any one or more of the following: user management, permission configuration, resource monitoring, log viewing, and alarm notification.

[0084] like Figure 3 The diagram shown is another structural schematic of the cross-domain community data unified semantic representation and sharing system provided in an embodiment of the present invention.

[0085] In this embodiment, the distributed graph database is equipped with a knowledge graph semantic hierarchy system; the semantic hierarchy system includes: a core semantic layer, a business semantic layer, and a sensitive semantic layer. Additionally, with... Figure 2 Compared with the embodiments shown, Figure 3The cross-domain unified semantic representation and sharing system 200 of the illustrated embodiment further includes any one or more of the following: a system management module 301, a federated learning management module 302, and an access control module 303. Wherein: The system management module 301 is used to perform incremental updates and online upgrades of the two-layer data model, the community governance dictionary, and the cross-domain knowledge graph.

[0086] The federated learning management module 302 is used to participate in federated learning with local data nodes of each department as participants and semantic representation results of the corresponding level in the cross-domain knowledge graph as the basic data for federated learning. Among them, core semantic layer data serves as the common foundation for federated computing of each department, business semantic layer data serves as the core basis for exclusive computing of each department, and sensitive semantic layer data is only calculated on local nodes and does not participate in global parameter sharing. The access control module 303 is used to configure access permissions according to the semantic hierarchy and to perform desensitization processing on the access data.

[0087] In some embodiments, the cross-domain data unified semantic representation and sharing system 200 may further include: an operation and maintenance management module (not shown), used to provide operation and maintenance management operations, the operation and maintenance management operations including any one or more of the following: user management, permission configuration, resource monitoring, log viewing, and alarm notification.

[0088] It should be noted that, without departing from the technical concept of this invention, semantic representation, knowledge graph construction, secure sharing and data expansion can also be achieved through equivalent semantic algorithms, alternative graph neural networks, other homomorphic encryption systems, hybrid federated learning architectures or distributed storage schemes; related model structures, permission rules and exchange protocols can also be extended and adapted according to actual application scenarios, and all the above-mentioned alternative solutions should be considered to fall within the protection scope of this invention.

[0089] The cross-domain community data unified semantic representation and sharing method and system provided in this invention adopts a two-layer data modeling, unified semantic representation, and distributed knowledge graph construction approach. With the goal of "standard unification, semantic interoperability, and knowledge association", it first performs distributed layered modeling on large-scale heterogeneous data from multiple departments in the community to form a unified data standard; then, it constructs a unified semantic space through semantic recognition, alignment, and fusion to eliminate cross-domain ambiguity; subsequently, it constructs a distributed cross-domain knowledge graph based on the semantic results to achieve deep association of multi-source knowledge, and then provides data sharing services to external parties based on the cross-domain knowledge graph.

[0090] Compared with the prior art, the solution of the present invention has the following beneficial effects: (1) It solves the problem of data standards and semantic ambiguity, and greatly improves the efficiency of fusion. It achieves standard unification through a two-layer data model, eliminates cross-domain ambiguity through unified semantic representation, and can process structured, semi-structured and unstructured data in a unified manner. The error rate of large-scale data mapping is significantly reduced, the degree of automation is greatly improved, and manual intervention is reduced.

[0091] (2) A balance between security and sharing is achieved. Sensitive data is fully controllable by adopting federated learning and homomorphic encryption to achieve "data is available but not visible, data does not leave the domain, and value can flow". Fine-grained semantic-level permission control avoids unauthorized access. Desensitization and auditing mechanisms reduce the risk of leakage and meet the requirements of data security and personal information protection regulations.

[0092] (3) It can perfectly adapt to PB-level large-scale data, and has strong engineering feasibility. The entire process adopts distributed, parallel, and incremental processing, supports high concurrency, high throughput, and low latency, and can stably support the real-time sharing of massive community data, solving the bottleneck of traditional solutions being unable to handle large-scale data.

[0093] (4) Enhance grassroots governance capabilities through deep knowledge association and intelligent reasoning. Construct cross-domain knowledge graphs to transform scattered data into related knowledge, support intelligent judgment, early warning and prediction, and collaborative handling, and upgrade from "data interoperability" to "knowledge collaboration", providing strong support for the modernization of grassroots governance.

[0094] (5) It is highly versatile, flexible in expansion, and has a wide range of applications. The model and interface are standardized and can be quickly deployed to different communities, streets, districts and counties. It is compatible with new departments and new businesses, and does not require repeated development, thus reducing construction and maintenance costs.

[0095] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0096] The present invention also provides a storage medium, which is a computer-readable storage medium storing a computer program thereon, the computer program being executable when it runs. Figure 1The method shown may include some or all of the steps. The storage medium may include read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc. The storage medium may also include non-volatile memory or non-transitory memory, etc.

[0097] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer program are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data provider to another website, computer, server, or data provider via wired or wireless means.

[0098] The embodiments of the present invention have been described in detail above. Specific implementation methods have been used to illustrate the present invention. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and apparatus of the present invention, and are only a part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention, and the content of this specification should not be construed as a limitation of the present invention. Therefore, any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for unified semantic representation and sharing of cross-domain community data, characterized in that, The method includes: By accessing data sources from various departments within the community, we obtained multi-source heterogeneous data. A two-layer data model is constructed, which includes a community core data model and departmental domain data models, and a bidirectional mapping rule base is established between the departmental domain data models and the community core data model. Based on the pre-established community governance dictionary and the two-layer data model, the multi-source heterogeneous data is mapped to a unified semantic space to obtain a standardized semantic library; the community governance dictionary contains community-specific vocabulary. A distributed cross-domain knowledge graph is constructed based on the standardized semantic library, and multiple subgraphs of the distributed cross-domain knowledge graph are stored in multiple storage nodes of the distributed graph database respectively. The distributed graph database is used to provide data sharing services to the outside world. The data sharing services include any one or more of the following: query services, shared access, and visualization.

2. The method for unified semantic representation and sharing of cross-domain community data according to claim 1, characterized in that, Building the core data model for the community includes: Identify the core entities that are universally applicable across the entire community; The attribute information and association rules of the core entities are determined to obtain the core data model of the community.

3. The method for unified semantic representation and sharing of cross-domain community data according to claim 2, characterized in that, The core entities include: residents, houses, grids, events, organizations, and devices; the resident entity is uniquely identified by a hash value. The attribute information includes: standard attributes, data types, encoding rules, and uniqueness constraints.

4. The method for unified semantic representation and sharing of cross-domain community data according to claim 1, characterized in that, Building data models for each department and domain includes: Based on the business characteristics of each department, the core community data model is extended with attributes to obtain the domain data models for each department.

5. The method for unified semantic representation and sharing of cross-domain community data according to claim 1, characterized in that, The bidirectional mapping rule base includes: an entity mapping rule base, an attribute mapping rule base, and a relationship mapping rule base.

6. The method for unified semantic representation and sharing of cross-domain community data according to claim 1, characterized in that, The process of mapping the multi-source heterogeneous data to a unified semantic space based on a pre-established community governance dictionary and the two-layer data model to obtain a standardized semantic library includes: Community instance information is extracted from the multi-source heterogeneous data according to the community governance dictionary. The community instance information includes: entity, attribute and relationship information. The similarity of community instance information from different departments is calculated based on the two-layer data model. Based on the calculation results, the multi-source heterogeneous data is mapped to a unified semantic space to obtain a standardized semantic library.

7. The method for unified semantic representation and sharing of cross-domain community data according to claim 1, characterized in that, The construction of a distributed cross-domain knowledge graph based on the standardized semantic library includes: Knowledge triples are extracted from the standardized semantic library, and the knowledge triples are represented in a structured form. The knowledge triples are fused to ensure their accuracy and generate target triples. Construct a distributed cross-domain knowledge graph based on the target triples.

8. The method for unified semantic representation and sharing of cross-domain community data according to any one of claims 1 to 7, characterized in that, The method further includes: A knowledge graph semantic hierarchy is established in the distributed graph database; the semantic hierarchy includes: a core semantic layer, a business semantic layer, and a sensitive semantic layer; Participating in federated learning according to the semantic hierarchy includes: using local data nodes of each department as participants, and using the semantic representation results of the corresponding level in the cross-domain knowledge graph as the basic data for federated learning; wherein, core semantic layer data serves as the common foundation for federated computation of each department, business semantic layer data serves as the core basis for exclusive computation of each department, and sensitive semantic layer data is only computed on local nodes and does not participate in global parameter sharing; and / or configuring access permissions according to the semantic hierarchy and de-identifying the accessed data.

9. The method for unified semantic representation and sharing of cross-domain community data according to any one of claims 1 to 7, characterized in that, The method further includes: Incremental updates and online upgrades are performed on the two-layer data model, the community governance dictionary, and the cross-domain knowledge graph.

10. A cross-domain community data unified semantic representation and sharing system, characterized in that, The system includes: The data access module is used to access data sources from various departments in the community to obtain multi-source heterogeneous data. The model and rule base construction module is used to construct a two-layer data model, which includes a community core data model and departmental domain data models, and to establish a bidirectional mapping rule base between the departmental domain data models and the community core data model. The standardization module is used to map the multi-source heterogeneous data to a unified semantic space based on the pre-established community governance dictionary and the two-layer data model, thereby obtaining a standardized semantic library; the community governance dictionary contains community-specific vocabulary. The knowledge graph generation module is used to construct a distributed cross-domain knowledge graph based on the standardized semantic library, and to store multiple subgraphs of the distributed cross-domain knowledge graph in multiple storage nodes of the distributed graph database respectively. The service providing module is used to provide data sharing services to the outside world using the distributed graph database. The data sharing services include any one or more of the following: query services, shared calls, and visualization displays.

11. The cross-domain community data unified semantic representation and sharing system according to claim 10, characterized in that, The standardization module includes: The information extraction unit is used to extract community instance information from the multi-source heterogeneous data according to the community governance dictionary. The community instance information includes: entity, attribute and relationship information. The mapping unit is used to calculate the similarity of community instance information from different departments based on the two-layer data model, and to map the multi-source heterogeneous data to a unified semantic space based on the calculation results to obtain a standardized semantic library.

12. The cross-domain community data unified semantic representation and sharing system according to claim 10, characterized in that, The knowledge graph generation module includes: The triple extraction unit is used to extract knowledge triples from the standardized semantic library, wherein the knowledge triples adopt a structured representation. The fusion processing unit is used to perform fusion processing on the knowledge triples to ensure the accuracy of the knowledge triples and generate target triples. The knowledge graph building unit is used to construct a distributed cross-domain knowledge graph based on the target triples.

13. The cross-domain community data unified semantic representation and sharing system according to claim 10, characterized in that, The system also includes: The operation and maintenance management module is used to provide operation and maintenance management operations, which include any one or more of the following: user management, permission configuration, resource monitoring, log viewing, and alarm notification.

14. The cross-domain community data unified semantic representation and sharing system according to any one of claims 10 to 13, characterized in that, The distributed graph database is equipped with a knowledge graph semantic hierarchy system. The semantic hierarchy includes: a core semantic layer, a business semantic layer, and a sensitive semantic layer; the system also includes: a federated learning management module and / or an access control module. The federated learning management module is used to participate in federated learning with local data nodes of each department as participants and the semantic representation results of the corresponding level in the cross-domain knowledge graph as the basic data for federated learning. Among them, the core semantic layer data serves as the common foundation for federated computing of each department, the business semantic layer data serves as the core basis for exclusive computing of each department, and the sensitive semantic layer data is only calculated on the local node and does not participate in global parameter sharing. The access control module is used to configure access permissions according to the semantic hierarchy and to perform anonymization processing on the accessed data.

15. The cross-domain community data unified semantic representation and sharing system according to any one of claims 10 to 13, characterized in that, The system also includes: The system management module is used to perform incremental updates and online upgrades of the two-layer data model, the community governance dictionary, and the cross-domain knowledge graph.

16. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is run by the processor, it executes the steps of the cross-domain community data unified semantic representation and sharing method according to any one of claims 1 to 9.