Cross-unit data management method

By standardizing the format and extracting semantic features of cross-unit data, building a semantic association network and dynamically generating query authorization, the problems of data semantic isolation and static permission management in the existing technology are solved, and efficient cross-unit data query and fine-grained permission control are realized.

CN119989418APending Publication Date: 2025-05-13贵州惠智电子技术有限责任公司

Patent Information

Application Number
CN202510468954.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing cross-unit data management technology has problems such as data semantic isolation, static permission management, low query efficiency and insufficient ability to resolve cross-unit data conflicts.

Method used

By standardizing the format and extracting semantic features of the data in the data pool of various agencies, building a semantic association network, dynamically generating query authorization and permission mapping tables, optimizing query routing using natural language processing and reinforcement learning, and introducing materialized views and query rewriting technologies for data fusion and conflict resolution.

Benefits of technology

It realizes business logic correlation analysis between data items, supports composite semantic queries, dynamically responds to data changes and scenario requirements, improves query efficiency and data consistency, and ensures fine-grained permission control and effective resolution of cross-unit data semantic conflicts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989418A_ABST
    Figure CN119989418A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to a cross-unit data management method, which comprises the following steps of: performing format standardization processing on data in a data pool of each organ and unit, converting the data into a uniform standard format, extracting semantic features to construct a semantic association network, and recording a semantic relationship of data items and a dynamic authorization and authority mapping mechanism. Then, according to the inquirer identity information, the unit attributes, the business requirements and the data sensitivity level, dynamically generating inquiry authorization and establishing an authority mapping table; and when a query condition is received, analyzing by utilizing a natural language processing technology, and generating a query route and optimizing a statement in combination with the semantic association network and the permission mapping table. And finally, calling data from each data pool according to the route, performing fusion processing, displaying the data to a querier in real time, and recording a query log. The invention aims to solve the problems of data semantic isolation, authority management staticization, low query efficiency and insufficient cross-unit data conflict resolution capability in the existing cross-unit data management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of data processing and relates to a cross-unit data management method. Background Art

[0002] With the continuous increase in the demand for government informationization and data sharing, cross-agency collaborative data management has become an important means to optimize public services and improve governance efficiency. At present, various agencies (such as environmental protection, industry and commerce, taxation, etc.) have generally established independent data pools to store a large amount of structured, semi-structured and unstructured data. However, due to problems such as inconsistent data formats, differences in semantic definitions, and decentralized authority management, cross-agency data integration and query face significant challenges. For example, the "enterprise pollution data" of the environmental protection department and the "enterprise registration information" of the industrial and commercial department may use different field names and unit standards, making it difficult to directly associate and use the data.

[0003] At present, the data formats of various units are not unified. For example, there are differences in text encoding and numerical precision, which requires a lot of manpower for format conversion. The data is semantically isolated and lacks business logic association. For example, when the data of a certain unit becomes sensitive due to policy adjustments, the system cannot automatically shrink the relevant query permissions, which poses a risk of data leakage. At the same time, the mapping relationship between permissions and data is extensive, making it difficult to achieve fine-grained control, and the existing query methods rely on preset routing strategies and cannot dynamically optimize the path based on data pool load, semantic associations, etc. For example, complex queries need to traverse multiple unit data pools, with a long response time and a lack of natural language processing support. Users need to use structured query statements, which have a high threshold and are prone to errors.

[0004] In response to the existing problems, the patent publication number is CN119377582A, and the patent name is A management method and system for multi-terminal engineering internal data. It discloses that the data is fused through spatiotemporal features to build an engineering information data set, realize the physical integration of multi-terminal data, and improve the efficiency of engineering data management. However, it only focuses on spatiotemporal features, does not build semantic associations between data items, and cannot support complex business logic queries; another example is a patent name A general data management method and system based on identity resolution technology, with the patent number CN119378532A. This solution uses identity resolution technology to standardize data formats, realize industry data aggregation, solve the problem of heterogeneous data formats, and support cross-system data exchange, but relies on predefined identifiers. rules, cannot dynamically expand semantic associations, and permission management is still based on fixed roles; and a hybrid parsing middleware and query method based on cross-type databases, patent number CN119396848A, which realizes unified access to multiple types of databases through middleware and maintains cross-database queries, reducing the complexity of application system development, however, the routing strategy is preset and static, and cannot optimize the query path according to data semantics and real-time status; and a data processing method and system, patent number CN118838968A, which realizes data synchronization through message middleware, supports distributed cluster data integration, improves data synchronization efficiency, and supports high concurrency scenarios, but does not resolve semantic conflicts and cannot meet the fine-grained control requirements of sensitive data.

[0005] Existing methods have made some progress in data integration, permission management and query efficiency, but they still have the following common defects: lack of business logic association analysis between data items and inability to support complex semantic queries; authorization mechanisms and query routing strategies rely on manual configuration and cannot respond to data changes and scenario requirements in real time; conflict resolution capabilities are weak; no effective cross-unit data semantic conflict resolution solution is provided, and data consistency is difficult to ensure. Summary of the invention

[0006] The present invention provides a cross-unit data management method to solve the problems of data semantic isolation, static authority management, low query efficiency and insufficient cross-unit data conflict resolution capability in existing cross-unit data management.

[0007] In order to solve the above problems, the technical solution adopted by the invention is: A cross-unit data management method comprises the following steps: S01 standardizes the format of data in the data pool of each agency, converts data of different formats into a unified standard format, extracts semantic features for each data item, builds a semantic association network, records the semantic association relationship between different data items through the semantic association network, and implements a dynamic authorization and permission mapping mechanism; S02 dynamically generates query authorization based on the inquirer's identity information, unit attributes, business requirements, and data sensitivity level, and establishes a permission mapping table to map the inquirer's permissions with data nodes in the semantic association network; S03 When receiving the query conditions, the query conditions are parsed using natural language processing technology, combined with the semantic association network and permission mapping table, and the dependency syntax analysis and intent recognition model are used to convert the natural language query into a structured query tree. Combined with the semantic association network and permission mapping table, the optimal data node path is selected based on reinforcement learning to generate the query route. At the same time, materialized views and query rewriting technology are introduced to decompose complex queries into sub-queries for parallel execution to optimize the query statements. The reward function of reinforcement learning is

[0008] in, is the permission compliance, the value range is [0,1], Indicates data processing delay in milliseconds. is the predicted value of conflict probability, and the parallelism is based on the formula, calculate, Refers to the total amount of data to be processed, in bytes; Node capacity The upper limit of the amount of data that a data node can carry, in bytes; Priority The priority of the query task, the value range is [1-10]; S04 retrieves corresponding data from the data pool of each agency unit according to the query route, integrates the retrieved data, and finally displays the integrated data to the inquirer in real time and records the query log.

[0009] The principles and advantages of this solution are: In the data preprocessing stage, the format of the data in the data pool of each agency unit is standardized, and semantic features are extracted and a semantic association network is constructed to record the semantic association relationship between data items. Then, query authorization is dynamically generated based on the multi-faceted information of the inquirer, and a permission mapping table is established to associate permissions with semantic network nodes to ensure accurate control of the inquirer's permissions. When the query conditions are received, the conditions are parsed using natural language processing technology, and query routes are generated and query statements are optimized in combination with the semantic association network and permission mapping table to efficiently locate the required data. Finally, data is retrieved according to the query route, displayed to the inquirer after fusion processing, and query logs are recorded for subsequent system optimization.

[0010] Compared with the existing technology, this solution has significant creativity and advantages. In terms of data integration, the existing technology focuses on the unification of data formats, while this solution not only standardizes the format, but also deeply mines the semantic information of the data and builds a semantic association network. For example, in the data integration of the environmental protection and meteorological departments, the "pollutant emission data" of the environmental protection department and the "air quality data" of the meteorological department can be logically linked through the semantic association network, so that the originally isolated data can form an organic whole, providing support for complex data analysis and decision-making. In terms of authority management, the existing technology usually adopts a static authorization method and cannot adapt to dynamic changes. This solution dynamically generates authorization based on the identity, business needs and data sensitivity level of the inquirer to achieve fine-grained authority control. If a certain inquirer needs to temporarily access sensitive data in a specific time period due to work, the system can dynamically adjust the authorization according to its business needs, meeting the needs while ensuring data security; first, the system will authenticate the inquirer. This can be achieved in many ways, such as username / password combination, digital certificate, biometric technology such as fingerprint recognition, face recognition. Through identity authentication, the basic identity information of the inquirer, such as name, department, position, etc., is determined. When the inquirer submits a query request, the system will analyze the business needs behind the request. This may involve natural language processing and semantic understanding of the query conditions. For example, when querying "get the sales department performance data of last month for market analysis report", the system can identify that the business need is market analysis, the data involved is the sales department performance data, and the time range is last month. The system will determine the data resources and operation types that need to be accessed based on pre-defined business rules and data associations; dynamic authorization can set the validity period of permissions. For example, when the inquirer needs to access specific sensitive data due to a temporary task, the system only grants access rights during the task period, such as one week. Once the validity period expires, the permission automatically expires to prevent abuse of permissions. In terms of query processing, the query routing and statement optimization capabilities of the existing technology are limited. This solution uses natural language processing technology to directly parse the user's natural language query conditions, combine the semantic association network and the permission mapping table to generate the optimal query route, and intelligently optimize the query statement. For example, when a user inputs "query the penalties for high-polluting enterprises in a certain area in the past month", the system can quickly and accurately retrieve relevant data from multiple government data pools, improving query efficiency and accuracy. In addition, the data fusion processing of this solution combines the semantic association network, which can effectively solve data conflicts and redundancy problems, and ensure that the data displayed to the inquirer is accurate, complete and consistent. At the same time, the recording and analysis of query logs helps to continuously optimize system performance and authority management strategies, and further improve the quality and efficiency of cross-unit data management.

[0011] Furthermore, in S01, semantic feature extraction uses BERT to vectorize data items, and simultaneously combines named entity recognition and relationship extraction technology to build a semantic association network, uses the language understanding ability of the pre-trained language model to convert data items into vector form, determines key entities in the data through named entity recognition, and uses relationship extraction technology to mine semantic relationships between entities, thereby building a semantic association network, through which the semantic association relationships between different data items are recorded; Based on cosine similarity, a metadata mapping relationship across data pools is established: for data A and data B, their semantic similarity is

[0012] in, and for BERT The generated vector, To obtain the strength of association between relationships, attribute-based access control is introduced to adjust permissions by combining timestamp, geographic location, and data update frequency. Sim(A,B) When it is greater than the set threshold θ, a metadata mapping relationship between data A and data B is established, and an attribute-based access control model is introduced, where the permission model is uth ,U is the user, D is the data, is the attribute weight, Including sensitivity level and unit trust; dynamically adjust user permissions through timestamp, geographic location, and data update frequency.

[0013] in, P(U,D) represents the final access right of user U to data D, is the attribute weight, including the sensitivity level weight of data D and the trust weight of the unit to which it belongs, and its value range is [0,1]; are attribute values ​​related to user U and data D, including timestamp-related attribute values, geographic location-related attribute values, and data update frequency-related attribute values, and their value range is [0,1].

[0014] The above scheme BERTAs a pre-trained language model, after training with a large-scale corpus, it has a strong language understanding ability. It can capture the deep semantic information in the data items. After converting the data items into vector form, these vectors can more accurately represent the semantic content of the data. NER can accurately identify key entities in the data, such as names of people, places, and organizational names, while the relationship extraction technology can mine the semantic relationship between these entities. By constructing a semantic association network, the semantic association relationship between different data items can be clearly recorded, which helps to better understand the context and relevance of data in cross-unit data management. For example, in business data involving multiple units, the data interaction relationship between different units can be discovered through the semantic association network, providing more comprehensive information for data analysis and decision-making. At the same time, because the semantic similarity takes into account the similarity between vectors and the strength of association between entities, it can avoid the mismatch problem that may occur in the traditional query method based on keyword matching. For example, when querying "a company's financial statements", not only can data containing the keywords "a company" and "financial statements" be found, but also other data related to the company's finances, such as financial analysis reports, can be found through the semantic association network. In addition, the permission model takes into account multiple attributes, such as sensitivity level and unit trust, and can perform personalized permission allocation according to the characteristics of different users and data. Different data may have different sensitivity levels, and different users may belong to different units with different trust levels. By weighting these attributes, the most appropriate permissions can be assigned to each user and data combination. For example, for data with high sensitivity levels, only users from high-trust units with corresponding permissions can access it. At the same time, as factors such as time, geographic location, and data update frequency change, user permissions can be adjusted in real time. For example, when the frequency of data updates increases, the system can automatically increase the operation frequency limit of users with data update permissions; when users leave a specific geographic location, the system can automatically limit their access to certain data. This real-time permission adjustment mechanism can adapt to the dynamically changing business environment and ensure the security and availability of data.

[0015] Further, in S02, the permission mapping table uses Neo4j The mapping relationship between the storage permission nodes and data nodes is expressed by edge weights, and the mapping relationship between permissions and data is expressed by graph databases. The mapping algorithm is: , where is the sensitive change threshold, recursively updates the relevant permissions, P represents the permissions, , is a function that recursively updates permissions.

[0016] In the above schemes, the mapping relationship between permissions and data is often complicated in cross-unit data management. Different users, roles, data resources and various permission combinations constitute a complex network structure. Neo4j The mapping relationship between storage permission nodes and data nodes can be displayed intuitively in the form of a graph. Permission nodes and data nodes are vertices in the graph, and the edges between them represent the mapping relationship. The weight of the edge can clearly represent the permission granularity, such as read, write, modify and other different operation permissions. This intuitive representation method helps administrators quickly understand and manage the permission system; in cross-unit data management, business needs and permission systems may change over time. Neo4j The permission mapping relationship is stored. When you need to add new permissions, data nodes, or modify the mapping relationship, you only need to add or modify the corresponding vertices and edges in the graph. The operation is simple and has little impact on the existing data structure. For example, when a new business department joins, you can easily create a new permission node for it and establish a mapping relationship with the relevant data nodes without making large-scale adjustments to the entire database structure.

[0017] Furthermore, in S04, corresponding data are retrieved from the data pool of each agency, and the differences of heterogeneous data are eliminated by practical ontology alignment and data cleaning technology. The confidence of multi-source data is fused based on evidence theory, and data fusion is performed according to the fusion formula.

[0018] ,in For the i The basic probability distribution of data sources, The credibility of the data source; Build a conflict detection rule base and use game theory negotiation model to resolve conflicts; ,in is the unit priority, is the utility value of the conflict resolution solution, Bel represents the confidence after data fusion, Resolve represents the conflict resolution function, Maximize Represents the maximum value function.

[0019] In the above scheme, in cross-unit data management, the same data may exist in multiple data sources, and the reliability and accuracy of each data source may be different. Evidence theory can comprehensively consider the information of multiple data sources and distribute the information through basic probability. and data source credibility To calculate the confidence level of the data

[0020] This method can make full use of the complementarity of multi-source data and improve the accuracy and reliability of data fusion. For example, different data sources may give different estimates of the probability of an event. Through evidence theory fusion, a more reasonable comprehensive estimate can be obtained. In cross-unit data management, different units may have different interests and priorities for data. The game theory negotiation model can regard these units as participants in the game. By considering the unit priority and the utility value of conflict resolution To find the best conflict resolution method. This method can fully consider the interests of all parties, make the conflict resolution results more fair and reasonable, and improve the acceptance of data fusion results by all parties. For example, when dealing with data ownership conflicts, the game theory negotiation model can find a solution acceptable to all parties based on the importance of each unit and the degree of demand for data. At the same time, the game theory negotiation model is dynamic and can continuously adjust the conflict resolution method according to the actual situation. When the priorities of the units change or new interests emerge, the model can recalculate the optimal solution to ensure that the conflict is always effectively resolved. This enables the conflict resolution mechanism to adapt to the ever-changing business environment and data situation.

[0021] Furthermore, the format standardization process includes unifying the encoding format of text data in different formats, unifying the measurement unit and precision of numerical data. The data sources of the same agency are extensive, and the text data may use multiple encoding formats, such as UTF-8, GBK, etc.; the measurement unit and precision of numerical data also vary greatly. After unifying the encoding format and measurement unit precision, these format differences can be eliminated, so that data from different units can be smoothly integrated. For example, in the data integration of the environmental protection department and the meteorological department, the monitoring report text of the environmental protection department may use GBK encoding, while the data text of the meteorological department uses UTF-8 encoding. After the unified encoding, the data of the two can be merged and processed in one system to build a comprehensive environmental information database. The unified encoding format and measurement unit precision can lower the threshold for data sharing, making the data of each unit easier to be understood and used by other units. For example, in the information sharing platform between government departments, data in a unified format can be directly called and analyzed by other departments, which improves the efficiency and effect of data sharing.

[0022] Furthermore, in the process of dynamically generating query authorization, the query frequency, query data type, and query time distribution information of the queryer within a certain period of time are counted through the log analysis system. Based on these statistical results, if the queryer frequently queries a certain type of data, more query permissions for this type of data will be provided when dynamically generating query authorization; if an abnormal query that differs greatly from the historical query pattern occurs, the system automatically conducts a strict review or restriction on the query authorization. By analyzing the queryer's recent query behavior pattern and historical query records, it is possible to gain an in-depth understanding of their actual needs. For example, for a queryer who has been paying attention to the atmospheric pollution data of the environmental protection department for a long time, when dynamically generating query authorization in the future, the system can provide them with more data query permissions related to atmospheric pollution, including detailed data of specific areas and specific time periods, so that the authorization is more in line with the actual work needs of the queryer; the historical query behavior of the queryer usually has a certain regularity. If an abnormal query behavior that does not conform to the historical pattern occurs, it may mean that there is a data security risk. For example, a queryer who usually only queries public data suddenly frequently requests sensitive data. By comparing its historical query records, the system can promptly discover this abnormality and conduct a strict review or restriction on the query authorization, thereby effectively preventing security issues such as data leakage and illegal access.

[0023] Further, in the S03, the data changes are monitored in real time. Once the sensitivity of the data changes, the permission settings of the corresponding data nodes in the permission mapping table are adjusted according to the preset sensitivity level rules. When the permission of the queryer changes, the mapping relationship between the corresponding queryer and the data node is directly modified in the permission mapping table. The permission mapping table uses the Neo4j graph database to store the mapping relationship between the permission node and the data node, and uses the edge weight to represent the permission granularity, and uses the characteristics of the graph database to achieve rapid updates. When the data update causes some originally public data to become sensitive data, the real-time update of the permission mapping table can timely limit the access rights of the queryer. For example, in the medical industry, some of the patient's examination data was originally open to specific departments within the hospital. If it is subsequently discovered that these data contain new sensitive information, the real-time update of the permission mapping table can prevent unauthorized personnel from continuing to access, effectively preventing data leakage; if the permission mapping table cannot be updated in real time, there may be a lag phenomenon that the queryer's permissions are inconsistent with the actual situation. This will cause the queryer to be blocked due to insufficient permissions when they need to access certain data, or they can still access data when the permissions have been revoked, affecting work efficiency. Real-time updates can promptly eliminate this lag, allowing inquirers to successfully obtain the required data and improve the smoothness of the workflow.

[0024] Furthermore, the natural language processing technology includes word segmentation, part-of-speech tagging, named entity recognition and semantic understanding. Through the semantic understanding of the query statement, the system can obtain more contextual information and user intentions, thereby making more intelligent decisions. For example, when a user queries "the environmental protection compliance status of enterprises in a certain industry in this city", the system can not only provide relevant compliance data, but also perform intelligent analysis based on the data, such as providing an overall environmental protection situation assessment of the industry, comparative analysis with other industries, etc., to provide users with more valuable decision support.

[0025] Further, in S03, the CPU usage rate, memory occupancy rate load information of each data pool, and data storage location information are obtained in real time through the monitoring system. When generating query routes, data pools that are closer to the query initiator or have a shorter data transmission path are given priority; when the data pool is in a high-load state, the query request is directed to the data pool with a lower load. The storage location of the data will affect the distance and time of data transmission. Considering the storage location of the data to generate query routes, data pools that are closer to the query initiator or have a shorter data transmission path can be given priority. For example, in cross-regional unit data management, if the inquirer is located in area A, and the relevant data is stored in data pool B that is closer to area A, the query route generated by the system will preferentially point to data pool B, thereby reducing the data transmission time in the network and speeding up the query response speed; at the same time, when some data pools are in a high-load state, the speed of processing queries will be significantly reduced. Considering the load of the data pool, the system can direct the query request to the data pool with a lower load. For example, at a certain moment, the CPU usage and memory occupancy of data pool C are very high, while data pool D is in a light load state. The system will route the query to data pool D to avoid query delays caused by waiting for the high-load data pool to process. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 is a flow chart of the method of the present invention; DETAILED DESCRIPTION Embodiment 1, as Figure 1 As shown, a cross-unit data management method includes the following steps: S01 standardizes the format of data in the data pool of each agency, converts data of different formats into a unified standard format, extracts semantic features for each data item, builds a semantic association network, records the semantic association relationship between different data items through the semantic association network, and implements dynamic authorization and permission mapping mechanism; S02 dynamically generates query authorization based on the inquirer's identity information, unit attributes, business requirements, and data sensitivity level, and establishes a permission mapping table to map the inquirer's permissions with data nodes in the semantic association network; S03 When receiving the query conditions, the query conditions are parsed using natural language processing technology, combined with the semantic association network and permission mapping table, and the dependency syntax analysis and intent recognition model are used to convert the natural language query into a structured query tree. Combined with the semantic association network and permission mapping table, the optimal data node path is selected based on reinforcement learning to generate the query route. At the same time, materialized views and query rewriting technology are introduced to decompose complex queries into sub-queries for parallel execution to optimize the query statements. The reward function of reinforcement learning is

[0027] in, is the permission compliance, the value range is [0,1], Indicates data processing delay in milliseconds. is the predicted value of conflict probability, and the parallelism is based on the formula, calculate, Refers to the total amount of data to be processed, in bytes; Node capacity The upper limit of the amount of data that a data node can carry, in bytes; Priority The priority of the query task, the value range is [1-10]; S04 retrieves corresponding data from the data pool of each agency unit according to the query route, integrates the retrieved data, and finally displays the integrated data to the inquirer in real time and records the query log.

[0028] In the data preprocessing stage, the format of the data in the data pool of each agency unit is standardized, and semantic features are extracted and a semantic association network is constructed to record the semantic association relationship between data items. Then, query authorization is dynamically generated based on the multi-faceted information of the inquirer, and a permission mapping table is established to associate permissions with semantic network nodes to ensure accurate control of the inquirer's permissions. When the query conditions are received, the conditions are parsed using natural language processing technology, and query routes are generated and query statements are optimized in combination with the semantic association network and permission mapping table to efficiently locate the required data. Finally, data is retrieved according to the query route, displayed to the inquirer after fusion processing, and query logs are recorded for subsequent system optimization.

[0029] In terms of data integration, existing technologies focus on the unification of data formats, while this solution not only standardizes the format, but also deeply mines the semantic information of the data and builds a semantic association network. For example, in the data integration of the environmental protection and meteorological departments, the "pollutant emission data" of the environmental protection department and the "air quality data" of the meteorological department can be logically linked through the semantic association network, so that the originally isolated data can form an organic whole, providing support for complex data analysis and decision-making. In terms of authority management, existing technologies usually adopt static authorization methods and cannot adapt to dynamic changes. This solution dynamically generates authorization based on the identity, business needs and data sensitivity level of the inquirer to achieve fine-grained authority control. If a certain inquirer needs to temporarily access sensitive data in a specific time period due to work, the system can dynamically adjust the authorization according to its business needs, while meeting the needs and ensuring data security; first, the system will authenticate the inquirer. This can be achieved in a variety of ways, such as username / password combination, digital certificate, biometric technology such as fingerprint recognition, face recognition. The basic identity information of the inquirer, such as name, department, position, etc., is determined through identity authentication. When the inquirer submits a query request, the system will analyze the business needs behind the request. This may involve natural language processing and semantic understanding of query conditions. For example, in the query "get the sales department performance data of last month for market analysis report", the system can identify that the business demand is market analysis, the data involved is the sales department performance data, and the time range is last month. The system will determine the data resources and operation types that need to be accessed based on pre-defined business rules and data associations; dynamic authorization can set the validity period of permissions. For example, when the inquirer needs to access specific sensitive data due to a temporary task, the system only grants access rights during the task period, such as one week. Once the validity period expires, the permission automatically expires to prevent abuse of permissions. In terms of query processing, the query routing and statement optimization capabilities of the existing technology are limited. This solution uses natural language processing technology to directly parse the user's natural language query conditions, combine the semantic association network and the permission mapping table to generate the optimal query route, and intelligently optimize the query statement. For example, when the user enters "query the punishment of high-polluting enterprises in a certain area in the past month", the system can quickly and accurately retrieve relevant data from multiple government data pools, improving query efficiency and accuracy. In addition, the data fusion processing of this solution combines the semantic association network, which can effectively solve data conflicts and redundancy problems, ensuring that the data displayed to the queryer is accurate, complete and consistent. At the same time, the recording and analysis of query logs helps to continuously optimize system performance and authority management strategies, and further improve the quality and efficiency of cross-unit data management.

[0030] In S01, semantic feature extraction is performed using BERTVectorize the data items and build a semantic association network by combining named entity recognition and relation extraction technology. Use the language understanding ability of the pre-trained language model to convert the data items into vector form, identify the key entities in the data through named entity recognition, and use relation extraction technology to mine the semantic relationship between entities, thereby building a semantic association network, which records the semantic association relationship between different data items. Based on cosine similarity, a metadata mapping relationship across data pools is established: for data A and data B, their semantic similarity is

[0031] in, and for BERT The generated vector, To obtain the strength of association between relationships, attribute-based access control is introduced to adjust permissions by combining timestamp, geographic location, and data update frequency. Sim(A,B) When it is greater than the set threshold θ, a metadata mapping relationship between data A and data B is established, and an attribute-based access control model is introduced, where the permission model is uth ,U is the user, D is the data, is the attribute weight, Including sensitivity level and unit trust; dynamically adjust user permissions through timestamp, geographic location, and data update frequency.

[0032] in, P(U,D) represents the final access right of user U to data D, is the attribute weight, including the sensitivity level weight of data D and the trust weight of the unit to which it belongs, and its value range is [0,1]; are attribute values ​​related to user U and data D, including timestamp-related attribute values, geographic location-related attribute values, and data update frequency-related attribute values, and their value range is [0,1].

[0033] As a pre-trained language model, BERT in the above scheme has strong language understanding ability after training with a large-scale corpus. It can capture the deep semantic information in the data items. After converting the data items into vector form, these vectors can more accurately represent the semantic content of the data. NER can accurately identify key entities in the data, such as names of people, places, and names of organizations, while the relationship extraction technology can mine the semantic relationship between these entities. By constructing a semantic association network, the semantic association relationship between different data items can be clearly recorded, which helps to better understand the context and relevance of data in cross-unit data management. For example, in business data involving multiple units, the data interaction relationship between different units can be discovered through the semantic association network, providing more comprehensive information for data analysis and decision-making. At the same time, because the semantic similarity considers the similarity between vectors and the strength of association between entities, it can avoid the mismatch problem that may occur in the traditional query method based on keyword matching. For example, when querying "financial statements of a certain company", not only can data containing the keywords "a certain company" and "financial statements" be found, but also other data related to the company's finances, such as financial analysis reports, can be found through the semantic association network. In addition, the permission model takes into account multiple attributes, such as sensitivity level and unit trust, and can perform personalized permission allocation according to the characteristics of different users and data. Different data may have different sensitivity levels, and different users may belong to different units with different trust levels. By weighting these attributes, the most appropriate permissions can be assigned to each user and data combination. For example, for data with high sensitivity levels, only users from high-trust units with corresponding permissions can access it. At the same time, as factors such as time, geographic location, and data update frequency change, user permissions can be adjusted in real time. For example, when the frequency of data updates increases, the system can automatically increase the operation frequency limit of users with data update permissions; when users leave a specific geographic location, the system can automatically limit their access to certain data. This real-time permission adjustment mechanism can adapt to the dynamically changing business environment and ensure the security and availability of data.

[0034] In S02, the permission mapping table uses Neo4j to store the mapping relationship between permission nodes and data nodes, represents the permission granularity through edge weights, and uses a graph database to represent the mapping relationship between permissions and data; the mapping algorithm is ,in is the sensitive change threshold, recursively updates the relevant permissions, P represents the permissions, A function that recursively updates permissions.

[0035] In the above scheme, in cross-unit data management, the mapping relationship between permissions and data is often intricate, and different users, roles, data resources, and various permission combinations constitute a complex network structure. Using Neo4j to store the mapping relationship between permission nodes and data nodes can intuitively display these relationships in the form of a graph. Permission nodes and data nodes are vertices in the graph, and the edges between them represent the mapping relationship. The weight of the edge can clearly represent the permission granularity, such as read, write, modify and other different operation permissions. This intuitive representation method helps administrators quickly understand and manage the permission system; in cross-unit data management, business needs and permission systems may change over time. Using Neo4j to store permission mapping relationships, when you need to add new permissions, data nodes, or modify mapping relationships, you only need to add or modify the corresponding vertices and edges in the graph. The operation is simple and has little impact on the existing data structure. For example, when a new business department joins, you can easily create a new permission node for it and establish a mapping relationship with the relevant data nodes without making large-scale adjustments to the entire database structure.

[0036] In S04, the corresponding data is retrieved from the data pool of each agency, and the differences of heterogeneous data are eliminated by practical ontology alignment and data cleaning technology. The confidence of multi-source data is fused based on evidence theory, and data fusion is performed according to the fusion formula.

[0037] ,in For the i The basic probability distribution of data sources, The credibility of the data source; Build a conflict detection rule base and use game theory negotiation model to resolve conflicts; ,in is the unit priority, is the utility value of the conflict resolution method, Bel represents the confidence after data fusion, and Resolve represents the conflict resolution function. Maximize Represents the maximum value function.

[0038] In the above scheme, in cross-unit data management, the same data may exist in multiple data sources, and the reliability and accuracy of each data source may be different. Evidence theory can comprehensively consider the information of multiple data sources and distribute the information through basic probability. and data source credibility To calculate the confidence level of the data

[0039] This method can make full use of the complementarity of multi-source data and improve the accuracy and reliability of data fusion. For example, different data sources may give different estimates of the probability of an event. Through evidence theory fusion, a more reasonable comprehensive estimate can be obtained. In cross-unit data management, different units may have different interests and priorities for data. The game theory negotiation model can regard these units as participants in the game. By considering the unit priority and the utility value of conflict resolution To find the best conflict resolution method. This method can fully consider the interests of all parties, make the conflict resolution results more fair and reasonable, and improve the acceptance of data fusion results by all parties. For example, when dealing with data ownership conflicts, the game theory negotiation model can find a solution acceptable to all parties based on the importance of each unit and the degree of demand for data. At the same time, the game theory negotiation model is dynamic and can continuously adjust the conflict resolution method according to the actual situation. When the priorities of the units change or new interests emerge, the model can recalculate the optimal solution to ensure that the conflict is always effectively resolved. This enables the conflict resolution mechanism to adapt to the ever-changing business environment and data situation.

[0040] The format standardization process includes unifying the encoding format of text data in different formats, unifying the measurement unit and precision of numerical data. The data sources of the same agency are wide, and the text data may use multiple encoding formats, such as UTF-8, GBK, etc.; the measurement unit and precision of numerical data also vary greatly. After unifying the encoding format and measurement unit precision, these format differences can be eliminated, so that data from different units can be smoothly integrated. For example, in the data integration of the environmental protection department and the meteorological department, the monitoring report text of the environmental protection department may use GBK encoding, while the data text of the meteorological department uses UTF-8 encoding. After the unified encoding, the data of the two can be merged and processed in one system to build a comprehensive environmental information database. The unified encoding format and measurement unit precision can lower the threshold for data sharing, making the data of each unit easier to be understood and used by other units. For example, in the information sharing platform between government departments, data in a unified format can be directly called and analyzed by other departments, which improves the efficiency and effect of data sharing.

[0041] In the process of dynamically generating query authorization, the query frequency, query data type, and query time distribution information of the inquirer within a certain period of time are counted through the log analysis system. Based on these statistical results, if the inquirer frequently queries a certain type of data, more query permissions for this type of data will be provided to the inquirer when dynamically generating query authorization; if an abnormal query that differs greatly from the historical query mode occurs, the system automatically conducts strict review or restriction on the query authorization. By analyzing the inquirer's recent query behavior pattern and historical query records, its actual needs can be deeply understood. For example, for an inquirer who has been paying attention to the atmospheric pollution data of the environmental protection department for a long time, when the query authorization is dynamically generated later, the system can provide it with more data query permissions related to atmospheric pollution, including detailed data of specific areas and specific time periods, so that the authorization is more in line with the actual work needs of the inquirer; the historical query behavior of the inquirer usually has a certain regularity. If an abnormal query behavior that does not conform to the historical pattern occurs, it may mean that there is a data security risk. For example, a inquirer who usually only queries public data suddenly frequently requests sensitive data. By comparing its historical query records, the system can promptly discover this abnormality and conduct strict review or restriction on the query authorization, thereby effectively preventing security issues such as data leakage and illegal access.

[0042] In the S03, the data changes are monitored in real time. Once the sensitivity of the data changes, the permission settings of the corresponding data nodes in the permission mapping table are adjusted according to the preset sensitivity level rules. When the permission of the queryer changes, the mapping relationship between the corresponding queryer and the data node is directly modified in the permission mapping table. The permission mapping table uses the Neo4j graph database to store the mapping relationship between the permission node and the data node, and uses the edge weight to represent the permission granularity, and uses the characteristics of the graph database to achieve rapid updates. When the data update causes some originally public data to become sensitive data, the real-time update of the permission mapping table can timely limit the access rights of the queryer. For example, in the medical industry, some of the patient's examination data was originally open to specific departments within the hospital. If it is subsequently discovered that these data contain new sensitive information, the real-time update of the permission mapping table can prevent unauthorized personnel from continuing to access, effectively preventing data leakage; if the permission mapping table cannot be updated in real time, there may be a lag phenomenon in which the queryer's permissions do not match the actual situation. This will cause the queryer to be blocked due to insufficient permissions when they need to access certain data, or they can still access data when the permissions have been revoked, affecting work efficiency. Real-time updates can promptly eliminate this lag, allowing inquirers to successfully obtain the required data and improve the smoothness of the workflow.

[0043] The natural language processing technology includes word segmentation, part-of-speech tagging, named entity recognition and semantic understanding. Through the semantic understanding of query statements, the system can obtain more contextual information and user intentions, thereby making more intelligent decisions. For example, when a user queries "the environmental protection compliance status of enterprises in a certain industry in this city", the system can not only provide relevant compliance data, but also perform intelligent analysis based on the data, such as providing an overall environmental protection situation assessment of the industry, comparative analysis with other industries, etc., to provide users with more valuable decision support.

[0044] In the S03, the CPU usage rate, memory occupancy rate load information of each data pool, and the data storage location information are obtained in real time through the monitoring system. When generating the query route, the data pool that is closer to the query initiator or has a shorter data transmission path is given priority; when the data pool is in a high-load state, the query request is directed to the data pool with a lower load. The storage location of the data will affect the distance and time of data transmission. Considering the storage location of the data to generate the query route, the data pool that is closer to the query initiator or has a shorter data transmission path can be given priority. For example, in cross-regional unit data management, if the inquirer is located in area A, and the relevant data is stored in data pool B that is closer to area A, the query route generated by the system will give priority to data pool B, thereby reducing the data transmission time in the network and speeding up the query response speed; at the same time, when some data pools are in a high-load state, the speed of processing queries will be significantly reduced. Considering the load of the data pool, the system can direct the query request to the data pool with a lower load. For example, at a certain moment, the CPU usage and memory occupancy of data pool C are very high, while data pool D is in a light load state. The system will route the query to data pool D to avoid query delays caused by waiting for the high-load data pool to process.

[0045] In actual use, 1. Data Preprocessing For the data in the data pools of different agencies, the format is first standardized. For text data, professional text encoding detection tools, such as the chardet library, are used to identify the encoding format. If the monitoring report text of the environmental protection department is encoded in GBK and the data text of the meteorological department is encoded in UTF-8, they are uniformly converted to UTF-8 to ensure the compatibility of the data in subsequent processing. For numerical data, its original measurement unit and precision are determined by analyzing the metadata information and business rules of the data. For example, when integrating the data of the environmental protection and meteorological departments, it was found that the pollutant concentration data recorded by the environmental protection department was in ppm, and the relevant data of the meteorological department was in mg / m³. According to scientific unit conversion rules, all numerical data are unified into mg / m³, and the precision is unified to two decimal places, eliminating the differences in measurement units and precision of numerical data, laying the foundation for data integration.

[0046] The BERT model based on deep learning is used to extract semantic features from the processed standardized data. Taking the monitoring data of the environmental protection department as an example, the data is input into the pre-trained BERT model, and the model will conduct in-depth analysis of the data content, context, and related business rules. At the same time, a semantic association network is constructed by combining named entity recognition (NER) and relationship extraction technology. The NER technology is used to accurately identify key entities in the data, such as names of people, places, names of organizations, names of pollutants, etc.; the semantic relationships between entities are mined through relationship extraction technology, such as the causal relationship between "pollutant emissions" and "air quality", and the relationship between "enterprises" and "pollutant emission data". The extracted semantic features are used as nodes, and the semantic association relationships are used as edges to construct a semantic association network, clearly recording the semantic association relationships between different data items, which is convenient for subsequent data analysis and query processing.

[0047] Based on cosine similarity, a metadata mapping relationship across data pools is established. For data A and data B, their semantic similarity is

[0048] in, and for BERT The generated vector, is the strength of the relationship; the threshold θ is set through a large number of experiments and business experience, for example, θ=0.75. Sim(A,B) When it is greater than the set threshold θ, the metadata mapping relationship between data A and data B is established. The attribute-based access control model is introduced, and the permission model is uth ,U is the user, D is the data, is the attribute weight, Including sensitivity level and unit trust; dynamically adjust user permissions through timestamp, geographic location, and data update frequency.

[0049] in, P(U,D) represents the final access right of user U to data D, is the attribute weight, including the sensitivity level weight of data D and the trust weight of the unit to which it belongs, and its value range is [0,1]; is the attribute value related to user U and data D, including timestamp-related attribute value, geographic location-related attribute value, and data update frequency-related attribute value, and its value range is [0,1]. For example, for highly sensitive data, only users from highly trusted units with corresponding permissions can access it; when the data update frequency increases, the system automatically increases the operation frequency limit for users with data update permissions; when a user leaves a specific geographic location, the system automatically limits his or her access to certain data.

[0050] 2. Dynamically generate query authorization and permission mapping Collect the identity information of the inquirer, such as name, unit, position, etc. The collection of the identity information of the inquirer must be agreed by the inquirer, and the business requirements are obtained through the business requirements form filled out by the user in the system. The data owner or administrator sets the sensitivity level of the data according to the data content and relevant laws and regulations. At the same time, with the help of the log analysis system, statistics are collected on the query frequency, query data type, query time period and other information of the inquirer in the past period of time, such as the past three months. If it is found that a certain inquirer has been paying attention to the air pollution data of the environmental protection department for a long time, when generating the query authorization, more data query permissions related to air pollution are provided to him, including detailed data of specific areas and specific time periods, so that the authorization is more in line with the actual work needs of the inquirer. If the inquirer has abnormal query behavior that does not conform to the historical pattern, such as the inquirer who usually only queries public data suddenly frequently requests sensitive data, the system will detect this abnormality in time by comparing its historical query records, and strictly review or restrict the query authorization, effectively preventing security issues such as data leakage and illegal access.

[0051] Use Neo4j to store the mapping relationship between permission nodes and data nodes, and establish a permission mapping table. Permission nodes and data nodes are vertices in the graph, and the edges between them represent the mapping relationship. The weight of the edge is used to clearly represent the permission granularity, such as the read permission weight is 0.3, the write permission weight is 0.5, and the modify permission weight is 0.7. The mapping algorithm of the permission mapping table is: ,in is the sensitive change threshold, recursively updates the relevant permissions, P represents the permissions, is a function that recursively updates permissions. For example, when a new business department joins, a new permission node is created for it in the Neo4j graph database, and a mapping relationship is established with the relevant data nodes. The operation is simple and has little impact on the existing data structure. The permission mapping table will be updated in real time according to the update of data and the change of the queryer's permissions. When the data update causes some originally public data to become sensitive data, the system monitors the data changes in real time and adjusts the permission settings of the corresponding data nodes in the permission mapping table according to the preset sensitivity level rules; when the queryer's permissions change, the mapping relationship between the corresponding queryer and the data node is directly modified in the permission mapping table.

[0052] Query Processing When the query conditions are received, they are parsed using natural language processing technology. Natural language processing technology covers word segmentation, part-of-speech tagging, named entity recognition, and semantic understanding. Taking the query sentence "Query the penalties for high-polluting enterprises in a certain area in the past month" as an example, the Jieba word segmentation tool is used for word segmentation to obtain words such as "query", "recent month", "certain area", "high-polluting enterprises", and "penalty situation"; the part of speech is marked for each word through the NLTK library, such as "query" (verb), "recent month" (time phrase), "certain area" (place noun), "high-polluting enterprises" (noun phrase), and "penalty situation" (noun phrase); the named entity recognition model based on deep learning is used to identify entities in the query sentence, such as "certain area" (place entity), "high-polluting enterprises" (organization entity); Transformer The architecture model performs an overall semantic understanding of the query statement and obtains the user's query intent.

[0053] Combining the semantic association network and permission mapping table, the natural language query is converted into a structured query tree using dependency syntax analysis and intent recognition model. The query route is generated by selecting the optimal data node path based on reinforcement learning. The reward function of reinforcement learning is:

[0054] in, is the permission compliance, the value range is [0,1], Indicates data processing delay in milliseconds. is the predicted value of conflict probability. Through this reward function, the optimal data node path is selected by comprehensively considering permission compliance, data processing delay and conflict probability.

[0055] (II) Query statement optimization Materialized views and query rewriting technology are introduced to decompose complex queries into subqueries for parallel execution to optimize query statements. The degree of parallelism is based on the formula: calculate, Refers to the total amount of data to be processed, in bytes; Node capacity The upper limit of the amount of data that a data node can carry, in bytes; Priority is the priority of the query task, and its value range is [1-10]. For example, for a complex query involving a large amount of data, the parallel execution resources can be reasonably allocated according to the data volume, data node carrying capacity and query task priority to improve query efficiency.

[0056] 3. Data pool selection The monitoring system can obtain the CPU usage rate, memory occupancy rate and other load information of each data pool in real time, as well as the data storage location information, such as IP address and geographic location. When generating query routes, data pools that are closer to the query initiator or have shorter data transmission paths are given priority; when some data pools are in a high-load state, query requests are directed to data pools with lower loads. For example, in cross-regional unit data management, if the inquirer is located in area A, and the relevant data is stored in data pool B, which is closer to area A, and the load of data pool B is lower, the query route generated by the system will give priority to data pool B, reducing the data transmission time in the network and speeding up the query response speed; if the CPU usage rate and memory occupancy rate of data pool C are both high, and data pool D is in a light-load state, the system will direct the query route to data pool D to avoid query delays caused by waiting for the high-load data pool to process.

[0057] 4. Data retrieval and fusion processing According to the generated query route, the corresponding data is retrieved from the data pool of each agency. For example, according to the instructions of the query route, data on the punishment of high-polluting enterprises in a certain area is obtained from the data pool of the environmental protection department and the data pool of relevant law enforcement departments.

[0058] Ontology alignment and data cleaning techniques are used to eliminate the differences in heterogeneous data. Ontology alignment unifies the data at the semantic level by establishing mapping relationships between ontologies of different data sources; data cleaning removes noise, duplicate data, and erroneous data. Based on the evidence theory, the confidence of multi-source data is fused according to the fusion formula:

[0059] ,in For the i The basic probability distribution of data sources, The credibility of the data source; For example, for a company's environmental protection data, there may be multiple data sources such as the environmental protection department and the company's own monitoring system. Through evidence theory, the information of each data source can be comprehensively considered to improve the accuracy and reliability of data fusion.

[0060] 3. Conflict Resolution Build a conflict detection rule base and use game theory negotiation model to resolve conflicts; ,in is the unit priority, is the utility value of the conflict resolution solution, Bel represents the confidence after data fusion, Resolve represents the conflict resolution function, MaximizeRepresents the maximum function. When dealing with data ownership conflicts, the game theory negotiation model finds a solution acceptable to all parties based on the importance of each unit and the degree of demand for data. For example, when the environmental protection department and the enterprise have a dispute over the ownership of a piece of environmental protection data, the model determines the final ownership and use of the data by considering the priorities of both parties and the utility values ​​of different solutions.

[0061] 4. Data display and log recording The fused data is displayed to the inquirer in real time in the form of intuitive charts, reports, etc., which is convenient for the inquirer to understand and use. For example, the punishment of high-polluting enterprises in a certain area is displayed in a table, including information such as enterprise name, punishment time, punishment reason, and punishment result. At the same time, the query log is recorded to record the identity information of the inquirer, query time, query conditions, retrieved data, and detailed information of the query results. Through the query log, once a data security incident occurs, the inquirer involved and the operation time can be quickly located, and the incident investigation and responsibility tracing can be carried out; the query log is analyzed to understand the user's usage habits and demand preferences for data, and the system functions can be improved and expanded in a targeted manner.

[0062] The above are only embodiments of the present invention. Common knowledge such as the known specific structures and characteristics in the scheme is not described in detail here. Ordinary technicians in the relevant field are aware of all the common technical knowledge in the technical field to which the invention belongs before the application date or priority date, can obtain all the existing technologies in the field, and have the ability to apply conventional experimental means before that date. Ordinary technicians in the relevant field can improve and implement this scheme in combination with their own abilities under the enlightenment given by this application. Some typical known structures or known methods should not become obstacles for ordinary technicians in the relevant field to implement this application. It should be pointed out that for those skilled in the art, several deformations and improvements can be made without departing from the structure of the present invention, which should also be regarded as the scope of protection of the present invention, and these will not affect the effect of the implementation of the present invention and the practicality of the patent. The scope of protection required by this application shall be based on the content of its claims, and the specific implementation methods and other records in the specification can be used to interpret the content of the claims.

Claims

1. A cross-unit data management method, characterized in that: The following steps are involved: S01 standardizes the format of data in the data pool of each agency, converts data of different formats into a unified standard format, extracts semantic features for each data item, builds a semantic association network, records the semantic association relationship between different data items through the semantic association network, and implements dynamic authorization and permission mapping mechanism; S02 dynamically generates query authorization based on the inquirer's identity information, unit attributes, business requirements, and data sensitivity level, and establishes a permission mapping table to map the inquirer's permissions with data nodes in the semantic association network; S03 When receiving the query conditions, the query conditions are parsed using natural language processing technology, combined with the semantic association network and the permission mapping table, and the natural language query is converted into a structured query tree using dependency syntax analysis and intent recognition models. Combined with the semantic association network and the permission mapping table, the optimal data node path is selected based on reinforcement learning to generate the query route. At the same time, materialized views and query rewriting technology are introduced to decompose the query into sub-queries for parallel execution to optimize the query statement. The reward function of reinforcement learning is in, is the permission compliance, the value range is [0,1], Indicates data processing delay in milliseconds. is the predicted value of conflict probability, and the parallelism is based on the formula, calculate, Refers to the total amount of data to be processed, in bytes; Nodecapacity The upper limit of the amount of data that a data node can carry, in bytes; Priority The priority of the query task, the value range is [1-10]; S04 retrieves corresponding data from the data pool of each agency unit according to the query route, integrates the retrieved data, and finally displays the integrated data to the inquirer in real time and records the query log.

2. A cross-unit data management method according to claim 1, characterized in that: In S01, semantic feature extraction uses BERT to vectorize data items, and combines named entity recognition and relationship extraction technology to build a semantic association network. The language understanding ability of the pre-trained language model is used to convert the data items into vector form, and the key entities in the data are determined by named entity recognition. The semantic relationship between entities is mined by using the relationship extraction technology, so as to build a semantic association network, through which the semantic association relationship between different data items is recorded; Based on cosine similarity, a metadata mapping relationship across data pools is established: for data A and data B, their semantic similarity is in, and The vector generated by BERT, To obtain the strength of association between relationships, attribute-based access control is introduced to adjust permissions by combining timestamp, geographic location, and data update frequency. Sim(A,B) When it is greater than the set threshold θ, a metadata mapping relationship between data A and data B is established, and an attribute-based access control model is introduced, where the permission model is uth ,U is the user, D is the data, is the attribute weight, Including sensitivity level and unit trust; dynamically adjust user permissions through timestamp, geographic location, and data update frequency. in, P(U,D) represents the final access right of user U to data D, is the attribute weight, including the sensitivity level weight of data D and the trust weight of the unit to which it belongs, and its value range is [0,1]; are attribute values ​​related to user U and data D, including timestamp-related attribute values, geographic location-related attribute values, and data update frequency-related attribute values, and their value range is [0,1].

3. The cross-unit data management method according to claim 1, characterized in that: In S02, the permission mapping table uses Neo4j The mapping relationship between the storage permission nodes and data nodes is expressed by edge weights, and the complex mapping relationship between permissions and data is expressed by graph databases. The mapping algorithm is: ,in, is the sensitive change threshold, recursively updates the relevant permissions, P represents the permissions, A function that recursively updates permissions.

4. The cross-unit data management method according to claim 1, characterized in that: In S04, the corresponding data is retrieved from the data pool of each agency, and the differences of heterogeneous data are eliminated by practical ontology alignment and data cleaning technology. The confidence of multi-source data is fused based on evidence theory, and data fusion is performed according to the fusion formula. , in For the i The basic probability distribution of data sources, The credibility of the data source; Build a conflict detection rule base and use game theory negotiation model to resolve conflicts; ,in is the unit priority, is the utility value of the conflict resolution solution, Bel represents the confidence after data fusion, Resolve represents the conflict resolution function, Maxmize Represents the maximum value function.

5. The cross-unit data management method according to claim 1, characterized in that: The format standardization process includes unifying the encoding format of text data in different formats and unifying the measurement unit and precision of numerical data.

6. A cross-unit data management method according to claim 1, characterized in that: In the process of dynamically generating query authorization in S02, the query frequency, query data type, and query time distribution information of the inquirer within a certain period of time are counted through the log analysis system. Based on these statistical results, if the inquirer frequently queries a certain type of data, more query permissions for this type of data are provided to the inquirer when dynamically generating query authorization; If an abnormal query occurs that differs significantly from the historical query pattern, the system will automatically conduct strict review or restriction on the query authorization.

7. The cross-unit data management method according to claim 1, characterized in that: The changes of data are monitored in real time in S03. Once the sensitivity of the data changes, the permission settings of the corresponding data nodes in the permission mapping table are adjusted according to the preset sensitivity level rules. When the permission of the queryer changes, the mapping relationship between the corresponding queryer and the data node is directly modified in the permission mapping table. The permission mapping table uses the Neo4j graph database to store the mapping relationship between the permission node and the data node, and uses the edge weight to represent the permission granularity, and uses the characteristics of the graph database to achieve rapid updates.

8. The cross-unit data management method according to claim 1, characterized in that: The natural language processing technology includes word segmentation, part-of-speech tagging, named entity recognition and semantic understanding.

9. The cross-unit data management method according to claim 1, characterized in that: In S03, the CPU usage rate, memory occupancy rate load information, and data storage location information of each data pool are obtained in real time through the monitoring system. When generating a query route, a data pool that is closer to the query initiator or has a shorter data transmission path is preferentially selected; when a data pool is in a high-load state, the query request is directed to a data pool with a lower load.

Citation Information

Patent Citations

  • Data processing method and system

    CN118838968A

  • Management method and system for multi-terminal engineering interior data

    CN119377582A

  • General data management method and system based on identification analysis technology

    CN119378532A

  • Mixed analysis middleware based on cross-type database and query method

    CN119396848A

  • Firewall attacked surface carding and security reinforcement method

    CN119276632A

Cited By

  • Intelligent query system and method for power transmission and distribution production data based on natural language interaction

    CN120256451A

  • Data anonymization adjustment method and system based on dynamic association risk analysis

    CN120449213A

  • Dynamic authority management system and method based on multi-source salary data integration

    CN120541825A

  • Enterprise data information collection authority management method

    CN120705232A

  • Data production and application method based on index management

    CN120910102A