Hash power network-based multi-source heterogeneous environment data asset construction method

By processing multi-source heterogeneous environmental data and intelligently constructing data warehouses, thematic data marts, and knowledge query graphs, the integration and management problems of multi-source heterogeneous environmental data have been solved, achieving efficient storage and analysis of environmental data and meeting the needs of environmental informatization and smart environmental protection.

CN115858829BActive Publication Date: 2026-02-13JINAN ENVIRONMENTAL RES INST (JINAN YELLOW RIVER BASIN ECOLOGICAL PROTECTION PROMOTION CENT)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211514810.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-30
Publication Date
2026-02-13
Estimated Expiration
2042-11-30

AI Technical Summary

Technical Problem

Existing technologies cannot effectively integrate and manage multi-source heterogeneous environmental data, resulting in large differences in data structure, lack of correlation and integration, low application level, and inability to meet the needs of environmental information and smart environmental protection.

Method used

This approach employs multi-source heterogeneous environment data processing, data warehouse classification and construction, thematic data mart construction, knowledge query graph construction, and cross-media retrieval. It combines MPP distributed relational database, Hadoop cluster, in-memory database, and Storm streaming computing framework to process different types of data and establish multi-source heterogeneous environment data assets.

Benefits of technology

It enables efficient storage, management and analysis of environmental data, meets the application needs of environmental information technology and smart environmental protection, improves the standardization and relevance of data, and enhances the efficiency of data utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115858829B_ABST
    Figure CN115858829B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of environmental management big data, and discloses a multi-source heterogeneous environmental data asset construction method based on a computing power network, comprising the following steps: S1, multi-source heterogeneous environmental data processing; S2, data warehouse classification construction; S3, theme data mart construction; S4, knowledge query graph construction; and S5, cross-media retrieval application.The present application solves the problems of poor general standardization of environmental management data, lack of data correlation and integration between different business departments, and low application degree in actual business applications by means of multi-source heterogeneous environmental data processing, data warehouse classification construction, theme data mart construction, knowledge query graph construction, and cross-media retrieval application, efficiently completes the preprocessing and accurate classification of various environmental management data on the computing power network, meets the current application requirements for massive data management in the informationization construction of environmental protection work, and realizes the transformation of environmental protection informationization to digital environmental protection and smart environmental protection.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of environmental management big data, and particularly relates to a multi-source heterogeneous environmental data asset construction method based on computing power network. BACKGROUND

[0002] With the popularization of application of new generation information technologies such as big data, artificial intelligence and cloud computing, the demand of society as a whole for data storage, calculation, transmission and application is greatly improved, and the requirement for computing power with strong industry penetration and wide social applicability is higher and higher. The computing power network taking computing power as a carrier has developed rapidly, and the radiation to the environmental protection industry is also not exceptional. The sources of environmental data include various business systems of environmental protection, remote sensing images, thematic maps, social information, GPS positioning tracking and many other sources. The differences of various data in data structure are obvious, which constitutes the characteristics of multi-source heterogeneous environmental data. The computing power network is the deepening and upgrading of environmental management cloud network integration. In the multi-source heterogeneous environmental data on the computing power network, the single-dimensional data resources in the fields of water, air, soil, pollution, and pollution sources cannot meet the application requirements of environmental protection personnel, cannot guarantee the computing requirements of the computing power network, and cannot meet the application requirements of environmental protection personnel. The originally disorganized data such as pollution, monitoring and evaluation must be efficiently integrated and managed, and orderly data assets such as scientific research service library, pollution control library and environmental monitoring library must be constructed. This is not only the premise of informatization construction in the field of ecological environment, but also the basis for the progress of environmental informatization to digital environmental protection and smart environmental protection.

[0003] The access amount and the calculation amount of the data center on the computing power network in the era of big data are unprecedented. When the data is big enough, anything can happen. With the continuous development of various informationization construction of environmental management, the future data center will not only face the access amount of a base point, but also the access amount of all environmental protection departments in the province and even the whole country. At this time, the pressure faced by the data center will be unimaginable. In addition, the content in the environmental database is not only much, but also the structure has changed. Not only the two-dimensional table standard structure is stored, but also the unstructured data standard specification is not unified, such as office documents, texts, pictures, audios and videos, XML, HTML, various reports and the like. All the data is large and grows rapidly. In the face of such a large amount of data, the reading efficiency will be lower and lower. Therefore, the standardization of environmental data and the data warehouse architecture are not suitable for the high-speed reaction of the computing power network. In addition, the current various data systems of environmental management are numerous, the data sources are different, the data structures of various data are obviously different, and the data volume is huge, reaching TB level or even PB level. However, the environmental data is generally poor in standardization, the data between different business departments lacks correlation and integration, and the application degree is low in actual business application. The construction of environmental data assets crosses many fields such as environment, computer science, software and hardware technology, and the current technology does not fully utilize the research and development achievements in the field of machine learning and the like. Therefore, the service efficiency of the data asset construction cannot fully meet the needs of the actual environmental protection work, and the intelligent level still has a lot of room for improvement. SUMMARY

[0004] The purpose of the present application is to provide a multi-source heterogeneous environmental data asset construction method based on a computing power network to solve the problems in the background art.

[0005] To achieve the above purpose, the present application provides the following technical scheme:

[0006] A multi-source heterogeneous environmental data asset construction method based on a computing power network, comprising the following steps:

[0007] S1, multi-source heterogeneous environmental data processing: mixed processing of multi-source heterogeneous original environmental data on the computing power network, wherein the multi-source heterogeneous original environmental data includes structured environmental data, unstructured environmental data and real-time environmental data;

[0008] S2, data warehouse classification construction: extracting, converting and cleaning the processed multi-source heterogeneous environmental data, and establishing a data warehouse based on content, correlation, production standardization and time label to realize the classification construction of the data warehouse; and then classifying the historical and newly entered structured and unstructured environmental data of the user into the warehouse for conversion processing;

[0009] S3, theme data mart construction: according to the classified construction data warehouse, through the environment management object extraction, conversion and loading tool, from the closely related environmental data, intelligently extract the key words, keywords, and construct the theme data mart to meet the environmental business management analysis and prediction decision demand;

[0010] S4, knowledge query graph construction: application of syntax tree and semantic triple extraction in semantic analysis, the user's query statement is applied to the data warehouse, and all environmental data words that meet certain relationship are combined to form a phrase, and the problem that different entity identifiers in multi-source heterogeneous environmental data represent the same entity object is solved; then, one or more relationship phrases are formed by using entity extraction, relationship extraction and attribute extraction, and a knowledge query graph is constructed by using the object entity and the relationship network as the core structure;

[0011] S5, cross-media retrieval application: through the semantic association and the bottom feature between the multi-source heterogeneous media data, the similarity of the multi-source heterogeneous environmental media is established; meanwhile, according to the characteristics of the environmental data, the local typical correlation analysis and the multi-view learning technology are applied, and the unified expression model of the multi-source heterogeneous media is realized by combining the similarity of the multi-source heterogeneous media on the semantic level; then, the multi-source heterogeneous media information sorting algorithm is applied to carry out the structured analysis and management of the multi-source heterogeneous media, and the seamless identification and retrieval between the multi-source heterogeneous environmental management media are realized.

[0012] As a further scheme of the present application: the processing method of the structured environmental data in the S1 step: MPP distributed relational database is used for storage, the MPP distributed relational database uses Share-Nothing architecture, controls the host, operating system, memory and storage respectively, and each Segment is interconnected through IP network; the structured environmental data includes database data and format report data.

[0013] As a further scheme of the present application: the processing method of the unstructured environmental data in the S1 step: Hadoop cluster is used for data storage and calculation, HDFS provides distributed storage function, Map-Reduce provides offline batch calculation function, parallel processing is used to speed up the processing speed, and the reliability of the platform is guaranteed; the unstructured environmental data includes picture data, audio data and video data.

[0014] As a further scheme of the present application: the processing method of the real-time environmental data in the S1 step: the memory database+Storm streaming calculation framework is used for processing; the message data or real-time interactive data can be processed with extremely low delay, and the processed results are saved to the persistent medium; the real-time environmental data includes time-space remote sensing data.

[0015] As a further scheme of the present application: in the S2 step, the corresponding data warehouse can also be customized and constructed for the user according to the user's preferences and use logic.

[0016] As a further scheme of the present application: in the S2 step, the data warehouse includes a basic data warehouse, a configuration data warehouse, a business data warehouse, and a material data warehouse.

[0017] As a further scheme of the present application: in the S3 step, the corresponding theme data mart can also be customized and constructed for the user according to the construction classification of the data warehouse, to form flexible control of environmental management activities, and to provide environmental management decision-making release to the user in the form of a theme mart.

[0018] As a further scheme of the present application: in the S3 step, the theme data mart includes an environmental management object data mart, an environmental management activity data mart, and an environmental management decision-making application data mart.

[0019] As a further scheme of the present application: in the S4 step, the solution to the problem that different entity identifiers in the multi-source heterogeneous environmental data represent the same entity object is as follows: the entity identifiers are fused through a data preprocessing, partition indexing, and feature matching entity alignment process, to realize identification and merging of the same entity, multi-dimensional similarity comparison, and entity equivalence determination; and new associations between entities are established through background calculation reasoning, starting from the existing entity relationship data in the data warehouse.

[0020] Compared with the prior art, the present application has the following advantages:

[0021] The present application processes multi-source heterogeneous environmental data, constructs a data warehouse, constructs a theme data mart, constructs a knowledge query graph, and applies cross-media retrieval, solves the problems that environmental management data is generally poorly standardized, data between different business departments lacks correlation and integration, and the application degree is low in actual business applications, efficiently completes preprocessing and accurate classification of various environmental management data on the computing power network, meets the current application needs of massive data management in the informationization construction of environmental protection work, and realizes the transformation of environmental protection informationization to digital environmental protection and smart environmental protection. BRIEF DESCRIPTION OF DRAWINGS

[0022] Fig. 1 It is a functional architecture diagram of a multi-source heterogeneous environmental data asset construction method based on a computing power network;

[0023] Fig. 2 It is a hardware topology diagram of a multi-source heterogeneous environmental data asset construction method based on a computing power network;

[0024] Fig. 3 It is a technical architecture diagram of a multi-source heterogeneous environmental data asset construction method based on a computing power network. Detailed Implementation

[0025] Please see Figs. 1-3 In this embodiment of the invention, a method for constructing multi-source heterogeneous environmental data assets based on computing power networks is provided, such as... Fig. 1 As shown, it includes the following steps:

[0026] S1. Multi-source Heterogeneous Environment Data Processing: This step involves hybrid processing of raw environment data from multiple heterogeneous sources on the computing network. The raw environment data includes structured environment data, unstructured environment data, and real-time environment data. The processing method for structured environment data in step S1 is as follows: An MPP distributed relational database is used for storage. The MPP distributed relational database adopts a share-nothing architecture, controlling the host, operating system, memory, and storage separately. Each segment is interconnected via an IP network. The structured environment data includes database data and formatted data. The processing method for unstructured environment data in step S1 is as follows: A Hadoop cluster is used for data storage and computation. HDFS provides distributed storage functionality. MapReduce provides offline batch computing capabilities, accelerating processing speed through parallel processing and ensuring platform reliability; unstructured environmental data includes image data, audio data, and video data; the processing method for real-time environmental data in step S1 is to use an in-memory database + Storm streaming computing framework; it can process message data or real-time interactive data with extremely low latency and save the processed results to persistent media; real-time environmental data includes spatiotemporal remote sensing data;

[0027] By intelligently selecting appropriate data processing modes for different environmental data, efficient storage, management, and analysis of environmental data on the computing network can be achieved.

[0028] S2. Data Warehouse Classification and Construction: This involves extracting, transforming, and cleaning the processed multi-source heterogeneous environmental data, and establishing a content-based, associative, production-standardized, and time-stamped data warehouse. This achieves classified data warehouse construction. The data warehouse includes a basic data warehouse, a configuration data warehouse, a business data warehouse, and a documentation data warehouse. Custom data warehouses can also be built for users based on their preferences and usage logic. The classified data warehouses then process the user's historical and newly entered structured and unstructured environmental data through categorized storage and transformation, providing users with intelligent data integration and facilitating subsequent retrieval, querying, access, sharing, and distribution of environmental data on the computing network.

[0029] S3, theme data mart construction: according to the classification of the constructed data warehouse, through the environment management object extraction, conversion and loading tool, from the closely related environmental data, intelligently extract the keywords, keywords, and construct the theme data mart to meet the environmental business management analysis and prediction decision demand; The theme data mart includes environmental management object data mart, environmental management activity data mart and environmental management decision application data mart; According to the classification of the construction of data warehouse, the corresponding theme data mart can be constructed for user self-definition, so as to form flexible control of environmental management activities, and provide environmental management decision release to users in the form of theme market, so as to provide customized intelligent services for users; Through the theme market, the whole life cycle management of environmental data on the computing power network can be effectively tracked and guaranteed, and standardized services or customized intelligent services can be provided for users;

[0030] S4, knowledge query graph construction: application of syntax tree and semantic triple extraction in semantic analysis, application of user query statement to data warehouse, and composition of short phrases of all environmental data words meeting certain relationship, and solution of the problem that different entity identifiers in multi-source heterogeneous environmental data represent the same entity object; Then use entity extraction, relationship extraction and attribute extraction to form one or several relationship phrases, and build a knowledge query graph with object entity and relationship network as the core structure through reference resolution; The knowledge query graph provides support for knowledge representation, business relationship analysis and relationship search of environmental management in the form of visualization, browsing and analysis;

[0031] S5, cross-media retrieval application: through the semantic association and underlying features among multi-source heterogeneous media data, the similarity of multi-source heterogeneous environmental media is established; At the same time, according to the characteristics of environmental data, local typical correlation analysis and multi-view learning technology are applied, combined with the similarity of multi-source heterogeneous media at the semantic level, a unified expression model of multi-source heterogeneous media is realized; Then, the multi-source heterogeneous media information sorting algorithm is applied to carry out the structured analysis and management of multi-source heterogeneous media, and realize the seamless identification and retrieval among multi-source heterogeneous environmental management media; So as to solve the rapid positioning and searching of environmental management full data information, and carry out multi-source heterogeneous data media identification analysis for environmental multimedia information; The solution to the problem that different entity identifiers in multi-source heterogeneous environmental data represent the same entity object is as follows: the entity identifiers are fused through the entity alignment process of data preprocessing, partition index and feature matching, realizing the identification and merging of the same entity, multi-dimensional similarity comparison and entity equivalence determination; Starting from the existing entity relationship data in the data warehouse, new associations between entities are established through background calculation and reasoning, so as to expand and enrich the knowledge network, discover new knowledge from existing knowledge, and help users understand the hierarchy of environmental data parsing syntax structure on the computing power network, and solve the "long distance dependency" problem in natural language processing.

[0032] As Fig. 2 shown, it is exhibited in a hardware deployment manner. The computing power service provided by the existing computing center, edge computing node, etc. is inefficient, and the emergence of computing power network can better coordinate environmental management data resources and provide better services. The hardware deployment of the computing network center is connected by the computing nodes distributed in "cloud, management, edge, and end" through a gigabit network switch, which can real-time perceive the state of environmental data computing resources and network resources, and then intelligently allocate and schedule various data computing and service applications, forming a whole computing resource perceivable, allocatable, and schedulable multi-source heterogeneous environmental data cloud computing network. And the end computing node is associated with the regional computing power transaction center, which provides environmental data sources for each jurisdictional area through computing power providers, consumers, and AI empowerment, and connects the environmental protection special network, provides corresponding environmental data sources for each jurisdictional area, and each sub-area can provide application services for environmental-related workers based on the computing network data resources. Through the regional control center, it can realize cross-regional mutual recognition and communication of environmental management data analysis, provide cross-regional business services, and also provide a safe and reliable environmental management data resource computing power service platform for users.

[0033] As Fig. 3 shown, it is exhibited in a hardware deployment manner. The computing power service provided by the existing computing center, edge computing node, etc. is inefficient, and the emergence of computing power network can better coordinate environmental management data resources and provide better services. The hardware deployment of the computing network center is connected by the computing nodes distributed in "cloud, management, edge, and end" through a gigabit network switch, which can real-time perceive the state of environmental data computing resources and network resources, and then intelligently allocate and schedule various data computing and service applications, forming a whole computing resource perceivable, allocatable, and schedulable multi-source heterogeneous environmental data cloud computing network. And the end computing node is associated with the regional computing power transaction center, which provides environmental data sources for each jurisdictional area through computing power providers, consumers, and AI empowerment, and connects the environmental protection special network, provides corresponding environmental data sources for each jurisdictional area, and each sub-area can provide application services for environmental-related workers based on the computing network data resources. Through the regional control center, it can realize cross-regional mutual recognition and communication of environmental management data analysis, provide cross-regional business services, and also provide a safe and reliable environmental management data resource computing power service platform for users.

[0034] The above merely describes preferred specific embodiments of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art, according to the technical solution and inventive concept of the present application, makes equivalent replacement or change within the technical range disclosed by the present application, which should be covered within the protection scope of the present application.

Claims

1. A method for constructing multi-source heterogeneous environment data assets based on a computing power network, characterized in that, The method comprises the following steps: S1, multi-source heterogeneous environment data processing: mixed processing of multi-source heterogeneous raw environment data on the computing network, wherein the multi-source heterogeneous raw environment data includes structured environment data, unstructured environment data and real-time environment data; S2, data warehouse classification construction: extracting, converting and cleaning the processed multi-source heterogeneous environment data, and establishing a data warehouse based on content, correlation, production standardization and time label to realize the classification construction of the data warehouse; and then classifying the user history and newly entered structured and unstructured environment data through the classified data warehouse; S3, theme data mart construction: according to the classified data warehouse, the environment management object extraction, conversion and loading tool is used to intelligently extract keywords and key words from closely related environment data to construct a theme data mart to meet the needs of environmental business management analysis and prediction decision-making; S4, knowledge query graph construction: applying syntax tree and semantic triple extraction in semantic analysis, applying the user's query statement to the data warehouse, and forming a phrase of all environment data words meeting certain relationship, and solving the problem that different entity identifiers in multi-source heterogeneous environment data represent the same entity object; then using entity extraction, relationship extraction and attribute extraction to form one or several relationship phrases, and constructing a knowledge query graph with object entity and relationship network as the core structure through reference resolution; S5, cross-media retrieval application: through the semantic association and underlying features between multi-source heterogeneous media data, the similarity of multi-source heterogeneous environment media is established; at the same time, according to the characteristics of environment data, local typical correlation analysis and multi-view learning technology are applied, combined with the similarity of multi-source heterogeneous media on the semantic level, a unified expression model of multi-source heterogeneous media is realized; then a multi-source heterogeneous media information sorting algorithm is applied to realize the structured analysis and management of multi-source heterogeneous media, and realize the seamless identification and retrieval between multi-source heterogeneous environment management media.

2. The method according to claim 1, wherein, The processing method of structured environment data in step S1: MPP distributed relational database is used for storage, MPP distributed relational database uses Share-Nothing architecture, controls host, operating system, memory and storage respectively, and each Segment is interconnected through IP network; structured environment data includes database data and format report data.

3. The method of claim 1, wherein, The processing method of unstructured environment data in step S1: Hadoop cluster is used for data storage and calculation, HDFS provides distributed storage function, Map-Reduce provides offline batch calculation function, parallel processing is used to speed up the processing speed and ensure the reliability of the platform; unstructured environment data includes picture data, audio data and video data.

4. The method according to claim 1, wherein, The processing method of the real-time environmental data in the step S1: the processing is performed by using a memory database and a Storm streaming calculation framework; message data or real-time interaction data is processed at a very low delay, and the processed result is saved to a persistent medium; the real-time environmental data includes space-time remote sensing data.

5. The method of claim 1, wherein, The step S2 further includes self-defining and constructing a corresponding data warehouse for a user according to the user's preferences and use logic.

6. The computing power network-based multi-source heterogeneous environment data asset construction method according to claim 1, characterized in that, The data warehouse in the step S2 includes a basic data warehouse, a configuration data warehouse, a business data warehouse and material data.

7. The method of claim 1, wherein, The step S3 further includes self-defining and constructing a corresponding theme data mart for a user according to the construction classification of the data warehouse, so as to form flexible management and control of environmental management activities, and provide environmental management decision-making release to the user in a theme mart manner.

8. The method according to claim 1, wherein, The theme data mart in the step S3 includes an environmental management object data mart, an environmental management activity data mart and an environmental management decision-making application data mart.

9. The method according to claim 1, wherein, The solution to the problem that different entity identifiers in the multi-source heterogeneous environmental data represent the same entity object in the step S4 is as follows: the entity identifiers are fused by using a data preprocessing, partition index and feature matching entity alignment process, so as to realize the identification and merging of the same entity, multi-dimensional similarity comparison and entity equivalence determination; and a new association between entities is established by starting from the existing entity relationship data in the data warehouse and through background calculation reasoning.

Citation Information

Patent Citations

  • Urban multi-dimensional space multivariate heterogeneous information data processing method

    CN113821702A

  • Multi-source heterogeneous graph data fusion method and system based on supercomputing

    CN114399006A