Method, apparatus, computer device and product for constructing a basin information data architecture

By constructing a gravitational model of the basin data and generating a basin information data asset catalog, the problems of low availability and quality of basin data are solved, data availability and decision-making quality are improved, and operational efficiency and data circulation are optimized.

CN119441546BActive Publication Date: 2025-06-24THREE GORGES GROUP IND DEVELOPMENT (BEIJING) CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510039743.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-06-24
Estimated Expiration
2045-01-10

AI Technical Summary

Technical Problem

The availability of basin data is low, making it difficult to ensure the quality of basin data, thereby reducing the quality of analytical decision-making of basin information models.

Method used

By obtaining multi-source watershed data, building a watershed data gravitational model, dividing the watershed data domain and data topics, generating a watershed information data asset catalog, and numbering it to form a watershed information data architecture.

Benefits of technology

It has improved the availability and storage and retrieval efficiency of multi-source watershed data, greatly improved the quality of watershed data and decision-making quality such as basin cascade scheduling, optimized the operational efficiency of power generation and shipping, and enhanced data circulation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119441546B_ABST
    Figure CN119441546B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of basin information management, and discloses a method, device, computer equipment and product for constructing a basin information data architecture. The method includes: obtaining multi-source basin data, and constructing a basin data gravity model based on the multi-source basin data; dividing a basin data domain and data themes based on the basin data gravity model to generate a basin information data asset catalog; numbering and storing the basin information data asset catalog to generate a basin information data architecture. The present invention improves the usability and storage and retrieval efficiency of multi-source basin data, greatly improves the quality of basin data and decision-making quality such as cascade scheduling of the basin, optimizes the operation efficiency such as power generation and shipping, and enhances data circulation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of watershed information management, and particularly relates to a method, device, computer device and product for constructing a watershed information data architecture. Background Art

[0002] The watershed information model can provide a design basis for estimating the outflow process of the watershed, the river channel flood routing process, and the construction of small and medium-sized reservoirs, the development of irrigation canals, and the establishment of hydropower stations of different scales; in addition, it can provide useful data for urban water supply, shipping and transportation, etc., and provide decision-making support for flood control and water transfer scheduling, etc.

[0003] The watershed information model encompasses various types of watershed information and their relationships, including diverse data forms, different types, and varying scales, resulting in low availability of watershed data, making it difficult to ensure the quality of watershed data, and thus reducing the analysis and decision-making quality of the watershed information model. Summary of the Invention

[0004] In view of this, the present invention provides a method, device, computer device and product for constructing a watershed information data architecture to solve the problem of low availability of watershed data and difficulty in ensuring the quality of watershed data.

[0005] In a first aspect, the present invention provides a method for constructing a watershed information data architecture, the method comprising:

[0006] Obtaining multi-source watershed data and constructing a watershed data gravity model based on the multi-source watershed data;

[0007] Dividing the watershed data domain and data themes based on the watershed data gravity model to generate a catalog of watershed information data assets;

[0008] Numbering and storing the catalog of watershed information data assets to generate a watershed information data architecture.

[0009] The method for constructing a watershed information data architecture provided in this embodiment constructs a watershed data gravity model based on multi-source watershed data, divides the watershed data domain and data themes based on the watershed data gravity model to generate a catalog of watershed information data assets; numbers and stores the catalog of watershed information data assets to generate a watershed information data architecture; wherein, a watershed data gravity model is constructed according to the characteristics of multi-modal watershed data, and then the watershed data gravity model is used to form a catalog of watershed data assets and storage management, construct a watershed information data architecture, improve the availability and storage and retrieval efficiency of multi-source watershed data, greatly enhance the quality of watershed data and the decision-making quality of watershed cascade scheduling, etc., optimize the operation efficiency of power generation and shipping, etc., and enhance data circulation.

[0010] In an alternative embodiment, constructing a watershed data gravity model based on multi-source watershed data includes:

[0011] Perform multi-modal named entity recognition on multi-source basin data to obtain entities and the number of times the entities appear;

[0012] Determine the interaction force between entities based on the entities and the number of times the entities appear;

[0013] Associate entities using the interaction force between entities to generate a basin data gravitational model.

[0014] A method for constructing a basin information data architecture provided in this embodiment determines the interaction force between entities through entities and the number of times the entities appear, associates entities using the interaction force between entities, generates a basin data gravitational model, generates a basin data gravitational model for characteristics such as multi-modal basin data, defines the gravitational force of basin information data, and lays a foundation for subsequent formation of a basin information data asset catalog and numbered storage.

[0015] In an alternative implementation, divide the basin data domain and data topics based on the basin data gravitational model to generate a basin information data asset catalog, including:

[0016] Construct a data gravitational graph based on the basin data gravitational model and calculate the attraction ability of vertices in the data gravitational graph;

[0017] Divide data topics based on the attraction ability of vertices to obtain the data topic names corresponding to the divided data topics;

[0018] Obtain the subject classification names and calculate the edit distance between the data topic names and the subject classification names;

[0019] Aggregate the basin data domains corresponding to the data topic names based on the edit distance to obtain the basin information data asset catalog.

[0020] A method for constructing a basin information data architecture provided in this embodiment constructs a data gravitational graph through the basin data gravitational model, transforms the problem of structured organization of basin data into the problem of graph node contribution degree, and then divides data topics and basin data domains through the attraction ability of vertices in the data gravitational graph, establishes a basin information data asset catalog, solves the problems of standardization and structured organization of basin data, realizes accurate division and clustering of data topics and basin data domains, improves the usability of data, and provides a standard data organization architecture.

[0021] In an alternative implementation, construct a data gravitational graph based on the basin data gravitational model and calculate the attraction ability of vertices in the data gravitational graph, including:

[0022] Use the entities in the basin data gravitational model as vertices, use the connected routes between entities as edges, determine the weights of the edges based on the interaction force between entities, and construct a data gravitational graph based on the vertices, edges, and edge weights respectively;

[0023] Determine the degree of vertices, the number of vertex occurrences, and the gravitational force of vertex connection routes based on the data gravity graph, and calculate the attraction ability of vertices based on the degree of vertices, the number of vertex occurrences, and the gravitational force of vertex connection routes.

[0024] A method for constructing a basin information data architecture provided in this embodiment constructs a data gravity graph through a basin data gravity model, and then calculates the attraction ability of vertices in the data gravity graph, providing a basis for subsequent division of data themes and basin data domains.

[0025] In an alternative implementation, divide data themes based on the attraction ability of vertices to obtain the data theme names corresponding to the divided data themes, including:

[0026] Determine the division ratio based on the attraction ability of vertices, and determine the division core vertices based on the division ratio;

[0027] Obtain the connection relationships between other vertices and the division core vertices, and divide data themes based on the connection relationships between other vertices and the division core vertices;

[0028] Perform keyword recognition on the entities corresponding to the divided data themes to obtain the data theme names.

[0029] A method for constructing a basin information data architecture provided in this embodiment determines the division core vertices through the attraction ability of vertices, and then divides data themes by using the connection relationships between other vertices and the division core vertices, realizing the accurate division of basin information data corresponding to different data themes, as well as the multi-source data integration, fusion, and structured organization of basin information data.

[0030] In an alternative implementation, number and store the basin information data asset catalog to generate a basin information data architecture, including:

[0031] Define the entity nodes corresponding to the basin information data asset catalog, number the entity nodes, and add labels to the entity nodes based on the numbers;

[0032] Use a graph database to determine the storage management information of the basin information data asset catalog;

[0033] Construct a basin information data architecture based on the basin information data asset catalog after adding labels and the storage management information.

[0034] A method for constructing a basin information data architecture provided in this embodiment adds tags by numbering the basin information data asset catalog, and uses a graph database for storage management, achieving standard data organization and storage, improving storage and retrieval efficiency, and further supporting basin data sharing, data analysis, and data research by constructing the basin information data architecture, which will greatly improve the quality of basin data, the quality of decision-making such as cascade scheduling of the basin, the operation efficiency of optimizing power generation and shipping, and enhance data circulation.

[0035] In a second aspect, the present invention provides a device for constructing a basin information data architecture, which includes:

[0036] A construction module for obtaining multi-source basin data and constructing a basin data gravity model based on the multi-source basin data;

[0037] A division module for dividing the basin data domain and data topics based on the basin data gravity model to generate a basin information data asset catalog;

[0038] A numbering and storage module for numbering and storing the basin information data asset catalog to generate a basin information data architecture.

[0039] In a third aspect, the present invention provides a computer device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the method for constructing a basin information data architecture according to the first aspect or any corresponding embodiment thereof.

[0040] In a fourth aspect, the present invention provides a computer-readable storage medium, on which computer instructions are stored, and the computer instructions are used to cause a computer to execute the method for constructing a basin information data architecture according to the first aspect or any corresponding embodiment thereof.

[0041] In a fifth aspect, the present invention provides a computer program product, including computer instructions, and the computer instructions are used to cause a computer to execute the method for constructing a basin information data architecture according to the first aspect or any corresponding embodiment thereof. Description of the Drawings

[0042] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required to be used in the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0043] Figure 1It is a schematic flowchart of a method for constructing a basin information data architecture according to an embodiment of the present invention;

[0044] Figure 2 It is a schematic flowchart of another method for constructing a basin information data architecture according to an embodiment of the present invention;

[0045] Figure 3 It is a schematic structural diagram of a data gravity map according to an embodiment of the present invention;

[0046] Figure 4 It is a schematic diagram of a basin information data asset catalog according to an embodiment of the present invention;

[0047] Figure 5 It is a schematic flowchart of yet another method for constructing a basin information data architecture according to an embodiment of the present invention;

[0048] Figure 6 It is a structural block diagram of a device for constructing a basin information data architecture according to an embodiment of the present invention;

[0049] Figure 7 It is a schematic hardware structure diagram of a computer device according to an embodiment of the present invention. Specific embodiments

[0050] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0051] The embodiments of the present invention provide a method for constructing a basin information data architecture, which generates a basin data gravity model for characteristics such as multi-modal basin data, forms a basin information data asset catalog and number storage based on the undirected graph of the gravity model, and constructs a basin information data architecture; among them, according to characteristics such as multi-modal basin data, the data gravity theory is proposed, a basin data gravity model is constructed, a basin data asset catalog is formed based on the data gravity model, and graph data is used for number storage to construct a basin information data architecture, providing standardized data organization and storage, improving data availability and storage and retrieval efficiency, greatly enhancing basin data quality, decision-making quality such as basin cascade scheduling, optimizing operation efficiency such as power generation and shipping, and enhancing data circulation.

[0052] An embodiment of the present invention provides a method for constructing a basin information data architecture. It should be noted that the execution subject of the method for constructing the basin information data architecture provided by the embodiment of the present invention can be a device for constructing the basin information data architecture. The device for constructing the basin information data architecture can be implemented as part or all of an electronic device through software, hardware, or a combination of software and hardware. Among them, the electronic device can be a server or a terminal. Among them, the server in the embodiments of the present application can be a single server or a server cluster composed of multiple servers. The terminal in the embodiments of the present application can be other intelligent hardware devices such as a smart phone, a personal computer, a tablet computer, a wearable device, and a smart robot. In the following method embodiments, the execution subject is taken as an electronic device as an example for illustration.

[0053] According to an embodiment of the present invention, an embodiment of a method for constructing a basin information data architecture is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And, although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0054] In this embodiment, a method for constructing a basin information data architecture is provided, which can be used for the above-mentioned electronic device. Figure 1 It is a flowchart of a method for constructing a basin information data architecture according to an embodiment of the present invention, as Figure 1 shown, the process includes the following steps:

[0055] Step S101, obtain multi-source basin data, and construct a basin data gravity model based on the multi-source basin data.

[0056] Specifically, the multi-source basin data includes structured data, among which data is collected and recorded through sensors, and the collected and recorded data includes 3D models, as well as unstructured data and semi-structured data such as text, tables, emails, pictures, audio, video, and web pages.

[0057] Further, according to the characteristics of multi-modal basin data, etc., a data gravity theory is proposed to construct a basin data gravity model.

[0058] Step S102, divide the basin data domain and data themes based on the basin data gravity model, and generate a catalog of basin information data assets.

[0059] Specifically, the basin data domain and data themes are divided based on the subject division standard to form a catalog of basin information data assets.

[0060] Step S103, number and store the catalog of basin information data assets to generate a basin information data architecture.

[0061] Specifically, based on the basin information data asset catalog, each node is numbered and tagged, and a graph database is applied for storage.

[0062] A method for constructing a basin information data architecture provided in this embodiment constructs a basin data gravity model based on multi-source basin data, divides the basin data domain and data themes based on the basin data gravity model, and generates a basin information data asset catalog; numbers and stores the basin information data asset catalog to generate a basin information data architecture. Among them, a basin data gravity model is constructed according to the characteristics of multi-modal basin data, and then the basin data gravity model is used to form a basin data asset catalog and storage management, construct a basin information data architecture, improve the availability and storage and retrieval efficiency of multi-source basin data, greatly improve the quality of basin data and the quality of decision-making such as cascade scheduling of the basin, optimize the operation efficiency such as power generation and shipping, and enhance data liquidity.

[0063] In this embodiment, a method for constructing a basin information data architecture is provided, which can be used in the above-mentioned electronic device. Figure 2 It is a flowchart of a method for constructing a basin information data architecture according to an embodiment of the present invention, as Figure 2 shown, and this process includes the following steps:

[0064] Step S201, obtain multi-source basin data, and construct a basin data gravity model based on the multi-source basin data.

[0065] Specifically, the above step S201 includes:

[0066] Step S2011, perform multi-modal named entity recognition on the multi-source basin data to obtain entities and the number of times the entities appear.

[0067] Specifically, the multi-modal named entity recognition (Multimodal Named Entity Recognition, abbreviated as MNER) method is used to identify entities in the multi-source basin data, take the type of the multi-source basin data as the attribute of the entity, and count the number of times each entity appears, as well as the number of times two entities appear simultaneously.

[0068] Step S2012, determine the mutual force between entities based on the entities and the number of times the entities appear.

[0069] Specifically, assume that the number of times entity n appears is , the number of times entity n and entity p appear together is , and define the mutual force between two directly associated entities as , and its calculation formula is as follows:

[0070] (1)

[0071] Among them, represents the occurrence times of entity , represents the occurrence times of entity , represents the occurrence times of entity and entity co - occurrence times.

[0072] Furthermore, define the interaction force between two non - directly associated entities (only in the case where two entities can be associated through an intermediate entity) as , and its calculation formula is as follows:

[0073] ( ) (2)

[0074] Among them, b is the intermediate node between entity a and c. If entity a has multiple branches connecting to entity c, the gravitational force between entity a and entity c is the sum of the gravitational forces of all paths not exceeding 3.

[0075] Step S2013, use the interaction force between entities to associate entities and generate a gravitational model of watershed data.

[0076] Specifically, define the gravitational model of watershed data as: entities interact with each other through gravity and are connected into an associated model body, and there can be multiple gravitational models of watershed data.

[0077] Step S202, based on the gravitational model of watershed data, divide the watershed data domain and data themes, and generate a catalog of watershed information data assets.

[0078] Specifically, the above - mentioned step S202 includes:

[0079] Step S2021, construct a data gravity graph based on the gravitational model of watershed data and calculate the attraction ability of vertices in the data gravity graph.

[0080] In some alternative embodiments, the above - mentioned step S2021 includes:

[0081] Step a1, use the entities in the gravitational model of watershed data as vertices, use the connected routes between entities as edges, and determine the weights of the edges based on the interaction force between entities. Then construct a data gravity graph based on vertices, edges, and the weights of the edges respectively.

[0082] Specifically, as Figure 3 shown, take each entity in the gravitational model of watershed data as a vertex, connect an edge between entities with a gravitational relationship, and use the sum of the gravitational forces of the connected routes as the weight of the edge to construct multiple weighted undirected graphs, called data gravity graph G.

[0083] Step a2: Determine the degree of vertices, the number of times vertices appear, and the gravitational force of vertex connection routes based on the data gravity graph, and calculate the attraction ability of vertices based on the degree of vertices, the number of times vertices appear, and the gravitational force of vertex connection routes.

[0084] Specifically, for each data gravity graph G, calculate the attraction ability of each vertex, and its calculation formula is as follows:

[0085] (3)

[0086] Among them, represents the attraction ability of vertex n, represents the gravitational force (or interaction force) of the connection route of vertex n. This connection route is a path with no more than 3 entities, represents the degree of vertex n, represents the sum of the number of times vertex n appears.

[0087] Step S2022: Divide data themes based on the attraction ability of vertices to obtain the data theme names corresponding to the divided data themes.

[0088] In some alternative embodiments, the above step S2022 includes:

[0089] Step b1: Determine the division ratio based on the attraction ability of vertices, and determine the division core vertices based on the division ratio.

[0090] Specifically, count the attraction ability of all vertices, set the division ratio q (q is generally less than 0.2), sort the vertices from largest to smallest according to the attraction ability of vertices, and take the top q vertices as the division core vertices.

[0091] Step b2: Obtain the connection relationship between other vertices and the division core vertices, and divide data themes based on the connection relationship between other vertices and the division core vertices.

[0092] Specifically, other vertices (that is, the remaining vertices in the data gravity graph except the division core vertices) have the following situations:

[0093] 1) If it is directly connected to one of the division core vertices and not directly connected to other division core vertices, then divide the data theme corresponding to the other division core vertices into the data theme of this division core vertex.

[0094] 2) If it is directly connected to one of the division core vertices and directly connected to other division core vertices, then calculate the gravitational value from this division core vertex to the connected division core vertices, and divide the data theme corresponding to the other division core vertices into the data theme corresponding to the maximum gravitational value.

[0095] 3) If not directly connected to the core vertex of the partition, calculate the paths from the directly connected vertices to each core vertex, and partition them into the data topics corresponding to the shortest paths.

[0096] Step b3, identify keywords for the entities corresponding to the partitioned data topics to obtain the data topic names.

[0097] Step S2023, obtain the subject classification names and calculate the edit distances between the data topic names and the subject classification names.

[0098] Specifically, the subject classification names can adopt the second- and third-level subject classification names in the subject classification standard, and calculate the edit distances ds and dt between the data topic names and the second- and third-level subject classification names.

[0099] Step S2024, aggregate the basin data domains corresponding to the data topic names based on the edit distances to obtain the basin information data asset catalog.

[0100] Specifically, if the dt of a certain data topic name is less than or equal to 2, then use ds as the data domain of this data topic name.

[0101] Furthermore, as Figure 4 shown, the basin information data asset catalog is divided into 3 levels, including the L1 data domain (i.e., the basin data domain), the L2 data topic, and the L3 concept entity. Among them, the integrated logical entity and the corresponding attributes (including files, documents, audio, video materials, etc. of multimodal data) are the L3-level concept entities.

[0102] Step S203, number and store the basin information data asset catalog to generate the basin information data architecture. For details, please refer to Figure 1 the steps of Embodiment shown in S103, which will not be elaborated here.

[0103] A method for constructing a basin information data architecture provided by this embodiment determines the interaction forces between entities through the entities and the number of entity occurrences, associates the entities using the interaction forces between entities to generate a basin data gravitational model, generates a basin data gravitational model for characteristics such as multimodality of basin data, defines the basin information data gravity, and lays a foundation for subsequent formation of the basin information data asset catalog and numbered storage; secondly, constructs a data gravity graph through the basin data gravitational model, transforms the problem of structured organization of basin data into the problem of vertex contribution degrees in the graph, and then divides the data topics and basin data domains through the attraction ability of the vertices in the data gravity graph, establishes the basin information data asset catalog, solves the problems of standardization and structured organization of basin data, realizes accurate partitioning and clustering of data topics and basin data domains, improves the availability of data, and provides a standard data organization architecture.

[0104] In this embodiment, a method for constructing a watershed information data architecture is provided, which can be used in the above-mentioned electronic device. Figure 5 It is a flowchart of a method for constructing a watershed information data architecture according to an embodiment of the present invention. As Figure 5 shown, the process includes the following steps:

[0105] Step S501, obtain multi-source watershed data, and construct a watershed data gravity model based on the multi-source watershed data. For details, please refer to Figure 2 step S201 of the embodiment shown, which will not be elaborated here.

[0106] Step S502, divide the watershed data domain and data theme based on the watershed data gravity model, and generate a watershed information data asset catalog. For details, please refer to Figure 2 step S202 of the embodiment shown, which will not be elaborated here.

[0107] Step S503, number and store the watershed information data asset catalog to generate a watershed information data architecture.

[0108] Specifically, the above step S503 includes:

[0109] Step S5031, define entity nodes corresponding to the watershed information data asset catalog, number the entity nodes, and add labels to the entity nodes based on the numbers.

[0110] Specifically, represent the conceptual entities in the watershed information data asset catalog as entity nodes, represent the gravity between entities as edges, and add attributes to each entity node.

[0111] Furthermore, number each level and entity node in the watershed information data asset catalog respectively. The number of each conceptual entity is: "data domain number - data theme number - conceptual entity number - logical entity number", and the numbers of the remaining levels are the successive parent node numbers of the entity node until the node itself number. Then, use the numbers as labels for each entity node to improve the subsequent query efficiency.

[0112] Step S5032, use a graph database to determine the storage management information of the watershed information data asset catalog.

[0113] Specifically, use a graph database (such as Neo4j, ArangoDB, Amazon Neptune, etc.) to store the watershed data.

[0114] Step S5033, construct a watershed information data architecture based on the watershed information data asset catalog with added labels and the storage management information.

[0115] A method for constructing a basin information data architecture provided by this embodiment adds tags by numbering the basin information data asset catalog, and stores and manages it using a graph database, achieving standard data organization and storage, improving storage and retrieval efficiency, and further supporting basin data sharing, data analysis, and data research by constructing a basin information data architecture, which will greatly improve the quality of basin data, the quality of decision-making such as cascade scheduling of the basin, optimize the operation efficiency such as power generation and shipping, and enhance data circulation.

[0116] In this embodiment, a device for constructing a basin information data architecture is also provided. This device is used to implement the above-mentioned embodiment and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that realizes a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0117] This embodiment provides a device for constructing a basin information data architecture, as Figure 6 shown, including:

[0118] A construction module 601, configured to obtain multi-source basin data and construct a basin data gravity model based on the multi-source basin data;

[0119] A division module 602, configured to divide a basin data domain and data topics based on the basin data gravity model, and generate a basin information data asset catalog;

[0120] A numbering and storage module 603, configured to number and store the basin information data asset catalog to generate a basin information data architecture.

[0121] In some optional implementation manners, the construction module 601 includes:

[0122] An identification unit, configured to perform multi-modal named entity recognition on the multi-source basin data to obtain entities and the number of times the entities appear;

[0123] A first determination unit, configured to determine the interaction force between entities based on the entities and the number of times the entities appear;

[0124] A generation unit, configured to associate entities using the interaction force between entities to generate a basin data gravity model.

[0125] In some optional implementation manners, the division module 602 includes:

[0126] A first calculation unit, configured to construct a data gravity graph based on the basin data gravity model and calculate the attraction ability of vertices in the data gravity graph;

[0127] A partitioning unit, configured to partition data topics based on the attraction ability of vertices, and obtain the data topic names corresponding to the partitioned data topics;

[0128] A second calculation unit, configured to obtain subject classification names and calculate the edit distance between the data topic names and the subject classification names;

[0129] An aggregation unit, configured to aggregate the basin data domains corresponding to the data topic names based on the edit distance, and obtain a basin information data asset catalog.

[0130] In some alternative embodiments, the first calculation unit includes:

[0131] A construction subunit, configured to use the entities in the basin data gravity model as vertices, the connection routes between the entities as edges, and determine the weights of the edges based on the mutual forces between the entities, and respectively construct a data gravity graph based on the vertices, edges, and the weights of the edges;

[0132] A calculation subunit, configured to determine the degree of vertices, the number of vertex occurrences, and the gravity of the vertex connection routes based on the data gravity graph, and calculate the attraction ability of the vertices based on the degree of vertices, the number of vertex occurrences, and the gravity of the vertex connection routes.

[0133] In some alternative embodiments, the partitioning unit includes:

[0134] A determination subunit, configured to determine a partitioning ratio based on the attraction ability of vertices, and determine partitioning core vertices based on the partitioning ratio;

[0135] A partitioning subunit, configured to obtain the connection relationships between other vertices and the partitioning core vertices, and partition data topics based on the connection relationships between other vertices and the partitioning core vertices;

[0136] An identification subunit, configured to perform keyword identification on the entities corresponding to the partitioned data topics to obtain data topic names.

[0137] In some alternative embodiments, the number storage module 603 includes:

[0138] A definition unit, configured to define entity nodes corresponding to the basin information data asset catalog, number the entity nodes, and add labels to the entity nodes based on the numbers;

[0139] A second determination unit, configured to use a graph database to determine the storage management information of the basin information data asset catalog;

[0140] A construction unit, configured to construct a basin information data architecture based on the basin information data asset catalog with added labels and the storage management information.

[0141] The further function descriptions of the above-mentioned various modules and units are the same as those in the corresponding above-mentioned embodiments, and will not be elaborated here.

[0142] The device for constructing a basin information data architecture in this embodiment is presented in the form of functional units. Here, the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and a memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0143] The embodiment of the present invention further provides a computer device having the above Figure 6 shown device for constructing a basin information data architecture.

[0144] Please refer to Figure 7 , Figure 7 which is a schematic structural diagram of a computer device provided by an optional embodiment of the present invention. As Figure 7 shown, the computer device includes: one or more processors 10, a memory 20, and an interface for connecting each component, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (such as an array of servers, a set of blade servers, or a multi-processor system). Figure 7 In

[0145] FIG. 1, a processor 10 is taken as an example.

[0146] The memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.

[0147] The memory 20 may include a program storage area and a data storage area. The program storage area may store an operating system and application programs required for at least one function. The data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 may optionally include a memory remotely provided with respect to the processor 10, and these remote memories may be connected to the computer device through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0148] The memory 20 may include a volatile memory, such as a random access memory. The memory may also include a non-volatile memory, such as a flash memory, a hard disk, or a solid-state drive. The memory 20 may also include a combination of the above types of memories.

[0149] The computer device further includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30, and the output device 40 may be connected through a bus or other means. Figure 7 Taking connection through a bus as an example.

[0150] The input device 30 may receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the computer device, such as a touch screen, a keypad, a mouse, a trackpad, a touchpad, a pointing stick, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 may include a display device, an auxiliary lighting device (such as an LED), and a tactile feedback device (such as a vibration motor), etc. The above-mentioned display device includes, but is not limited to, a liquid crystal display, a light-emitting diode, a display, and a plasma display. In some alternative embodiments, the display device may be a touch screen.

[0151] Embodiments of the present invention also provide a computer-readable storage medium. The method according to the embodiments of the present invention can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented as computer code that is originally stored in a remote storage medium or a non-transitory machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored as such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method shown in the above embodiments is implemented.

[0152] A part of the present invention can be applied as a computer program product, such as computer program instructions, which when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should be able to understand that the forms of existence of computer program instructions in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Herein, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to the computer.

[0153] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A method for constructing a watershed information data architecture, characterized in that: The method comprises: Acquire multi-source watershed data, and construct a watershed data gravity model based on the multi-source watershed data; Divide the watershed data domains and data themes based on the watershed data gravity model, and generate a watershed information data asset catalog; The watershed information data asset directory is numbered and stored to generate a watershed information data architecture; The constructing of a watershed data gravity model based on the multi-source watershed data includes: Performing multimodal named entity recognition on the multi-source watershed data to obtain entities and entity occurrence counts; Determine the interaction force between entities based on the entities and the number of occurrences of the entities; assuming that the number of occurrences of entity n is , the number of times entity n and entity p appear together is , the interaction force between two directly associated entities is defined as , the calculation formula is as follows: in, Representing Entities Number of occurrences, Representing Entities Number of occurrences, Representing Entities and entities co-occurrence count; The interaction force between two entities that are not directly related is defined as , the calculation formula is as follows: ( ) Where b is the intermediate node between entities a and c. If entity a has multiple branches connected to entity c, the gravitational force between entity a and entity c is the sum of the gravitational forces of all paths that do not exceed 3. Associating the entities by using the interaction forces between the entities to generate the watershed data gravity model; The dividing of the watershed data domains and data themes based on the watershed data gravity model to generate a watershed information data asset directory includes: Constructing a data gravity graph based on the watershed data gravity model, and calculating the attraction capacity of vertices in the data gravity graph; Dividing the data topics based on the attraction capabilities of the vertices to obtain data topic names corresponding to the divided data topics; Obtain a subject classification name, and calculate the edit distance between the data subject name and the subject classification name; The watershed data domains corresponding to the data subject names are aggregated based on the edit distance to obtain the watershed information data asset directory.

2. The method according to claim 1, characterized in that The step of constructing a data gravity graph based on the watershed data gravity model and calculating the attraction capacity of vertices in the data gravity graph includes: The entities in the watershed data gravity model are taken as vertices, the connection routes between the entities are taken as edges, and the weights of the edges are determined based on the interaction forces between the entities, and the data gravity graph is constructed based on the vertices, the edges and the weights of the edges respectively; The degree of the vertex, the number of vertex occurrences and the gravity of the vertex connection path are determined based on the data gravity graph, and the attraction ability of the vertex is calculated based on the degree of the vertex, the number of vertex occurrences and the gravity of the vertex connection path.

3. The method according to claim 1, characterized in that The dividing of the data topics based on the attraction ability of the vertices to obtain data topic names corresponding to the divided data topics includes: Determine a division ratio based on the attraction ability of the vertices, and determine a division core vertex based on the division ratio; Acquire the connection relationship between other vertices and the partition core vertex, and partition the data subject based on the connection relationship between the other vertices and the partition core vertex; Keyword recognition is performed on entities corresponding to the divided data topics to obtain the data topic names.

4. The method according to claim 1, characterized in that: The method of storing the watershed information data asset directory by numbers to generate a watershed information data architecture includes: Define the entity nodes corresponding to the watershed information data asset directory, number the entity nodes, and add labels to the entity nodes based on the numbers; Determining storage management information of the watershed information data asset directory using a graph database; The watershed information data architecture is constructed based on the labeled watershed information data asset directory and the storage management information.

5. A device for constructing a watershed information data architecture, characterized in that: The device comprises: A construction module, used for acquiring multi-source watershed data and constructing a watershed data gravity model based on the multi-source watershed data; A division module, used to divide the watershed data domains and data themes based on the watershed data gravity model, and generate a watershed information data asset directory; A number storage module, used for numbering and storing the watershed information data asset directory to generate a watershed information data architecture; The building blocks include: The recognition unit is used to perform multimodal named entity recognition on multi-source watershed data to obtain entities and entity occurrence counts; The first determination unit is used to determine the interaction force between entities based on the entity and the number of entity occurrences; assuming that the number of entity n occurrences is , the number of times entity n and entity p appear together is , the interaction force between two directly associated entities is defined as , the calculation formula is as follows: in, Representing Entities Number of occurrences, Representing Entities Number of occurrences, Representing Entities and entities co-occurrence count; The interaction force between two entities that are not directly related is defined as , the calculation formula is as follows: ( ) Where b is the intermediate node between entities a and c. If entity a has multiple branches connected to entity c, the gravitational force between entity a and entity c is the sum of the gravitational forces of all paths that do not exceed 3. A generating unit, used for associating entities by using the interaction forces between entities to generate a gravity model of watershed data; The division modules include: A first calculation unit is used to construct a data gravity graph based on the watershed data gravity model and calculate the attraction capacity of vertices in the data gravity graph; A partitioning unit, used for partitioning the data topics based on the attraction ability of the vertices, and obtaining data topic names corresponding to the partitioned data topics; The second calculation unit is used to obtain the subject classification name and calculate the edit distance between the data subject name and the subject classification name; The aggregation unit is used to aggregate the watershed data domains corresponding to the data subject names based on the edit distance to obtain the watershed information data asset directory.

6. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the method for constructing a watershed information data architecture according to any one of claims 1 to 4 by executing the computer instructions.

7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the method for constructing a watershed information data architecture according to any one of claims 1 to 4.

8. A computer program product, characterized in that It includes computer instructions, and the computer instructions are used to enable a computer to execute the watershed information data architecture construction method described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Village development potential evaluation and village classification method and device based on multi-source data

    CN113610346A

  • Multi-dimensional data base plate construction method and system for flood control of drainage basin and medium

    CN118331944A