A structured data modeling method, modeling system and management method
By constructing a knowledge processing link model and a knowledge label matrix model, the shortcomings of traditional database modeling in the application of knowledge graphs are solved, automatic construction and intelligent retrieval of knowledge links are realized, and data management and application efficiency of shale gas exploration and development are improved.
Patent Information
- Application Number
- CN202210667322.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-13
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2042-06-13
AI Technical Summary
Traditional database modeling methods cannot effectively support the application of knowledge graphs, cannot automatically form a knowledge label system, cannot intelligently search and knowledge correlation, and the data processing volume in shale gas exploration and development is large, and the deep model management function is lacking.
Using structured data modeling methods, a knowledge processing link model is built, a knowledge label matrix model is built, and instantiated judgment is carried out and the data set entity home is established, a knowledge label system is established, and the knowledge link is automatically constructed.
It realizes an efficient processing link from data to information and then to knowledge, automatically completes the construction of the knowledge label system, lays a solid foundation for subsequent knowledge search and intelligent recommendation, and improves the level of data management and application.
Smart Images

Figure CN117271784B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of shale gas exploration and development, and specifically relates to a structured data modeling method, a modeling system, and a management system. Background Art
[0002] In recent years, shale gas extraction has flourished, and major oil and gas companies have invested significant manpower and capital in shale gas exploration and development. This is particularly true in the development of data resources. The collection, storage, management, and application of shale gas-related data are key areas of focus for oil and gas companies in their information technology efforts. Domestic oil and gas companies have accumulated over two decades of experience in building conventional oil and gas databases and have achieved significant success. However, their approaches to database construction remain mired in traditional models: focused primarily on data management, with technical considerations focused on issues such as data redundancy and conformance to the appropriate paradigm. With the continuous emergence of new IT technologies and approaches in recent years, oil and gas company managers and researchers have increasingly demanded more data applications. Improving data management and application capabilities in this new environment is a challenge facing oil and gas companies.
[0003] Knowledge graph is a new technology that has emerged in recent years to serve knowledge management and application. It is a modern theory that achieves the purpose of multidisciplinary integration by combining the theories and methods of disciplines such as applied mathematics, graphics, information visualization technology, and information science with methods such as quantitative citation analysis and co-occurrence analysis, and using visual graphs to vividly display the core structure, development history, frontier fields, and overall knowledge architecture of the discipline.
[0004] In the field of oil exploration and development, the application of knowledge graphs can provide practical and valuable references for research. This is especially true for unconventional oil and gas development, such as shale gas, which has only just begun in recent years and is still in its infancy, making it particularly suitable for the implementation of knowledge graphs. However, the problem is that traditional database modeling is completely unsuitable for knowledge graph applications: First, traditional database modeling focuses solely on data management, and model results are limited to "data tables," "data items," and "relationships between tables." Subsequent application development requires extensive processing. Second, database construction in the oil and gas field exploration and development sector focuses on the refined management of business data. Over 90% of this data is structured basic data, which serves only as the foundation for knowledge processing rather than the knowledge itself. Traditional database modeling focuses solely on "data" and fails to clearly reflect the "data-information-knowledge" framework, resulting in a high workload for subsequent knowledge processing. Third, traditional database modeling often uses table and item notes as labels. These labels contain little content and cannot be automatically organized into a system, making them completely incapable of supporting intelligent retrieval and knowledge association. Fourth, traditional database modeling often relies on commercial modeling software, which focuses on model creation but has limited functionality for model management, especially deep processing. Therefore, based on the above reasons, new methods and tools for shale gas structured data modeling for knowledge graph applications should be adopted. Summary of the Invention
[0005] The present invention aims to provide a structured data modeling method, modeling system, and management system that effectively handles the processing chain from data to information and then to knowledge, and automatically completes the construction of a knowledge tag system and the construction of knowledge chains during the modeling process, laying a solid data model foundation for the subsequent development of applications such as knowledge search and intelligent recommendation. To achieve the above objectives, the present invention provides the following technical solutions:
[0006] A structured data modeling method, the method comprising:
[0007] Based on structured data, build a knowledge processing link model;
[0008] Based on the knowledge processing link model, a knowledge label matrix model is constructed;
[0009] Perform instantiation judgment on the knowledge label matrix;
[0010] The structured data and the knowledge label matrix after instantiation judgment are used to reposition the data set entities and complete the structured data modeling.
[0011] Preferably, the structured data includes basic entity data, information set entities and knowledge entity values, wherein:
[0012] Basic entity data, including basic entity data of all businesses;
[0013] The information set entity is formed by extracting and fusing multiple basic entity data;
[0014] The knowledge entity value is obtained according to the information set entity.
[0015] Preferably, a knowledge processing link model is constructed based on structured data, specifically,
[0016] Mark the basic entity data and build a basic entity data resource pool;
[0017] According to the information set entity of the basic entity data, a basic data set entity data resource pool is constructed;
[0018] Build a knowledge system based on the knowledge entities of basic entity data;
[0019] Analyze the processing chain relationship between the knowledge system and the data set entity resource pool and the basic entity resource pool, and build a knowledge processing chain model.
[0020] Preferably, each data of the basic entity data refers to one and only one basic entity, and is a single basic entity data.
[0021] Preferably, the identity tag is a unique identity tag of the basic entity data.
[0022] Preferably, the basic data set entity data resource pool is constructed based on the information set entity of the basic entity data, including:
[0023] Divide data set entities by the finest-grained business content, without any business overlap;
[0024] For each divided dataset entity, select a combination element from the basic entity data resource pool to establish a one-to-many relationship between the dataset entity and the basic entity data;
[0025] Uniquely tag dataset entities that have a one-to-many relationship with basic entity data to build a basic dataset entity data resource pool.
[0026] Preferably, the processing chain relationship between the analysis knowledge system and the data set entity resource pool and the basic entity resource pool is constructed to construct a knowledge processing chain model, including:
[0027] List the basic data entities in the dataset entity resource pool or basic entity resource pool used by each knowledge entity in the computational knowledge system;
[0028] Solidify the mapping relationship between each knowledge entity in the knowledge system and the dataset entity resource pool and basic entity resource pool, and build a knowledge processing link model.
[0029] Preferably, the constructing of the knowledge label matrix model based on the knowledge processing link model includes establishing a multi-dimensional knowledge label matrix according to the business field in which the knowledge processing link model is located;
[0030] Performing main labeling on the dimensions of the knowledge label matrix;
[0031] For each fixed knowledge main tag, the knowledge sub-tags are divided step by step to obtain the minimum granularity knowledge sub-tags and complete the construction of the knowledge tag matrix model.
[0032] Preferably, in the multi-dimensional knowledge label matrix, each dimension represents a knowledge label system, each data set entity corresponds to an element in the matrix, the knowledge labels of all dimensions of the element are the labels of the data set, and the data set and the knowledge label elements are in an "N to 1" relationship.
[0033] Preferably, the multi-dimensional label matrix includes:
[0034] There is no business overlap between any two dimensions; the content of each dimension can be disassembled or subdivided level by level;
[0035] The broken-down content within each dimension has a clear and fixed business logic relationship.
[0036] Preferably, the minimum granularity sub-tags cannot be further subdivided from the perspective of business understanding, or further subdivision is stopped when the requirements of subsequent knowledge application construction are met.
[0037] Preferably, the instantiation determination of the knowledge label matrix includes:
[0038] In each knowledge tag dimension, the knowledge tags are graded, where the value of the instance element knowledge tag of the lower gear is the same as the value of the instance element knowledge tag of the upper gear;
[0039] The instance element is the intersection of the basic entity data in N dimensions.
[0040] Preferably, the knowledge tag setting level is specifically:
[0041] Any two-by-two combinations of N label dimensions can reduce the N-dimensional matrix into "1+2+...(N-1)" 2-dimensional matrices.
[0042] For each instance element in the 2D matrix, determine whether the corresponding instances of the first-level gear label combination belong to the same level. If they belong to the same level, automatically remove all second-level gear label combination instances that do not belong to the same level in the first-level gear label combination. In this way, judge and automatically remove them level by level to complete the level marking of all instance elements in each 2D matrix.
[0043] After completing the level determination of all instance elements of the 2-dimensional matrix, the level evaluation of all instance elements of N dimensions is automatically achieved.
[0044] Preferably, the structured data and the knowledge label matrix after instantiation judgment are subjected to data set entity placement, specifically,
[0045] Select a dataset entity, select its corresponding label in each label dimension, and perform dataset entity placement; wherein the corresponding label is the label with the smallest granularity, and each dataset entity has one and only one corresponding label in each label dimension.
[0046] A structured data modeling system, the modeling system comprising:
[0047] The model building unit is used to build a knowledge processing link model based on structured data; and to build a knowledge label matrix model based on the knowledge processing link model;
[0048] A judgment unit, used for performing instantiation judgment on the knowledge label matrix;
[0049] The homing unit homing the structured data and the knowledge label matrix after instantiation judgment performs data set entity homing to complete the structured data modeling.
[0050] Preferably, the construction of the knowledge processing link model is specifically as follows:
[0051] The model building unit performs identity tagging on basic entity data and builds a basic entity data resource pool;
[0052] The model building unit constructs a basic data set entity data resource pool based on the information set entity of the basic entity data;
[0053] The model building unit constructs a knowledge system based on the knowledge entities of the basic entity data;
[0054] The model building unit analyzes the processing chain relationship between the knowledge system and the data set entity resource pool and the basic entity resource pool, and constructs a knowledge processing link model.
[0055] Preferably, the knowledge tag matrix model is constructed based on the knowledge processing link model, including:
[0056] The model building unit builds a multi-dimensional knowledge label matrix based on the business domain of the knowledge processing link model;
[0057] The model building unit performs main label archiving on the dimensions of the knowledge label matrix;
[0058] The model building unit divides each fixed knowledge main tag into knowledge sub-tags step by step, obtains the minimum granularity knowledge sub-tag, and completes the construction of the knowledge tag matrix model.
[0059] Preferably, the instantiation determination of the knowledge label matrix includes:
[0060] In the dimension of each knowledge tag, the knowledge tags are divided into levels, wherein the value of the instance element knowledge tag of the lower level is the same as the value of the instance element knowledge tag of the upper level.
[0061] Preferably, the structured data and the knowledge label matrix after instantiation judgment are subjected to data set entity placement, specifically,
[0062] Select a dataset entity, select a corresponding label in each label dimension, and perform dataset entity placement, wherein the corresponding label is the label with the smallest granularity, and each dataset entity has one and only one corresponding label in each label dimension.
[0063] A structured data management method, which manages the modeling method described in claims 1-14.
[0064] Preferably, the management method includes management of the knowledge link model, management of the knowledge tag matrix model, management of the instantiation of the knowledge tag matrix, and management of the entity placement of the data set.
[0065] Preferably, the management of the knowledge link model includes the management of the basic entity resource pool, the management of the data set entity resource pool, the knowledge system management and the knowledge graph association algorithm service.
[0066] Management of the basic entity resource pool, including registration of basic entities into the pool and maintenance of basic entity naming and unique identity identification;
[0067] The management of the dataset entity resource pool includes registering the dataset entity into the pool, maintaining the dataset entity name and unique identity, and maintaining the relationship between the dataset entity and the basic entity;
[0068] Knowledge system management, including maintaining knowledge naming and unique identity, managing knowledge hierarchy, and managing the relationship between knowledge and dataset entities + basic entities;
[0069] The management of knowledge graph association algorithm services includes analyzing the structured data relationships between basic entities, dataset entities, and knowledge systems, and building knowledge processing link models.
[0070] Preferably, the management of the knowledge tag matrix model includes:
[0071] Manage the content and relationships of knowledge tag dimensions,
[0072] According to the content and relationship of knowledge tag dimensions, main tags are determined and sub-tags are divided.
[0073] Preferably, the instantiation management of the knowledge tag matrix includes:
[0074] By combining two dimensions and labeling positions within any dimension, all instance elements in N dimensions are marked.
[0075] Determine whether the same-level label combination in the 2D matrix belongs to the same level, determine and automatically eliminate it level by level, and mark the level value of all instance elements in each 2D matrix;
[0076] According to the level values of all instance elements of all 2D matrices, the level value of each element of the N-dimensional matrix is automatically determined and fed back.
[0077] Preferably, the management of the entity placement of the data set includes:
[0078] Mark and store N-dimensional labels for each dataset entity;
[0079] Automatically check and determine whether the N-dimensional labels of each dataset entity belong to the same level.
[0080] Preferably, the management method further includes an intelligent search and recommendation engine, including:
[0081] Based on the input labels, it automatically locates labels that fully match or are similar in the label matrix, or directly recommends dataset entities corresponding to the label combination.
[0082] The automatically recommended data set entity can automatically associate related knowledge systems or data foundation entities according to the knowledge link model.
[0083] The technical effects and advantages of the present invention are as follows:
[0084] The structured data modeling method, modeling system and management system of the present invention well handle the processing link from data to information and then to knowledge, and automatically complete the construction of the knowledge label system and the construction of the knowledge link during the modeling process, laying a solid data model foundation for the subsequent research and development of applications such as knowledge search and intelligent recommendation. Other features and advantages of the present invention will be explained in the subsequent description, and part of them will become obvious from the description, or can be understood by practicing the present invention. The objects and other advantages of the present invention can be realized and obtained through the structures indicated in the description, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0085] Figure 1 is a flow chart of the modeling method of the present invention;
[0086] Figure 2 A knowledge processing link diagram from the basic entity to the indicator knowledge entity of the present invention;
[0087] Figure 3 Modeling flow chart for the knowledge processing link of the present invention;
[0088] Figure 4 This is the knowledge label matrix system diagram of the present invention. DETAILED DESCRIPTION
[0089] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0090] To address the deficiencies of the prior art, the present invention discloses a structured data modeling method, such as Figure 1 As shown, the method includes constructing a knowledge processing link model based on structured data; constructing a knowledge label matrix model based on the knowledge processing link model; performing instantiation judgment on the knowledge label matrix; and returning the structured data and the knowledge label matrix after instantiation judgment to the data set entity to complete structured data modeling.
[0091] The structured data is established through the knowledge graph, and several common knowledge representation methods of the knowledge graph are: selecting an appropriate data representation form and determining the core data structure of the knowledge graph. For example, in the process of building the knowledge graph, for text data, it is necessary to combine NLP (Natural Language Processing) technology to extract basic entity data from the text, and it is also possible to reversely annotate the text based on the basic entity data; using the RDF (WIDETABLE, Property Graph, logical storage solution) graph model, the basic entity data of different fields, different structures, and different formats are integrated to form a data set entity; based on the combination of data set entities, domain knowledge and business calculations, reasoning, machine learning (for example, K-nearest neighbor algorithm, HMM (Markov) model, CRF (conditional random field) model, time model, currency model and organizational structure, etc.) and network analysis and other logical calculations are performed on the knowledge graph to obtain knowledge entities.
[0092] Further, combined with Figure 2 As can be seen, the structured data includes basic entity data, information set entities, and knowledge entity values. Basic entity data includes basic entity data for all businesses; information set entities are formed by extracting multiple basic entity data; and knowledge entity values are obtained through logical calculations based on the information set entities. The logical calculations vary depending on the specific application domain. Exemplary methods include, but are not limited to, the K-nearest neighbor algorithm, the HMM (Markov model), and the CRF (Conditional Random Field) model. A knowledge processing chain is established from basic entities to knowledge entities. In a specific embodiment of the present invention, the primary function of structured data in the oil and gas exploration and development field is to form a basis for analyzing exploration and development patterns through multiple processing, aggregation, and analysis of basic data. This is primarily manifested in the form of "knowledge," which can be broadly categorized into management knowledge, technical knowledge, and economic knowledge. Specifically, this knowledge content includes knowledge on annual production completion rates and single-well productivity.
[0093] When analyzing structured data from the perspective of the knowledge system, the value stored in a single basic entity is just "data", such as "daily gas production = 1000"; the data information set entity formed by the combination of several basic data entities can be called "information", such as "W1 well + yyyymmdd + daily gas production = 1000"; the knowledge entity value formed by using multiple basic data entities in multiple data set entities to perform specific purpose algorithm processing reflects complex business logic and can support certain expert opinions, which can be called "knowledge".
[0094] Further, combined with Figure 3It can be seen that constructing a knowledge processing chain model based on structured data includes the following steps: constructing a knowledge system based on the knowledge entities of the basic entity data; performing identity tagging on the basic entity data in the knowledge system to construct a basic entity data resource pool; constructing a basic dataset entity data resource pool based on the information set entities of the basic entity data; and analyzing the processing chain relationship between the knowledge system and the dataset entity resource pool and the basic entity resource pool to construct a knowledge processing chain model. Each piece of basic entity data is referred to by one and only one basic entity, and is a single basic entity data.
[0095] Furthermore, knowledge system construction: knowledge entities should be determined according to business needs. The selection and division of knowledge entities should follow the following rules: Quantifiable: the content corresponding to the knowledge entity must be a specific numerical value, which can be automatically compared by computer on the data, and cannot be qualitative; Formulated: it can be automatically calculated based on the basic data entity using clear formulas, and no human intervention is required in the process; Systematized: each piece of knowledge is not completely independent, and the business and hierarchical relationship must be clarified, such as which are top-level indicator knowledge, which are lower-level indicator knowledge, which are core knowledge, and which are the influencing factors of knowledge.
[0096] Taking the research-oriented knowledge of shale gas exploration, development and production as an example, the following indicator knowledge entities (partial) can be established:
[0097] Drilling engineering knowledge: number of wells started / number of wells completed / well depth / horizontal section length / drilling cycle;
[0098] Fracturing engineering knowledge: number of fracturing wells / fracturing section length / average fluid volume per well / average sand volume per well / number of fracturing sections / fluid intensity / sand addition intensity;
[0099] Gas well testing knowledge: number of test wells / test production / test pressure;
[0100] Comprehensive knowledge of production and construction: number of wells put into production / waiting period for fracturing / fracturing period / waiting period for testing / testing period / well construction period / average daily production in the first year / average daily production per 100-meter section in the first year / first-year production decline rate.
[0101] Furthermore, the main purpose of building a basic entity resource pool is to establish a complete set of basic entities, including the following:
[0102] Extracting basic entities: Extract and define basic entities from the entire business scope. The primary principle of identifying basic entities is to be "basic" and avoid "reprocessing" entities as much as possible.
[0103] Granularity-based splitting of basic entities: Each basic entity must be "single" and avoid "composite entities" as much as possible. When encountering such composite entities, appropriate splitting is required to separate them into multiple basic entities.
[0104] Deduplication of basic entities: After extraction and splitting, multiple basic entities need to be compared and deduplicated based on business rules and scopes to ensure that a piece of data refers to only one basic entity.
[0105] Unique identity tag of basic entities: Since there are a large number of basic entities and they need to be frequently referenced in subsequent processing, the basic entities need to be uniquely identified.
[0106] Furthermore, the basic data set entity data resource pool is constructed based on the information set entity of the basic entity data, including dividing the data set entity according to the finest-grained business content without business overlap; for each divided data set entity, a combination element is selected from the basic entity data resource pool to establish a one-to-many relationship between the data set entity and the basic entity data; because the data set entity is frequently referenced in subsequent knowledge processing and label matrix repositioning, it is necessary to uniquely identify the data set entity, that is, to uniquely identify the data set entity in a one-to-many relationship with the basic entity data, and construct a basic data set entity data resource pool.
[0107] Physically, each "dataset entity" is a collection and combination of multiple basic entities within a certain business scope. Dataset entities only reference basic entities and should not be created repeatedly. For example, the two basic entities representing date (year-month-day) and well number only exist in one copy. Date + well number + production can constitute a "single well daily production data set entity", and date + well number + casing pressure can constitute a "single well daily status data set entity". Only one "single well daily production" is retained. If multiple dataset entities include a "single well daily production basic entity", they should all point to the same basic entity. This process can ensure the uniqueness of the basic data source.
[0108] Furthermore, the purpose of knowledge chain construction is to establish a mapping relationship between the indicator knowledge system and the data set entity resource pool and the basic entity resource pool. The processing chain relationship between the analysis knowledge system and the data set entity resource pool and the basic entity resource pool is constructed to construct a knowledge processing chain model, including:
[0109] List the basic data entities in the dataset entity resource pool or basic entity resource pool used to calculate each knowledge entity in the knowledge system, that is, analyze and list in detail which data entities are needed to support the calculation of each indicator knowledge, that is, the processing chain from data to knowledge (the specific algorithm belongs to the software code implementation and is not considered in the modeling stage). For example: the calculation of drilling cycle indicator knowledge requires two basic entities, namely the start-up date and completion date, in the drilling basic dataset entity; solidify the mapping relationship between each knowledge entity in the knowledge system and the dataset entity resource pool and the basic entity resource pool, and build a knowledge processing chain model.
[0110] Furthermore, through the construction of the knowledge processing chain model described above, all business data (which can be structured) in shale gas exploration and development is guaranteed to be stored in corresponding "dataset entities" with granularity accurate to each "basic entity," and a knowledge chain from basic data to "knowledge" is established. However, the current data model cannot be directly used in knowledge graph applications because it lacks "labels."
[0111] In this invention, the labeling process for constructing the knowledge graph is completed through the "knowledge label matrix", such as Figure 4 In the knowledge label matrix system shown, label dimension 1, label dimension 2, label dimension 3, label dimension 4, label dimension 5, and so on, and label dimension n, form an N-dimensional matrix, with each dimension containing several labels. For example, taking one label from each of the N dimensions defines a "matrix element," which is essentially the collection of N labels. Labels within the same dimension must be marked with "inclusion" relationships, depending on the granularity of the divisions; and "sequence" relationships must be marked based on the business logic relationships embodied by the labels. Each "dataset entity" corresponds to a corresponding "matrix element," and the knowledge labels of the N dimensions of the "matrix element" constitute the dataset's label.
[0112] Furthermore, the knowledge processing link model is based on which a knowledge label matrix model is constructed, including establishing a multi-dimensional knowledge label matrix according to the business field in which the knowledge processing link model is located; filing the dimensions of the knowledge label matrix with main labels; dividing each filed knowledge main label into knowledge sub-labels step by step to obtain the minimum granularity knowledge sub-label and complete the construction of the knowledge label matrix model.
[0113] Defining the dimension of the label matrix involves defining a number of necessary knowledge label systems based on the business domain of the model, for knowledge retrieval and application. The number of systems determines the dimensionality of the label matrix. This work must adhere to the following principles: Independence: No two dimensions should have significant or significant business overlap; Refinability: The content of each dimension must be decomposable and subdividable; Strong Relationship: The decomposed content within each dimension must have clear and fixed business logic relationships (whether inclusion, sequential, or parallel).
[0114] In a specific embodiment of the present invention, for shale gas exploration and development, this technical solution pre-sets a five-dimensional label matrix, including the following label dimensions: Business dimension: for example, exploration-development-production, a sequential relationship. The shale gas exploration and development chain is very long, which means that the business must be divided into multiple business stages, and the main businesses have a sequential relationship; Work dimension: for example, mineral rights reserves-planning plan-exploration evaluation-experimental testing-development plan-development dynamics-ecological and environmental protection, a parallel relationship. The work dimension often corresponds to the division of departments or positions in a unit; Application dimension: for example, management-oriented-research-oriented-financial analysis-oriented, a parallel relationship; Target dimension: for example, company-block-well area-platform-well, an inclusion relationship. Each record in the data set must describe a certain "target"; Time frequency dimension: for example, year-quarter-month-week-ten-day, an inclusion relationship.
[0115] Furthermore, the main label is defined as all elements of the full and accurate label matrix. Each dimension is first divided into several first-level levels to form the main label. For example, the first level of the business dimension is divided into exploration, development, and production.
[0116] Furthermore, subtags are divided into the smallest granularity for each level. This minimum granularity subtag stops being subdivided when it is no longer feasible from a business perspective, or when it meets the requirements of subsequent knowledge application development. Based on the results of the master tag leveling, subtags are divided into sub-levels within each level. This process is carried out step by step, gradually reducing the granularity of the tags. The specific level of subtag refinement can be determined by referring to the following suggestions: gradually reducing the granularity until it is no longer feasible from a business perspective; or stopping subdivision when it meets the requirements of subsequent knowledge application development.
[0117] The result of constructing a knowledge label matrix model is the specific content of each knowledge label dimension. However, when it comes to a specific data model, not every combination of two dimensions is considered reasonable or can correspond to a dataset entity. For example, "experimental test (work dimension) + well (target dimension)" can generate "analysis and testing data for a single well sample," which is considered reasonable. However, "experimental test (work dimension) + company (target dimension)" does not generate data and is therefore considered unreasonable.
[0118] The task of instantiating the knowledge label matrix involves assessing the rationality of all intersections (referred to as "instance elements") across the five dimensions of the knowledge label matrix, based on the actual circumstances of shale gas exploration and development. Assuming 100 labels for each dimension, there would be 1,000,000,000,000 instances. It's impossible to assess the rationality of each instance individually, requiring scientific methods and algorithms, as described in detail below.
[0119] The work content and steps of instantiating the knowledge label matrix are as follows: on the dimension of each knowledge label, the knowledge label is assigned a level, where the value of the knowledge label of the instance element of the lower gear is the same as the value of the knowledge label of the instance element of the upper gear, and the calculation formula is defined as follows: when the label instance of the directly upper gear is 0, the lower label instances are all 0, and the formula is marked as F (10) ; The instance element is the intersection of the basic entity data in N dimensions.
[0120] Any two-by-two combinations of N label dimensions can reduce the N-dimensional matrix into "1+2+...(N-1)" 2-dimensional matrices.
[0121] For each instance element in the 2D matrix, determine whether the corresponding instance of the first-level gear label combination belongs to the same level. If they belong to the same level, automatically remove all the second-level gear label combination instances that do not belong to the same level in the first-level gear label combination. In this way, judge and automatically remove them level by level to complete the level marking of all instance elements in each 2D matrix; for example, X 11 Y 11 =0, X 11 Y 12 =1, and so on. When the first gear is marked, apply the above formula F (10) , you can automatically remove X 11 Y 11 All the secondary gear label combination instances under it are judged and automatically eliminated level by level, which greatly reduces the workload of marking rationality. Finally, the rationality of all instance elements in each 2D matrix is marked (0 or 1);
[0122] After completing the level judgment of all instance elements of the 2D matrix, the level evaluation of all instance elements of N dimensions is automatically realized. For example, if the two dimensions of a matrix are X and Y, first judge the rationality of the corresponding instance from the first-level gear label combination, for example, X 11 Y 11 =0, X 11 Y 12 =1, and so on. When the first gear is marked, apply the above formula F (10) After completing the rationality judgment of all instance elements in the above 2-dimensional matrix, any label will form a combination with any label in other dimensions and be marked as 0 or 1; here, the formula F is defined (n) The purpose is to automatically realize the rationality assessment of all intersections of N dimensions (referred to as "instance elements"). The algorithm is: the instance element corresponds to two combinations of n labels. When all the combination values in the 2-bit matrix are 1, it is marked as 1, otherwise it is marked as 0, indicating that this instance element is unreasonable.
[0123] Each data set entity is matched with an element in the "knowledge label matrix", that is, an association relationship between structured data and the knowledge label system is established. The structured data is matched with the knowledge label matrix after instantiation and judgment to perform data set entity placement. Specifically, select a data set entity, select its corresponding label in each label dimension, and perform corresponding association; wherein, the corresponding label is the label with the smallest granularity, and each data set entity has one and only one corresponding label in each label dimension. (n) Represents the label value of N dimensions of a dataset entity, and defines the formula S (n) , use E (n) Bring it into the knowledge label matrix instance and automatically locate the corresponding instance element. If the instance element = 0, it means that the label selected for this dataset entity is unreasonable. At this point, the above structured data modeling work is completed.
[0124] A structured data modeling system includes a model building unit for constructing a knowledge processing link model based on structured data; constructing a knowledge label matrix model based on the knowledge processing link model; a judgment unit for performing instantiation judgment on the knowledge label matrix; and a corresponding association unit for correspondingly associating the structured data with the knowledge label matrix after instantiation judgment to complete structured data modeling.
[0125] The construction of the knowledge processing link model is specifically as follows:
[0126] The model building unit performs identity tagging on basic entity data and builds a basic entity data resource pool;
[0127] The model building unit constructs a basic data set entity data resource pool based on the information set entity of the basic entity data;
[0128] The model building unit constructs a knowledge system based on the knowledge entities of the basic entity data;
[0129] The model building unit analyzes the processing chain relationship between the knowledge system and the data set entity resource pool and the basic entity resource pool, and constructs a knowledge processing link model.
[0130] The present application also proposes a structured data management method, which manages the modeling method of the structured data.
[0131] The management method includes a knowledge link model management unit, a knowledge tag matrix model management unit, a knowledge tag matrix instantiation management unit, a data set entity placement management unit and an intelligent search and recommendation engine service management unit, wherein:
[0132] The knowledge link model management unit is used to develop methods and data in the process of building a knowledge processing link model based on structured data; build a knowledge label matrix model based on the knowledge processing link model; perform instantiation judgment on the knowledge label matrix; and associate the structured data with the knowledge label matrix after instantiation judgment to complete the structured data modeling. The management of the knowledge link model includes the management of the basic entity resource pool, the management of the data set entity resource pool, the knowledge system management, and the knowledge graph association algorithm service.
[0133] Management of the basic entity resource pool, including registration of basic entities into the pool, maintenance of basic entity naming and unique identity identification; management of the dataset entity resource pool, including registration of dataset entities into the pool, maintenance of dataset entity naming and unique identity identification, and maintenance of the relationship between dataset entities and basic entities;
[0134] Knowledge system management, including knowledge system creation, maintenance of knowledge system naming and unique identity, management of knowledge hierarchy, and management of the relationship between knowledge and dataset entities + basic entities;
[0135] The management of knowledge graph association algorithm services includes analyzing the structured data relationships between basic entities, dataset entities, and knowledge systems, and building knowledge processing link models.
[0136] The knowledge graph association algorithm service implements "master-slave entity association", "parallel basic entity business association", "parallel basic entity knowledge association", and "parallel data set entity knowledge association" based on the knowledge link relationship between basic entities, data set entities, and knowledge.
[0137] "Master-slave entity association": Specify any indicator knowledge and automatically associate the various dataset entities and basic entities that support the indicator knowledge calculation; specify any dataset and automatically associate the basic entities.
[0138] "Parallel Basic Entity Business Association": Specify any two basic entities and verify the association by automatically identifying whether they support the same dataset entity. The algorithm result is "Associated" or "Not Associated".
[0139] "Parallel basic entity indicator knowledge association": Specify any two basic entities and verify the association relationship by automatically counting the number of "simultaneously supported indicator knowledge". The algorithm result is a specific number, representing the number of simultaneously supported indicator knowledge.
[0140] "Parallel Dataset Entity Indicator Knowledge Association": Specify any two dataset entities and automatically count the number of basic entities that "support the same indicator knowledge" to reflect the association relationship. The algorithm result is a specific number.
[0141] The knowledge label matrix model management is used to manage the content and relationship of the knowledge label dimensions, and provides the operational functions of main label scheduling and sub-label division. The management of the knowledge label matrix model includes managing the content and relationship of the knowledge label dimensions, scheduling the main label and dividing the sub-label according to the content and relationship of the knowledge label dimensions. The instantiation management of the knowledge label matrix includes marking all instance elements of N dimensions through two-dimensional combinations and label positions within any dimension; judging whether the same-level label combination in the 2-dimensional matrix belongs to the same level, judging and automatically eliminating it level by level, marking the level values of all instance elements in each 2-dimensional matrix; automatically judging the level values of each element of the N-dimensional matrix and feeding back the level values of all instance elements marked in all 2-dimensional matrices.
[0142] Furthermore, it manages the rationality marking management function of all intersections of N dimensions (referred to as "instance elements"); provides the function of managing two-dimensional combinations; provides the function of managing label gears in any dimension; provides the rationality marking function of the same-level label combination in a 2-dimensional matrix; after completing the rationality marking function of the same-level label combination in the 2-dimensional matrix, provides the function of automatically judging and automatically eliminating label combinations at each level, and realizes the automatic calculation function of the rationality of all intersections of N dimensions (referred to as "instance elements"), that is, based on the marking results of all 2-dimensional matrices, automatically calculates the rationality value of each element of the N-dimensional matrix. It also provides the function of retrieving any N-dimensional matrix element and feeding back the rationality value.
[0143] The management of dataset entity placement includes marking and storing the N-dimensional label of each dataset entity; and automatically verifying and determining whether the N-dimensional label of each dataset entity belongs to the same hierarchy. Furthermore, the dataset entity placement management unit is used to implement the labeling and storage functions of each dataset entity's N-dimensional label; and to implement the function of automatically verifying the rationality of the dataset entity's N-dimensional label: based on the N labels marked for a certain entity, the corresponding element in the "Knowledge Label Matrix Instantiation" result is automatically found, and the rationality value (0 or 1) is automatically retrieved. If the rationality value is 0, an automatic alarm is prompted.
[0144] The intelligent search and recommendation engine automatically locates fully matching or similar tags in the tag matrix based on the input tags, or directly recommends the data set entities corresponding to the tag combination; the automatically recommended data set entities can automatically associate the relevant knowledge system or data foundation entities based on the knowledge link model.
[0145] Furthermore, intelligent search and recommendation engine services include:
[0146] Keyword intelligent search: The "Knowledge Label Matrix" is equivalent to labeling each data set. When the user enters a keyword for search, the search engine provided by this tool will automatically locate the exact or similar tags in the label matrix;
[0147] When a user enters multiple keywords in parallel, the search engine will automatically locate the labels in N different dimensions and automatically remove all elements with a rationality of 0 based on the instantiated matrix. It will then recommend all reasonable N-dimensional label combinations to the user for continued precise search, or directly recommend the dataset entity corresponding to the label combination.
[0148] For automatically recommended entities, other knowledge or basic entities can be automatically linked up and down based on the knowledge link model.
[0149] Finally, it should be noted that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or make equivalent substitutions for some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A structured data modeling method, characterized in that: The method comprises, Based on structured data, a knowledge processing link model is constructed; the structured data is established through a knowledge graph; in the knowledge graph construction process, for text data, it is necessary to combine NLP technology to extract basic entity data from the text; Based on the knowledge processing link model, a knowledge label matrix model is constructed; Perform instantiation judgment on the knowledge label matrix; The structured data and the knowledge label matrix after instantiation judgment are used to return the data set entities to complete the structured data modeling; Based on structured data, a knowledge processing link model is constructed, specifically, Mark the basic entity data and build a basic entity data resource pool; According to the information set entity of the basic entity data, a basic data set entity data resource pool is constructed; Build a knowledge system based on the knowledge entities of basic entity data; Analyze the processing chain relationship between the knowledge system and the data set entity resource pool and the basic entity resource pool, and build a knowledge processing chain model; The knowledge tag matrix model is constructed based on the knowledge processing link model, including: Establish a multi-dimensional knowledge label matrix according to the business field in which the knowledge processing link model is located; and assign main labels to the dimensions of the knowledge label matrix; For each designated knowledge main tag, divide the knowledge sub-tags step by step, obtain the minimum granularity knowledge sub-tags, and complete the knowledge tag matrix model construction; The instantiation determination of the knowledge label matrix includes: In each knowledge tag dimension, the knowledge tags are graded, where the value of the instance element knowledge tag of the lower gear is the same as the value of the instance element knowledge tag of the upper gear; The instance element is the intersection of the basic entity data in N dimensions; The structured data and the knowledge label matrix after instantiation judgment are used to perform data set entity placement, specifically, Select a dataset entity, select its corresponding label in each label dimension, and perform dataset entity placement; wherein the corresponding label is the label with the smallest granularity, and each dataset entity has one and only one corresponding label in each label dimension.
2. A structured data modeling method according to claim 1, characterized in that: The structured data includes basic entity data, information set entities and knowledge entity values, wherein: Basic entity data, including basic entity data of all businesses; The information set entity is formed by extracting and fusing multiple basic entity data; The knowledge entity value is obtained according to the information set entity.
3. A structured data modeling method according to claim 1, characterized in that: Each data of the basic entity data refers to one and only one basic entity, and is a single basic entity data.
4. A structured data modeling method according to claim 1 or 3, characterized in that: The identity tag is a unique identity tag of the basic entity data.
5. A structured data modeling method according to claim 1 or 3, characterized in that: The basic data set entity data resource pool is constructed based on the information set entity of the basic entity data, including: Divide data set entities by the finest-grained business content, without any business overlap; For each divided dataset entity, select a combination element from the basic entity data resource pool to establish a one-to-many relationship between the dataset entity and the basic entity data; Uniquely tag dataset entities that have a one-to-many relationship with basic entity data to build a basic dataset entity data resource pool.
6. A structured data modeling method according to claim 1, characterized in that: The processing chain relationship between the analysis knowledge system and the data set entity resource pool and the basic entity resource pool is constructed to construct a knowledge processing chain model, including: List the basic data entities in the dataset entity resource pool or basic entity resource pool used by each knowledge entity in the computational knowledge system; Solidify the mapping relationship between each knowledge entity in the knowledge system and the dataset entity resource pool and basic entity resource pool, and build a knowledge processing link model.
7. A structured data modeling method according to claim 1, characterized in that: In the multi-dimensional knowledge label matrix, each dimension represents a knowledge label system, each data set entity corresponds to an element in the matrix, and the knowledge labels of all dimensions of the element are the labels of the data set. The data set and the knowledge label elements are in an "N to 1" relationship.
8. A structured data modeling method according to claim 1 or 7, characterized in that: The multi-dimensional knowledge label matrix includes: There is no business overlap between any two dimensions; the content of each dimension can be disassembled or subdivided level by level; The broken-down content within each dimension has a clear and fixed business logic relationship.
9. A structured data modeling method according to claim 1, characterized in that: The minimum granularity knowledge sub-tags are no longer subdivided when they meet the requirements of business understanding or subsequent knowledge application construction.
10. A structured data modeling method according to claim 1, characterized in that: The knowledge tag setting level is specifically: Any two-by-two combinations of N label dimensions can be used to reduce the N-dimensional matrix into "1+2+…(N-1)" 2-dimensional matrices. For each instance element in the 2D matrix, determine whether the corresponding instances of the first-level gear label combination belong to the same level. If they belong to the same level, all second-level gear label combination instances that do not belong to the same level in the first-level gear label combination are automatically eliminated. This is done by judging and automatically eliminating them level by level, marking the levels of all instance elements in each 2D matrix. After completing the level determination of all instance elements of the 2-dimensional matrix, the level evaluation of all instance elements of N dimensions is automatically achieved.
11. A structured data modeling system, characterized in that: The modeling system includes, The model building unit is used to build a knowledge processing link model based on structured data; the structured data is built through the knowledge graph; in the knowledge graph building process, for text data, it is necessary to combine NLP technology to extract basic entity data from the text; based on the knowledge processing link model, a knowledge label matrix model is built; wherein, based on the structured data, the knowledge processing link model is built, specifically, Mark the basic entity data and build a basic entity data resource pool; According to the information set entity of the basic entity data, a basic data set entity data resource pool is constructed; Build a knowledge system based on the knowledge entities of basic entity data; Analyze the processing chain relationship between the knowledge system and the data set entity resource pool and the basic entity resource pool, and build a knowledge processing chain model; The knowledge tag matrix model is constructed based on the knowledge processing link model, including: Establish a multi-dimensional knowledge label matrix according to the business field in which the knowledge processing link model is located; and assign main labels to the dimensions of the knowledge label matrix; For each designated knowledge main tag, divide the knowledge sub-tags step by step, obtain the minimum granularity knowledge sub-tags, and complete the knowledge tag matrix model construction; A determination unit is used to perform instantiation determination on the knowledge label matrix; wherein, the instantiation determination on the knowledge label matrix includes: In each knowledge tag dimension, the knowledge tags are graded, where the value of the instance element knowledge tag of the lower gear is the same as the value of the instance element knowledge tag of the upper gear; The instance element is the intersection of the basic entity data in N dimensions; The homing unit performs data set entity homing on the structured data and the knowledge label matrix after instantiation judgment to complete the structured data modeling, wherein the homing on the data set entity on the structured data and the knowledge label matrix after instantiation judgment is specifically as follows: Select a dataset entity, select its corresponding label in each label dimension, and perform dataset entity placement; wherein the corresponding label is the label with the smallest granularity, and each dataset entity has one and only one corresponding label in each label dimension.
12. A structured data management method, characterized in that: The management method manages the modeling method described in claims 1-10.
13. A structured data management method according to claim 12, characterized in that: The management method includes management of the knowledge link model, management of the knowledge label matrix model, management of the instantiation of the knowledge label matrix, and management of the entity placement of the data set.
14. A structured data management method according to claim 13, characterized in that: The management of the knowledge link model includes the management of the basic entity resource pool, the management of the data set entity resource pool, the management of the knowledge system and the knowledge graph association algorithm service. Management of the basic entity resource pool, including registration of basic entities into the pool and maintenance of basic entity naming and unique identity identification; The management of the dataset entity resource pool includes registering the dataset entity into the pool, maintaining the dataset entity name and unique identity, and maintaining the relationship between the dataset entity and the basic entity; Knowledge system management, including maintaining knowledge naming and unique identity, managing knowledge hierarchy, and managing the relationship between knowledge and dataset entities + basic entities; The management of knowledge graph association algorithm services includes analyzing the structured data relationships between basic entities, dataset entities, and knowledge systems, and building knowledge processing link models.
15. A structured data management method according to claim 13, characterized in that: The management of the knowledge tag matrix model includes: Manage the content and relationships of knowledge tag dimensions, According to the content and relationship of knowledge tag dimensions, main tags are determined and sub-tags are divided.
16. A structured data management method according to claim 13, characterized in that: The instantiation management of the knowledge tag matrix includes: By combining two dimensions and labeling positions within any dimension, all instance elements in N dimensions are marked. Determine whether the same-level label combination in the 2D matrix belongs to the same level, determine and automatically eliminate it level by level, and mark the level value of all instance elements in each 2D matrix; According to the level values of all instance elements of the entire 2-dimensional matrix, the level value of each element of the N-dimensional matrix is automatically determined and fed back.
17. A structured data management method according to claim 13, characterized in that: The management of the entity placement of the data set includes: Mark and store N-dimensional labels for each dataset entity; Automatically check and determine whether the N-dimensional labels of each dataset entity belong to the same level.
18. A structured data management method according to claim 12, characterized in that: The management method also includes an intelligent search and recommendation engine, including, Based on the input labels, it automatically locates labels that fully match or are similar in the label matrix, or directly recommends dataset entities corresponding to the label combination. The automatically recommended dataset entities can automatically associate related knowledge systems or data foundation entities based on the knowledge link model.
Citation Information
Patent Citations
Method for carrying out information interaction by utilizing academic resource interaction platform of social network
CN103034728A
Data classification and modelling based application compliance analysis
US20210224298A1