Structured database system for local water pollutant discharge standards based on data mining
By constructing a hierarchical relational network based on domain knowledge graphs and a weighted document semantic model, the problem that existing technologies cannot express complex semantic logic and hierarchical constraints is solved, enabling in-depth mining and intelligent application of local water pollutant discharge standards and improving the system's intelligent decision-making capabilities.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINESE RES ACAD OF ENVIRONMENTAL SCI
- Filing Date
- 2025-12-30
- Publication Date
- 2026-07-31
AI Technical Summary
Existing technologies cannot effectively express the complex semantic logic and hierarchical constraints in local water pollutant discharge standards, making it difficult to support in-depth data mining and intelligent applications. In particular, their application value is limited in advanced scenarios such as precise compliance diagnosis and intelligent decision comparison.
Through modules such as data acquisition, preprocessing, knowledge extraction, anomaly detection, and database management, a hierarchical relational network based on domain knowledge graph is constructed. Semantic vectorization is performed using a weighted document semantic model, and structured and intelligent analysis of standard text is achieved through temporal database storage and intelligent application services.
It enables in-depth analysis and intelligent application of local water pollutant discharge standards, supports dynamic reasoning and compliance review in advanced scenarios, and improves the efficiency and accuracy of information retrieval and intelligent decision-making.
Smart Images

Figure CN121786096B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data mining technology, and in particular to a structured database system for local water pollutant discharge standards based on data mining. Background Technology
[0002] In the process of environmental management informatization, digital management of water pollutant discharge standards in various regions is an important requirement. Existing technologies mostly adopt standard query systems based on keyword retrieval and relational databases. Through manual or rule templates, key information in standard documents is extracted and stored in a preset database, realizing electronic storage and conditional query of standard information, and providing users with basic information retrieval convenience.
[0003] There are core limitations in supporting deep data mining and intelligent applications, namely, insufficient ability to express complex semantic logic and hierarchical constraints in standard texts. Local emission standards contain a large number of conditional rules and multi-level associations, and the existing flat data models are unable to effectively capture and structure such complex relationships. As a result, the system can only provide static queries and cannot support dynamic reasoning under different needs. It is difficult to realize the inherent inconsistency of standard provisions and compliance review, which limits its application value in advanced scenarios such as accurate compliance diagnosis and intelligent decision comparison. Summary of the Invention
[0004] In view of the aforementioned existing problems, the present invention is proposed.
[0005] Therefore, this invention provides a structured database system for local water pollutant discharge standards based on data mining, which solves the core problem that existing technologies cannot express the complex semantic logic and hierarchical constraints in standard texts due to the flattened data model, thus restricting deep data mining and intelligent applications.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: This invention provides a structured database system for local water pollutant discharge standards based on data mining, which includes a data acquisition module that collects text of local water pollutant discharge standards from multiple locations and performs preprocessing to obtain structured standard text data. The knowledge extraction module extracts key information entities and relationships from structured standard text data based on domain knowledge graphs, constructs a hierarchical association relationship network with associated weight coefficients, and obtains the first structured data. The processing module uses a weighted document semantic model to perform semantic vectorization on the complete document text corresponding to the first structured data, and obtains a standard semantic vector set. The anomaly detection module performs anomaly detection on the first structured data and assigns anomaly weights to obtain high-quality structured data with weight labels. The database management module defines the physical storage mode of the temporal database through a hierarchical association network of weight coefficients, and stores high-quality structured data with weight labels and standard semantic vector sets into the temporal database. The intelligent application service module provides intelligent application services based on a temporal database. By calling the correlation weight coefficient and the anomaly weight coefficient, it prioritizes the recommendation results and provides hierarchical prompts for anomaly information.
[0007] As a preferred embodiment of the structured database system for local water pollutant discharge standards based on data mining described in this invention, the following steps are included: collecting local water pollutant discharge standard texts from multiple locations and preprocessing them to obtain structured standard text data: We collected texts of local water pollutant discharge standards from multiple regions and converted them to a unified format. The local water pollutant discharge standards text, after being uniformly converted in format, undergoes character encoding cleaning before the document's logical structure is parsed. Standard text data for chapter and paragraph identifiers is output based on the document's logical structure parsing.
[0008] As a preferred embodiment of the structured database system for local water pollutant discharge standards based on data mining as described in this invention, the extraction of key information entities and relationships from structured standard text data based on domain knowledge graphs includes the following steps: Identifying key information entities from structured standard text data based on domain knowledge graphs; Based on domain knowledge graphs, key information relationships between entities are extracted from structured standard text data and then integrated.
[0009] As a preferred embodiment of the structured database system for local water pollutant discharge standards based on data mining as described in this invention, the following steps are included in constructing a hierarchical association network with associated weight coefficients to obtain the first structured data: Establish a hierarchical relationship network for key information entities and relationships based on the level of explicitness in standard text data; Based on the location and clarity of the key information entities in the standard text data, each relationship in the hierarchical relationship network is assigned a relationship weight coefficient to obtain the hierarchical relationship with the attached relationship weight coefficient. The hierarchical association network based on the associated weight coefficients constitutes the first structured data.
[0010] As a preferred embodiment of the structured database system for local water pollutant discharge standards based on data mining as described in this invention, the following steps are included: Semantic vectorization of the complete document text corresponding to the first structured data is performed using a weighted document semantic model to obtain a set of standard semantic vectors: Assign computational weights to different parts of the complete document text corresponding to the first structured data; The expression for calculating the weights is: ; in, For the first in the document The weights of each text segment are calculated. For the first part of the complete document text A partial text fragment, As a moderating factor for the semantic role rating item, Analysis of Semantic Role Labeling Technology in Natural Language Processing The grammatical structure of each sentence in the text. For normalization function, For term frequency-inverse document frequency based on a domain terminology dictionary, Indexing partial text fragments; We utilize a weighted document semantic model to process the complete document text, assigning and calculating weights accordingly. The weighted document semantic model converts each complete local water pollutant discharge standard document into a high-dimensional semantic vector; A standard semantic vector set is constructed based on high-dimensional semantic vectors.
[0011] As a preferred embodiment of the local water pollutant discharge standard structured database system based on data mining described in this invention, the following steps are included: Anomaly detection and anomaly weight assignment are performed on the first structured data to obtain high-quality structured data with weighted labels: The first structured data is subjected to legality conflict checks, internal consistency checks, and ambiguity detection to detect anomalies; Assign anomaly weights to the detected anomalies based on their anomaly types, and then mark the first structured data with assigned anomaly weights as high-quality structured data with weighted labels.
[0012] As a preferred embodiment of the structured database system for local water pollutant discharge standards based on data mining as described in this invention, the physical storage mode of the temporal database is defined through a hierarchical association network of weighted coefficients, including the following steps: Based on the entities, relations, and association weight coefficients defined in the hierarchical association relationship network with associated association weight coefficients, the table structure of the temporal database is obtained; Based on the entities, relations, and association weight coefficients defined in the hierarchical association relationship network with attached association weight coefficients, foreign key associations between temporal database tables are obtained; Based on the entities, relationships, and association weight coefficients defined in the hierarchical association relationship network with associated association weight coefficients, a dedicated field for storing association weight coefficients is established in the temporal database table. The physical storage model of a temporal database consists of its table structure, foreign key relationships, and special fields.
[0013] As a preferred embodiment of the data mining-based structured database system for local water pollutant discharge standards described in this invention, the storage of high-quality structured data with weighted labels and a set of standard semantic vectors in a temporal database includes the following steps: High-quality structured data with weighted labels is mapped to the corresponding table in the temporal database, and the standard semantic vector set is stored in the dedicated vector field of the temporal database, resulting in the persistent storage of high-quality structured data with weighted labels and the standard semantic vector set in the temporal database.
[0014] As a preferred embodiment of the structured database system for local water pollutant discharge standards based on data mining as described in this invention, the system provides intelligent application services based on a temporal database, including the following steps: Provides similarity-based recommendation services based on a set of standard semantic vectors in a temporal database; This service provides anomaly alerts based on high-quality, weighted, structured data in a temporal database.
[0015] As a preferred embodiment of the structured database system for local water pollutant discharge standards based on data mining as described in this invention, the following steps are included: prioritizing the recommendation results and providing graded prompts for abnormal information by calling the association weight coefficient and the anomaly weight coefficient: In the similarity-based recommendation service, the association weight coefficient is used to sort the initial recommendation results; The association weight coefficients are used to assign priority order to the sorted recommendation results, thus obtaining the priority order of the recommendation results; In the abnormal information prompting service, the severity level of each abnormal message is determined by calling the abnormal weight coefficient, and the abnormal messages of different severity levels are then graded and prompted accordingly.
[0016] The beneficial effects of this invention are as follows: A data acquisition module acquires and preprocesses standard text to generate structured standard text data; a knowledge extraction module extracts key entities and relationships based on a domain knowledge graph, constructs a hierarchical association network with associated weight coefficients, and forms the first structured data; a processing module uses a weighted document semantic model to semantically vectorize the complete text, generating a set of standard semantic vectors; an anomaly detection module performs anomaly verification on the preliminary data and assigns anomaly weights, outputting high-quality structured data with weight labels; a database management module defines the physical storage mode of the temporal database based on the weighted network to achieve persistent data storage; and an intelligent application service module, by calling association weights and anomaly weights, implements priority ranking of recommendation results and hierarchical prompts for anomaly information, solving the problems of existing standard data management systems, which, due to their flat models, struggle to express complex semantic logic and hierarchical constraints, and cannot support deep mining and intelligent reasoning. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a schematic diagram of a structured database system for local water pollutant discharge standards based on data mining.
[0019] Figure 2 This is a flowchart for anomaly detection.
[0020] Figure 3 Flowchart for constructing a hierarchical network.
[0021] Figure 4 This is a flowchart of the weighted document semantic vectorization process. Detailed Implementation
[0022] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0023] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0024] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0025] Reference Figures 1-4 As one embodiment of the present invention, this embodiment provides a structured database system for local water pollutant discharge standards based on data mining, including the following steps: The data acquisition module collects local water pollutant discharge standards from multiple locations and preprocesses them to obtain structured standard text data.
[0026] We collected texts of local water pollutant discharge standards from multiple regions and converted them to a unified format.
[0027] Furthermore, a target area list is established, and based on this list, access is made to the official public information platforms of ecological and environmental authorities at all levels to obtain local water pollutant discharge standards and related policy interpretations and drafting instructions. For the original files obtained in different formats, such as portable document formats with fixed layouts, easily editable document formats, and Hypertext Markup Language (HTML) pages, a format parsing library is used for decoding. Portable document formats need to be distinguished between text and image types. Text types are directly extracted from the character stream, while image types need to be converted into text using optical character recognition (OCR) technology. Document format files are parsed to open document structures to extract text and basic formatting information. HTML pages are cleaned of tags and the main text is retained. All documents are ultimately converted into a unified structured text format, such as JavaScript object notation, which can retain text content and mark metadata such as the original source format, eliminating parsing obstacles caused by initial format differences.
[0028] The local water pollutant discharge standards text, after being uniformly converted in format, undergoes character encoding cleansing before the document's logical structure is parsed.
[0029] Furthermore, the text data, after undergoing standardized format conversion, undergoes deep character encoding cleansing to resolve potential character set conflicts arising from diverse document sources. Non-standard character encodings are converted to a universal character encoding standard, filtering out control characters, garbled text, and irrelevant formatting symbols. The clean text content is then subjected to document logical structure analysis. Based on text segmentation algorithms and rules from natural language processing, the algorithm identifies visual and semantic cues such as title keywords, numbering patterns, and paragraph spacing to divide the continuous text stream into logically meaningful units, such as chapters, clauses, and paragraphs, and establishes hierarchical relationships between them. For example, text beginning with "Chapter 1" or "Section 1" is identified as chapter titles, and the subordinate content is categorized as the main text of the corresponding chapter.
[0030] Standard text data for chapter and paragraph identifiers is output based on the document's logical structure parsing.
[0031] Furthermore, after parsing the document's logical structure and constructing a hierarchical index of chapters and paragraphs, the structural information is bound and serialized with the corresponding text content. Each identified logical unit is assigned a unique identifier, and the hierarchical relationship is recorded. The entire document is identified as the root node, the next chapter as a child node, and the paragraphs under each chapter as grandchild nodes. Each node's data content includes title text and body text. All this text data with complete chapter and paragraph identification information is organized and output in a machine-readable structured data format. This output is standard text data identifying chapters and paragraphs.
[0032] The knowledge extraction module extracts key information entities and relationships from structured standard text data based on domain knowledge graphs, constructs a hierarchical association relationship network with associated weight coefficients, and obtains the first structured data.
[0033] Identifying key information entities from structured standard text data based on domain knowledge graphs.
[0034] Furthermore, core entity categories are defined in the knowledge graph of water pollutant discharge standards, such as pollutant names, industry classifications, discharge monitoring locations, and control measure types. For structured standard text data that has undergone logical structure parsing, named entity recognition technology based on a pre-trained language model is used for sequence labeling. This scans the text content and identifies phrases or terms belonging to the aforementioned predefined categories. For example, from the text fragment "Chemical Oxygen Demand (COD) emission limit is 50 milligrams per liter," COD can be identified as a pollutant name entity, and 50 milligrams per liter as an emission limit entity. Utilizing the entity dictionary and contextual semantic rules provided by the domain knowledge graph, ambiguities of general terms in a specific domain are eliminated, ensuring the accuracy of extraction and achieving a precise mapping from natural language descriptions to structured domain concepts.
[0035] Based on domain knowledge graphs, key information relationships between entities are extracted from structured standard text data and then integrated.
[0036] Furthermore, relying on predefined relation types in the domain knowledge graph, relation extraction techniques based on dependency parsing and semantic role labeling are used to analyze the grammatical structure of sentences or adjacent sentence groups containing multiple entities to determine whether predefined relations exist between entities. For example, from the text "The direct emission limit of chemical oxygen demand in the textile dyeing and finishing industry is 80 milligrams per liter," an execution relation between the textile dyeing and finishing industry and chemical oxygen demand can be extracted, while a limit relation between chemical oxygen demand and 80 milligrams per liter can be extracted. The scattered extracted relation triples are integrated to eliminate duplicate relations pointing to the same fact and form a consistent relation set.
[0037] Based on the level of explicitness in standard text data, a hierarchical relationship network is established for key information entities and relationships.
[0038] Furthermore, analyzing the logical structure of structured standard text data reveals the underlying document structure. For example, standard documents typically have a hierarchical relationship of standard-industry / watershed-pollutant-emission destination-emission limit. Based on this hierarchy, entities are arranged, establishing connections between high-level entities (such as a local standard) and low-level entities (such as the industry / watershed to which the standard applies). Low-level entities are then connected to even finer-grained entities, thus forming a tree-like or network-like hierarchical structure. The construction of this hierarchical network of relationships not only records the relationships between entities but, more importantly, restores the regulatory framework and scope of application of the standard text itself.
[0039] Based on the location and clarity of the key information entities in the standard text data, each relationship in the hierarchical relationship network is assigned a relationship weight coefficient, thus obtaining a hierarchical relationship network with associated relationship weight coefficients.
[0040] Furthermore, a correlation weight coefficient is calculated for each edge in the hierarchical relationship network. The calculation process comprehensively considers two main factors: the occurrence position quantification value and the explicitness of the description quantification value. The occurrence position quantification value is determined by the position of the text fragment upon which the relationship is based in the document. For example, a relationship appearing in a core table of the main text is given a higher base weight than a relationship appearing in an appendix description. The explicitness of the description quantification value is determined by the explicitness of the text describing the relationship. For example, a relationship using a mandatory execution statement is given a higher explicitness weight than a relationship using a referential execution statement. ,in It is a fuzzy inference-based function used to comprehensively process the quantized values of the location. and express explicit quantifiable values Input the data to obtain a basic confidence level; This represents the global structural importance of the relationship within the entire network; It is an adjustment coefficient that combines the explicit features of the text with the structural features of the network, so that the association weight coefficient can more comprehensively and reasonably reflect the actual importance and reliability of each relationship in the standard.
[0041] The expression for the association weight coefficient is: ; in, For entities at higher levels The weighting coefficients for relationships with lower-level entities. For higher-level entities, For entities of lower hierarchy, The gain coefficient of the adaptive factor. For entities at higher levels Global importance in the network of relationships of lower-level entities. To produce a position quantization value, To clarify the quantification value, This is fuzzy reasoning.
[0042] The hierarchical association network based on the associated weight coefficients constitutes the first structured data.
[0043] Furthermore, the hierarchical relational network with associated weight coefficients, along with all nodes (key information entities), edges (relationships), and the weights on the edges (association weight coefficients), is serialized and encapsulated. The encapsulated data set is the first structured data, which fully expresses the semantic knowledge, hierarchical structure, and quantitative importance assessment of each knowledge point extracted from the original standard text. This first structured data is no longer a simple text copy or database record, but a knowledge network rich in semantic associations and importance measurements.
[0044] It should be noted that fuzzy reasoning is a soft computing technique that mimics human approximate reasoning. It transforms precise input values (such as positional quantifications and explicit expression quantifications) into high, medium, and low fuzzy linguistic terms by defining membership functions. Then, it performs reasoning based on a rule base containing if-then rules (e.g., high positional importance and high explicit expression lead to high base confidence). Finally, through a defuzzification process, the fuzzy outputs obtained from the reasoning are aggregated into a precise numerical output for computation. The processing module uses a weighted document semantic model to perform semantic vectorization on the complete document text corresponding to the first structured data, resulting in a standard set of semantic vectors.
[0045] Calculation weights are assigned to different parts of the complete document text corresponding to the first structured data.
[0046] Furthermore, the complete local water pollutant discharge standard document text is logically divided into multiple text fragments. For each text fragment, the word frequency-inverse document frequency (IF-IVF) value and semantic role score are calculated based on the domain terminology dictionary. The IF-IVF value calculation focuses on the occurrence density of predefined core terms in the water pollutant discharge standard domain within the text fragment, reflecting the richness of the fragment's professional content. The semantic role score analyzes the grammatical structure of each sentence in the text fragment using semantic role annotation technology in natural language processing, identifying and counting the number of phrases expressing key semantic roles of norms, constraints, and obligations, reflecting the strength of the fragment's normative expression. After adding the IF-IVF value and the semantic role score based on the domain terminology dictionary, the result is processed through a normalization function to make the calculated weights of all text fragments comparable. Specifically, the method of allocating calculation weights combines statistical features and deep semantic function analysis, which can dynamically and finely distinguish between prescriptive core content and explanatory auxiliary content in a document, ensuring that key information is fully emphasized in subsequent vectorization.
[0047] The expression for calculating the weights is: ;
[0048] in, For the first in the document The weights of each text segment are calculated. For the first part of the complete document text A partial text fragment, As a moderating factor for the semantic role rating item, Analysis of Semantic Role Labeling Technology in Natural Language Processing The grammatical structure of each sentence in the text. For normalization function, For term frequency-inverse document frequency based on a domain terminology dictionary, This is an index for partial text fragments.
[0049] We use a weighted document semantic model to process the complete document text and assign calculated weights.
[0050] Furthermore, the complete document text, with its computational weights assigned, is input into a weighted document semantic model. When processing the text sequence, the weighted document semantic model adjusts the contribution of the words or subwords in each text segment to the model's attention mechanism based on the final computational weight assigned to that segment. Ultimately, the text segments with higher computational weights contain words that have a greater impact on the model's computation of contextual representations, thus occupying a more important position in the overall representation of the generated document.
[0051] The weighted document semantic model converts each complete local water pollutant discharge standard document into a high-dimensional semantic vector.
[0052] Furthermore, a pre-trained language model based on the Transformer architecture encodes the weighted complete document text. Through a multi-layer self-attention mechanism and a feedforward neural network, it captures the complex contextual dependencies between words in the text, gathers information from the entire sequence, and generates a fixed-dimensional, dense floating-point vector to represent the semantic core of the document. During the processing, it is guided by weights based on domain term density and semantic role rating. The generated high-dimensional semantic vector can more prominently reflect the key semantic information of the document, such as core regulations, limit requirements, and constraints in pollutant emission control.
[0053] A standard semantic vector set is constructed based on high-dimensional semantic vectors.
[0054] Furthermore, the high-dimensional semantic vectors corresponding to each local water pollutant discharge standard document processed by the weighted document semantic model are aggregated to form a standard semantic vector set. This set constitutes a semantic space, where each vector represents a specific local standard document. In the semantic space, standard documents with similar semantics will have their vectors closer together. The standard semantic vector set organizes the deep semantic information of multiple local water pollutant discharge standard documents in a structured and computable form.
[0055] The anomaly detection module performs anomaly detection on the first structured data and assigns anomaly weights to obtain high-quality structured data with weighted labels.
[0056] The first structured data is subjected to legality conflict checks, internal consistency checks, and ambiguity detection to detect anomalies.
[0057] Furthermore, by comparing the local water pollutant emission standard limits recorded in the first structured data with the corresponding national water pollutant emission standard limits, when a local standard limit is found to be more lenient than a national standard limit, it is marked as a legality conflict anomaly. Internal consistency verification is performed by scanning the first structured data to check whether there are different limit provisions for the same pollutant under the same applicable conditions in the same standard document, or whether there are inconsistent definitions of the same key term. Such contradictions are marked as internal inconsistency anomalies. The ambiguity detection uses natural language processing technology to analyze the words used to describe the provisions or conditions in the standard text, identify appropriate, necessary, and in principle lacking objective quantitative indicators in vague expressions, and mark them as ambiguity anomalies. This achieves automated and multi-angle review of the quality of the standard text content, discovering potential compliance, logical rigor, and clarity of expression issues in the standard itself, and providing a basic guarantee for data quality.
[0058] Assign anomaly weights to the detected anomalies based on their anomaly types, and then mark the first structured data with assigned anomaly weights as high-quality structured data with weighted labels.
[0059] Furthermore, each anomaly type is pre-defined with a severity baseline. For example, legality conflict anomalies are assigned the highest severity baseline because they directly violate higher-level laws; internal inconsistency anomalies are assigned the second-highest severity baseline because they cause self-contradictions in standards; and ambiguous statements are assigned a relatively low severity baseline because they affect the certainty of execution. For each specific anomaly instance, an impact range factor is multiplied by the severity baseline of the corresponding anomaly type. The factor assesses the size of the data range affected by this anomaly instance. For example, the impact range factor of an anomaly instance that affects the entire industry standard is greater than that of an anomaly instance that only affects a specific process condition.
[0060] Specifically, by calculating the anomaly weight as equal to the severity base value multiplied by the impact range factor, the quantitative severity of each anomaly instance is obtained. The first structured data is then associated and labeled with these anomalies and their weights to form high-quality structured data with weight labels. This transforms qualitative anomalies into quantitative risk indicators, thus achieving anomaly management.
[0061] The expression for assigning abnormal weights is: ; in, For the first One abnormal weight, For severity base function, It is an exception type. To evaluate the first The size of the data range affected by each anomaly For the first One abnormal instance, Index for abnormal instances.
[0062] The database management module defines the physical storage mode of the temporal database through a hierarchical association network of weight coefficients, storing high-quality structured data with weight labels and standard semantic vector sets into the temporal database.
[0063] Based on the entities, relationships, and association weight coefficients defined in the hierarchical association relationship network with associated association weight coefficients, the table structure of the temporal database is obtained.
[0064] Furthermore, the analysis of entity types in the hierarchical association network with associated weight coefficients maps each entity type to an independent data table in the temporal database. For example, the local standard entity type in the network is mapped to a local standard information table, and the pollutant entity type is mapped to a pollutant information table. The structure of each data table is defined by the attributes of the corresponding entity type. For example, the local standard information table includes fields such as standard name, issuing authority, and effective date. This ensures that the database table structure can fully carry the entity information in the knowledge network, realizing a direct and lossless conversion from the semantic model to the physical model, and laying a structural foundation for accurate data storage and efficient querying.
[0065] Foreign key associations between temporal database tables are obtained based on the entities, relationships, and association weight coefficients defined by the hierarchical association relationship network with attached association weight coefficients.
[0066] Furthermore, the relationships between entities in the hierarchical association network with associated weight coefficients are analyzed, and each relationship type is mapped to foreign key constraints between temporal database tables. For example, if there is an applicable relationship between local standard entities and industry entities in the network, then a foreign key field pointing to the primary key of the local standard information table and a foreign key field pointing to the primary key of the industry information table are set in the industry applicable relationship table, so that the semantic association between entities is strictly maintained at the database level through foreign key relationships.
[0067] Based on the entities, relationships, and association weight coefficients defined in the hierarchical association relationship network with associated association weight coefficients, a dedicated field for storing association weight coefficients is established in the temporal database table.
[0068] Furthermore, in the data table storing the relationship between entities, a dedicated field is added to each record to store the association weight coefficient corresponding to that relationship. For example, in the industry applicability relationship table, in addition to the foreign key fields pointing to the local standard and the industry, an association weight field is added to record the weight value of the relationship that the local standard applies to the industry. This makes the importance measurement of the relationship an intrinsic part of the database schema, solidifying the deep insights obtained from data mining into the storage layer.
[0069] The physical storage model of a temporal database consists of its table structure, foreign key relationships, and special fields.
[0070] Furthermore, the collection of data tables, the network of relationships established between tables through foreign keys, and the dedicated fields in the tables used to store the weight coefficients of the relationships, together form a logically rigorous overall design scheme. At this point, the physical storage mode of the database not only specifies the storage format of the data, but also defines the logical connections and quantitative attributes between the data. The physical storage mode is not a simple data container, but a deeply structured data model rich in semantic relationships and importance measurements. It can efficiently support complex and intelligent application services at the upper level and is a key infrastructure for realizing the leap from data storage to knowledge services.
[0071] The intelligent application service module provides intelligent application services based on a temporal database. By calling the correlation weight coefficient and the anomaly weight coefficient, it prioritizes the recommendation results and provides hierarchical prompts for anomaly information.
[0072] This provides similarity-based recommendation services using a set of standard semantic vectors from a temporal database.
[0073] Furthermore, when a user queries or browses a document on water pollutant discharge standards for a specific region, the high-dimensional semantic vector corresponding to the document is used as the query vector. By calculating the similarity between the query vector and all other vectors in the standard semantic vector set of the temporal database, such as using cosine similarity, the set of candidate recommended standard documents that are semantically closest can be found. The similarity comparison based on high-dimensional semantic vectors can capture the similarity between standard documents at the abstract level, such as management ideas, scope of application, and technical focus. This overcomes the limitations of traditional keyword matching and realizes intelligent recommendation based on the overall semantic content of the document. It can discover standard documents with different titles but highly related or complementary content.
[0074] This service provides anomaly alerts based on high-quality, weighted, structured data in a temporal database.
[0075] Furthermore, when a user accesses a specific clause of a particular standard, the service queries the temporal database for pre-detected and marked anomalies and their weights associated with that clause. These anomalies include legality conflicts, internal inconsistencies, and ambiguities in expression. By displaying these anomalies and their weights in association with the specific clause, the static standard query is transformed into a dynamic risk assessment process, providing users with in-depth compliance insights and risk warnings to assist in making more rigorous decisions.
[0076] In the similarity-based recommendation service, the association weight coefficient is used to sort the initial recommendation results.
[0077] Furthermore, after obtaining a preliminary set of candidate recommendation standard documents through semantic vector similarity, the average association weight coefficient of each candidate recommendation standard document stored in the temporal database is further invoked. The average association weight coefficient reflects the clarity and binding strength of the internal regulations of the candidate standard documents. Through a comprehensive scoring expression, semantic similarity and the average association weight coefficient are combined to calculate the final score of each candidate document, ensuring that semantic similarity is the basis for ranking, while the average association weight coefficient, as a moderating factor, adds points to standard documents with more explicit and mandatory regulations. This elevates the recommendation standard from content relevance to a level that is not only content-related but also of higher quality and with greater reference value, thereby improving the practicality and intelligence of the recommendation results.
[0078] The overall score expression is: ; in, Recommended standard document The overall score, To search for documents, As candidate documents, The standard document currently being viewed With candidate recommendation criteria document Semantic similarity between them As an influence moderating factor for association weights, Recommended documents for candidates The average association weight coefficient.
[0079] exist Phase: The calculation begins by finding all values related to the temporal database. Similar documents, These are merely potential candidate documents; therefore, the most accurate description is a candidate recommendation standard document, emphasizing its candidate status and indicating that it has not yet been finalized for recommendation.
[0080] exist Stage: Final scores have been calculated for candidate documents, and a final recommendation list is generated based on the scores. It has already completed the transformation from candidate to about to be recommended, so it can be called a recommendation criterion document. The recommendation describes an outcome state.
[0081] In It is the object of semantic comparison, and In It is the object of quality assessment and ranking. In The semantic comparison focuses on content features. In The focus of quality assessment and ranking is on intrinsic attributes.
[0082] The association weight coefficients are used to assign priority order to the sorted recommendation results, thus obtaining the priority order of the recommendation results.
[0083] Furthermore, based on the final score of each candidate recommended standard document, all candidate documents are sorted in descending order, with documents with higher scores appearing at the top of the recommendation list and receiving higher priority; documents with lower final scores are placed at the bottom and receive lower priority. The process of sorting based on a quantitative comprehensive score involves calling the associated weight coefficients to assign priority to the sorted recommendation results, generating a well-organized recommendation list with clear distinctions in value. This allows users to prioritize and refer to standard documents that are not only relevant in content but also more authoritative and operational in terms of regulations, greatly improving information acquisition efficiency.
[0084] In the abnormal information prompting service, the severity level of each abnormal message is determined by calling the abnormal weight coefficient, and the abnormal messages of different severity levels are then graded and prompted accordingly.
[0085] Furthermore, before notifying users of abnormal information, the system calls the abnormal weight coefficient corresponding to each abnormal message. Based on the analysis of historical abnormal data samples, the continuous value range of the abnormal weight coefficient is divided into several intervals corresponding to different severity levels, and threshold intervals are set. According to the threshold intervals, the abnormal weight coefficient is mapped to different severity levels, such as high, medium, and low risk levels. In the user interface or feedback report, different prompts are used according to different severity levels, such as color, icon, or text emphasis to distinguish the display. This quantifies and classifies the severity of abnormal information, realizes refined risk management and intuitive visualization, guides users to prioritize high-risk issues, and optimizes the decision-making process.
[0086] In summary, this invention acquires and preprocesses standard text through a data acquisition module to generate structured standard text data; a knowledge extraction module extracts key entities and relationships based on a domain knowledge graph, constructs a hierarchical association network with associated weight coefficients, and forms the first structured data; a processing module uses a weighted document semantic model to semantically vectorize the complete text, generating a set of standard semantic vectors; an anomaly detection module performs anomaly verification on the preliminary data and assigns anomaly weights, outputting high-quality structured data with weight labels; a database management module defines the physical storage mode of the temporal database based on the weighted network to achieve persistent data storage; and an intelligent application service module uses association weights and anomaly weights to prioritize recommendation results and provide hierarchical prompts for anomaly information. This invention solves the problems of existing standard data management systems, which, due to their flat models, struggle to express complex semantic logic and hierarchical constraints and cannot support deep mining and intelligent reasoning.
[0087] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A structured database system for local water pollutant discharge standards based on data mining, characterized in that: include, The data acquisition module collects local water pollutant discharge standard documents from multiple locations and preprocesses them to obtain structured standard text data. The knowledge extraction module extracts key information entities and relationships from structured standard text data based on domain knowledge graphs, constructs a hierarchical association relationship network with associated weight coefficients, and obtains the first structured data. The processing module uses a weighted document semantic model to perform semantic vectorization on the complete document text corresponding to the first structured data, and obtains a standard semantic vector set. The anomaly detection module performs anomaly detection on the first structured data and assigns anomaly weights to obtain high-quality structured data with weight labels. The database management module defines the physical storage mode of the temporal database through a hierarchical association network of weight coefficients, and stores high-quality structured data with weight labels and standard semantic vector sets into the temporal database. The intelligent application service module provides intelligent application services based on a temporal database. By calling the correlation weight coefficient and the anomaly weight coefficient, it prioritizes the recommendation results and provides hierarchical prompts for anomaly information. The complete document text corresponding to the first structured data is semantically vectorized using a weighted document semantic model to obtain a standard semantic vector set, including the following steps: Assign computational weights to different parts of the complete document text corresponding to the first structured data; The expression for calculating the weights is: ; in, For the first in the document The weights of each text segment are calculated. For the first part of the complete document text A partial text fragment, As a moderating factor for the semantic role rating item, Analysis of Semantic Role Labeling Technology in Natural Language Processing The grammatical structure of each sentence in the text. For normalization function, For term frequency-inverse document frequency based on a domain terminology dictionary, Indexing partial text fragments; We utilize a weighted document semantic model to process the complete document text, assigning and calculating weights accordingly. The weighted document semantic model converts each complete local water pollutant discharge standard document into a high-dimensional semantic vector; A standard semantic vector set is constructed based on high-dimensional semantic vectors; Anomaly detection and anomaly weighting are performed on the first structured data to obtain high-quality structured data with weighted labels, including the following steps: The first structured data is subjected to legality conflict checks, internal consistency checks, and ambiguity detection to detect anomalies; Assign anomaly weights to the detected anomalies based on their anomaly types, and then mark the first structured data with assigned anomaly weights as high-quality structured data with weighted labels.
2. The structured database system for local water pollutant discharge standards based on data mining as described in claim 1, characterized in that: Local water pollutant discharge standards from multiple regions were collected and preprocessed to obtain structured standard text data, including the following steps: We collected texts of local water pollutant discharge standards from multiple regions and converted them to a unified format. The local water pollutant discharge standards text, after being uniformly converted in format, undergoes character encoding cleaning before the document's logical structure is parsed. Standard text data for chapter and paragraph identifiers is output based on the document's logical structure parsing.
3. The structured database system for local water pollutant discharge standards based on data mining as described in claim 2, characterized in that: Extracting key information entities and relationships from structured standard text data based on domain knowledge graphs includes the following steps: Identifying key information entities from structured standard text data based on domain knowledge graphs; Based on domain knowledge graphs, key information relationships between entities are extracted from structured standard text data and then integrated.
4. The structured database system for local water pollutant discharge standards based on data mining as described in claim 3, characterized in that: Constructing a hierarchical association network with associated weight coefficients to obtain the first structured data includes the following steps: Establish a hierarchical relationship network for key information entities and relationships based on the level of explicitness in standard text data; Based on the location and clarity of the key information entities in the standard text data, each relationship in the hierarchical relationship network is assigned a relationship weight coefficient, thus obtaining a hierarchical relationship network with associated relationship weight coefficients. The hierarchical association network based on the associated weight coefficients constitutes the first structured data.
5. The structured database system for local water pollutant discharge standards based on data mining as described in claim 4, characterized in that: The physical storage model of the temporal database is defined through a hierarchical association network of weighted coefficients, including the following steps: Based on the entities, relations, and association weight coefficients defined in the hierarchical association relationship network with associated association weight coefficients, the table structure of the temporal database is obtained; Based on the entities, relations, and association weight coefficients defined in the hierarchical association relationship network with attached association weight coefficients, foreign key associations between temporal database tables are obtained; Based on the entities, relationships, and association weight coefficients defined in the hierarchical association relationship network with associated association weight coefficients, a dedicated field for storing association weight coefficients is established in the temporal database table. The physical storage model of a temporal database consists of its table structure, foreign key relationships, and special fields.
6. The structured database system for local water pollutant discharge standards based on data mining as described in claim 5, characterized in that: Storing high-quality structured data with weighted labels and a standard set of semantic vectors into a temporal database includes the following steps: High-quality structured data with weighted labels is mapped to the corresponding table in the temporal database, and the standard semantic vector set is stored in the dedicated vector field of the temporal database, resulting in the persistent storage of high-quality structured data with weighted labels and the standard semantic vector set in the temporal database.
7. The structured database system for local water pollutant discharge standards based on data mining as described in claim 6, characterized in that: Providing intelligent application services based on temporal databases includes the following steps: Provides similarity-based recommendation services based on a set of standard semantic vectors in a temporal database; This service provides anomaly alerts based on high-quality, weighted, structured data in a temporal database.
8. The structured database system for local water pollutant discharge standards based on data mining as described in claim 7, characterized in that: By invoking the correlation weight coefficient and the anomaly weight coefficient, the recommendation results are prioritized and the anomaly information is displayed in a tiered manner, including the following steps: In the similarity-based recommendation service, the association weight coefficient is used to sort the initial recommendation results; The association weight coefficients are used to assign priority order to the sorted recommendation results, thus obtaining the priority order of the recommendation results; In the abnormal information prompting service, the severity level of each abnormal message is determined by calling the abnormal weight coefficient, and the abnormal messages of different severity levels are then graded and prompted accordingly.