Recommendation method and device for investigation and classified supervision of soil pollution risk enterprise
By integrating multi-source data through a knowledge graph and large language model, the method addresses inefficiencies in identifying soil and groundwater pollution risks, providing precise enterprise classification for effective risk management.
Patent Information
- Application Number
- GB2025007979
- Authority / Receiving Office
- GB · GB
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-24
- Filing Date
- 2025-05-22
- Publication Date
- 2025-11-26
AI Technical Summary
Existing methods for identifying soil and groundwater pollution risks in industrial enterprises are inefficient and inaccurate due to the lack of standardized enterprise data and missing correlations between basic enterprise data and pollution risks, leading to incomplete risk identification and management.
A method and device utilizing a knowledge graph and large language model to integrate multi-source data, reason enterprise production characteristics, and calculate the probability of volatile and migratory pollutants, combined with spatial analysis to create a classified supervision list.
Accurately identifies and classifies enterprises prone to soil pollution, enhancing the precision of risk management and supporting targeted supervision and remediation efforts.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
The present application relates to the technical field of soil pollution risk control and supervision of enterprises, and more particularly, to a recommendation method and device for investigation and classified supervision of soil pollution risk enterprises. BACKGROUND The emission of toxic and hazardous substances in the production and operation activities of industrial enterprises may easily lead to soil and groundwater pollution. If the pollution is allowed to migrate and spread without supervision and risk control, it will affect the safety of human living environment and ecological environment. It is necessary to promptly investigate the plots with pollution risks and carry out targeted supervision. Relevant management departments are trying various methods to conduct pollution risk inspections and dynamically update supervision lists, such as the list of key soil pollution supervision units and the list of priority supervision plots. Soil and groundwater pollution has significant spatiotemporal heterogeneity, and pollution risk investigation requires multidimensional data information as a basic support. However, enterprise land archives management in the early stage is not standardized, basic data and historical information are lacking, and more manpower and material resources need to be invested for data collection and analysis, and the investigation efficiency is low and there is a possibility of omission. How to quickly and effectively identify enterprises with pollution risks and carry out targeted classified management has become a technical problem that needs to be solved urgently. In the existing technologies, basic enterprise data, such as enterprise names, industry categories, etc., are used to preliminarily identify suspected polluted enterprise plots through black box algorithms such as machine learning. However, the basic enterprise data is not strongly correlated with the indication of soil and groundwater pollution risks of the plots, and key characteristics such as enterprise production characteristics and characteristic pollutants are easily missed. At the same time, it is also impossible to specifically distinguish risk types, which affects the accuracy of the identification results and the precision of classified management. SUMMARY In order to solve the above technical problems, the present application provides a recommendation method and device for investigation and classified supervision of soil pollution risk enterprises. The following technical solutions are adopted: A recommendation method for investigation and classified supervision of soil pollution risk enterprises includes the following steps: step 1: starting from a full-aperture enterprise list, screening and integrating a list of key industrial enterprises to which need to be paid attention according to a specific rule as an investigation base; step 2: according to availability and convenience of data, gathering multi-source data such as enterprise production characteristics and environmental distribution characteristics to screen enterprise basic attribute indicators with high correlation indicating risks of soil and groundwater pollution in plots, the enterprise basic attribute indicators comprising an enterprise production scale, a production and operation activity time, a three wastes (wastewater, waste gas and solid waste) emission management level, enterprise main product information, a ground anti-seepage level of an enterprise location, soil permeability, a groundwater depth, and a surrounding sensitive point density; step 3: focusing on the production characteristics of the enterprise in risk investigation, by establishing a knowledge graph, reasoning from known information whether the enterprise is involved in volatile and migratory pollutants, where in the reasoning process, a large language model is combined to achieve efficient data identification and batch matching to screen out enterprises involved in volatile and migratory pollutants by calculation; where risk investigation focuses primarily on the production characteristics of enterprises, and screens out enterprises involved in production, use, storage, disposal or discharge of toxic and hazardous substances, with special attention paid to volatile and migratory pollutants among toxic and hazardous substances; however, much direct information on volatile and migratory pollutants obtained from enterprises is missing; however, there is a logical chain of association between the production process, raw materials and auxiliary materials and products, and pollutants of enterprises; by establishing the knowledge graph, it is possible to reason whether the enterprise is involved in volatile and migratory pollutants from known information; in the reasoning process, the large language model is further combined to achieve efficient data identification and batch matching to screen out enterprises involved in volatile and migratory pollutants by calculation; step 4: determining a possibility of the enterprise pollutants entering the soil, based on pollution characteristic statistics of supervised enterprises, integrating characteristic thresholds of various pollution indicators, screening out enterprises that are prone to soil pollution from those involving volatile and migratory pollutants; and step 5: distinguishing a degree of impact of the migration and spread of the enterprise pollution on the surrounding environment again, setting characteristic thresholds based on hydrogeological conditions of pollution spread and the distribution characteristics of the surrounding environment by combining prior knowledge to screen out enterprises that are prone to pollution spread, enterprises that are highly sensitive in the surrounding areas, and enterprises that are subject to key risk investigation in sequence from enterprises that are prone to foil pollution, and forming a classified supervision list; and identifying key focus areas according to the spatial distribution characteristics of each type of enterprises. Optionally, a specific method of step 1 includes: obtaining full-aperture yellow page enterprises and industrial and commercial enterprise data in batches, forming a list of key industrial enterprises according to a screening rule of key soil pollution industries, obtaining spatial coordinates of the enterprises based on a map service API interface, and merging duplicate and redundant enterprises; and comparing an enterprise list that have been included in supervision by a management department, marking supervision statuses of the enterprises, and distinguishing the enterprises that have been included in supervision and the enterprises that have not been included in supervision. Optionally, in step 2, the basic attribute indicators of the enterprise are standardized using the following method: for the production scale, performing numerical classification based on the number of employees and annual turnover distribution of the enterprise; for the time of production and operation activities, calculating an interval between registration and closure of the enterprise; for the three wastes discharge management level, calculating the distribution density of nearby enterprises that have been supervised; and for the main product information of the enterprise, performing batched structured recognition by the large language model; the enterprise spatial data indicators are standardized using the following method: superimposing an enterprise distribution layer with a surface impermeable surface layer, a soil layer property layer, a groundwater depth layer, and a sensitive point distribution layer through spatial information coordinate conversion to extract the spatial characteristic data of the corresponding enterprise location, including a ground anti-seepage ratio, soil permeability, a groundwater depth, and a surrounding sensitive point density. Optionally, in step 4, based on reasoning by the knowledge graph, the probability of the relationship between enterprise products and volatile and migratory pollutants is calculated, and when product-related data of the enterprise to be evaluated is inputted, the probability of the enterprise involved in volatile and migratory pollutants is outputted immediately. Optionally, a specific method for calculating the probability of the relationship between enterprise products and volatile and migratory pollutants includes: for a purpose of knowledge reasoning, defining a knowledge graph ontology in combination with enterprise production, pollution emission characteristics and pollutant attributes, and expert industry knowledge of pollutant attributes, including entities, entity attributes and relationships between the entities; designing a triple structure data table to store the above entities, entity attributes and relationships between the entities; corresponding to the entities, the entity attributes, and relationship fields between the entities, extracting entity data and attributes and relationships there of one by one using pollution discharge permit registration data of representative sample enterprises, , storing the same according to the triple structure data table, and carrying out data cleaning and integration to achieve mapping to the knowledge graph ontology; analyzing the relationship probability between the entities to assign weights to key node relationships, determining a connection condition probability of each edge in the knowledge graph, and filling this conditional probability into the relationship attributes of the relationship table in the triple table; establishing a visual graph database based on the filled triple structure data table, and using a graph database for storage, where when new enterprise triple data is loaded and updated, the entities, the entity attributes, the relationships between the entities, and corresponding entity relationship conditional probabilities in the relevant triple structure data table are also updated; based on the conditional probability between entity relationships in the graph database, calculating the probability of the enterprise to be evaluated involving volatile and migratory pollutants according to a reasoning formula; if the product-related data of the enterprise to be evaluated is semi-structured, calling the large language model to perform batched structural recognition of product entities, and performing batched matching and reasoning calculations with the corresponding entities in the existing knowledge graph library; and setting a probability threshold for a reasoning result and screening out enterprises involved in volatile and migrating pollutants from the key industry enterprises. Optionally, a comprehensive probability of volatile and migratory pollutants generated by the product is calculated using the following formula: n m zrh i-l J-l where i is a path number, j is the number of the edge on the path i, and m is the number of edges on the path i; if the product of the enterprise to be evaluated, a production process or raw materials and auxiliary materials information used by the product are available at the same time, when calculating the above comprehensive probability, only an input edge weight of the production process or raw materials and auxiliary materials needs to be set to 1, and input edge weights of the remaining production processes or raw materials and auxiliary materials need to be set to 0. Optionally, in step 5, the pollution indicator characteristics of the supervised soil pollution enterprises are statistically analyzed, and the pollution indicator data of the enterprises to be evaluated are screened accordingly, such that each indicator value is not lower than a lower quartile value of the corresponding indicator of the supervised enterprises, and the enterprises that are prone to soil pollution are comprehensively screened out. Optionally, in step 5, the indicator thresholds are set hierarchically according to the expert prior knowledge, and enterprises with underground soil permeability higher than a specific value and groundwater depth lower than a specific value are screened according to the pollution diffusion conditions; and based on the impact of pollution on the surrounding environment, enterprises with a density of sensitive points around locations thereof higher than a certain value are screened to achieve rapid screening of specific types of enterprises. A recommendation device for investigation and classified supervision of soil pollution risk enterprises includes: an integrated data loading and storage module, an attribute data processing module, a spatial data processing module, a knowledge graph reasoning module, a pollution characteristic statistics module, a classification screening and list recommendation module, and a spatial display module, where the device, when executing a computing program, is configured to implement the method described in any one of claims 1 to 5, and where the data loading and storage module is configured to classify, mark, and store acquired multisource data, and support data import and update; the attribute data processing module is configured to integrate and clean list attribute data, perform standardized processing according to a rule, and support batched export; the spatial data processing module is configured to perform standardized processing according to the rule using ArcGIS mapping analysis software and support batched export; the knowledge graph reasoning module is configured to load the filled triple structure data table, establish a graph database, and support data update; and call the large language model to perform batched structural processing on main product information of the enterprise to be evaluated, and calculate a probability of the enterprise involved in volatile and migratory pollutants according to the rule, and. screen out the enterprises involved in volatile and migratory pollutants after set the probability threshold; the pollution characteristic statistics module is configured to automatically identify and count the relevant pollution indicator data of supervised enterprises and provide pollution indicator screening thresholds; the hierarchical screening and list recommendation module is configured to automatically export the recommended list according to default indicator screening thresholds, and support manual update of the indicator screening thresholds; and the spatial display module is configured to display the spatial distribution of various recommended enterprises in a classified manner, distinguish the enterprises by color and legend marks, and support identification of key focusing areas. Optionally, the integrated data loading and storage module, the attribute data processing module, the spatial data processing module, the knowledge graph reasoning module, the pollution characteristic statistics module, the classification screening and list recommendation module, and the spatial display module are implemented based on a computer server. In summary, the present application includes at least one of the following beneficial technical effects: The present application can provide recommendation method and device for investigation and classified supervision of soil pollution risk enterprises. The method includes: gathering multisource data such as enterprise production characteristics and environmental distribution characteristics, and integrating efficient data processing technologies such as the knowledge graph, the large language model, and spatial analysis, accurately focusing on enterprises that are prone to soil pollution from the full-aperture enterprise list, paying further attention to the impact of pollution spread on the surrounding environment, and distinguishing enterprise risk types in combination with geographical environmental conditions of pollution spread, automatically calculating and recommending the classified supervision list to support relevant managers to supplement and improve the supervision list and promote subsequent risk control and remediation activities. BRIEF DESCRIPTION OF THE DRAWINGS Fig. lisa schematic flowchart of a recommendation method for investigation and classified supervision of soil pollution risk enterprises according to the present application; Fig. 2 is a schematic flowchart of integrating the list of key soil pollution industries and standardizing multi-source data according to the present application; Fig. 3 is a schematic diagram of a process of establishing a knowledge graph and knowledge reasoning for enterprise production and pollution identification according to the present application; Fig. 4 is a schematic structural diagram of a knowledge graph ontology according to the present application; Fig. 5 is a schematic diagram of a recommendation device for investigation and classified supervision of soil pollution risk enterprises according to the present application; and Fig. 6 is an effect diagram of spatial classification display of a specific embodiment of the present application. DETAILED DESCRIPTION The present application will be described in further detail below in conjunction with the accompanying drawings. Embodiments of the present application discloses a recommendation method and device for investigation and classified supervision of soil pollution risk enterprises. Referring to Figs. 1 to 6, in a first aspect, starting from a full- aperture enterprise list, a list of key industrial enterprises that need attention is screened and integrated according to a specific rule as an investigation base. Taking a certain city as an example, more than 60,000 yellow pages of enterprises and more than 400,000 pieces of industrial and commercial enterprise list data of the city are obtained in batches for data cleaning and field integration. According to the key industrial screening rules in Appendix B of the Technical Guidelines for Investigation of Soil Pollution Status of Construction Land (HJ 25.1-2019), non-key industrial enterprises are eliminated to form a list of key industrial enterprises; enterprise spatial coordinates are obtained based on a map service API interface and enterprises are merged with duplicate and redundant addresses. After screening and integration, more than 10,000 key industrial enterprises in the city are obtained, which serve as the investigation base. The comprehensive enterprise names, spatial coordinates, and industrial category information are compared with the list of enterprises supervised by the management department, and the supervision status of each enterprise in the list is marked to distinguish enterprises that have been supervised and enterprises that have not been unsupervised. In a second aspect, based on the availability and convenience of data, highly correlated indicators indicating the risk of soil and groundwater pollution in plots are screened, including (1) basic attribute data: an industry category to which the enterprise belongs, a production scale, a production and operation activity period, a three wastes (wastewater, waste gas and solid waste) emission management level, and a main product information of the enterprise; (2) spatial data: a ground anti-seepage level, soil permeability, a groundwater depth, and a density of surrounding sensitive points at the enterprise location. In the embodiment, the corresponding indicator data in the list of key industrial enterprises in the city are collected in batches and standardized. In the basic attribute data, for the industrial category indicators, the industry category data in the list are uniformly processed according to the standard format of the National Economic Industry Classification (GB / T4754-2017); for the production scale indicator, the number of employees and annual turnover of the enterprise are comprehensively counted, and each data is divided into 13 levels from 1 to 13, and the highest level of the two types of data is taken as the standardized data, which is a discrete numerical field; for the production and operation activity time indicator, the annual interval from enterprise registration to closure is calculated as standardized data, which is a discrete numeric field; for the three wastes emission management level indicator, ArcGIS is used to conduct a kernel density analysis on the land parcels that have been included in the supervision, and the kernel density values of the land parcels included in the supervision at the location of each enterprise are used as the standardized data, which is a continuous numerical field; and for the information of the main products of the enterprise, batched structured recognition is performed through the large language model (Described in detail in the third aspect). In the spatial data, for the ground anti-seepage level, soil permeability, groundwater depth, and surrounding sensitive point density indicator, the impermeable surface grid data (accuracy 10 m), soil layer property grid data (1:100,000), groundwater depth point data, and human settlement sensitive point POI data (schools, residential areas, hospitals, public management and public service places) of the city are collected in batches on the corresponding public source data website. The enterprise distribution layer is superimposed on the above layers after spatial information coordinate conversion, and the spatial characteristic data of the corresponding enterprise location is extracted, including: for the ground anti-seepage level, the ground anti-seepage ratio corresponding to each enterprise distribution point is classified as 1 or 0, which is a discrete numerical field; for soil permeability, the soil property classification corresponding to each enterprise distribution point (1 heavy clay, 2 silty clay, 3 clay, 4 silty clay loam, 5 clay loam, 6 silt, 7 silt loam, 8 sandy clay, 9 loam, 10 sandy clay loam, 11 sandy loam, 12 loamy sand, 13 sand) is used as standardized attribute data, which is discrete numerical fields; for the groundwater depth, ArcGIS inverse distance interpolation is used to rasterize the point data, and the depth values corresponding to the distribution points of each enterprise are used as the standardized attribute data as continuous numerical fields; ArcGIS is used to perform kernel density analysis on the density of surrounding sensitive points. The density of sensitive points corresponding to the distribution points of each enterprise is normalized and used as standardized attribute data, which is continuous numerical fields. In a third aspect, the risk investigation focuses primarily on the production characteristics of enterprises and screens out enterprises involved in the production, use, storage, disposal or discharge of toxic and hazardous substances, with special attention paid to volatile and migratory pollutants among toxic and hazardous substances. For the fields in the enterprise list after cleaning and integration, the data coverage varies. For example, the enterprise product information is relatively comprehensive, but much information on the volatile and migrating pollutants of the enterprise is missing, affecting the accuracy of enterprise screening. However, there is a logical chain of association between the fields of enterprise production processes, raw materials and auxiliary materials and products, and volatile and migratory pollutants. Through deductive reasoning of complex relationships, the probability of the relationship between enterprise products and volatile and migratory pollutants can be screened out. Based on the above foundation, the present application establishes a method for reasoning whether an enterprise is involved in volatile and easily migrating pollutants based on the knowledge graph. The general framework is as follows: based on the purpose of this knowledge graph deductive reasoning, a top-down approach is used to first define the knowledge graph ontology, then map the entities, entity attributes, and entity relationships in the enterprise sample data to the defined ontology, and assign weights to key node relationships based on the parsed probability of the sample entity relationships to construct a knowledge graph, and on this basis calculate and reason the probability that the enterprise to be evaluated is involved in volatile and migrating pollutants. A specific implementation process includes: First, the knowledge graph ontology is defined based on expert and industry knowledge such as enterprise production, pollution generation, and discharge characteristics, and pollutant attributes, including entities, entity attributes, and relationships between the entities. The ontology structure relationship is shown in Fig. 4, where the entities include eight categories: enterprises, products, industries, production processes, raw materials and auxiliary materials, intermediate products, toxic and harmful pollutants, and volatile and easily migrating pollutants, which are represented by nodes in the knowledge graph; the entity attributes refer to the properties of the entities themselves, that is, node attributes. For example, the attributes of raw materials and auxiliary materials may include pollutants. The relationship between the entities refers to the mapping relationship between the entities. In the knowledge graph, edges are used to represent the relationship between entity nodes. For example, the production process nodes and the intermediate 5 product nodes are connected by the "produce" edges. A triple structure data table is set to store the above entities, entity attributes and relationships between the entities, and each row of labels reflects the relationship between the entities. In the embodiment, a triple structure data table (Table 1) between the entity 1 enterprise and the entity 2 products is provided, including the enterprise name, the three products produced by the enterprise 10 and the main attributes of the products; in another embodiment, a triple structure data table (Table 2) between the production process of the entity 3 and the raw materials and auxiliary materials of the entity 4 of the same enterprise is provided, including three types of production processes involved in the enterprise, the raw materials and auxiliary materials corresponding to each production process and their pollutant properties. In this embodiment, there is a one-to-many 15 relationship between the entities, and all are one-to-one correspondence through row labels; the triple structure data tables between other entities are similar to the above structures, as shown in Table 1 and Table 2. Table 1 Triple structure data table example 1 Entity 1 Entity 2 Relationship Entity Enterprise Name Entity Product Name Product Attributes Relationship Relationship Name Enterprise Enterprise A Product Hardware Main Products Have Produce Enterprise Enterprise A Product Jewelry Pieces Non-main products Have Produce Enterprise Enterprise A Product Hardware Main Products Have Produce Table 2 Triple structure data table example 2 Entity 3 Entity 5 Relationship Entity Process Name Entity Names of raw materials and auxiliary materials Properties of pollutants in raw materials and auxiliary materials Relationship Relationship Name Production process Galvanizing A Raw materials and auxiliary materials Chromic anhydride Chromium Have Need Production process Galvanizing A Raw materials and auxiliary materials Sulfuric acid Sulfuric acid Have Need Production process Galvanizing B Raw materials and auxiliary materials Sodium hydroxide Sodium hydroxide Have Need Production process Galvanizing B Raw materials and auxiliary materials Zinc sheet Zinc Have Need Production process coppering Raw materials and auxiliary materials Phosphorus copper particles Copper Have Need Production process coppering Raw materials and auxiliary materials Copper sulfate Copper Have Need Production process Coppering Raw materials and auxiliary materials Hydrochloric acid Hydrochloric acid Have Need Secondly, according to the entities and entity attributes as well as the relationships between the entities in the ontology, entity data and its attributes and relationships are extracted directly one by one from the pollution discharge permit registration data of representative sample enterprises by using a Python script and are stored in the triple structure data table to achieve 5 mapping to the defined ontology. In the embodiment, the pollutant discharge permit registration forms of 50 representative enterprises in the metal surface treatment and thermal processing industry are selected, and the enterprises, products, production processes, raw materials and auxiliary materials, intermediate product entity data and attributes thereof and corresponding relationships are extracted one by one by using the Python script and automatically filled in the corresponding triple structure data table; for the entity attributes of toxic and hazardous substances and volatile and migratable pollutants, the pollutant property parameter dictionary library that is currently more mature in related industry applications is automatically loaded. The style may refer to Appendix B of HJ25.3. The automatically extracted and filled data of each enterprise may contain duplication, redundancy, errors or relationship contradictions. In order to ensure the accuracy of the representative sample analysis data, manual review and correction are carried out on the data such as "product pollutant attributes", "raw material pollutant attributes", and "intermediate product pollutant attributes", such as eliminating the pollutant attributes with synonyms or near synonyms in the same entity, thereby completing data cleaning and integration to form a relatively standardized and unified structural data table to improve the accuracy of subsequent relationship weighting and probability reasoning. Once more, the relationship probability between the entities is analyzed to assign weights to key node relationships, determine the connection condition probability of each edge in the knowledge graph, and fill this conditional probability into the relationship attribute of the relationship table in the triple table. The present application mainly targets three types of entity relationships: product and production process, production process and raw materials and auxiliary materials, and production process and intermediate products. Based on traversal retrieval and statistics of the frequency of co-occurrence between two entities, the probability of each connection condition is determined based on its proportion in the total sample. The experience of industry experts may be further introduced to correct the statistical conditional probability. For example, through traversal retrieval statistics, it is found that the conditional probability of the product "nut" and the production process "galvanizing" is 81%, the conditional probability of the product "nut" and the production process "nickel plating" is 16%, and other production processes include "chrome plating" and "spray painting". Finally, a visual graph database is established based on the filled triple structure data table (including the relationship conditional probability attributes calculated above), which is generally stored in the graph database neo4j to facilitate query and quality assessment. For example, when the product "nuts" is queried in the graph database, the relationship indicator is indicatored to all the production processes, raw materials and auxiliary materials, intermediate products, toxic and hazardous substances, volatile and migratory pollutants and conditional probability situations involved in its production; this may be used to verify whether the queried relationship conforms to the defined ontology structure or whether there is a missing relationship. If so, it is returned to modify the original triple structure data table. At the same time, when the representative enterprise sample is further expanded, the entities, entity attributes, entity relationships and corresponding entity relationship conditional probabilities in the triple structure data table may be updated respectively. Next, based on the conditional probabilities between entity relationships in the graph database, the probability of the enterprise to be evaluated involved in volatile and migratory pollutants is reasoned. Specifically, a traversal search is performed in the graph database to find all paths (assuming there are n paths) with the starting point being the product entity of the enterprise to be evaluated and with the end point being the volatile and mobile pollutant entity. The weights of the adjacent node edges on each path are retrieved (i is the path number, j is the number of the edge on the path i), and the comprehensive probability of the volatile and mobile pollutants generated by the product is calculated by the following formula (m is the number of edges on the path i): n m [=1 j=l If there is the product of the enterprise to be evaluated, the production technology or the information of the raw materials and auxiliary materials adopted by the product at the same time, when the above-mentioned comprehensive probability is calculated, only the input edge weight of the production technology or the raw materials and auxiliary materials needs to be set to 1, and the input edge weights of the remaining production technologies or the raw materials and auxiliary materials are set to 0. When multiple products are involved in the enterprise , the maximum value of the probability of each product involving volatile and migratory pollutants is taken; when it comes to information on non-primary product attributes, the probability is discounted by 20%, as shown in Table 3. Table 3 Probability results of enterprises involving volatile and migratory pollutants Enterprise Name Product Name Involving volatile and migratory pollutants Enterprise A Metalware 94.20% Enterprise A Accessories 82.10% Finally, based on the above reasoning rule, the probability of volatile and migratory pollutants involved in the batch of enterprises to be evaluated summarized in the second aspect is calculated. If the product-related data of the enterprise to be evaluated is semi-structured, the Python script is used to call the large language model to perform batched recognition of products, raw materials and auxiliary materials, and production process entities thereof, and match them with the corresponding entities in the existing knowledge graph library in batches. In a specific embodiment, the large language model prompt is set as follows: You are an assistant for extracting the main product and business scope of an enterprise, and now you want to extract the information of "products", "production technology", and "r raw materials and auxiliary materials "from the original text: (1) “Products” refers to the items produced by the enterprise for sale, such as “leather gloves”, “metalware”, and the like; (2) “Production process” refers to the process required to produce products, such as “galvanizing”;(3) “Raw materials and auxiliary materials” refer to the raw materials and auxiliary materials necessary for the production of products. Please extract the above information from the following information, if not, return "none" and return it in the form of JSON: <Original text information to be inputted>. <Original text information to be inputted> is replaced with the specific text information to be extracted when used, such as metal parts and galvanizing, and the structured result returned is as follows: {"product": "metal parts", "Production process": "galvanizing", "raw materials and auxiliary materials": "none"} The identified entities are semantically matched with the entities in the existing knowledge graph library. For example, the above results are matched with the product entity "metalware" and the production process "galvanizing" in the knowledge graph library, and then the probability of the enterprise involved in volatile and migratory pollutants is obtained according to the reasoning formula. In the embodiment, the Python script is used to import semi-structured data such as the main product and business scope of the enterprise to be evaluated summarized in the second aspect in batches. After structured processing and entity matching by the large language model, batch reasoning is used to calculate the probability of each enterprise involving migratory pollutants. Those with a probability of not less than 0.9 are screened as enterprises involving volatile and migratory pollutants, and there are more than 2,000 such enterprises involved in this city. In a fourth aspect, the pollution indicator characteristics of the supervised soil polluting enterprises are statistically analyzed, and the pollution indicator data of the enterprises to be evaluated are screened accordingly, such that each indicator value is not lower than the lower quartile value of the corresponding indicator of the supervised enterprises, and the enterprises that are prone to soil pollution are comprehensively screened out. In the embodiment, 150 samples of soil polluting enterprises that have been supervised are selected from more than 2000 enterprises involving volatile and migratory pollutants, and the lower quartile values of the indicator data of production and operation activity time, production scale, three wastes emission management level, and ground anti-seepage level thereof are respectively counted. The indicators of the enterprise to be evaluated are screened accordingly, such that each indicator value is not less than the lower quartile value of the corresponding indicator of the supervised enterprise, that is, the production and operation activity time is not less than 17 years, the production scale is not less than 10, and the three wastes emission management level is not less than 0.1, thereby screening out more than 100 enterprises to be included in the supervision that are prone to soil pollution (excluding those already included in the supervision). According to this indicator data limit, all enterprises that have been included in the supervision of the city are verified, and the enterprises screened out are all soil polluting enterprises. This method is based on the screening of pollution characteristics of existing enterprises in the city. The pollution characteristics of enterprises in different cities are different, so the data screening values may be different; different screening quantile thresholds may be selected based on the supervisory intensity of different cities, such as deciles and medians. In a fifth aspect, the conditions for the diffusion of soil pollution in enterprises and the possibility of impact thereof on the surrounding areas are analyzed, and the indicator thresholds are set in layers according to the expert prior knowledge to screen out enterprises with good pollution migration and diffusion conditions and high surrounding sensitivity. In the embodiment, in the enterprises prone to soil pollution, based on the pollution migration and diffusion conditions, enterprises (more than 20) with underground soil permeability not less than 11 and groundwater depth less than 6 are selected and marked as enterprises of concern for pollution diffusion; in terms of surrounding sensitivity, enterprises (more than 10) with a relative density of surrounding sensitive points higher than 0.6 are selected and marked as surrounding sensitive enterprises; enterprises (several) that meet both types of screening criteria will be marked as key risk screening enterprises. Based on this rule, rapid screening of enterprises with specific risk types is achieved. Different indicator thresholds may be set based on the hydrogeological conditions and supervisory intensity of different cities, such as groundwater depth less than 2. In a sixth aspect, an automatic recommendation computing device is constructed, which includes a data loading and storage module, an attribute data processing module, a spatial data processing module, a knowledge graph reasoning module, a pollution characteristic statistics module, a classified screening and list recommendation module, and a spatial display module. In the embodiment, the data loading and storage module imports the source data such as the full-aperture enterprise list, the supervised enterprise list, the enterprise spatial coordinates, the soil layer property grid data, the groundwater depth point data, the sensitive point POI data, and the like collected in batches, and classifies and marks the storage, and may be imported and replaced multiple times according to the data update situation. In the attribute data processing module, list integration is clicked, and automatic calculation is performed according to the integration rule described in the first aspect to unify the enterprise list, and the duplicate enterprises that may not be automatically identified are manually corrected. Attribute data cleaning is clicked, and according to the attribute data cleaning rule described in the second aspect, the corresponding field data in the data loading and storage module is identified, automatic standard structured processing is performed, and it is stored in the enterprise list field table; the data reported as errors in automatic processing may be manually corrected in batches. In the spatial data processing module, according to the spatial data processing rule described in the second aspect, the corresponding spatial layer data in the data loading and storage module is identified, and the ArcGIS mapping analysis software is set to be called, and the analysis rule is embedded into the software interface to automatically extract the corresponding spatial characteristic data of the corresponding enterprise location, and match and write it into the enterprise list field table. In the knowledge graph reasoning module, the filled triple structure data table described in the third aspect is loaded to establish a visual graph database; it supports loading and updating of new enterprise triple data, and the entities, entity attributes, entity relationships and corresponding entity relationship conditional probabilities in the relevant triple structure data table are updated immediately. The main product fields of enterprises in the data loading and storage module are automatically identified, to support calling of the large language model to perform data structure processing and entity matching; the knowledge reasoning formula described in the third aspect is embedded to calculate and derive the probability of each enterprise involved in volatile and migratory pollutants in batches. After the probability thresholds are set, the enterprises involved in volatile and migratory pollutants are screened and derived. In the pollution characteristic statistics module, according to the soil pollution enterprise characteristic statistics rule described in the fourth aspect, the enterprise list field in the data loading and storage module is automatically identified, the corresponding pollution indicator screening threshold is calculated, and the manual modification of the pollution indicator screening threshold is supported in the interface; after the threshold is confirmed, the enterprises that are prone to soil pollution will be screened out and included in the supervision. In the classification screening and list recommendation module, according to the description in the fifth aspect, the threshold for manually filling in the indicator screening is set in the interface, and the enterprise list field in the data loading and storage module is automatically identified, and pollution spread concern enterprises, surrounding sensitive concern enterprises, and key risk investigation enterprises are respectively exported according to the threshold screening. In the spatial display module, ArcGIS mapping analysis software is set to be called, and the coordinates of the aforementioned multi-level screening recommended enterprises are automatically accessed, distinguished by color and legend marks, and the spatial distribution of the enterprises in each recommended list is classified and displayed, as shown in Fig. 6, to support visual identification of key focus areas. The above is preferred embodiments of the present application, but is not intended to limit the scope of protection of the present application. Therefore, Any equivalent changes made according to the structure, shape, and principle of the present application should be included in the protection scope of the present application.
Claims
1. A recommendation method for investigation and classified supervision of soil pollution risk enterprises, characterized by the recommendation method being implemented a computer device, and comprising the following steps:step 1: starting from a full-aperture enterprise list, screening and integrating a list of key industrial enterprises to which need to be paid attention according to a specific rule as an investigation base;step 2: according to availability and convenience of data, gathering multi-source data such as enterprise production characteristics and environmental distribution characteristics to screen enterprise basic attribute indicators with high correlation indicating risks of soil and groundwater pollution in plots, the enterprise basic attribute indicators comprising an enterprise production scale, a production and operation activity time, a three wastes (wastewater, waste gas and solid waste) emission management level, enterprise main product information, a ground anti-seepage level of an enterprise location, soil permeability, a groundwater depth, and a surrounding sensitive point density;step 3: first focusing on the production characteristics of the enterprise in risk investigation, by establishing a knowledge graph, reasoning from known information whether the enterprise is involved in volatile and migratory pollutants, wherein in the reasoning process, a large language model is combined to achieve efficient data identification and batch matching to screen out enterprises involved in volatile and migratory pollutants by calculation;step 4: determining a possibility of the enterprise pollutants entering the soil, based on pollution characteristic statistics of supervised enterprises, integrating characteristic thresholds of various pollution indicators, screening out enterprises that are prone to soil pollution from enterprises involving volatile and migratory pollutants; andstep 5: distinguishing a degree of impact of the migration and spread of the enterprise pollution on the surrounding environment again, setting characteristic thresholds based on hydrogeological conditions of pollution spread and the distribution characteristics of the surrounding environment by combining prior knowledge to screen out enterprises that are prone to pollution spread, enterprises that are highly sensitive in the surrounding areas, and enterprises that are subject to key risk investigation in sequence from enterprises that are prone to foilpollution, and forming a classified supervision list; and identifying key focus areas according to the spatial distribution characteristics of each type of enterprises;step 6: generating a pollution control scheme based on the classified supervision list and the key focus areas, and sending the pollution control scheme to key industrial enterprises.
2. The recommendation method for investigation and classified supervision of soil pollution risk enterprises according to claim 1, characterized in that a specific method of step 1 comprises: obtaining full-aperture yellow page enterprises and industrial and commercial enterprise data in batches, forming a list of key industrial enterprises according to a screening rule of key soil pollution industries, obtaining spatial coordinates of the enterprises based on a map service API interface, and merging duplicate and redundant enterprises; andcomparing an enterprise list that have been included in supervision by a management department, marking supervision statuses of the enterprises, and distinguishing the enterprises that have been included in supervision and the enterprises that have not been included in supervision.
3. The recommendation method for investigation and classified supervision of soil pollution risk enterprises according to claim 1, characterized in that in step 2, the basic attribute indicators of the enterprise are standardized using the following method: for the production scale, performing numerical classification based on the number of employees and annual turnover distribution of the enterprise; for the time of production and operation activities, calculating an interval between registration and closure of the enterprise; for the three wastes discharge management level, calculating the distribution density of nearby enterprises that have been supervised; and for the main product information of the enterprise, performing batched structured recognition by the large language model; andthe enterprise spatial data indicators are standardized using the following methods: superimposing an enterprise distribution layer with a surface impermeable surface layer, a soil layer property layer, a groundwater depth layer, and a sensitive point distribution layer through spatial information coordinate conversion to extract the spatial characteristic data of the corresponding enterprise location, comprising a ground anti-seepage ratio, soil permeability, a groundwater depth, and a surrounding sensitive point density.
4. The recommendation method for investigation and classified supervision of soil pollutionrisk enterprises according to claim 1, characterized in that in step 4, based on reasoning by the knowledge graph, the probability of the relationship between enterprise products and volatile and migratory pollutants is calculated, and when product-related data of the enterprise to be evaluated is inputted, the probability of the enterprise involved in volatile and migratory pollutants is outputted immediately.
5. The recommendation method for investigation and classified supervision of soil pollution risk enterprises according to claim 4, characterized in that a specific method for calculating the probability of the relationship between enterprise products and volatile and migratory pollutants comprises: for a purpose of knowledge reasoning, defining a knowledge graph ontology in combination with enterprise production, pollution emission characteristics and pollutant attributes, and expert industry knowledge of pollutant attributes, including entities, entity attributes and relationships between the entities; designing a triple structure data table to store the above entities, entity attributes and relationships between the entities;corresponding to the entities, the entity attributes, and relationship fields between the entities, extracting entity data and attributes and relationships there of one by one using pollution discharge permit registration data of representative sample enterprises, storing the same according to the triple structure data table, and carrying out data cleaning and integration to achieve mapping to the knowledge graph ontology;analyzing the relationship probability between the entities to assign weights to key node relationships, determining a connection condition probability of each edge in the knowledge graph, and filling this conditional probability into the relationship attributes of the relationship table in the triple table;establishing a visual graph database based on the filled triple structure data table, and using a graph database for storage, wherein when new enterprise triple data is loaded and updated, the entities, the entity attributes, the relationships between the entities, and corresponding entity relationship conditional probabilities in the relevant triple structure data table are also updated;based on the conditional probability between entity relationships in the graph database, calculating the probability of the enterprise to be evaluated involving volatile and migratory pollutants according to a reasoning formula; if the product-related data of the enterprise to be evaluated is semi-structured, calling the large language model to perform batched structuralrecognition of product entities, and performing batched matching and reasoning calculations with the corresponding entities in the existing knowledge graph library; andsetting a probability threshold for a reasoning result and screening out enterprises involved in volatile and migrating pollutants from the key industry enterprises.
6. The recommendation method for investigation and classified supervision of soil pollution risk enterprises according to claim5, characterized in that a comprehensive probability of volatile and migratory pollutants generated by the product is calculated using the following formula:n mi— 1 j—1wherein i is a path number, j is the number of the edge on the path i, and m is the number of edges on the path I; if the product of the enterprise to be evaluated, a production process or raw materials and auxiliary materials information used by the product are available at the same time, when calculating the above comprehensive probability, only an input edge weight of the production process or raw materials and auxiliary materials needs to be set to 1, and input edge weights of the remaining production processes or raw materials and auxiliary materials need to be set to 0.
7. The recommendation method for investigation and classified supervision of soil pollution risk enterprises according to claim 1, characterized in that in step 5, the pollution indicator characteristics of the supervised soil pollution enterprises are statistically analyzed, and the pollution indicator data of the enterprises to be evaluated are screened accordingly, such that each indicator value is not lower than a lower quartile value of the corresponding indicator of the supervised enterprises, and the enterprises that are prone to soil pollution are comprehensively screened out.
8. The recommendation method for investigation and classified supervision of soil pollution risk enterprises according to claim 1, characterized in that:in step 5, the indicator thresholds are set hierarchically according to the expert prior knowledge, and enterprises with underground soil permeability higher than a specific value and groundwater depth lower than a specific value are screened according to the pollution diffusion conditions; and based on the impact of pollution on the surrounding environment, enterprises with a density of sensitive points around locations thereof higher than a certain value arescreened to achieve rapid screening of specific types of enterprises.
Citation Information
Cited By
Soil pollution area division method and system
CN121686370A