Enterprise nitrogen and phosphorus pollution discharge flux calculation method based on multi-source heterogeneous data fusion
Through multi-source heterogeneous data fusion and machine learning methods, enterprise attributes and spatial data are integrated, and the inconsistency and lack of enterprise nitrogen and phosphorus emission flux data are solved, efficient and accurate pollution emission flux calculation is achieved, and reliable technical support is provided for environmental supervision.
Patent Information
- Application Number
- CN202510880986.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-27
AI Technical Summary
At present, the acquisition and integration of nitrogen and phosphorus discharge flux of enterprises' pollutant discharge data faces the problems of dispersion, heterogeneity, inconsistency and lack of data sources, which affects the accuracy and completeness of data, and lacks effective data fusion and calculation methods.
Multi-source heterogeneous data fusion method is adopted, combined with map API, general large language model and batch scanning tools, and integrated enterprise attribute information, spatial location and remote sensing image data, and used random forest algorithms to complete the missing data to calculate the enterprise's nitrogen and phosphorus pollution emission flux.
It significantly improves data coverage and calculation dimensions, improves data processing efficiency and accuracy, realizes high-precision calculation of enterprise nitrogen and phosphorus discharge flux, and provides scientific support for environmental supervision.
Smart Images

Figure CN120373675A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of environmental data processing and pollution monitoring, and particularly relates to a method for calculating the nitrogen and phosphorus sewage discharge flux of enterprises based on the fusion of multi-source heterogeneous data. Background Art
[0002] Currently, the acquisition and integration of nitrogen and phosphorus sewage discharge flux data of enterprises still face technical challenges: on the one hand, the data sources are scattered and heterogeneous. Enterprise information is obtained from the industrial and commercial registration system and enterprise credit database and stored in tabular form. While the pollution emission data mainly comes from the sewage discharge permit system and is mostly stored in the form of pictures. At the same time, the geographical location information depends on the map API for acquisition, and spatial data such as remote sensing images and land use data are scattered on different data platforms and stored in raster form. These data have different formats and problems such as missing and inconsistent data, which affect the accuracy and integrity of the nitrogen and phosphorus sewage discharge flux data of enterprises. On the other hand, there are missing pollution emission data for some enterprises. Due to the imperfect data collection mechanism, it is impossible to accurately evaluate the pollution emission situation of enterprises. With the development of large language model technology, it performs excellently in image understanding and text recognition, especially suitable for processing unstructured image data such as watermarked and complex background images, breaking through the limitations of traditional optical character recognition in heterogeneous data processing, and providing support for the standardization and intelligent analysis of pollution emission data.
[0003] Therefore, how to effectively fuse multi-source data, utilize the picture recognition ability of large language models to process heterogeneous data, construct a complete and reliable enterprise nitrogen and phosphorus sewage discharge flux data set, and at the same time use advanced data filling and calculation methods to improve data quality has become an urgent problem to be solved in the current field of environmental data processing. Summary of the Invention
[0004] To solve the above problems, the present invention proposes a method for calculating the nitrogen and phosphorus sewage discharge flux of enterprises based on the fusion of multi-source heterogeneous data. The method aims to fuse multi-source heterogeneous data, construct an enterprise nitrogen and phosphorus sewage discharge flux data set, and use the random forest algorithm to complement missing data and predict emission situations, improving data integrity and accuracy, and providing scientific support for environmental supervision.
[0005] To achieve the above object, the present invention adopts the following technical solutions: A method for calculating the nitrogen and phosphorus sewage discharge flux of enterprises based on the fusion of multi-source heterogeneous data, comprising the following steps: S1. Collect the multi-source attribute information of enterprises and conduct unified classification according to industry standards; S2. Call multiple map APIs, obtain and standardize the longitude and latitude coordinates of enterprises, and realize the fusion of attribute data and spatial location information; S3. Integrate multi-source spatial data, match it with enterprise location information, and extract regional environmental characteristics. Among them, the multi-source spatial data includes remote sensing image data, land use map data, and administrative division data. The specific process of step S3 is as follows: S31. Obtain the publicly available building roof data, uniformly convert the original coordinate system to the WGS84 coordinate system, so that the building roof data and the enterprise longitude and latitude coordinate data adopt the same spatial reference system. S32. Load the converted building roof data and enterprise vector point data in the ArcMap software, and use the spatial overlay analysis tool for spatial matching processing. S33. Use the function of extracting raster values at points to extract the roof raster value information corresponding to each enterprise vector point. S34. Conduct discriminant analysis on the extraction results to identify the enterprise vector point data that can accurately fall on the building roof. S35. Eliminate the enterprise vector point data that fails to accurately fall on the building roof to exclude positioning errors caused by spatial coordinate conversion errors or data offsets. S4. Use the general large language model and batch scanning tool to identify and structurally extract heterogeneous format data respectively, and fuse, compare, and correct the extraction results to obtain enterprise nitrogen and phosphorus emission permit data. S5. Integrate the attribute information, spatial characteristics of the enterprise, and enterprise nitrogen and phosphorus emission permit data, use machine learning methods to complete the missing information, calculate the enterprise nitrogen and phosphorus pollution discharge flux, and perform rasterization processing. The calculation formula is: Input the enterprise data to be measured, and define as follows: , where is the enterprise data; X1 is the establishment date; X2 is the registered capital; X3 is the enterprise category; S K is the one-hot encoded variable of the industry scale; Use the trained random forest regression model for prediction. The calculation formula is: , where is the predicted value of the final nitrogen and phosphorus pollutant emissions; M is the number of decision trees in the random forest regression model; m is the decision tree number; is the predicted output of the m-th decision tree; Correspond the total nitrogen, total phosphorus, and ammonia nitrogen of the three nitrogen and phosphorus pollution discharge fluxes to the raster.
[0006] Preferably, the attribute information in step S1 includes company name, registration status, unified social credit code, legal representative, available phone number, registered address, affiliated district or county, establishment date, approval date, industry classification, national standard industry category, national standard industry major category, national standard industry medium category, enterprise scale, registered capital, paid-in capital, business term, affiliated province, affiliated city, company type, former name, taxpayer identification number, registration number, organization code, number of insured persons, annual report to which the number of insured persons belongs, address of the latest annual report, communication address, website, email, and business scope.
[0007] Preferably, the map APIs in step S2 include Baidu Map API and Amap API; the specific process of obtaining and standardizing the enterprise longitude and latitude coordinates in step S2 is as follows: S21. Filter non-numeric or out-of-range coordinate data; S22. Use the offset correction method to convert the coordinates obtained by the map API into Mars coordinates; S23. Convert the Mars coordinates into WGS84 coordinates through a geographic transformation algorithm to obtain enterprise vector point data.
[0008] Preferably, in step S4, the specific process of using a general large language model to identify and structurally extract heterogeneous format data is as follows: S41A. Identify the image file and screen the content containing the target keywords; S42A. Set the root directory path where the image to be processed is located and specify the supported image formats; the image formats include PNG, JPG, JPEG, BMP, WEBP, TIF, and TIFF; S43A. Set the output file path for saving the recognition result text and define the keyword set for screening; the keywords include total nitrogen, total phosphorus, and ammonia nitrogen; S44A. Traverse all subfolders and image files under the root directory. For each image file that meets the format requirements, read the image content and convert it into the Base64 encoding form, and construct a standard data URL format string; S45A. Call the language model interface with image recognition capabilities and submit a request containing image data and recognition instructions; the recognition instructions are used to extract the water quality index information in the table in the image and limit to only retain the relevant row data containing the keywords; S46A. After obtaining the recognition result, parse the returned text content line by line, screen out the lines containing the target keywords, and if there are lines that meet the conditions, write them together with the image path information into the output file; if there is no matching content, record it as not recognizing the relevant lines; S47A. Set a fixed time interval between image recognition tasks until a task completion prompt is output after all image processing is completed, and save all recognition results.
[0009] Preferably, in step S4, the specific process of using a batch scanning tool to identify and structurally extract heterogeneous format data is as follows: S41B. Data organization and automatic traversal: Set the root directory for storing pictures to be processed, and the system automatically traverses all subfolders under this directory; S42B. Optical character recognition text recognition and preprocessing: Initialize the PaddleOCR tool library, and enable the angle classification function to optimize the recognition of tilted text; Process eligible images one by one, and call the optical character recognition model to perform text recognition; The recognition results include text content, text box coordinates, and confidence levels; S43B. Anomaly detection and error handling: If the recognition result of optical character recognition is empty or no text is detected, automatically record the image path and mark that no text is detected; S44B. Result storage and output: Store the recognition results of optical character recognition in a unified format. The recognition results include the image path, text content, text box coordinates, and confidence levels; The results are saved to a specified output text file, and different image recognition results are separated by delimiters; And use the append write mode to ensure that the processed results will not be lost if the program is interrupted during long-term operation; S45B. Parse the recognized text data: Read the text file storing the recognition results, and parse the picture path, text content, text coordinates, and confidence levels in the text file, and then store the parsed data in a structured manner; S46B. Intelligent grouping by ordinate: Use the picture path as the primary key to split the data so that the data of different pictures will not be mixed; Calculate the average ordinate value of the text box, and set the ordinate tolerance threshold to determine whether the text belongs to the same line; Through the proximity analysis method, automatically determine whether the recognized text belongs to the same visual line, and merge the text with a distance less than the threshold into the same group; S47B. Generate structured output: In the grouped data, retain the picture path of each group of text for subsequent traceability; Unify the storage method so that each group contains the text content, confidence score, and corresponding picture path of the text in the line where it is located, and generate a clear structured data table.
[0010] Preferably, in step S5, the spatial feature is enterprise plot positioning; The nitrogen and phosphorus pollution discharge fluxes include total nitrogen pollution discharge flux, total phosphorus pollution discharge flux, and ammonia nitrogen pollution discharge flux.
[0011] Preferably, the training steps of the random forest regression model in step S5 are as follows: S51. Data processing: The incorporation date is converted into the duration of the company's existence as of the scheduled date; The industry scale variable is converted into an industry scale coefficient through a mapping function to enhance the impact of categorical variables on the prediction results. The mapping function is defined as: , where is the industry scale coefficient; is the mapping function; C is the industry scale variable; f is the pollution discharge coefficient corresponding to the industry in the Second National Pollution Source Census Bulletin; One-hot encoding transformation is performed on the industry scale variable. The transformation formula is: OHE(S) =
S1, S2, ……, S K
[0012] After adopting the above technical solution, the present invention has the following beneficial effects: 1. The present invention adopts multi-source data deep fusion. By integrating multi-source heterogeneous data such as enterprise attribute data, spatial location information (latitude and longitude coordinates), remote sensing images, land use maps, administrative divisions, etc., it breaks through the limitations of a single data source, significantly improves data coverage and calculation dimensions, and provides comprehensive and multi-dimensional data support for sewage discharge flux analysis.
[0013] 2. The present invention adopts standardization and automation processing. It classifies enterprise attribute information uniformly based on industry standards and uses the map API to achieve the standardization processing of spatial data, effectively solving the problem of incompatible heterogeneous data formats, reducing manual intervention errors, and improving the efficiency and consistency of data processing.
[0014] 3. The present invention adopts intelligent data extraction and correction. By combining a general large language model (LLM) with a batch scanning tool, it automatically identifies, structurally extracts, and fuses and compares unstructured enterprise emission permit data, greatly improving the accuracy of data extraction, and correcting data deviations through a cross-validation mechanism to ensure the reliability of emission data.
[0015] 4. The present invention adopts dynamic spatial feature association. By matching enterprise location information with multi-source spatial data (such as remote sensing images, land use types), it accurately extracts regional environmental features, realizes the dynamic association analysis of sewage discharge flux and geographical environment, and enhances the spatio-temporal representativeness of calculation results.
[0016] 5. The present invention adopts machine learning optimization and completion. It uses machine learning methods to fuse multi-dimensional data and intelligently complete missing information, reduces the estimation deviation caused by incomplete data, improves the integrity and accuracy of sewage discharge flux calculation, and is especially suitable for data-sparse scenarios. In addition, the present invention can perform visualization and decision support. By rasterizing the sewage discharge flux results, it generates a visually intuitive spatial distribution visualization output, which is convenient for environmental supervision departments to quickly locate key sewage discharge areas and provides a scientific basis for differential control and precise pollution treatment.
[0017] In summary, through multi-source data fusion, intelligent processing, and machine learning optimization, the present invention realizes the high-precision and efficient calculation of the nitrogen and phosphorus sewage discharge flux of enterprises, providing a reliable technical means for environmental pollution supervision and treatment. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is the flow chart of the present invention; DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below in conjunction with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0020] As Figure 1 shown, a method for calculating the nitrogen and phosphorus sewage discharge flux of enterprises based on multi-source heterogeneous data fusion includes the following steps: S1. Collect the multi-source attribute information of enterprises and classify it uniformly according to industry standards; The attribute information described in step S1 includes company name, registration status, unified social credit code, legal representative, available telephone, registered address, district or county of affiliation, establishment date, approval date, industry classification, national standard industry category, national standard industry major category, national standard industry medium category, enterprise scale, registered capital, paid-in capital, business term, province of affiliation, city of affiliation, company type, former name, taxpayer identification number, registration number, organization code, number of insured persons, annual report to which the number of insured persons belongs, address of the latest annual report, communication address, website, email, and business scope; In step S1, classification is carried out in accordance with the principles of matching the primary industry with land distribution and land use data, subdividing the secondary industry, and merging the tertiary industry by combining the national standard industry category, national standard industry major category, and national standard industry medium category. The specific classification criteria are as follows: Agriculture: Enterprises engaged in the planting and production of crops such as grain, vegetables, and fruits.
[0021] Forestry: Enterprises engaged in forest cultivation, timber harvesting, forest product processing, and forest resource management.
[0022] Animal Husbandry: Enterprises engaged in the raising of livestock and poultry and the production and processing of livestock products.
[0023] Fishery: Enterprises engaged in aquaculture, fishing, and aquatic product processing.
[0024] Mining: Enterprises engaged in the extraction of coal, metal ores, non-metal ores, and other mineral resources.
[0025] Food Manufacturing: Enterprises that produce various foods, beverages, tobacco products, etc. for human consumption.
[0026] Textile Industry: Enterprises engaged in spinning, weaving, dyeing and finishing, clothing, and textile product production.
[0027] Chemical Products Manufacturing: Enterprises that produce chemical raw materials, chemical products, pharmaceuticals, chemical fibers, etc.
[0028] Metal Manufacturing: Enterprises engaged in the smelting, rolling processing of ferrous and non-ferrous metals, and the production of metal products.
[0029] All types of equipment manufacturing industries: enterprises that produce mechanical equipment, special equipment, general equipment, and transportation equipment.
[0030] Electronic manufacturing industry: enterprises that manufacture electronic products such as electronic components, electronic devices, communication equipment, and computers.
[0031] Other manufacturing industries: covering manufacturing industries not classified into other categories, such as furniture, toys, stationery production enterprises, etc.
[0032] Electric power, heat, gas, and water production and supply industry: enterprises that provide production, supply, and related services of electric power, heat, gas, and water.
[0033] Construction industry: enterprises engaged in the construction, installation, repair, and decoration of various construction projects.
[0034] Tertiary industry: covering wholesale and retail, transportation, warehousing, postal services, accommodation and catering, finance, education, medical care, culture, entertainment, and all other service industries.
[0035] S2. Call multiple map APIs, obtain and standardize the longitude and latitude coordinates of enterprises, and realize the integration of attribute data and spatial location information; The map APIs mentioned in step S2 include Baidu Map API and Gaode Map API; the specific process of obtaining and standardizing the longitude and latitude coordinates of enterprises in step S2 is as follows: S21. Filter non-numeric or out-of-range coordinate data; S22. Use the offset correction method to convert the coordinates obtained by the map API into Mars coordinates; S23. Convert the Mars coordinates into WGS84 coordinates through a geographic transformation algorithm to obtain enterprise vector point data; S3. Integrate multi-source spatial data, match it with the enterprise location information, and extract regional environmental characteristics; among them, the multi-source spatial data includes remote sensing image data, land use map data, and administrative division data; The specific process of step S3 is as follows: S31. Obtain public building roof data, uniformly convert the original coordinate system into the WGS84 coordinate system, so that the building roof data and the enterprise longitude and latitude coordinate data adopt the same spatial reference system; among them, the building roof data is a dataset of building roof distributions with a resolution of 2.5 for each year from 2016 to 2021, generated from Sentinel-2 images from 2016 to 2021 using super-resolution technology; S32. Load the converted building roof data and enterprise vector point data in ArcMap software, and use the spatial overlay analysis tool for spatial matching processing; S33. Use the function of extracting raster values at points to extract the roof raster value information corresponding to each enterprise vector point; S34. Conduct discriminant analysis on the extraction results to identify the enterprise vector point data that can accurately fall on the building roof; S35. Eliminate the enterprise vector point data that fails to accurately fall on the building roof to exclude positioning errors caused by spatial coordinate conversion errors or data offsets; S4. Use a general large language model and a batch scanning tool to respectively identify and structurally extract heterogeneous format data, and fuse and compare the extraction results and correct the data to obtain enterprise nitrogen and phosphorus emission permit data; In step S4, the specific process of using a general large language model to identify and structurally extract heterogeneous format data is as follows: S41A. Identify the image file and screen the content containing the target keywords; S42A. Set the root directory path where the image to be processed is located and specify the supported image formats; the image formats include PNG, JPG, JPEG, BMP, WEBP, TIF, and TIFF; S43A. Set the output file path for saving the recognition result text and define the keyword set for screening; the keywords include total nitrogen, total phosphorus, and ammonia nitrogen; S44A. Traverse all subfolders and image files under the root directory. For each image file that meets the format requirements, read the image content and convert it into the Base64 encoding form to construct a standard data URL format string; S45A. Call the language model interface with image recognition capabilities and submit a request containing image data and recognition instructions; the recognition instructions are used to extract the water quality index information in the table in the image and limit to only retain the relevant row data containing the keywords; S46A. After obtaining the recognition result, parse the returned text content line by line, screen out the lines containing the target keywords, and if there are lines that meet the conditions, write them into the output file together with the image path information; if there is no matching content, record that no relevant lines are recognized; S47A. Set a fixed time interval between image recognition tasks until all image processing is completed, output a task completion prompt, and save all recognition results; In step S4, the specific process of using a batch scanning tool to identify and structurally extract heterogeneous format data is as follows: S41B. Data organization and automatic traversal: Set the root directory where the pictures to be processed are stored, and the system automatically traverses all subfolders under this directory; S42B. Optical Character Recognition Text Recognition and Preprocessing: Initialize the PaddleOCR toolkit and enable the angle classification function to optimize the recognition of skewed text; process eligible images one by one, and call the optical character recognition model for text recognition; the recognition results include text content, text box coordinates, and confidence levels. S43B. Anomaly Detection and Error Handling: If the recognition result of optical character recognition is empty or no text is detected, automatically record the image path and mark that no text is detected. S44B. Result Storage and Output: Store the recognition results of optical character recognition in a unified format. The recognition results include the image path, text content, text box coordinates, and confidence levels; save the results to a specified output text file, and use a delimiter to distinguish the recognition results of different images; and use the append write mode to ensure that the processed results will not be lost if the program is interrupted during long-term operation. S45B. Parse Recognized Text Data: Read the text file storing the recognition results, parse the image path, text content, text coordinates, and confidence levels in the text file, and then store the parsed data in a structured manner. S46B. Intelligent Grouping by Ordinate: Use the image path as the primary key to split the data so that the data of different images will not be mixed; calculate the average ordinate value of the text box and set the ordinate tolerance threshold to determine whether the text belongs to the same line; through the proximity analysis method, automatically determine whether the recognized text belongs to the same visual line, and merge the text with a distance less than the threshold into the same group. S47B. Generate Structured Output: In the grouped data, retain the image path of each group of text for subsequent traceability; unify the storage method so that each group contains the text content, confidence score of the text in the line, and the corresponding image path, and generate a clear structured data table. S5. Integrate the enterprise's attribute information, spatial features, and enterprise nitrogen and phosphorus emission permit data, use machine learning methods to complete missing information, calculate the enterprise's nitrogen and phosphorus pollution discharge flux, and perform rasterization processing. The calculation formula is: Input the enterprise data to be measured and define as follows: , where is the enterprise data; X1 is the establishment date; X2 is the registered capital; X3 is the enterprise category; S K is the one-hot encoded variable of the industry scale; Use the trained random forest regression model for prediction. The calculation formula is: , where is the predicted value of the final nitrogen and phosphorus pollutant emissions; M is the number of decision trees in the random forest regression model; m is the decision tree number; is the predicted output of the m-th decision tree; Correspond the total nitrogen, total phosphorus, and ammonia nitrogen sewage discharge fluxes to the grid; In step S5, the spatial feature is the enterprise plot location; the nitrogen and phosphorus sewage discharge fluxes include the total nitrogen sewage discharge flux, the total phosphorus sewage discharge flux, and the ammonia nitrogen sewage discharge flux; The training steps of the random forest regression model described in step S5 are as follows: S51. Data processing: Convert the establishment date to the duration of the company's existence up to a predetermined date; The industry scale variable is converted into an industry scale coefficient through a mapping function to enhance the influence of the categorical variable on the prediction result. The mapping function is defined as: , where is the industry scale coefficient; is the mapping function; C is the industry scale variable; f is the sewage discharge coefficient corresponding to the industry in the Second National Pollution Source Census Bulletin; Perform one-hot encoding conversion on the industry scale variable. The conversion formula is: OHE(S)=
S1,S2,……,S K
[0036] As described above, the above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A method for calculating the nitrogen and phosphorus sewage discharge flux of enterprises based on multi-source heterogeneous data fusion, characterized in that It includes the following steps: S1. Collect the multi-source attribute information of enterprises and conduct unified classification according to industry standards; S2. Call multiple map APIs, obtain and standardize the longitude and latitude coordinates of enterprises, and realize the integration of attribute data and spatial location information; S3. Integrate multi-source spatial data, match it with the enterprise location information, and extract regional environmental characteristics; among them, the multi-source spatial data includes remote sensing image data, land use map data and administrative division data; The specific process of step S3 is: S31. Obtain the publicly available building roof data, uniformly convert the original coordinate system to the WGS84 coordinate system, so that the building roof data and the enterprise longitude and latitude coordinate data adopt the same spatial reference system; S32. Load the converted building roof data and enterprise vector point data in the ArcMap software, and use the spatial overlay analysis tool for spatial matching processing; S33. Adopt the function of extracting raster values from points to extract the roof raster value information corresponding to each enterprise vector point; S34. Conduct discriminant analysis on the extraction results to identify the enterprise vector point data that can accurately fall on the building roof; S35. Eliminate the enterprise vector point data that fails to accurately fall on the building roof to exclude positioning errors caused by spatial coordinate conversion errors or data offsets; S4. Use the general large language model and batch scanning tool to identify and structurally extract heterogeneous format data respectively, and fuse, compare and correct the extraction results to obtain the enterprise nitrogen and phosphorus emission permit data; S5. Integrate the attribute information, spatial characteristics and enterprise nitrogen and phosphorus emission permit data of the enterprise, use machine learning methods to complete the missing information, calculate the enterprise nitrogen and phosphorus pollution discharge flux and conduct rasterization processing. The calculation formula is: Input the enterprise data to be measured, defined as follows: , where is the enterprise data; X1 is the establishment date; X2 is the registered capital; X3 is the enterprise category; S K is the one-hot encoding variable of the industry scale; Use the trained random forest regression model for prediction, and the calculation formula is: , where is the predicted value of the final nitrogen and phosphorus pollutant emissions; M is the number of decision trees in the random forest regression model; m is the decision tree number; is the predicted output of the m-th decision tree; Correspond the total nitrogen, total phosphorus and ammonia nitrogen of the three nitrogen and phosphorus pollution discharge fluxes to the raster.
2. The method for calculating the enterprise nitrogen and phosphorus sewage discharge flux based on multi-source heterogeneous data fusion according to claim 1, wherein: The attribute information described in step S1 includes company name, registration status, unified social credit code, legal representative, available telephone, registered address, affiliated district or county, establishment date, approval date, industry classification, national standard industry category, national standard industry major category, national standard industry medium category, enterprise scale, registered capital, paid-in capital, business term, affiliated province, affiliated city, company type, former name, taxpayer identification number, registration number, organization code, number of insured persons, annual report to which the number of insured persons belongs, latest annual report address, communication address, website, email and business scope.
3. The method for calculating the enterprise nitrogen and phosphorus sewage discharge flux based on multi-source heterogeneous data fusion according to claim 1, characterized in that, The map APIs described in step S2 include Baidu Map API and Amap API; the specific process of obtaining and standardizing the longitude and latitude coordinates of enterprises in step S2 is: S21. Filter non-numerical or out-of-range coordinate data; S22. Use the offset correction method to convert the coordinates obtained by the map API to Mars coordinates; S23. Convert the Mars coordinates to WGS84 coordinates through a geographic transformation algorithm to obtain enterprise vector point data.
4. The enterprise nitrogen and phosphorus sewage flux calculation method based on multi-source heterogeneous data fusion according to claim 1, characterized in that In step S4, the specific process of using the general large language model to identify and structurally extract heterogeneous format data is: S41A. Identify the image file and screen the content containing the target keywords; S42A. Set the root directory path where the image to be processed is located and specify the supported image formats. The image formats include PNG, JPG, JPEG, BMP, WEBP, TIF, and TIFF. S43A. Set the output file path to save the recognition result text and define the keyword set for filtering. The keywords include total nitrogen, total phosphorus, and ammonia nitrogen. S44A. Traverse all subfolders and image files under the root directory. For each image file that meets the format requirements, read the image content and convert it into the Base64 encoding form, and construct a standard data URL format string. S45A. Call the language model interface with image recognition capabilities and submit a request containing image data and recognition instructions. The recognition instructions are used to extract the water quality index information in the table in the image and limit to retain only the relevant row data containing the keywords. S46A. After obtaining the recognition result, parse the returned text content line by line, filter out the lines containing the target keywords, and if there are qualified lines, write them together with the image path information into the output file. If there is no matching content, record it as not recognizing the relevant lines. S47A. Set a fixed time interval between image recognition tasks until all image processing is completed, then output a task completion prompt and save all recognition results.
5. The enterprise nitrogen and phosphorus sewage discharge flux calculation method based on multi-source heterogeneous data fusion according to claim 1, wherein, In step S4, the specific process of using the batch scanning tool to recognize and structurally extract heterogeneous format data is as follows: S41B. Data organization and automatic traversal: Set the root directory where the pictures to be processed are stored, and the system automatically traverses all subfolders under this directory. S42B. Optical character recognition text recognition and preprocessing: Initialize the PaddleOCR tool library and enable the angle classification function to optimize the recognition of inclined text. Process each qualified image one by one, and call the optical character recognition model to perform text recognition. The recognition result includes the text content, text box coordinates, and confidence level. S43B. Anomaly detection and error handling: If the recognition result of the optical character recognition is empty or no text is detected, automatically record the image path and mark that no text is detected. S44B. Result storage and output: Store the recognition results of the optical character recognition in a unified format. The recognition results include the image path, text content, text box coordinates, and confidence level. The results are saved to the specified output text file, and the recognition results of different images are separated by delimiters. And adopt the append write mode to ensure that if the program is interrupted during long-term operation, the processed results will not be lost. S45B. Parse the recognized text data: Read the text file that saves the recognition results, and parse the picture path, text content, text coordinates, and confidence level in the text file, and then store the parsed data in a structured manner. S46B. Perform intelligent grouping by the vertical coordinate: Use the picture path as the primary key to split the data so that the data of different pictures will not be mixed. Calculate the average vertical coordinate value of the text box and set the vertical coordinate tolerance threshold to determine whether the text belongs to the same line. Using the proximity analysis method, automatically determine whether the recognized text belongs to the same visual line, and merge the text with a distance less than the threshold into the same group; S47B. Generate a structured output: In the grouped data, retain the image path of each group of text for subsequent traceability; unify the storage method so that each group contains the text content, confidence score, and corresponding image path of the text in the line where it is located, and generate a clear structured data table.
6. The method for calculating the enterprise nitrogen and phosphorus sewage discharge flux based on multi-source heterogeneous data fusion according to claim 1, wherein, In step S5, the spatial feature is the enterprise plot location; the nitrogen and phosphorus pollution discharge fluxes include the total nitrogen pollution discharge flux, the total phosphorus pollution discharge flux, and the ammonia nitrogen pollution discharge flux.
7. The method for calculating the enterprise nitrogen and phosphorus sewage discharge flux based on multi-source heterogeneous data fusion according to claim 1, characterized in that The training steps of the random forest regression model in step S5 are as follows: S51. Data processing: Convert the establishment date into the duration of the company's existence as of the predetermined date; The industry scale variable is converted into an industry scale coefficient through a mapping function, which is used to enhance the impact of the categorical variable on the prediction result. The mapping function is defined as: , where is the industry scale coefficient; is the mapping function; C is the industry scale variable; f is the pollutant discharge coefficient corresponding to the industry in the Second National Pollution Source Census Bulletin; Perform one-hot encoding transformation on the industry scale variable, and the transformation formula is: OHE(S)=【S1,S2,……,S K 】, , , where is the one-hot encoding transformation result; S1,S2,……,S K are the one-hot encoding variables of the first k industry scales; S i =1 indicates that the enterprise belongs to the i-th category of industry scale; S i =0 indicates that the enterprise does not belong to the corresponding category; i is the enterprise number; k is the specific enterprise scale, including large, medium, small, and micro; Standardize the input data to make the order of magnitude of different feature values consistent. The calculation formula is as follows: , where x is the input data; μ is the mean of the enterprise numerical features; σ is the standard deviation of the enterprise numerical features; Divide the standardized data into a training set and a test set; S52. Use the training set data to train a random forest regression model, input 、 and the enterprise's attribute information to calculate the predicted value of nitrogen and phosphorus pollutant emissions; S53. Calculate the predicted mean squared error using the test set data. The calculation formula is as follows: , where MSE is the predicted mean squared error; N is the number of training samples; j is the sample number; y j is the true value of nitrogen and phosphorus pollutants; is the predicted value of nitrogen and phosphorus pollutants; S54. Splitting rule of decision tree: During the construction of a decision tree, at each decision node, a partial subset of input features is randomly selected as the candidate splitting feature to find the optimal feature x j and the corresponding threshold t to minimize the error of the left and right subtrees. The optimization objective is defined as: , where is the total number of samples in the data set of the current node; and are the number of samples included in the left and right child nodes respectively after splitting according to the optimal ; and are the mean squared errors of the samples in the left and right child nodes respectively.
Citation Information
Patent Citations
Steel industry carbon emission monitoring method based on electric power data driving
CN117689078A
Regional energy carbon emission rapid accounting method based on space-time correlation
CN119006013A
Flue gas sulfur dioxide concentration prediction method based on variational mode decomposition and ensemble learning
CN119357910A
Carbon emission fusion model modeling method and system based on multi-dimensional geographic data
CN119988820A
Automatic urban land identification system integrating business big data with building form
US20210217117A1