A machine learning-based chemical industrial park target ground object identification method

By dividing the chemical industrial park into sub-regions and combining texture and spectral features, a random forest classifier and an active learning optimization model are used to solve the problems of accuracy and wide applicability in chemical industrial park identification, thus achieving efficient identification of chemical industrial parks.

CN119559525BActive Publication Date: 2025-12-30INST OF GEOGRAPHICAL SCI & NATURAL RESOURCE RES CAS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411733725.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-29
Publication Date
2025-12-30
Estimated Expiration
2044-11-29

AI Technical Summary

Technical Problem

Existing technologies lack accurate identification methods for chemical industrial parks, and the acquisition and processing of high-resolution remote sensing images limit the ability to create large-scale, timely, and accurate maps, especially in the context of widespread application within chemical industrial parks.

Method used

By dividing the area to be identified into multiple sub-regions, a classification model for chemical industrial parks is constructed using machine learning algorithms. Combining texture features and spectral features, a random forest classifier is used for identification. Furthermore, the model is optimized through active learning and spatial clustering to adapt to complex environments.

Benefits of technology

It has achieved accurate identification of large-scale chemical industrial parks, improved the generalization and accuracy of the model, reduced intra-class differences, and enhanced the accuracy and adaptability of chemical industrial park identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119559525B_ABST
    Figure CN119559525B_ABST
Patent Text Reader

Abstract

The embodiment of the application discloses a chemical industry park target ground object recognition method based on machine learning, and relates to the technical field of data processing. Wherein, the method comprises: acquiring satellite image data and chemical industry park interest point vector site data of a region to be recognized; dividing the region to be recognized into multiple sub-regions in units of cities; extracting features from the satellite image data of each sub-region, and using the feature data and the interest point vector site data of each sub-region to construct a chemical industry park classification model of each sub-region, wherein each classification model is based on a machine learning algorithm, and the feature data of a grid in each sub-region is used as input, and whether the grid belongs to a chemical industry park is used as output; using the classification model of each sub-region to recognize the feature data in the full range of each sub-region, and obtaining the distribution of the chemical industry park in each sub-region. The embodiment realizes accurate recognition of a chemical industry park in a large range of regions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing, and more particularly to a method for identifying target features in chemical industrial parks based on machine learning. Background Technology

[0002] Currently, deep learning-based accurate feature identification of high-resolution remote sensing imagery is becoming increasingly popular. However, the training process often requires a large number of labeled samples; otherwise, the model can easily overfit to a limited number of training samples, resulting in poor performance when predicting new, unknown datasets. There is a lack of sample datasets covering general features of chemical industrial parks. Most studies use labeled samples of storage tanks for training and identification, which cannot be applied to the entire park in practice. Furthermore, considering the acquisition and processing of high-resolution imagery, the identification scope of chemical industrial parks is often focused on smaller areas such as city centers, limiting the ability to create large-scale, timely, and accurate maps.

[0003] In comparison, machine learning is a more mature application in remote sensing, with increasingly clear and detailed identification of ground features and a wider range of applications, making it an important methodological support for research such as land use monitoring and identification. However, there is still no accurate method for identifying chemical industrial parks, a specific type of man-made landscape. Summary of the Invention

[0004] This invention provides a machine learning-based method for identifying target features in chemical industrial parks to solve the aforementioned technical problems.

[0005] In a first aspect, embodiments of the present invention provide a method for identifying target features in a chemical industrial park based on machine learning, including:

[0006] Acquire satellite imagery data of the area to be identified and vector station data of points of interest in the chemical industrial park;

[0007] The region to be identified is divided into multiple sub-regions, with cities as the unit.

[0008] Features are extracted from satellite imagery data of each sub-region. Using the feature data and point of interest vector station data of each sub-region, a classification model for chemical industrial parks in each sub-region is constructed. Each classification model is based on machine learning algorithms and takes the feature data of the grid in each sub-region as input and whether the grid belongs to a chemical industrial park as output.

[0009] The distribution of chemical industrial parks in each sub-region is obtained by using the classification model of each sub-region to identify the feature data of the entire sub-region.

[0010] In a second aspect, embodiments of the present invention provide an electronic device, the electronic device comprising:

[0011] One or more processors;

[0012] Memory, used to store one or more programs.

[0013] When the one or more programs are executed by the one or more processors, the one or more processors implement the machine learning-based target feature identification method for chemical industrial parks as described in any embodiment.

[0014] Thirdly, embodiments of the present invention also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the machine learning-based target feature identification method for chemical industrial parks as described in any embodiment.

[0015] In summary, this invention provides a machine learning-based method for identifying target features in chemical industrial parks. By dividing the area to be identified into multiple sub-regions to achieve partitioned modeling, intra-class differences are reduced, making the basic framework based on feature data and RF classifier applicable to the entire large area. At the same time, the model parameters reflect the differences between sub-regions, balancing the generalization and accuracy of the model, and achieving accurate identification of chemical industrial parks in a large area. Attached Figure Description

[0016] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0017] Figure 1 This is a flowchart of a machine learning-based target feature identification method for chemical industrial parks, provided by an embodiment of the present invention.

[0018] Figure 2 This is a flowchart of another machine learning-based target feature identification method for chemical industrial parks provided in an embodiment of the present invention;

[0019] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0021] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0022] In the description of this invention, it should also be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0023] Figure 1 This is a flowchart illustrating a machine learning-based method for identifying target features in chemical industrial parks, as provided in an embodiment of the present invention. This method is suitable for identifying chemical industrial parks across a large area spanning multiple cities and is executed by electronic equipment. Figure 1 As shown, the method specifically includes:

[0024] S110. Acquire satellite imagery data of the area to be identified and vector station data of points of interest in the chemical industrial park.

[0025] In this embodiment, the area to be identified is a large area spanning multiple cities, such as a large area along the Yangtze River. This step first acquires satellite imagery data of the area, as well as POI (Points of Interest) vector station data related to the chemical industrial park, as the data source for the entire method.

[0026] Optionally, satellite imagery data can be Sentinel-2 remote sensing imagery data; POI vector station data can be obtained from the Internet using web crawling technology, such as obtaining POI vector station data along the Yangtze River in a certain year, including POI vector station data of CIP (Chemical Industrial Parks) and POI vector station data of non-CIP.

[0027] S120. Divide the region to be identified into multiple sub-regions, taking cities as units.

[0028] Chemical industrial parks in different regions often exhibit significant differences in appearance. Therefore, this embodiment divides the large area into multiple sub-regions, and subsequent modeling of each sub-region will reduce intra-class variations. Specifically, cities are selected as the basic unit for sub-region division.

[0029] S130. Extract features from satellite imagery data of each sub-region.

[0030] After partitioning, specific data features are extracted from satellite imagery data for each sub-region, serving as the basis for identifying chemical industrial parks (CIPs). Specifically, since the factory facilities of CIPs result in texture features in their remote sensing images that are completely different from other types of land cover considered as construction land, this embodiment uses the texture features of the grid as a key feature to improve classification performance. Optionally, a gray-level co-occurrence matrix can be used to calculate the texture features.

[0031] In one specific implementation, taking any sub-region as an example, the window size for calculating texture features is first determined. This window size significantly affects the texture feature extraction effect. It can be determined manually based on the terrain scale and the scale of equipment features within the chemical industrial park, or it can be calculated based on the clustering characteristics of texture features and the minimum footprint of the chemical industrial park. Optionally, the number of grids in the satellite image data can be used as the unit of measurement for the window size. The calculation of the window size can then include the following steps:

[0032] S1-1. Extract multiple positive sample grids belonging to the chemical industrial park and multiple negative sample grids not belonging to the chemical industrial park from the POIs data of the current sub-region. Optionally, multiple preliminary locations of chemical industrial parks and non-chemical industrial parks can be extracted from the POIs data of the current sub-region. Then, preliminary grid labeling is performed based on each preliminary location. From these preliminary locations, multiple typical grids belonging to the chemical industrial park are selected as positive samples, and multiple typical grids belonging to non-chemical industrial parks are selected as negative samples. For example, grids that are close to the preliminary location of a chemical industrial park and fall within the area of ​​that chemical industrial park can be considered as typical grids belonging to the chemical industrial park; similarly, grids that are close to the preliminary location of a non-chemical industrial park and fall within the area of ​​that non-chemical industrial park can be considered as typical grids belonging to non-chemical industrial parks.

[0033] S1-2. Set the initial window size and use it as the current window size. Optionally, the initial window size can be set to one remote sensing data grid.

[0034] S1-3. Extract the texture features of each positive and negative sample grid and cluster the texture features. Specifically, use the K-centroid clustering method, taking K=2, to cluster the texture features of all sample grids into two classes.

[0035] S1-4. Using the positive and negative sample grid sets as the actual clustering results, calculate the Land coefficients of the two clusters. Specifically, the larger the Land coefficient, the more closely the two clustering results match.

[0036] S1-5. If the Rand coefficient is greater than the coefficient threshold, then the current window size is used as the window size determined based on the clustering characteristics of the texture features; otherwise, the current window size is increased by a certain step size, and S1-3 is returned based on the new current window size, until the final Rand coefficient is greater than the coefficient threshold. This embodiment sets a relatively large coefficient threshold for the Rand coefficient. When the Rand coefficient is greater than this threshold, it indicates that the clustering results in S1-3 accurately distinguish between positive and negative samples. That is, the texture features of the current window are large enough to make positive and negative samples show significant differences, thus playing a stronger distinguishing role in the identification of chemical industrial parks.

[0037] In practice, as the window size increases, the Land coefficient tends to rise first and then fall, or rise first and then stabilize. This is because windows that are too large or too small are not conducive to extracting texture features commensurate with the size of the ground features. To avoid the subjectivity of manually selecting the window size or setting the coefficient threshold, a wider range of window size can be pre-selected, allowing the Land coefficient obtained under different window sizes within this range to exhibit a pattern of rising first and then falling, or rising first and then stabilizing. Then, the turning point where the Land coefficient rises to fall or from rise to stabilization is selected as the window size determined based on the clustering characteristics of texture features.

[0038] Finally, after determining the window size based on the clustering characteristics of the texture features, it is then considered whether the window size is smaller than the minimum footprint of the chemical industrial park. If it is smaller than the minimum footprint, or smaller than a set percentage of the minimum footprint, then the window size can be used as the final window size.

[0039] Once the window size is determined, the texture features of each grid are calculated based on that window size. Simultaneously, the spectral features of each grid are extracted, including the normalized difference vegetation index (NDVI), soil-adjusted vegetation index (SAVI), corrected normalized water index (MNDI), and normalized building index (NDI). These spectral and texture features together constitute the feature space of the grid.

[0040] S140. Using the feature data and point of interest vector station data of each sub-region, construct a classification model for the chemical industrial park in each sub-region. Each classification model is based on a machine learning algorithm, and takes the feature data of the grid in each sub-region as input and whether the grid belongs to the chemical industrial park as output.

[0041] This step involves rigorous visual labeling of the grid samples to construct an accurate sample set. This sample set is then used to train the machine learning model in different regions, resulting in classification models for chemical industrial parks in each sub-region.

[0042] In one specific implementation, firstly, based on the preliminary locations of each chemical industrial park and non-chemical industrial park obtained from the aforementioned POIs vector site data, the grids covering each chemical industrial park and non-chemical industrial park are rigorously visually labeled; the feature data and labeling information of each grid constitute a sample set. Optionally, negative samples can cover various land cover types (such as forests, grasslands, farmland, water bodies, and bare land) as well as non-CIPs areas similar to but different from chemical industrial parks (such as steel mills, cement plants, manufacturing plants, and conventional residential areas) to reduce the occurrence of false positives.

[0043] Next, select a classifier. Optionally, a Random Forest (RF) classifier can be chosen as the primary classification tool. Reasons include: RF has multilinear feature modeling capabilities, effectively handling complex relationships in multidimensional feature spaces, making it suitable for feature modeling in chemical industrial parks; RF is robust, exhibiting strong resistance to noise and outliers through the integration of multiple decision trees, better adapting to uncertainties in remote sensing data; RF has simple parameters, and compared to other classifiers (such as support vector machines), RF parameter tuning is relatively simple, requiring only adjustments to the number of decision trees, the maximum number of splits per tree, and the number of features used for training.

[0044] Once the classifier is determined, a Random Forward (RF) classifier can be constructed for any sub-region using sample data within that sub-region. Optionally, the RF classifier can be built using the Google Earth Engine (GEE) API, and the `Classifier.randomForest()` function can be called to train the model. During training, the grid search method can be used to fine-tune different combinations of parameters (including the number of decision trees, the maximum number of splits per tree, and the number of features used for training) to achieve the best classification results.

[0045] Furthermore, in the initial stage, a Radio Frequency (RF) classifier can be trained using labeled training data, followed by predictions on sub-region datasets. Errors and unlabeled data in the predictions will be returned to experts for further annotation. New labeled data will be used to retrain the model, focusing on missed CIPs pixels and false positive predictions, ensuring the model can progressively adapt and correct errors. False positive predictions refer to misclassifying non-CIPs as CIPs; this error is reduced by collecting more representative negative samples. Missed CIPs pixels refer to pixels that were not correctly classified as belonging to the chemical industrial park; these need to be relabeled to improve recognition accuracy.

[0046] In one specific implementation, for the missed CIPs pixels (i.e., the missed positive samples), the positive samples can be re-labeled based on the area and location of the chemical industrial parks within the positive samples. Specifically, due to the window size when calculating texture features, chemical industrial parks with areas smaller than the window size may be missed. In this case, the positive samples of these chemical industrial parks can be re-labeled and the training can be repeated.

[0047] The aforementioned iterative learning process can also be called active learning. If the model's test metrics still fail to meet the standards after multiple active learning iterations, the misclassified grids (referred to as misclassified grids) can be spatially clustered according to their geographical location to explore whether the misclassified grids exhibit regional clustering characteristics in geographic space. Spatial clustering refers to clustering the misclassified grids based on their geographical location. Optionally, there are initially no clusters; during clustering, each misclassified grid can be traversed once or multiple times, with new clusters added continuously as the traversal progresses. Specifically, in each iteration, for each encountered error grid, the following operations are performed: If the error grid does not belong to any existing cluster, for each existing cluster, determine whether the distance between the error grid and the cluster center of the existing cluster is less than a preset threshold. If so, the error grid is assigned to that cluster; otherwise, a new cluster is created with the error grid as the cluster center. If the error grid belongs to an existing cluster, for other clusters besides the one to which the error grid belongs, determine whether the distance between the error grid and the cluster center of its own cluster is greater than the distance between the error grid and the cluster centers of those other clusters. If so, the error grid is assigned to those other clusters. In this way, densely spaced error grids can be automatically clustered into one group, and the number of clusters is not fixed in advance but adaptively formed based entirely on the spatial distribution of the error grids.

[0048] After clustering is complete, clusters with weights greater than a set threshold are selected, and the density of misclassified grids and the density of correctly classified grids within these clusters are calculated. Specifically, the weight of each cluster can be determined based on the number of misclassified grids it contains. Optionally, the weight of a cluster can be the proportion of the number of misclassified grids within that cluster to the total number of misclassified grids in the current sub-region. If the weight of a cluster is greater than a set threshold, it indicates that the misclassified grids exhibit significant spatial clustering within that cluster, and it is necessary to further consider whether to model the area covered by that cluster separately.

[0049] Optionally, this embodiment calculates the density of erroneous grids and the density of correct grids within the cluster as a basis for determining whether to model them separately. Specifically, the density of erroneous grids is obtained by dividing the number of erroneous grids within the cluster by the area of ​​the block covered by the cluster; the density of correct grids is obtained by dividing the number of correct grids within the cluster by the area of ​​the block covered by the cluster. For ease of distinction and description, this embodiment refers to these two densities as the first density and the second density, respectively.

[0050] If the first density is greater than the first threshold and the second density is less than the second threshold (for example, the first threshold is 0.8 and the second threshold is 0.2), that is, the first density is large enough and the second density is small enough, it indicates that the current classification model exhibits significant poor performance in this block and is not suitable for most samples in this block. In this case, it is necessary to separate the block from the current sub-region for modeling.

[0051] At this point, the system continues to check whether the block covered by the cluster is connected to other sub-regions. If the block is connected to other sub-regions outside the current sub-region, it indicates that the block is located on the boundary of a neighboring city and may be influenced by the neighboring city, exhibiting similar chemical industrial park characteristics. In this case, the block can be merged into the other sub-regions, and the merged sub-regions can then be modeled. If the block is not connected to any sub-region outside the current sub-region, it can be modeled as an independent sub-region.

[0052] Furthermore, after modeling the blocks covered by the cluster separately, the test indicators of the existing model can be recalculated for other continuous blocks in the current sub-region other than the cluster. If the new test indicators meet the requirements, the existing model can be used as the chemical industrial park classification model for the other blocks. If the new test indicators do not meet the requirements, the other blocks can also be modeled as an independent sub-region.

[0053] In this way, this embodiment can promptly identify regional classification differences within the same sub-region during the modeling process, separating unsuitable blocks from the current sub-region for modeling; it can also break down barriers between sub-regions, allowing boundary blocks of sub-regions to participate in the overall modeling of adjacent sub-regions; ultimately making both partitioning and modeling more accurate.

[0054] S150. Using the classification model of each sub-region, the feature data of the entire range of each sub-region are identified to obtain the distribution of chemical industrial parks in each sub-region.

[0055] After all modeling is completed, the final sub-regions may differ from those divided in S120. Each final sub-region corresponds to a final classification model. For any given final sub-region, the feature data of all grids within its coverage area are input into the corresponding classification model for processing, which yields the identification result of whether each grid belongs to the chemical industrial park distribution. Optionally, grids belonging to the chemical industrial park are marked as 1, and grids not belonging to the chemical industrial park are marked as 0, thus obtaining the chemical industrial park distribution mask for that sub-region. After performing the above operation on all sub-regions, the chemical industrial park distribution mask for the entire area to be identified can be obtained.

[0056] In summary, this embodiment provides a machine learning-based method for identifying target features in chemical industrial parks, which can achieve the following beneficial effects:

[0057] 1. By adjusting the calculation window of texture features, the texture features of the chemical industrial park can be better captured, so that the feature data can reflect the significant characteristics of the chemical industrial park and achieve more accurate land cover identification.

[0058] 2. By combining active learning with RF classifiers, the identification performance of chemical industrial parks can be gradually improved. This not only enhances the ability to accurately identify chemical industrial parks in complex environments, but also provides a feasible solution for future remote sensing data processing. This process can be continuously iterated to improve classification accuracy.

[0059] 3. By dividing the region to be identified into multiple sub-regions to achieve partitioned modeling, intra-class differences are reduced, making the basic framework based on texture features, spectral features, and RF classifier applicable to the entire large area. At the same time, the differences between sub-regions are reflected through model parameters, taking into account both the generalization and accuracy of the model.

[0060] 4. By using cities as basic partitioning units and leveraging the spatial clustering characteristics of erroneous grids in model building, regional differences within basic partitioning units can be identified in a timely manner. Regions that are clearly unsuitable for the current model can be modeled separately. At the same time, the barriers between sub-regions can be broken down, allowing the boundary blocks of sub-regions to participate in the overall modeling of adjacent sub-regions, ultimately improving the accuracy of partitioning, modeling, and chemical industrial park classification.

[0061] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention, such as... Figure 3 As shown, the device includes a processor 60, a memory 61, an input device 62, and an output device 63; the number of processors 60 in the device can be one or more. Figure 3 Taking a processor 60 as an example; the processor 60, memory 61, input device 62, and output device 63 in the device can be connected via a bus or other means. Figure 3Taking the example of a connection between China and Israel via a bus.

[0062] The memory 61, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the machine learning-based target feature identification method for chemical industrial parks in this embodiment of the invention. The processor 60 executes various functional applications and data processing of the device by running the software programs, instructions, and modules stored in the memory 61, thereby realizing the aforementioned machine learning-based target feature identification method for chemical industrial parks.

[0063] The memory 61 may primarily include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a given function; the data storage area may store data created based on terminal usage. Furthermore, the memory 61 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory, or other non-volatile solid-state storage device. In some instances, the memory 61 may further include memory remotely located relative to the processor 60, which can be connected to the device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0064] Input device 62 can be used to receive input digital or character information, and to generate key signal inputs related to user settings and function control of the device. Output device 63 may include display devices such as a display screen.

[0065] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the machine learning-based target feature identification method for chemical industrial parks according to any embodiment.

[0066] The computer storage medium of this invention can be any combination of one or more computer-readable media. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0067] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.

[0068] Program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.

[0069] Computer program code for performing the operations of this invention can be written in one or more programming languages ​​or a combination thereof. Programming languages ​​include object-oriented programming languages—such as Java, Smalltalk, and C++—as well as conventional procedural programming languages—such as C or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0070] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the technical solutions of the embodiments of the present invention.

Claims

1. A method for identifying target ground objects in a chemical industrial park based on machine learning, characterized in that, The method comprises the following steps: acquiring satellite image data of a to-be-identified region and chemical industry park interest point vector site data; dividing the to-be-identified region into a plurality of sub-regions in units of cities; extracting features from the satellite image data of each sub-region, and respectively constructing a chemical industry park classification model for each sub-region by using the feature data of each sub-region and the interest point vector site data, wherein each classification model is based on a machine learning algorithm, and takes the feature data of a grid in each sub-region as input and takes whether the grid belongs to a chemical industry park as output; identifying the feature data in the full range of each sub-region by using the classification model of each sub-region to obtain the distribution of the chemical industry park in each sub-region; wherein the step of respectively constructing a chemical industry park classification model for each sub-region by using the feature data of each sub-region and the interest point vector site data comprises: constructing a sample set according to the feature data and the interest point vector site data of any sub-region, and training and testing the chemical industry park classification model of the any sub-region by using the sample set; if the test index of the model does not meet the standard, spatially clustering the grids classified incorrectly according to geographical location; selecting a cluster with a weight greater than a set threshold, and calculating a first density of the grids classified incorrectly and a second density of the grids classified correctly in the cluster; if the first density is greater than a first threshold and the second density is less than a second threshold, detecting whether the block covered by the cluster is connected to other sub-regions, wherein the first threshold is greater than the second threshold; if connected, merging the block into the other sub-regions and modeling the merged sub-regions; if not connected, modeling the block as an independent sub-region.

2. The method of claim 1, wherein, after the step of detecting whether the block covered by the cluster is connected to other sub-regions if the first density is greater than a first threshold and the second density is less than a second threshold, the method further comprises: recalculating the test index of the model for other blocks in the any sub-region except the cluster; if the new test index meets the standard, using the existing model as the chemical industry park classification model of the other blocks; if the new test index does not meet the standard, modeling the other blocks as independent sub-regions.

3. The method of claim 1, wherein, the step of constructing a sample set according to the feature data and the interest point vector site data of any sub-region comprises: obtaining the preliminary positions of a plurality of chemical industry parks and non-chemical industry parks according to the interest point vector site data of any sub-region; visually labeling the grids covered by each chemical industry park and non-chemical industry park according to the preliminary positions of each chemical industry park and non-chemical industry park; constructing a sample set from the feature data and labeling information of each grid.

4. The method of claim 1, wherein, the step of extracting features from the satellite image data of each sub-region comprises: determining the window size for calculating texture features in units of grids of the satellite image data; calculating the texture features of each grid according to the window size.

5. The method of claim 4, wherein, the step of determining the window size for calculating texture features in units of grids of the satellite image data comprises: Determine the window size for calculating the texture feature according to the clustering characteristics of the texture feature and the minimum area of the chemical industrial park.

6. The method of claim 5, wherein, The determining the window size for calculating the texture feature according to the clustering characteristics of the texture feature and the minimum area of the chemical industrial park comprises: S1-1, extract a plurality of positive sample grids belonging to the chemical industrial park and a plurality of negative sample grids not belonging to the chemical industrial park from the chemical industrial park interest point vector site data; S1-2, set an initial window size and take the initial window size as a current window size; S1-3, extract the texture features of each positive sample grid and negative sample grid within the current window size, and cluster each texture feature into two clusters; S1-4, take the positive sample grid set and the negative sample grid set as the true clustering result, and calculate the Land coefficient of the two clusters; S1-5, if the Land coefficient is greater than a coefficient threshold, take the current window size as the window size determined according to the clustering characteristics of the texture feature; otherwise, increase the current window size by a certain step, and return to S1-3 according to the new current window size until the final Land coefficient is greater than the coefficient threshold.

7. The method of claim 1, wherein, The method comprises: According to the feature data and the interest point vector site data of each sub-region, construct a chemical industrial park classification model for each sub-region respectively; According to the feature data and the interest point vector site data of any sub-region, construct a sample set, and train and test the chemical industrial park classification model of the any sub-region; Extract the omitted grids which are not correctly classified as chemical industrial parks, and re-label the omitted grids according to the area and position of the chemical industrial parks in the omitted grids; 8. An electronic device, comprising: Continue to train the chemical industrial park classification model by using the re-labeled omitted grids. The method comprises: One or more processors; Memory for storing one or more programs, 9. A computer-readable storage medium, characterized in that, When the one or more programs are executed by the one or more processors, the one or more processors implement the machine learning-based chemical industrial park target ground object recognition method of any one of claims 1-7. A computer program is stored thereon, which is executed by a processor to implement the machine learning-based chemical industrial park target ground object recognition method of any one of claims 1-7.

Citation Information

Patent Citations

  • Fine urban land use type identification model training method

    CN116484266A

  • Urban functional area identification method based on remote sensing-crowd source semantic deep clustering

    CN117274650A