Urban inland inundation ponding point rapid positioning method and device

By using the HanLP deep learning model to extract place names and combined with the optimization treatment of the digital surface model, the problem of inaccurate positioning of urban waterlogging waterlogging points is solved, and higher positioning accuracy is achieved.

CN119988649AInactive Publication Date: 2025-05-13SUN YAT SEN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510129588.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-05
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, place names are prone to ambiguity in extraction, resulting in the extracted place names being inaccurate enough, the location of the location and the actual water accumulation location are large, and the location of urban water accumulation points is not accurate enough.

Method used

The place name extraction model trained based on the HanLP deep learning model is used to extract address information of water accumulation related text information in social media data. The initial water accumulation point is optimized based on the surrounding elevation expansion to obtain the target water accumulation point.

Benefits of technology

By eliminating the ambiguity of place names, improving the accuracy of place names extraction, and combining with the optimization of digital surface models, the positioning accuracy of urban waterlogging waterlogging sites has been significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988649A_ABST
    Figure CN119988649A_ABST
Patent Text Reader

Abstract

The invention discloses a method and a device for quickly positioning an urban waterlogging water accumulation point, which are used for solving the problems that the extracted place name is not accurate enough, the deviation between the positioning place and the actual water accumulation position is large and the positioning of the urban water accumulation point is not accurate enough because the place name extraction is ambiguous in the related technology. The method comprises the steps of obtaining social media data needing waterlogging ponding point positioning, and extracting ponding related text information from the social media data; performing address information extraction on the ponding related text information through a place name extraction model trained based on a HanLP deep learning model to obtain ponding point address information; converting the address information of the waterlogging point into longitude and latitude information, and visualizing the longitude and latitude information as an initial waterlogging waterlogging point; and performing lowest point optimization processing based on surrounding elevation expansion on the initial waterlogging ponding point in combination with the digital earth surface model to obtain a target waterlogging ponding point.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of urban waterlogging monitoring, and in particular to a method for quickly locating waterlogging points in urban waterlogging, a device for quickly locating waterlogging points in urban waterlogging, an electronic device and a storage medium. Background Art

[0002] Urban flooding refers to local or large-scale flooding caused by the inability to drain water in a city in a timely manner due to heavy rainfall, poor drainage systems or other factors. This disaster not only affects the quality of life of residents, but may also cause property losses and casualties, posing a threat to the sustainable development of the city.

[0003] Under the influence of climate change and human activities, urban waterlogging occurs frequently. Urban waterlogging disasters have become a prominent problem that restricts the safe development of cities. Due to many factors, historical waterlogging information in most cities is seriously missing. In addition, most studies on obtaining disaster data from social media focus on public opinion analysis based on social media data and disaster analysis based on check-in data, resulting in low accuracy of urban waterlogging warning and forecasting. In addition, when extracting place names from social media text information, some place names are ambiguous, resulting in inaccurate positioning of waterlogging points. Summary of the invention

[0004] The present invention provides a method for quickly locating waterlogging points in urban waterlogging, a device for quickly locating waterlogging points in urban waterlogging, an electronic device and a storage medium, which are used to solve or partially solve the technical problems in the related technology that ambiguity is prone to occur in place name extraction, resulting in the extracted place names being not accurate enough, and at the same time, the deviation between the located location and the actual waterlogging location is large, resulting in the inaccurate positioning of urban waterlogging points.

[0005] The present invention provides a method for quickly locating waterlogging points in urban areas, the method comprising:

[0006] Acquire social media data for waterlogging locations that require location, and extract waterlogging-related text information from the social media data;

[0007] The address information of the waterlogging point is extracted from the waterlogging-related text information by using a place name extraction model trained based on the HanLP deep learning model;

[0008] Converting the waterlogging point address information into longitude and latitude information, and visualizing the longitude and latitude information as the initial waterlogging point;

[0009] The initial waterlogging point is optimized based on the lowest point of the surrounding elevation extension in combination with the digital surface model to obtain the target waterlogging point.

[0010] Optionally, the place name extraction model includes a word segmentation model and a named entity recognition model; the place name extraction model trained based on the HanLP deep learning model extracts address information from the waterlogging-related text information to obtain the address information of the waterlogging point, including:

[0011] Firstly, the waterlogging-related text information is reversely screened through a keyword database to filter out irrelevant text in the waterlogging-related text information to obtain processed text information;

[0012] Inputting the processed text information into the word segmentation model to perform Chinese part-of-speech word segmentation processing to obtain a word segmentation result;

[0013] Inputting the word segmentation result into the named entity recognition model to perform place name recognition to obtain a place name recognition result;

[0014] When the place name recognition result indicates that the processed text information contains at least one place name, at least one place name information is extracted from the processed text information as the water accumulation point address information.

[0015] Optionally, the process of Chinese part-of-speech segmentation processing includes:

[0016] Performing Chinese word segmentation on the processed text information to convert sentences in the processed text information into multiple Chinese part-of-speech text sequences;

[0017] According to the sentence context information, each of the Chinese part-of-speech text sequences is tagged with a part-of-speech tag to obtain a tagged part-of-speech text sequence corresponding to each of the Chinese part-of-speech text sequences;

[0018] performing a stop word removal process on all the annotated part-of-speech text sequences to select at least one target part-of-speech text sequence from all the annotated part-of-speech text sequences;

[0019] Construct a word vector for each of the target part-of-speech text sequences to convert each of the target part-of-speech text sequences into a feature value to obtain a final word segmentation result.

[0020] Optionally, the waterlogging point address information corresponds to at least one initial waterlogging point; for each of the initial waterlogging points, the initial waterlogging point is optimized based on the lowest point of surrounding elevation extension by combining the digital surface model to obtain a target waterlogging point, including:

[0021] Taking the initial waterlogging point as the center, expand the preset distance values ​​upward, downward, left and right in a point expansion manner to obtain multiple initial waterlogging expansion points;

[0022] Using a digital surface model as a ground elevation model, based on a value extraction to point function, extracting the elevation of each of the initial waterlogging expansion points;

[0023] Automatically delete all points with null elevation values ​​in the initial waterlogging expansion points to obtain multiple valid waterlogging expansion points;

[0024] Determine a target waterlogging extension point corresponding to the lowest elevation value from the multiple effective waterlogging extension points;

[0025] Based on the initial waterlogging point and the target waterlogging expansion point, a target waterlogging point is determined.

[0026] Optionally, the determining a target waterlogging point based on the initial waterlogging point and the target waterlogging expansion point includes:

[0027] Comparing the elevation values ​​of the initial waterlogging point and the target waterlogging expansion point;

[0028] When the elevation value of the initial waterlogging point is greater than the elevation value of the target waterlogging extension point, the target waterlogging extension point is used as the target waterlogging point;

[0029] When the elevation value of the initial waterlogging point is less than the elevation value of the target waterlogging expansion point, the initial waterlogging point is used as the target waterlogging point;

[0030] When the elevation value of the initial waterlogging point is equal to the elevation value of the target waterlogging extension point, the initial waterlogging point and / or the target waterlogging extension point are used as the target waterlogging point.

[0031] Optionally, extracting waterlogging-related text information from the social media data includes:

[0032] Based on the pre-selected keywords, a keyword search is performed on the social media data by a web crawler to obtain crawled information;

[0033] An entity object is created for the crawled information to obtain text information related to water accumulation.

[0034] Optionally, converting the waterlogging point address information into longitude and latitude information, and visualizing the longitude and latitude information as an initial waterlogging point includes:

[0035] At least one place name information in the address information of the waterlogging point is converted into corresponding longitude and latitude information through the map API, and each longitude and latitude information is visualized as an initial waterlogging point in the integrated geographic space platform.

[0036] The present invention also provides a device for quickly locating waterlogging points in urban areas, comprising:

[0037] A waterlogging-related text information extraction module is used to obtain social media data for locating waterlogging points and extract waterlogging-related text information from the social media data;

[0038] An address information extraction module is used to extract address information from the waterlogging-related text information through a place name extraction model trained based on a HanLP deep learning model to obtain address information of waterlogging points;

[0039] An initial waterlogging point generation module is used to convert the waterlogging point address information into longitude and latitude information, and visualize the longitude and latitude information as the initial waterlogging point;

[0040] The lowest point optimization processing module is used to optimize the lowest point of the initial waterlogging point based on the surrounding elevation expansion in combination with the digital surface model to obtain the target waterlogging point.

[0041] The present invention also provides an electronic device, the device comprising a processor and a memory:

[0042] The memory is used to store program code and transmit the program code to the processor;

[0043] The processor is used to execute the method for quickly locating urban waterlogging waterlogging points as described in any one of the above items according to the instructions in the program code.

[0044] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium is used to store program codes, and the program codes are used to execute the method for quickly locating waterlogging points in urban waterlogging as described in any one of the above items.

[0045] It can be seen from the above technical solutions that the present invention has the following advantages:

[0046] A method for quickly locating urban waterlogging waterlogging points is provided. First, social media data for waterlogging waterlogging points that need to be located is obtained, and waterlogging-related text information is extracted from the social media data; then, the address information of waterlogging-related text information is extracted through a place name extraction model trained based on a HanLP deep learning model to obtain the address information of the waterlogging point, so that the HanLP deep learning model can be used to identify place names according to semantic named entities during place name extraction, accurately identify complex place names in the text, eliminate most place name ambiguities, improve the accuracy of place name extraction, and further improve the accuracy of urban waterlogging point location; then, the waterlogging point address information is converted into longitude and latitude information, and the longitude and latitude information is visualized as the initial waterlogging waterlogging point; finally, the initial waterlogging waterlogging point is optimized based on the lowest point of the surrounding elevation extension in combination with a digital surface model to obtain the target waterlogging waterlogging point, so that the extracted information is optimized based on the lowest point of the surrounding elevation extension in combination with the digital surface model terrain data, so that the waterlogging point can be spatially located and the positioning accuracy can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0048] Figure 1 A flowchart of a method for quickly locating waterlogging points in urban areas;

[0049] Figure 2 A schematic diagram of web page location and corresponding XPath acquisition in a web crawler;

[0050] Figure 3 This is a schematic diagram of an interface for optimizing the lowest point of the initial waterlogging point based on the surrounding elevation expansion;

[0051] Figure 4 It is a schematic diagram of the overall process of a method for quickly locating waterlogging points in urban areas;

[0052] Figure 5 The present invention is a structural block diagram of a device for quickly locating waterlogging points in cities. DETAILED DESCRIPTION

[0053] The embodiments of the present invention provide a method for quickly locating waterlogging points in urban waterlogging, a device for quickly locating waterlogging points in urban waterlogging, an electronic device and a storage medium, which are used to solve or partially solve the technical problems in the related technology that ambiguity is prone to occur in place name extraction, resulting in the extracted place names being not accurate enough, and at the same time, the deviation between the located location and the actual waterlogging location is large, resulting in the inaccurate positioning of urban waterlogging points.

[0054] In order to make the purpose, features and advantages of the present invention more obvious and easy to understand, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described below are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0055] As an example, under the influence of climate change and human activities, urban flooding occurs frequently. Urban flooding disasters have become a prominent problem that restricts the safe development of cities. At present, the use of social media to obtain information has begun to be used in disaster research. However, the uncertainty of data sources will lead to problems such as inaccurate and incomplete data. Methods based on deep learning can process data more efficiently, but the model still has the problem of difficulty in training complex events.

[0056] On the one hand, foreign scholars started research on using data crowdsourcing (social media) to obtain disaster information earlier and focused on real-time detection and trend prediction of disaster events. Although China started late, it has also achieved some results in recent years. For example, based on the weighted model of user activity at the check-in point, disaster mapping is performed on the processed data, which is close to the actual situation. At present, there are few studies on social media data in urban waterlogging disasters. Most of them focus on public opinion analysis based on social media data, disaster analysis based on check-in data, and the use of machine image intelligent recognition technology to identify urban waterlogging. Less attention is paid to the collation of historical waterlogging data and the extraction of accurate place names from social media data.

[0057] In terms of disaster information extraction, extracting place name information from social media, such as microblog (hereinafter referred to as Weibo) text or pictures, belongs to Chinese named entity recognition. Compared with English named entity recognition, it is more difficult to achieve due to problems such as word segmentation ambiguity and polysemy. At present, there are four main methods and combinations of Chinese named entity recognition research: dictionary-based, rule-based, machine learning and deep learning. With the introduction of recurrent neural networks and long short-term memory networks (LSTM), deep learning-based methods can establish connections between previous and subsequent data and capture dependencies over longer distances. This makes it more efficient in processing sequence data, making it gradually become the mainstream method for named entity recognition. For example, compared with traditional machine learning methods, a convolutional neural network (CNN) model is used to classify upcoming events for situational awareness. However, studies have shown that during a single event setting, the convolutional neural network model used can learn each label well. However, since multiple events have closely related labels, the convolutional neural network model has difficulty learning for complex event settings.

[0058] Therefore, in practical applications, place name extraction is prone to ambiguity and the extracted information is not accurate enough. For example, in the text, "XX Road YY North Road Intersection" will be extracted into two place names, "XX Road" and "YY North Road Intersection", and "AA Avenue BB Section" will be extracted into two invalid place names, "AA Avenue" and "BB Road".

[0059] In addition, when using the common map API (MAP Application Programming Interface, an interface used to integrate map functions into applications) to extract the longitude and latitude of place names, due to the wide range covered by a place name (such as a road or a national highway), it is easy for the positioning to deviate greatly from the actual waterlogging location, and the positioning of urban waterlogging points is not accurate enough.

[0060] Therefore, one of the core invention points of the embodiment of the present invention is: in view of the shortcomings of the current technology, a method for quickly locating urban waterlogging points is provided. First, the historical waterlogging text information of data crowdsourcing (such as Weibo crowdsourcing) is obtained through web crawler technology, and then the HanLP (Hankcs Advanced Natural Language Processing, an open source natural language processing (NLP) toolkit for Chinese processing) deep learning model is used to identify place names according to semantic named entities when extracting place names, which can accurately identify complex place names in the text, eliminate most place name ambiguities, and improve the accuracy of place name extraction, thereby improving the accuracy of urban waterlogging point positioning. Then, the extracted information is optimized based on the lowest point of the surrounding elevation expansion in combination with the digital surface model (DSM) terrain data to achieve spatial positioning of waterlogging points and improve positioning accuracy. Therefore, by using the digital surface model data, according to the principle that water flows to low places, the terrain low point is combined with the data crowdsourcing waterlogging text information, and then the urban waterlogging point is spatially positioned, which can directly improve the accuracy of urban waterlogging point positioning. At the same time, the data acquisition difficulty of digital surface models is low, and the terrain conditions can be directly obtained through analysis, which makes the data acquisition requirements for improving accuracy lower.

[0061] Reference Figure 1 , shows a flowchart of a method for quickly locating a waterlogging point in an urban area provided by an embodiment of the present invention, which may specifically include the following steps:

[0062] Step 101, obtaining social media data for locating waterlogging points, and extracting waterlogging-related text information from the social media data;

[0063] This step mainly realizes data collection. In practical applications, a web crawler is a script or program that can automatically capture web page content and related information. With the development of crawler technology, web crawler technology based on Python (a high-level programming language) has become more mature. Among them, the more typical web crawler screening technologies include regular expressions, XPath (XML Path Language, XML (eXtensible Markup Language) path language), BeautifulSoup (a python library for parsing HTML (HyperText Markup Language) and XML documents), etc.

[0064] In the embodiment of the present invention, Python web crawler technology based on XPath path language is adopted. By using keyword search method, according to the selected waterlogging keyword, social media data such as the text information of Weibo, text information in pictures (if equipped with pictures), posting time, etc. are crawled through the parsing of web page data to obtain waterlogging related information. By creating entity objects for subsequent operations, waterlogging related text information for subsequent processing is obtained. Exemplarily, web page positioning and corresponding XPath acquisition in web crawler are shown as follows: Figure 2 shown.

[0065] In a specific implementation, waterlogging-related text information can be extracted from social media data by: first, based on pre-selected waterlogging keywords, keyword search is performed on social media data through a web crawler to obtain crawled information; then, entity objects are created for the crawled information to obtain waterlogging-related text information.

[0066] The text data (crowdsourcing information) used for locating waterlogging points can be obtained not only from the Weibo channel used in the above example, but also from other social media that people use daily. It is understandable that the present invention is not limited to this.

[0067] Step 102, extracting address information from the waterlogging-related text information through a place name extraction model trained based on the HanLP deep learning model to obtain address information of the waterlogging point;

[0068] HanLP provides a wealth of functions, ranging from basic word segmentation and part-of-speech tagging to advanced tasks such as named entity recognition and dependency syntax analysis. In the embodiment of the present invention, address information is extracted from the acquired waterlogging-related text information based on the HanLP deep learning model.

[0069] Address information extraction (also known as place name information extraction) belongs to Chinese named entity recognition. Compared with English named entity recognition, it is more difficult to implement due to problems such as word segmentation ambiguity and polysemy. The purpose of named entity recognition is to determine the boundaries of entities in the text and divide entity names into different types, such as place names. In an embodiment of the present invention, after the text information has been processed by reverse screening, Chinese word segmentation, part-of-speech tagging and other data processing, the processed text information obtained has the preliminary conditions for named entity recognition.

[0070] Traditional methods will not make part-of-speech tags relevant, and may output part-of-speech tags that do not conform to natural language grammar rules, such as multiple verbs appearing consecutively. However, the algorithm in the HanLP model of this solution can constrain its output tags and obtain an optimal labeling sequence of adjacent tags.

[0071] In addition, when using the current technology, for example, "XX Road YY North Road Intersection" in the text will be extracted into two place names "XX Road" and "YY North Road Intersection", and "AA Avenue BB Road Section" will be extracted into two invalid place names "AA Avenue" and "BB Road", etc. The embodiment of the present invention can eliminate most similar place name ambiguity problems by using the HanLP model for Chinese part-of-speech segmentation and named entity recognition, thereby improving the word segmentation ambiguity problem of the current technology.

[0072] Before the model is officially put into use, model construction and training steps are required. In the embodiment of the present invention, a deep learning framework based on HanLP is adopted, MSRA (Microsoft Research Asia, an entity corpus in the field of natural language processing) is used as training data and entity annotation is completed, and a word segmentation model and a named entity recognition model are trained as the actual place name extraction model.

[0073] In the subsequent positioning process, the HanLP algorithm can be used in the word segmentation model to complete the word segmentation in a pipeline manner, and the word segmentation results can be used as the model input of the named entity recognition model to complete the named entity recognition. If the microblog text contains a place name, the extracted place name is returned. If there is no place name, a null value is returned.

[0074] Combined with the above discussion, in a specific implementation, the address information of waterlogging-related text information is extracted through a place name extraction model trained based on the HanLP deep learning model, and the process of obtaining the address information of the waterlogging point can mainly include the following sub-steps S01 to S04:

[0075] Step S01: reversely screening the waterlogging-related text information through a keyword database to filter out irrelevant text in the waterlogging-related text information to obtain processed text information;

[0076] Reverse screening refers to using a keyword database to filter out irrelevant texts such as "hydrocephalus" and "hydronephrosis".

[0077] Step S02: input the processed text information into the word segmentation model to perform Chinese part-of-speech word segmentation processing to obtain the word segmentation result;

[0078] Furthermore, the process of Chinese part-of-speech segmentation processing may include the following sub-steps S021 to S024:

[0079] Step S021: performing Chinese word segmentation on the processed text information to convert sentences in the processed text information into multiple Chinese part-of-speech text sequences;

[0080] Chinese word segmentation refers to converting a Chinese sentence into a word sequence. Specifically, the coarse segmentation mode in the HanLP model can be used to convert sentences in the text into word sequences.

[0081] Step S022: According to the sentence context information, perform part-of-speech tagging on each Chinese part-of-speech text sequence to obtain the tagged part-of-speech text sequence corresponding to each Chinese part-of-speech text sequence;

[0082] Part-of-speech tagging refers to giving each word in a sentence a correct part-of-speech tag according to the sentence context information. For example, a word indicating an action is tagged as a verb, and a word for description is tagged as an adjective, etc.

[0083] Step S023: Perform stop-word removal processing on all tagged part-of-speech text sequences to screen out at least one target part-of-speech text sequence from all tagged part-of-speech text sequences;

[0084] Stop-word removal refers to removing words that often appear in natural language but do not necessarily represent the actual semantics of the sentence. Such as "de", "di", "de", etc.

[0085] Step S024: Construct word vectors for each target part-of-speech text sequence to convert each target part-of-speech text sequence into feature values and obtain the final word segmentation result.

[0086] Word vector construction refers to converting the words in the text into feature values that can be calculated by a computer for the next step of named entity recognition.

[0087] Step S03: Input the word segmentation result into the named entity recognition model for place name recognition to obtain the place name recognition result;

[0088] Step S04: When the place name recognition result indicates that the processed text information contains at least one place name, extract at least one place name information from the processed text information as the waterlogging point address information.

[0089] It can be understood that there may be more than one location where urban waterlogging occurs in the same text information. Then, after information extraction, when the place name recognition result indicates that the processed text information contains at least one place name, extract at least one place name information from the processed text information as the waterlogging point address information.

[0090] It should be pointed out that in addition to using HanLP for Chinese word segmentation and place name recognition, other suitable deep learning models can also be used to improve the accuracy of place name recognition. For example, for Chinese word segmentation, the precise word segmentation mode of the python jieba (a Chinese word segmentation library) package can be used to convert sentences into word sequences. For place name recognition, the BiLSTM-CRF (Bidirectional Long Short-Term Memory-Conditional Random Field) algorithm can be used to implement named entity recognition, etc. It is understandable that the present invention is not limited to this.

[0091] Step 103, converting the waterlogging point address information into longitude and latitude information, and visualizing the longitude and latitude information as the initial waterlogging point;

[0092] Specifically, the address information of the waterlogging point is converted into longitude and latitude information, and the longitude and latitude information is visualized as the initial waterlogging point. This can be done by: using the currently commonly available map API, converting at least one place name information (street structured address) in the address information of the waterlogging point into corresponding longitude and latitude information, and visualizing each longitude and latitude information as the initial waterlogging point in the comprehensive geographic spatial platform.

[0093] Step 104 , optimizing the initial waterlogging point based on the lowest point of surrounding elevation extension in combination with the digital surface model to obtain the target waterlogging point.

[0094] In this step, the initial waterlogging points are optimized in combination with the Digital Surface Model (DSM) to obtain accurate urban waterlogging points.

[0095] Among them, the digital surface model can be regarded as a ground elevation model that includes the height information of surface buildings, bridges, trees, etc. The digital surface model not only includes the elevation information of the ground, but also includes the height information of all objects on the ground, such as buildings, vegetation, bridges, etc.

[0096] Compared with the digital surface model, the digital elevation model (DEM) commonly used in current positioning only contains the elevation information of the terrain, but not other surface information. It can be understood that the digital surface model is based on the digital elevation model and further covers the elevation of other surface information (including buildings, forests and other objects) besides the ground.

[0097] Combined with the above content, it can be known that the waterlogging point address information can correspond to at least one initial waterlogging point. For each initial waterlogging point, the initial waterlogging point is optimized based on the surrounding elevation extension in combination with the digital surface model to obtain the target waterlogging point. This process can be achieved by executing the following sub-steps S11 to S15:

[0098] Step S11: Taking the initial waterlogging point as the center, expand the preset distance values ​​upward, downward, left and right in a point expansion manner to obtain multiple initial waterlogging expansion points;

[0099] In actual situations, rainwater tends to gather in low-lying areas of terrain after heavy rains, forming urban waterlogging. Based on the previous steps, the longitude and latitude information is visualized as waterlogging points by using a comprehensive geographic spatial platform. However, the longitude and latitude converted by the map API are likely to have certain errors. Therefore, the lowest point in the square area with a side length of n meters centered on the longitude and latitude point of the initial waterlogging point can be found as the real waterlogging point. Among them, the values ​​of n are compared to 200m, 500m, 1000m, and 1600m. When n is selected as 500m, it can include an area large enough to find the low point, and the intersection is small.

[0100] Therefore, the first step of the lowest point optimization based on the surrounding elevation expansion is to expand the waterlogging point. Specifically, the python script is used to expand the initial waterlogging point upward, downward, left, and right by 250 meters. The expansion step is automatically adjusted according to the maximum number (corresponding to the maximum number of rows in Excel is 65536 rows).

[0101] Step S12: using the digital surface model as the ground elevation model, based on the value extraction to point function, extracting the elevation of each initial waterlogging expansion point;

[0102] Then read the point elevation. Using the digital surface model as the ground elevation model, use the value extraction to point function in the integrated geospatial platform, and use the terrain data of the digital surface model to extract the elevation of each initial waterlogging expansion point.

[0103] Step S13: automatically deleting all points with null elevations among the initial waterlogging extension points, and obtaining multiple valid waterlogging extension points;

[0104] Step S14: determining a target waterlogging extension point corresponding to the lowest elevation value from a plurality of effective waterlogging extension points;

[0105] For example, Figure 3 A schematic diagram of an interface for optimizing the lowest point of an initial waterlogging point based on surrounding elevation expansion is shown.

[0106] The green points correspond to the initial waterlogging points. The black points correspond to the initial waterlogging expansion points. The red points correspond to the lowest points within 500m determined by the terrain data of the digital surface model.

[0107] Step S15: Determine the target waterlogging point based on the initial waterlogging point and the target waterlogging expansion point.

[0108] Furthermore, based on the initial waterlogging point and the target waterlogging expansion point, the target waterlogging point can be determined by: comparing the elevation values ​​of the initial waterlogging point and the target waterlogging expansion point; and determining the final target waterlogging point based on the elevation value comparison result.

[0109] As a case, when the elevation value of the initial waterlogging point is greater than the elevation value of the target waterlogging extension point, the target waterlogging extension point is used as the target waterlogging point.

[0110] As another case, when the elevation value of the initial waterlogging point is less than the elevation value of the target waterlogging expansion point, the initial waterlogging point is used as the target waterlogging point.

[0111] Although the possibility is small, it is possible that the elevation values ​​of the initial waterlogging point and the target waterlogging extension point are exactly equal. When the elevation value of the initial waterlogging point is equal to the elevation value of the target waterlogging extension point, the initial waterlogging point and / or the target waterlogging extension point are used as the target waterlogging point. That is, one or both points can be selected according to the actual situation, or both points can be used as the target waterlogging point.

[0112] Generally speaking, the number of target waterlogging points in the city's administrative area that is finally obtained corresponds to the number of initial waterlogging points. However, in actual situations, the extracted initial waterlogging points may not be within the city's administrative area. At this time, for the city's administrative area, the initial waterlogging point is a null point (i.e., there is no elevation data), and it is necessary to delete the null point before optimizing the positions of other initial waterlogging points. Therefore, it can be concluded that the number of target waterlogging points in the city's administrative area is the number of initial waterlogging points minus the number of null points. When the number of null points is 0, the number of target waterlogging points is equal to the number of initial waterlogging points.

[0113] In an embodiment of the present invention, a method for quickly locating urban waterlogging points is provided. First, the historical waterlogging text information of data crowdsourcing (such as Weibo crowdsourcing) is obtained through web crawler technology, and then the HanLP deep learning model is used to identify the place names according to the semantic named entities when extracting place names, which can accurately identify the complex place names in the text, eliminate most of the ambiguity of place names, and improve the accuracy of place name extraction, thereby improving the accuracy of locating urban waterlogging points. Then, the extracted information is optimized based on the lowest point of the surrounding elevation expansion in combination with the digital surface model terrain data to achieve spatial positioning of the waterlogging point and improve the positioning accuracy. Therefore, by using the digital surface model data, according to the principle that water flows to the lower place, the terrain low point is combined with the data crowdsourcing waterlogging text information, and then the urban waterlogging point is spatially located, which can directly improve the accuracy of locating the urban waterlogging point. At the same time, the digital surface model data is easy to obtain, and the terrain conditions can be directly obtained through analysis, so that the data acquisition requirements for improving the accuracy are relatively low.

[0114] For better explanation, refer to Figure 4 , showing an overall flow diagram of a method for quickly locating waterlogging points in urban waterlogging provided by an embodiment of the present invention. It should be noted that this embodiment only briefly describes the general process of quickly locating waterlogging points in urban waterlogging. The specific implementation process of each step can be understood by referring to the relevant content in the aforementioned embodiment. It will not be repeated here. It can be understood that the present invention does not limit this.

[0115] Step 401: obtaining social media data for locating waterlogging points, extracting waterlogging-related text information from the social media data based on a web crawler, and performing reverse screening on the waterlogging-related text information through a keyword database to obtain processed text information;

[0116] Step 402: input the processed text information into a word segmentation model trained based on the HanLP deep learning model to perform Chinese part-of-speech word segmentation processing, and input the obtained word segmentation result into a named entity recognition model trained based on the HanLP deep learning model to perform place name recognition, thereby obtaining a place name recognition result;

[0117] Step 403: when the place name recognition result indicates that the processed text information contains at least one place name, extract at least one place name information from the processed text information as the water accumulation point address information;

[0118] Step 404: convert at least one place name information in the address information of the waterlogging point into corresponding longitude and latitude information through the map API, and visualize each longitude and latitude information as an initial waterlogging point in the integrated geospatial platform;

[0119] Step 405: for each initial waterlogging point, taking the initial waterlogging point as the center, expand the preset distance values ​​upward, downward, left and right in a point expansion manner to obtain multiple initial waterlogging expansion points;

[0120] Step 406: using the digital surface model as the ground elevation model, based on the value extraction to point function, extracting the elevation of each initial waterlogging expansion point;

[0121] Step 407: automatically deleting the points with null elevation values ​​in all the initial waterlogging extension points, obtaining multiple valid waterlogging extension points, and determining the target waterlogging extension point corresponding to the lowest elevation value from the multiple valid waterlogging extension points;

[0122] Step 408: Based on the initial waterlogging point and the target waterlogging expansion point, determine the target waterlogging point, and visualize the target waterlogging expansion point on the integrated geographic space platform.

[0123] Reference Figure 5 , shows a structural block diagram of a device for quickly locating waterlogging points in urban areas provided by an embodiment of the present invention, which may specifically include:

[0124] The waterlogging-related text information extraction module 501 is used to obtain social media data for locating waterlogging points and extract waterlogging-related text information from the social media data;

[0125] An address information extraction module 502 is used to extract address information from the waterlogging-related text information through a place name extraction model trained based on a HanLP deep learning model to obtain address information of waterlogging points;

[0126] An initial waterlogging point generating module 503 is used to convert the waterlogging point address information into longitude and latitude information, and visualize the longitude and latitude information as the initial waterlogging point;

[0127] The lowest point optimization processing module 504 is used to optimize the lowest point of the initial waterlogging point based on the surrounding elevation extension in combination with the digital surface model to obtain the target waterlogging point.

[0128] In an optional embodiment, the place name extraction model includes a word segmentation model and a named entity recognition model; the address information extraction module 502 includes:

[0129] A reverse screening module, used to reversely screen the waterlogging-related text information through a keyword database to filter out irrelevant text in the waterlogging-related text information to obtain processed text information;

[0130] A Chinese part-of-speech word segmentation processing module is used to input the processed text information into the word segmentation model to perform Chinese part-of-speech word segmentation processing to obtain a word segmentation result;

[0131] A place name recognition module, used for inputting the word segmentation result into the named entity recognition model to perform place name recognition and obtain a place name recognition result;

[0132] The place name information extraction module is used to extract at least one place name information from the processed text information as the water accumulation point address information when the place name recognition result indicates that the processed text information contains at least one place name.

[0133] In an optional embodiment, the Chinese part-of-speech word segmentation processing module includes:

[0134] A Chinese word segmentation module, used for performing Chinese word segmentation on the processed text information to convert sentences in the processed text information into multiple Chinese part-of-speech text sequences;

[0135] A part-of-speech tagging module is used to perform part-of-speech tagging on each of the Chinese part-of-speech text sequences according to the sentence context information, and obtain a tagged part-of-speech text sequence corresponding to each of the Chinese part-of-speech text sequences;

[0136] A stop word removal processing module is used to remove stop words from all the annotated part-of-speech text sequences, so as to select at least one target part-of-speech text sequence from all the annotated part-of-speech text sequences;

[0137] The word vector construction module is used to construct the word vector of each target part-of-speech text sequence to convert each target part-of-speech text sequence into a feature value to obtain the final word segmentation result.

[0138] In an optional embodiment, the waterlogging point address information corresponds to at least one initial waterlogging point; the lowest point optimization processing module 504 includes:

[0139] A waterlogging point expansion module is used to expand the preset distance values ​​upward, downward, left and right in a point expansion manner with the initial waterlogging point as the center to obtain multiple initial waterlogging expansion points;

[0140] An extension point elevation extraction module is used to extract the elevation of each of the initial waterlogging extension points based on a value extraction to point function by using a digital surface model as a ground elevation model;

[0141] A null value deletion module is used to automatically delete all points with null elevation values ​​in the initial waterlogging expansion points to obtain multiple valid waterlogging expansion points;

[0142] A target waterlogging extension point determination module is used to determine a target waterlogging extension point corresponding to a minimum elevation value from the plurality of valid waterlogging extension points;

[0143] The target waterlogging point determination module is used to determine the target waterlogging point based on the initial waterlogging point and the target waterlogging expansion point.

[0144] In an optional embodiment, the target waterlogging point determination module is specifically used to:

[0145] Comparing the elevation values ​​of the initial waterlogging point and the target waterlogging expansion point;

[0146] When the elevation value of the initial waterlogging point is greater than the elevation value of the target waterlogging extension point, the target waterlogging extension point is used as the target waterlogging point;

[0147] When the elevation value of the initial waterlogging point is less than the elevation value of the target waterlogging expansion point, the initial waterlogging point is used as the target waterlogging point;

[0148] When the elevation value of the initial waterlogging point is equal to the elevation value of the target waterlogging extension point, the initial waterlogging point and / or the target waterlogging extension point are used as the target waterlogging point.

[0149] In an optional embodiment, the waterlogging-related text information extraction module 501 includes:

[0150] An information crawling module, used to perform keyword search on the social media data through a web crawler based on pre-selected keywords, to obtain crawled information;

[0151] The entity object creation module is used to create an entity object for the crawled information and obtain text information related to water accumulation.

[0152] In an optional embodiment, the initial waterlogging point generating module 503 is specifically used for:

[0153] At least one place name information in the address information of the waterlogging point is converted into corresponding longitude and latitude information through the map API, and each longitude and latitude information is visualized as an initial waterlogging point in the integrated geographic space platform.

[0154] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the aforementioned method embodiment.

[0155] An embodiment of the present invention further provides an electronic device, the device comprising a processor and a memory:

[0156] The memory is used to store the program code and transmit the program code to the processor;

[0157] The processor is used to execute the method for quickly locating urban waterlogging waterlogging points according to any embodiment of the present invention according to the instructions in the program code.

[0158] An embodiment of the present invention further provides a computer-readable storage medium, which is used to store program codes, and the program codes are used to execute the method for quickly locating waterlogging points in urban waterlogging according to any embodiment of the present invention.

[0159] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0160] In the several embodiments provided by the present invention, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0161] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0162] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0163] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc., various media that can store program codes.

[0164] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for quickly locating urban waterlogging waterlogging points, characterized in that: include: Acquire social media data for waterlogging locations that require location, and extract waterlogging-related text information from the social media data; The address information of the waterlogging point is extracted from the waterlogging-related text information by using a place name extraction model trained based on the HanLP deep learning model; Converting the waterlogging point address information into longitude and latitude information, and visualizing the longitude and latitude information as the initial waterlogging point; The initial waterlogging point is optimized based on the lowest point of the surrounding elevation extension in combination with the digital surface model to obtain the target waterlogging point.

2. The method for quickly locating urban waterlogging points according to claim 1 is characterized in that: The place name extraction model includes a word segmentation model and a named entity recognition model; the place name extraction model trained based on the HanLP deep learning model extracts address information from the waterlogging related text information to obtain the address information of the waterlogging point, including: Firstly, the waterlogging-related text information is reversely screened through a keyword database to filter out irrelevant text in the waterlogging-related text information to obtain processed text information; Inputting the processed text information into the word segmentation model to perform Chinese part-of-speech word segmentation processing to obtain a word segmentation result; Inputting the word segmentation result into the named entity recognition model to perform place name recognition to obtain a place name recognition result; When the place name recognition result indicates that the processed text information contains at least one place name, at least one place name information is extracted from the processed text information as the water accumulation point address information.

3. The method for quickly locating urban waterlogging points according to claim 2 is characterized in that: The process of Chinese part-of-speech segmentation processing includes: Performing Chinese word segmentation on the processed text information to convert sentences in the processed text information into multiple Chinese part-of-speech text sequences; According to the sentence context information, each of the Chinese part-of-speech text sequences is tagged with a part-of-speech tag to obtain a tagged part-of-speech text sequence corresponding to each of the Chinese part-of-speech text sequences; performing a stop word removal process on all the annotated part-of-speech text sequences to select at least one target part-of-speech text sequence from all the annotated part-of-speech text sequences; Construct a word vector for each of the target part-of-speech text sequences to convert each of the target part-of-speech text sequences into a feature value to obtain a final word segmentation result.

4. The method for quickly locating waterlogging points in urban waterlogging according to claim 2 is characterized in that: The waterlogging point address information corresponds to at least one initial waterlogging point; for each of the initial waterlogging points, the initial waterlogging point is optimized based on the surrounding elevation extension by combining the digital surface model to obtain the target waterlogging point, including: Taking the initial waterlogging point as the center, expand the preset distance values ​​upward, downward, left and right in a point expansion manner to obtain multiple initial waterlogging expansion points; Using a digital surface model as a ground elevation model, based on a value extraction to point function, extracting the elevation of each of the initial waterlogging expansion points; Automatically delete all points with null elevation values ​​in the initial waterlogging expansion points to obtain multiple valid waterlogging expansion points; Determine a target waterlogging extension point corresponding to the lowest elevation value from the multiple effective waterlogging extension points; Based on the initial waterlogging point and the target waterlogging expansion point, a target waterlogging point is determined.

5. The method for quickly locating urban waterlogging points according to claim 4 is characterized in that: The step of determining a target waterlogging point based on the initial waterlogging point and the target waterlogging expansion point includes: Comparing the elevation values ​​of the initial waterlogging point and the target waterlogging expansion point; When the elevation value of the initial waterlogging point is greater than the elevation value of the target waterlogging extension point, the target waterlogging extension point is used as the target waterlogging point; When the elevation value of the initial waterlogging point is less than the elevation value of the target waterlogging expansion point, the initial waterlogging point is used as the target waterlogging point; When the elevation value of the initial waterlogging point is equal to the elevation value of the target waterlogging extension point, the initial waterlogging point and / or the target waterlogging extension point are used as the target waterlogging point.

6. The method for quickly locating urban waterlogging waterlogging points according to any one of claims 2 to 5, characterized in that: The step of extracting waterlogging related text information from the social media data includes: Based on the pre-selected keywords, a keyword search is performed on the social media data by a web crawler to obtain crawled information; An entity object is created for the crawled information to obtain text information related to water accumulation.

7. The method for quickly locating urban waterlogging points according to claim 6, characterized in that: The step of converting the waterlogging point address information into longitude and latitude information, and visualizing the longitude and latitude information as an initial waterlogging point includes: At least one place name information in the address information of the waterlogging point is converted into corresponding longitude and latitude information through the map API, and each longitude and latitude information is visualized as an initial waterlogging point in the integrated geographic space platform.

8. A device for quickly locating waterlogging points in urban areas, characterized in that: include: A waterlogging-related text information extraction module is used to obtain social media data for locating waterlogging points and extract waterlogging-related text information from the social media data; An address information extraction module is used to extract address information from the waterlogging-related text information through a place name extraction model trained based on a HanLP deep learning model to obtain address information of waterlogging points; An initial waterlogging point generation module is used to convert the waterlogging point address information into longitude and latitude information, and visualize the longitude and latitude information as the initial waterlogging point; The lowest point optimization processing module is used to optimize the lowest point of the initial waterlogging point based on the surrounding elevation expansion in combination with the digital surface model to obtain the target waterlogging point.

9. An electronic device, characterized in that: The device comprises a processor and a memory: The memory is used to store program codes and transmit the program codes to the processor; The processor is used to execute the method for quickly locating urban waterlogging waterlogging points according to any one of claims 1-7 according to the instructions in the program code.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to store program codes, and the program codes are used to execute the method for quickly locating urban waterlogging waterlogging points as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Water accumulation point intelligent identification method and device, electronic equipment and storage medium

    CN119271810A