Flood warning method, system, device and product based on multi-source social perception and swarm intelligence

CN122528751APending Publication Date: 2026-08-07CHINA THREE GORGES UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA THREE GORGES UNIV
Filing Date
2026-06-09
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0007]本申请的目的是提供一种基于多源社会感知与群体智能的洪水预警方法、系统、设备及产品,如何解决难以满足高精度、全覆盖的洪水智能化快速预警的问题

Benefits of technology

[0020]本申请提供了一种基于多源社会感知与群体智能的洪水预警方法及系统,通过采集并预处理社交媒体数据与手机基于位置的服务(Location-Based Services,LBS)数据,构建时空对齐的多源融合数据集,为模型模拟与风险评估提供高质量统一数据基础,解决多源数据融合不充分的问题,同时突破传统监测站点稀疏限制,实现流域全域无盲区覆盖;利用自然语言处理(Natural Language Processing,NLP)和空间聚类算法生成标准化洪水事件信息包,准确识别洪水事件并划分危险等级,精确定位高风险聚集区域;结合土壤与水评估工具(Soil and Water Assessment Tool,SWAT),即水文模型模拟与智能算法,提高洪水模拟精度;通过构建三维风险指标体系,利用组合权重和阈值设定,解决风险评估模糊问题,有效突出风险等级,提供精准预警信息;根据风险等级、用户位置生成差异化预警信息和推送渠道的不同选择,执行分级分类精准快速预警,有效解决预警一刀切以及响应效率低的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122528751A_ABST
    Figure CN122528751A_ABST
Patent Text Reader

Abstract

The application discloses a flood warning method and system based on multi-source social perception and swarm intelligence, and relates to the technical field of water informatics, social perception computing and swarm intelligence cross. The method solves the problem of insufficient multi-source data fusion by constructing a multi-source fusion dataset aligned in time and space, breaks through the limitation of traditional sparse monitoring sites, and realizes full-coverage of the basin without blind area. The method generates a standardized flood event information package by using NLP and spatial clustering algorithm, accurately identifies flood events and divides the danger level, and accurately locates the high-risk aggregation area. The method improves the flood simulation accuracy by combining the hydrological model SWAT simulation and intelligent algorithm. The method solves the fuzzy problem of risk assessment by constructing a three-dimensional risk index system, using combined weights and threshold settings, and providing accurate early warning information. According to the risk level, user location and push channel, the method performs hierarchical classification and accurate and rapid early warning, effectively solving the problem of one-size-fits-all early warning and low response efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the interdisciplinary fields of hydroinformatics, social sensing computing, and swarm intelligence, and in particular to a flood early warning method, system, device, and product based on multi-source social sensing and swarm intelligence. Background Technology

[0002] Floods are characterized by their suddenness, destructive power, wide impact, and long chains of secondary disasters. Once they occur, they can easily cause casualties, collapsed houses, damaged infrastructure, crop failures, and significant economic losses, seriously threatening the safety of people's lives and property and regional social stability. In the research of flood early warning technology, developed countries, represented by the United States, the European Union, and Japan, have built highly automated flood early warning systems based on high-density monitoring networks, advanced numerical weather prediction, and hydrodynamic models. They have mature experience in multi-source data fusion, real-time forecast correction, and global early warning services. After years of development, China has formed an integrated air-ground monitoring and forecasting system. With its independent hydrological model as the core, it integrates multi-source data from radar, satellites, and BeiDou, gradually realizing distributed hydrological simulation, hydrodynamic coupling, and intelligent early warning. It is also rapidly developing in areas such as digital twin river basins, short-term forecasts, and multi-level coordinated early warning.

[0003] In summary, current flood early warning systems in China and abroad mainly rely on traditional monitoring methods such as meteorological stations, hydrological stations, radar, and satellites, combined with hydrological and hydrodynamic models to achieve forecasting and early warning. However, systemic technical deficiencies still exist in practical applications. Traditional monitoring data suffers from insufficient coverage and limited real-time performance: Existing early warning technologies mainly rely on data collected by traditional monitoring equipment such as meteorological stations and hydrological monitoring stations. While this data is accurate, its coverage is limited, the monitoring network is sparse, and the update frequency is relatively low, resulting in delayed early warnings and an inability to comprehensively monitor large blind spots in urban-rural fringe areas, remote regions, and small and medium-sized river basins.

[0004] The warning information lacks refinement and personalization: Traditional warnings mostly adopt a unified release mode across the entire region, without distinguishing user location, risk level and degree of impact. This results in warnings not being prominent in high-risk areas and information overload in low-risk areas, low public response efficiency, untimely evacuation and relocation, and an inability to accurately provide targeted warning information and effectively push risk level warnings.

[0005] Insufficient fusion of multi-source data and inconsistent spatiotemporal benchmarks: Social media text data and mobile phone location data are out of sync in time, mismatched in space, and inconsistent in format, making it impossible to form an effective complement. This leads to inaccurate identification of flood events, large deviations in risk assessment, insufficient credibility of early warnings, and difficulty in supporting accurate early warning decisions.

[0006] The aforementioned deficiencies directly result in problems such as low accuracy, slow response, and insufficient coverage in the existing flood early warning system, making it difficult to fully meet the urgent need for high-precision, full-coverage, intelligent, and timely early warning of rainstorms and floods in modern river basin flood control and disaster reduction. Summary of the Invention

[0007] The purpose of this application is to provide a flood early warning method, system, device and product based on multi-source social perception and swarm intelligence, and to solve the problem of difficulty in meeting the requirements of high-precision and full-coverage intelligent and rapid flood early warning.

[0008] To achieve the above objectives, this application provides the following solution.

[0009] Firstly, this application provides a flood early warning method based on multi-source social perception and swarm intelligence, including: Collect and preprocess social media data and mobile LBS data to construct a spatiotemporally aligned multi-source fusion dataset.

[0010] Based on the multi-source fusion dataset, standardized flood event information packages are generated using NLP and spatial clustering algorithms.

[0011] The standardized flood event information package is input into the SWAT hydrological model to simulate floods, construct a three-dimensional risk index system, and classify risk levels.

[0012] Based on the aforementioned risk information, the user's real-time location, and the push channel, a tiered and categorized early warning system will be implemented.

[0013] Secondly, this application provides a flood early warning system based on multi-source social perception and swarm intelligence, comprising: The multi-source fusion dataset construction module collects and preprocesses social media data and mobile LBS data to build a spatiotemporally aligned multi-source fusion dataset.

[0014] The standardized flood event information package determination module generates standardized flood event information packages based on the multi-source fusion dataset using NLP and spatial clustering algorithms.

[0015] The risk level classification module inputs the standardized flood event information package into the SWAT hydrological model to simulate floods, construct a three-dimensional risk indicator system, and classify risk levels.

[0016] The early warning push module performs tiered and categorized early warnings based on the risk level, user location, and push channel.

[0017] Thirdly, this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the flood early warning method based on multi-source social perception and swarm intelligence as described above.

[0018] Fourthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the flood early warning method based on multi-source social perception and swarm intelligence as described above.

[0019] Based on the specific embodiments provided in this application, the following technical effects are disclosed.

[0020] This application provides a flood early warning method and system based on multi-source social perception and swarm intelligence. By collecting and preprocessing social media data and location-based services (LBS) data from mobile phones, a spatiotemporally aligned multi-source fusion dataset is constructed, providing a high-quality unified data foundation for model simulation and risk assessment. This addresses the problem of insufficient multi-source data fusion and overcomes the limitations of traditional sparse monitoring stations, achieving comprehensive coverage of the entire watershed without blind spots. Standardized flood event information packages are generated using Natural Language Processing (NLP) and spatial clustering algorithms to accurately identify flood events, classify hazard levels, and precisely locate high-risk clusters. The accuracy of flood simulation is improved by combining hydrological model simulation with intelligent algorithms. A three-dimensional risk index system is constructed, and combined weights and threshold settings are used to solve the problem of ambiguity in risk assessment, effectively highlighting risk levels and providing accurate early warning information. Differentiated early warning information and different push channels are generated based on risk level and user location, enabling hierarchical and categorized accurate and rapid early warning, effectively solving the problems of one-size-fits-all early warnings and low response efficiency. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 A flowchart illustrating the overall process of a flood early warning method based on multi-source social perception and swarm intelligence provided in this application embodiment. Figure 2A schematic diagram of the structure of a flood early warning method based on multi-source social perception and swarm intelligence provided in an embodiment of this application; Figure 3 This is a schematic diagram of four-dimensional structured information domain parsing provided in an embodiment of this application; Figure 4 This is a schematic diagram of the internal structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0023] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0024] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0025] This application discloses a flood early warning method based on multi-source social perception and swarm intelligence, belonging to the interdisciplinary fields of hydroinformatics, social perception computing and swarm intelligence. Specifically, it involves a method that integrates multi-source social perception data such as social media and mobile LBS, combined with hydrological model simulation and intelligent algorithms, to address the technical shortcomings of existing flood early warning monitoring, such as insufficient coverage, poor real-time performance and poor early warning push effect, and to achieve real-time perception, risk assessment and hierarchical classification of watershed rainstorms and floods.

[0026] like Figures 1-2 As shown in the embodiments of this application, the flood early warning method based on multi-source social perception and swarm intelligence includes: Step 1: Collect and preprocess social media data and mobile LBS data to construct a spatiotemporally aligned multi-source fusion dataset.

[0027] Step 2: Based on the multi-source fusion dataset, generate standardized flood event information packages using NLP and spatial clustering algorithms.

[0028] Step 3: Input the standardized flood event information package into the SWAT hydrological model to simulate floods, construct a three-dimensional risk index system, and classify risk levels.

[0029] Step 4: Based on the risk information, the user's real-time location, and the push channel, execute a tiered and categorized early warning system.

[0030] In an exemplary embodiment, step 1 involves collecting and preprocessing social media data and anonymized mobile LBS data to unify the spatiotemporal benchmark and construct a spatiotemporally aligned multi-source fusion dataset, specifically including the following steps.

[0031] Step 101: Obtain the social media data through API interfaces or compliant web crawlers, and obtain mobile LBS data through telecom operators or relevant location service providers, and complete anonymization and desensitization processing; the processed mobile LBS data includes: location coordinate data (latitude, longitude, altitude, etc.), timestamp information (location time), movement trajectory data (historical location sequence), and context information (speed, direction, stop point, etc.); the two types of data are deduplicated, cleaned, desensitized, and unified with spatiotemporal reference to construct a consistent WGS-84 spatiotemporal reference dataset to ensure data quality and spatiotemporal consistency.

[0032] Step 102: Convert the timestamps of the social media data into Unix timestamp format.

[0033] Step 103: Using the dynamic time window method combined with time error, the converted social media data and the processed mobile LBS data are synchronized and aligned in time.

[0034] Step 104: Using geocoding mapping and spatial membership, the converted social media data is spatially synchronized with the mobile LBS data collection.

[0035] Step 105: Construct a spatiotemporally aligned multi-source fusion dataset based on time-synchronized and spatially synchronized social media data and mobile LBS data.

[0036] As an optional implementation, in step 101, multi-source social perception data is collected and preprocessed through API interfaces or web crawling technology, relying on cooperation with telecommunications operators or relevant location service providers. The specific method is as follows.

[0037] (1) Obtain social media data and mobile LBS data.

[0038] ① Obtain social media awareness data based on API interfaces. Register a developer account on the target social media platform, complete the entity qualification verification, create an application in the developer console, and obtain core access credentials as the authentication basis for API calls. Study the API interface documentation published by the platform to clarify the interface categories, interface endpoints, request methods, parameter constraints, call quotas, and data return formats. Construct a structured request parameter set and configure request headers according to the interface documentation requirements. Use an HTTP client tool to initiate interface requests and transmission control, query relevant awareness data, including qualitative descriptions, scenario evidence, and spatiotemporal matching flood model information. Receive the HTTP response messages returned by the interface, prioritizing status code verification. For successful responses, extract the response body data and perform deserialization. Convert string-formatted data into structured data objects recognizable by the programming language, completing the mapping from raw data to program-operable data, providing a foundation for subsequent data use.

[0039] ② Obtain social media awareness data using web crawling technology. Set up a basic Python environment, install web crawling dependencies, confirm compliance requirements, clarify the rules of the target platform, and limit the use of data. Based on the provided code, modify the core parameters to adapt for flood information capture in the target monitoring area, including modifying keywords, target links, and setting output files. Save the configured web crawler code and start the crawler to execute commands. Check the command line output to confirm whether valid information has been captured; if a "request failed" message appears, check the network or adjust the request headers. After execution, find the generated CSV file in the same folder, which contains information such as user nicknames, posting times, and post content. Obtain the raw information, remove duplicates and content without flood keywords, manually or using simple tools, extract flood simulation information data from the posts, and standardize the cleaned data.

[0040] ③ Obtain user mobile phone LBS data by relying on telecom operators or relevant location service providers. Determine the required data type, spatiotemporal resolution, and coverage area. Based on relevant personal information protection clauses and telecom industry data security standards, the partner will anonymize and desensitize the original user data, extracting core basic positioning data including user anonymity identifiers, globally unique base station identifiers, latitude and longitude coordinates, second-level timestamps, and signal interaction status. Simultaneously, through cloud edge computing, derived feature analysis data covering population spatiotemporal distribution characteristics, movement trajectory fragments, and regional communication activity will be generated to form a standardized dataset to support subsequent business operations.

[0041] (2) Integrate multi-source social perception data.

[0042] ①Unify the spatiotemporal benchmark for data.

[0043] Set the social media dataset as ,in For the total amount of social media data, each data sample Carry timestamp (Accuracy down to the second, formatted as a Unix timestamp) and spatial tags Mobile LBS data set is The total amount of LBS data, per sample Carry timestamp and precise geographic coordinates Geographic coordinates are obtained using latitude and longitude in the WGS-84 coordinate system.

[0044] ②Alignment time.

[0045] After uniformly converting the timestamps of social media data to the Unix timestamp format consistent with LBS data, let the time conversion function be denoted as . The converted social media data timestamps are To ensure consistent time accuracy, the conversion error must meet the following requirements. ,in, For LBS data Unix timestamps.

[0046] A dynamic time window method is adopted, which adaptively adjusts the width of the time alignment window according to the data acquisition frequency. The value ranges from 5 to 30 seconds. For any LBS data sample Constructing a temporal neighborhood Iterate through the social media data samples, if they meet the following conditions: Then it is believed Meeting the time alignment criteria constitutes a candidate spatiotemporal alignment pair. To optimize alignment accuracy, a time synchronization error correction model is introduced. Let the time reference deviation between the two source data be denoted as... The optimal value of the deviation is obtained by using the least squares method.

[0047] .

[0048] in, This represents the initial number of candidate alignment pairs. For the first Social media timestamps on candidate data For the first LBS timestamps for candidate data. Corrected social media timestamps are... , will be updated to.

[0049] .

[0050] Sample pairs that meet the time alignment criteria will be included in the spatiotemporal alignment candidate set. ③ Alignment space.

[0051] Using geocoding mapping function Spatial labeling of social media data Transform to geographic coordinates under LBS data, i.e., spatial labels of social media data. The spatial labels are text-based, while the geographic coordinates in the LBS data are labeled as latitude and longitude rectangular areas. The mapping error is measured using the area of ​​the region corresponding to the geographic coordinates in the LBS data, maintaining a geometrical correspondence between the two. For example, text-based spatial labels (such as "Hongshan District, Wuhan City") can be converted into latitude and longitude rectangular areas. ,in, These are the minimum and maximum longitudes of the region, respectively. These are the minimum and maximum latitudes of the region, respectively. The mapping error must satisfy the requirement of the region area. Approximately 111km 2 Here, 0.001 refers to the upper limit of the product of latitude and longitude differences, which is adapted to the spatial accuracy requirements of cities.

[0052] Based on spatiotemporal alignment candidate set Each pair of candidate samples Calculate LBS coordinates With rectangular area Spatial membership is defined using a distance-weighted model, which is used to define the membership function.

[0053] .

[0054] In the formula, ρ is the shortest distance from the geographic coordinate point to the rectangular area. If the geographic coordinate point is within the area, then ρ = 0; otherwise, it is the Euclidean distance from the geographic coordinate point to the boundary of the area. This is the distance attenuation coefficient, with a custom value of 500m to suit the data collection density within the city.

[0055] Set spatial alignment membership threshold ,when At that time, the judgment Spatial alignment was completed, and the data was incorporated into the final spatiotemporally aligned dataset. .

[0056] Table 1 shows the social media data, mobile LBS data, and data preprocessing results for a certain region.

[0057] Table 1. Social media data, mobile LBS data, and data preprocessing results for a certain region. To further optimize accuracy, the spatiotemporal alignment dataset was optimized. Spatial error correction is employed to achieve accurate unification of the spatial reference between the two data sources. The corrected LBS coordinates are as follows.

[0058] .

[0059] in, , ,in These are the minimum and maximum longitudes of the region, respectively. These are the minimum and maximum latitudes of the region, respectively.

[0060] In another exemplary embodiment, step 2 generates a standardized flood event information package based on the multi-source fusion dataset using natural language processing (NLP) technology and spatial clustering algorithms, specifically including the following steps.

[0061] Step 201: Based on the improved TF-IDF algorithm in the NLP, the matching coefficient of the core dictionary in the flood domain is introduced to optimize the weight of the text keywords in the social media data and determine the optimized keywords.

[0062] Step 202: Based on the BERT entity recognition model in the NLP, extract flood event elements from social media text according to the optimized keywords.

[0063] Step 203: Based on the spatial clustering algorithm, the DBSCAN algorithm or K-means algorithm is used to process the mobile phone LBS data to obtain the population gathering areas within the watershed; and the silhouette coefficient is used to evaluate the spatial clustering effect of the population location data.

[0064] Step 204: Based on the flood event elements and the population gathering areas within the basin, construct a spatial correlation model between flood events and population gathering areas, and generate a standardized flood event information package to provide structured data support for subsequent risk assessment.

[0065] The flood event elements include time, location, severity of the disaster, and rescue needs. The standardized flood event information package is encapsulated in JSON format and includes a basic information field, a spatiotemporal characteristic field, a disaster impact field, and an emergency association field. It uses a three-level unique encoding of "administrative division code-date-time sequence number" and adds a traceability identifier to enable reverse tracing of the original data.

[0066] As an optional implementation method, the method described in step 2, which uses intelligent algorithms to extract flood information based on socially perceived data and construct a standardized flood event information package, is as follows.

[0067] (1) Use natural language processing (NLP) technology to process social perception data.

[0068] ① Dataset based on spatiotemporal alignment Social media data This enables the structured extraction of flood-related information.

[0069] Define an unstructured text collection The total number of texts is The word sequence after word segmentation of a single text is as follows: ,in For the first The number of words in each text segment.

[0070] Constructing a core dictionary for the field of flood Including floods, waterlogging, water levels, and levee breaches. One core keyword, introducing the matching coefficient of the core dictionary in the flood field. :like ,but ,otherwise The improved weighting formula for TF-IDF is as follows.

[0071] .

[0072] In the formula, , for Number of times it appears in the text To determine the inverse document frequency, add 1 to the numerator and denominator to avoid zero values.

[0073] Set weight threshold Filter out The words used as core keywords constitute a single-text keyword set. Simultaneously, edit distance of word sequences is used. Calculate the similarity between keyword phrases and merge phrases with similar meanings to avoid a large amount of semantic repetition. For example, "flood" and "flood spread" can be merged into "flood diffusion".

[0074] The keyword extraction and optimization results of the selected samples are shown in Table 2.

[0075] Table 2 shows the keyword extraction and optimization results of the selected samples. ② Construct a sentiment-risk level mapping model to quantify text sentiment into risk values.

[0076] Following the same steps as above, construct a flood-related sentiment dictionary. The keyword sentiment was labeled as negative. (1 = slight impact, 5 = extreme risk), calculate the average negative tendency of a single text.

[0077] .

[0078] Introducing text credibility Adjust the negative bias mean to construct the final risk assessment model.

[0079] .

[0080] In the formula, A is the score for the accuracy of key information, such as the core information in the text and the similarity between official reports and monitoring data; B is the score for content consistency, such as the consistency between what this text says and the content posted by others in the same time and area; the values ​​of A and B are both in the range of 0 to 1.

[0081] Set hazard level thresholds, Low risk Medium risk For high risk, output a single text hazard level corresponding to the hazard level threshold. .

[0082] The sentiment-risk mapping results of the selected samples are shown in Table 3.

[0083] Table 3 shows the sentiment-risk mapping results of the selected samples. ③ Extract the core elements of flood events, construct an event correlation model, and explore the spatiotemporal and semantic relationships between different events.

[0084] Define the set of flood event elements (Time, location, disaster severity, rescue needs), extract elements based on the BERT entity recognition model, and define the element matching degree.

[0085] .

[0086] In the formula, For the first The first text extracted Class elements, For the first Domain entity set of class elements, element matching degree When the elements are valid, an event is constituted. and events .

[0087] Building event correlation It integrates the dimensions of time, space and semantics.

[0088] .

[0089] In the formula, Ei and Ej refer to two existing valid events; , This is the time similarity weight, with a value of 0.3. This represents the spatial similarity weight, with a value of 0.4. This is the semantic similarity weight for keywords, with a value of 0.3. For time similarity, the smaller the time difference, the closer the value is to 1. For spatial similarity, the closer the locations of two events are, the higher the event relevance score. For keyword semantic similarity, the more similar the keywords of two events, the higher the event relevance score. When When the two events are determined to be strongly correlated, it is determined that they are related.

[0090] The results of flood event element extraction and correlation analysis of the selected samples are shown in Table 4.

[0091] Table 4. Results of flood event element extraction and correlation analysis for the selected samples. (2) Spatial clustering algorithm is used to process mobile phone LBS data.

[0092] ① For the collected mobile LBS positioning data, coordinate calibration preprocessing is first performed to remove abnormal positioning points such as signal drift and invalid coordinates, ensuring data quality. Based on the preprocessed mobile LBS data, K-means and DBSCAN spatial clustering algorithms are used, relying on the spatial proximity characteristics of user location coordinates, to divide all user positioning points in the watershed into several independent spatial clustering regions. ,in This represents the total number of cluster regions. Representing the Each cluster region corresponds to a specific geographical area.

[0093] To ensure clustering accuracy, the silhouette coefficient is introduced as a clustering effectiveness evaluation index to quantify the compactness and separation of the clustering results. The formula is as follows.

[0094] .

[0095] In the formula: for All internal positioning points to The average distance to other points within the area; for The average distance from all localization points within the cluster to their nearest neighbors within the cluster region. Profile coefficient. The closer the value is to 1, the better the clustering effect. The clustering parameters include the number of clusters K in the K-means algorithm, the spatial neighborhood radius ϵ and the minimum number of samples in the DBSCAN algorithm. By adaptively adjusting the clustering parameters to maximize the silhouette coefficient, high-precision structured aggregation of user spatial distribution within the watershed is achieved, laying a data foundation for subsequent event spatial correlation analysis and population feature mining.

[0096] ②Based on the core elements of social media flood events extracted by NLP, spatial correlation analysis is performed with the spatial clustering results of mobile LBS data to construct a quantitative spatial correlation model, thereby achieving precise coupling between flood events and areas where people gather.

[0097] Define flood events Clustering regions Spatial correlation model for.

[0098] .

[0099] In the formula, For spatial matching degree weight, As a weight for the risk level, To prioritize ensuring spatial matching, the following settings are configured. , ; For flood events Location of occurrence With densely populated areas within the watershed Spatial matching degree, if Falling If the value is within the specified range, then take 1; otherwise, apply the distance decay formula. calculate, for arrive The shortest distance to the boundary; The number of users in a densely populated area. This represents the total number of users within the watershed. Reflects the population density in the clustered region. The danger level of the flood event.

[0100] By using a spatial correlation model to statistically analyze the number of associated flood events and the total correlation within each cluster region, high-correlation cluster regions were identified. The population distribution characteristics provide core technical support for basin-wide risk assessment, disaster impact scope definition, and precise allocation of rescue resources.

[0101] The clustering results of mobile LBS data space using the spatial correlation model are shown in Table 5.

[0102] Table 5 Clustering results of mobile LBS data space (3) Construct standardized flood event information packages.

[0103] ① Establish a unified metadata standard, adopt a three-level unique coding system of "administrative division code-date-time sequence number", define required and optional fields, unify the JSON format and WGS-84 spatiotemporal benchmark, add traceability identifiers to each information sub-item, complete the structured construction and standardized encapsulation of standardized flood event information packages, and realize the reverse tracing of original data.

[0104] ②For example Figure 3 As shown, based on the aforementioned metadata specifications, a four-dimensional structured information domain is constructed, comprising a basic information domain, a spatiotemporal feature domain, a disaster impact domain, and an emergency correlation domain. The basic information domain primarily includes event identifiers, a list of data source types and proportions, and information credibility based on multi-source fusion results. The spatiotemporal feature domain primarily includes the event's spatiotemporal range, evolution trajectory, and water level changes. The disaster impact domain primarily includes the impact on personnel, infrastructure, and loss estimation. The emergency correlation domain primarily includes real-time monitoring data from associated hydrological or rainfall stations, nearby flood control material reserve points, and the location of rescue teams.

[0105] A case study of constructing a standardized flood event information package for selected samples is shown in Table 6.

[0106] Table 6 Examples of Standardized Flood Event Information Packages In another exemplary embodiment, step 3 involves inputting the standardized flood event information package into the SWAT hydrological model to simulate floods, constructing a three-dimensional risk indicator system, and classifying risk levels. This step specifically includes the following steps.

[0107] Step 301: Input the standardized flood event information package into the hydrological model (Soil and Water Assessment Tool, SWAT), correct the model parameters and boundary conditions. The model parameters include five specific parameters: actual hydrology, soil, topography, vegetation, and river channel. The boundary conditions include four types of boundary parameters: spatial, temporal, hydrological, and meteorological. Simulate the complete flood process within the basin, output hydrological elements, and construct a three-dimensional risk indicator system. The complete flood process within the basin includes: rainfall, interception, infiltration, runoff generation, and confluence. The hydrological elements include: inundation range, inundation depth, and duration. The three-dimensional risk indicator system includes: hydrological inundation characteristics, risk of casualties, and risk of economic loss.

[0108] Step 302: Based on the three-dimensional risk index system, the comprehensive risk index is calculated using a combined weighting method of AHP (Analytic Hierarchy Process) and entropy weighting.

[0109] Step 303: Based on the comprehensive risk index, and in conjunction with industry standards and expert consultation, determine the threshold, classify the risk levels, and achieve a quantitative and reproducible risk assessment. The risk levels include low risk, medium risk, and high risk.

[0110] Step 304: Based on the comprehensive risk index, combined with industry standards and expert opinions, the risk level is classified.

[0111] As an optional implementation method, step 3 combines flood model simulation to construct a three-dimensional risk index system and classify risk levels. The specific method is as follows.

[0112] (1) Simulate the complete process of floods in the basin using a hydrological model.

[0113] ① Combining the characteristics of the watershed, the SWAT model was selected as the sole hydrological simulation tool. Based on the concept of distributed hydrological models, the spatiotemporal watershed (i.e., the spatiotemporal range within the spatiotemporal feature domain, processed LBS data) was divided into several sub-watersheds and hydrological response units (HRUs) according to DEM topographic data. Specifically, relying on ArcGIS spatial analysis functions, underlying surface data such as topography, soil, and land use were integrated as input to the SWAT model. Based on DEM water system analysis, the watershed boundary was extracted and divided into several sub-watersheds. Based on the homogeneous computational units generated by superimposing land use, soil type, and slope, the HRUs were divided. The complete hydrological process of "rainfall-interception-infiltration-runoff generation-confluence" within the watershed was simulated. By combining physical mechanisms with empirical parameters, the natural laws of flood formation were restored. This method is suitable for flood simulation in small and medium-sized watersheds and areas with strong underlying surface heterogeneity, and is highly adaptable to rainstorm flood simulation and water conservancy project impact analysis.

[0114] a. The interception process.

[0115] .

[0116] In the formula: This is net rainfall. This represents the measured rainfall at HRU. This represents the amount retained by the canopy.

[0117] b. Infiltration-flow process (SCS curve method).

[0118] .

[0119] In the formula: This represents the actual surface runoff. The maximum retention capacity of the watershed ( CN2 is the number of curves.

[0120] c. Convergence process.

[0121] .

[0122] In the formula: This represents the total outflow rate of the cross-section. It is surface runoff on the slope. For lateral groundwater recharge, It is the base current.

[0123] ② After inputting the standardized flood event information package into the SWAT model, the key input parameters of the model (hydrological, underlying surface and meteorological parameters) are calibrated with field measured data, and the spatial differentiation optimization of the parameters is achieved through deviation quantification.

[0124] Standardized flood event information packages were imported into the ArcGIS platform, and spatial overlay analysis was performed on sub-basin vector layers and HRU vector layers to establish a spatial mapping relationship between flood event data and SWAT model calculation units. First-level validity information was prioritized and matched with the corresponding HRU information of the sub-basin. Among them, the first-level validity information (location information with spatial accuracy ≤100m, event time with temporal accuracy ≤1 hour, and feature matching degree ≥0.6) is high-confidence flood event data extracted from multi-source social perception data and cross-validated by multiple sources, ensuring the quality of model input data and ensuring the scientificity and effectiveness of deviation rate calculation and model optimization.

[0125] The deviation quantification formula is used as the basis for correcting the parameters of the SWAT hydrological model, and the deviation rate δ is .

[0126] .

[0127] In the formula: For parameter deviation rate, This is the measured value of the information packet. For model simulation values; when When the threshold is exceeded, parameters are iteratively adjusted based on standardized flood event information packages. After adjustment, the validity level of the updated information packages is fed back, forming a closed-loop optimization. The adjustment process can be viewed in real time on an ArcGIS map to observe changes in the spatial distribution of parameters, adapting to spatially heterogeneous requirements.

[0128] (2) Construct a three-dimensional risk indicator system.

[0129] Based on the "ontology-impact-loss" logical relationship of flood risk, and combining the output of the SWAT model with the standardized flood event information package, a three-level hierarchical structure is constructed, and the consistency of the three levels is judged. The consistency test formula is as follows.

[0130] .

[0131] In the formula, For consistency ratio, Consistent at different time levels; As a consistency indicator, , To determine the largest eigenvalue of a matrix, The number of criterion layers; The average random consistency index, .

[0132] The first layer is the target layer, where the comprehensive flood risk index is calculated.

[0133] .

[0134] In the formula, α is the weight of hydrological inundation characteristics, β is the weight of casualty risk, γ is the weight of economic loss risk, and α+β+γ=1. α, β, and γ are determined by a combination weighting method using the AHP (Analytic Hierarchy Process) combined with the entropy weighting method.

[0135] The second layer is the criteria layer, which includes three core criteria, all of which are connected to the model output and information package data. Hydrological inundation characteristic criteria. Corresponding spatiotemporal feature domain information package module and model flooding simulation data; personnel casualty risk criteria This includes information packages related to disaster impact and emergency response, as well as simulation data on personnel exposure in the model; and economic loss risk criteria. This corresponds to the disaster impact domain information package module and the model loss simulation data.

[0136] The third layer is the indicator layer, which breaks down quantifiable and reproducible sub-indicators and determines the quantification formulas. Each indicator is derived from the fused data and model output. For example, relative inundation range. Classified inundation depth index Population risk index Estimated total economic losses wait.

[0137] Among them, the relative inundation range for.

[0138] .

[0139] In the formula, The total inundated area is represented by data from SWAT model simulations. To assess the total area of ​​the region.

[0140] Classified flood depth index for.

[0141] .

[0142] In the formula, , The inundation depth is classified into four levels: k=1 (inundation depth: 0-0.5m), k=2 (inundation depth: 0.5-1m), k=3 (inundation depth: 1-2m), and k=4 (inundation depth: >2m). The value of the k-th level inundation depth is: 0.25 for k=1, 0.75 for k=2, 1.5 for k=3, and 2.5 for k=4. This represents the percentage of the area submerged at level k. The inundation area of ​​level k is based on the simulation results of the SWAT model. The total inundated area is given in the formula for the relative inundation range. Consistent.

[0143] Population Risk Index for.

[0144] .

[0145] In the formula, This is a population risk index, ranging from 0 to 1. The higher the value, the greater the risk of flooding to the population. The population risk exposure zone is divided into three categories: low agglomeration zone when m=1, medium agglomeration zone when m=2, and high agglomeration zone when m=3. The number of people in the flooded area of ​​the m-th type of population agglomeration area is derived from the information package module and is directly counted from population perception data without distinguishing age. The weight of the m-th type of population agglomeration area is assigned according to the degree of population agglomeration: 0.2 for low agglomeration area, 0.3 for medium agglomeration area, and 0.5 for high agglomeration area. The higher the degree of agglomeration, the greater the weight, which is consistent with the pattern of flood risk. This represents the total population of the flooded area; the data is sourced from the information package module.

[0146] Estimated total economic loss for.

[0147] .

[0148] In the formula, Total economic loss; For direct economic losses, ,in, Based on asset type, it is divided into 3 categories: residential buildings at t=1, farmland at t=2, and industrial facilities at t=3. The inundation area represents the area of ​​asset type t, and the data comes from the information package module. The value per unit area of ​​asset class t is derived from the information package module. The damage rate of asset class t is determined based on the flooding depth simulated by the SWAT model: Within the range of 0 to 0.5m, =0.1, Within the range of 0.5 to 1m, =0.3, Within the range of 1 to 2 meters, =0.6, Within the range of >2m =0.9; As indirect economic losses, ,in, This is the indirect loss coefficient, assigned a value based on the regional industry type: agricultural areas, =0.2; Industrial zone, =0.4; Commercial area, =0.5, Duration of flood impact.

[0149] The combined weighting method, which combines the Analytic Hierarchy Process (AHP) with the entropy weighting method, can avoid the limitations of a single weight. By adjusting the weights based on the accuracy of the model output and the reliability of the information packet data, the calculation process and basis are clearly defined.

[0150] The core risk indicators are provided, as shown in Table 7.

[0151] Table 7 Core Risk Indicators (3) Classify risk levels.

[0152] ①Based on the comprehensive risk index With the core as the reference, the range is 0 to 1. The threshold is determined by combining the hydrological inundation simulation results output by the model, the risk of casualties and economic losses, and industry standards, and is determined by expert demonstration.

[0153] ② Quantitative definition of risk level.

[0154] When the comprehensive risk index R ∈ [0, 0.3) and simultaneously meets the first condition in the industry standard, the current risk level is determined to be low risk level after expert evaluation; the first condition includes the simulated inundation depth < 0.5m and the relative inundation range. <10%, Population Risk Index <0.2 and estimated total economic loss If the percentage is less than 5%, there is no need to activate a Level I emergency response.

[0155] When the comprehensive risk index R ∈ [0.3, 0.7) and simultaneously meets the second condition in the industry standard, the current risk level is determined to be medium risk level after expert evaluation; the second condition includes the simulated inundation depth of 0.5–1.5 m and the relative inundation range. Population risk index and estimated total economic losses If the proportion is 5% to 20%, a Level II emergency response needs to be activated.

[0156] When the comprehensive risk index R ∈ [0.7, 1] and simultaneously meets the third condition in the industry standard, the current risk level is determined to be high risk level after expert evaluation; the third condition includes the simulated inundation depth ≥ 1.5m and the relative inundation range ≥30%, Population Risk Index ≥0.6 and estimated total economic loss If the proportion is ≥20%, a Level I emergency response needs to be initiated.

[0157] The risk levels and their defined scope are provided, as shown in Table 8.

[0158] Table 8 Risk Levels and Their Scope ③ Adjust the thresholds based on the model simulation results and the scenario adaptability of the information package to ensure applicability to different scenarios, and directly connect the risk level classification results with flood warnings and emergency resource allocation.

[0159] In another exemplary embodiment, step 4 performs hierarchical and categorized precise early warning push based on risk information, user real-time location and push channel.

[0160] Step 401: Generate differentiated early warning information based on the risk information, crowd density, and user's real-time location. The risk information includes: basic risk information, risk level information, and risk impact information.

[0161] Step 402: Adjust the content length of the push channel according to the differentiated early warning information. At the same time, establish a hierarchical push mechanism based on the comprehensive risk index, which includes multiple strong reminders for high-risk early warnings and regular coverage for medium and low-risk situations. Implement hierarchical and classified early warnings to achieve accurate, timely, and comprehensive delivery of early warning information, thereby improving public response efficiency and emergency response capabilities.

[0162] As an optional implementation method, step 4, based on risk assessment and user location information, customizes early warning information and constructs a hierarchical and classified early warning push mechanism, the specific method of which is as follows.

[0163] (1) Early warning information generation and customization.

[0164] The customization of early warning information follows the core processing principle of "standardized modules + channel adaptation". It relies on risk assessment reports to extract key related elements and completes the customization according to unified specifications to ensure that the information is highly consistent with the assessment results, while adapting to the dissemination characteristics of different push channels.

[0165] ① Clearly define the core relevant elements of risk assessment. Using the risk assessment report as the core support, extract key relevant elements as the basis for customizing early warning information, ensuring a high degree of consistency between the information and the assessment results. Specifically, this includes...

[0166] a. Basic Risk Information: Clearly state the risk name and source, and match the risk identification results in the assessment report, so that the recipient can quickly understand the core attributes of the risk.

[0167] b. Risk level information: directly adopts the comprehensive risk index. and threshold The defined risk level is marked simultaneously. The specific values ​​should be clearly defined to indicate the urgency and severity of the risk.

[0168] c. Risk impact information: Combining inundation simulation results, , Data such as these are used to clarify the specific geographical scope of the risk impact, the size of the affected groups, the core business processes, and the scope of assets, while also... Quantify and estimate the extent of loss.

[0169] d. Handling guidance information: Based on mobile phone LBS location data and risk response suggestions in the assessment report, extract concise and operable preliminary handling measures, clarify the responsible parties and key handling points, and avoid leaving the recipient at a loss.

[0170] e. Warning period information: Based on the risk duration determined in the risk assessment. and the rate of change of the risk index Define the effective time of the warning information Expiration time , At the same time, according to The absolute value dynamically determines the real-time early warning update frequency for disaster-stricken areas. The derivation is as follows.

[0171] .

[0172] ②Adjust the length based on the push channel and adopt a differentiated processing strategy. For SMS channels, strictly limit to 100 characters, prioritizing the retention of four key information categories: warning title, risk level, core handling guidelines, and warning period, while simplifying the description of quantitative parameters; for email or announcement channels, supplement with complete quantitative derivation basis, flooding simulation results, and expert opinions, etc., to ensure the integrity and traceability of information.

[0173] Table 9 provides examples of warning text messages corresponding to different risk levels.

[0174] Table 9 Examples of early warning text messages corresponding to different risk levels (2) Construct a tiered early warning and push mechanism.

[0175] A tiered early warning and notification mechanism based on a comprehensive risk index Using this as the core quantitative basis, combined with quantitative monitoring data and expert evaluation results, we can accurately define the risk level and then formulate tiered push rules to ensure the targeting and timeliness of the push.

[0176] An early warning push mechanism is provided, as shown in Table 10.

[0177] Table 10 Early Warning Push Mechanism ① Implement tiered and differentiated push notification processing, adopting corresponding push modes based on the urgency of the risk level to ensure accurate information delivery.

[0178] High-risk warning information adopts a multi-pronged push mode of "SMS + mobile application pop-up + telephone notification", and is set up to trigger a second telephone or SMS reminder if the user does not read it within 5 minutes, to ensure the timeliness and reach of emergency information.

[0179] For medium- and low-risk information: Based on the channel compatibility score, prioritize convenient channels such as SMS and mobile applications, while simultaneously covering public channels such as radio and television to achieve comprehensive early warning coverage, balancing push efficiency and cost.

[0180] ② Based on risk assessment results and push notification effectiveness data, construct a mechanism for continuous learning and dynamic updating to ensure the rigor and sustainability of the tiered early warning push mechanism. Continuously update and process data, establishing a multi-source data real-time collection and regular update mechanism to continuously collect various relevant data such as social media data, mobile LBS data, and actual flood measurement data.

[0181] To address the problems of existing technologies, such as delayed perception, low accuracy, poor targeting, slow response, and insufficient coverage, this application provides a flood early warning method based on multi-source social perception and swarm intelligence. This method enables early perception, rapid identification, and accurate location of flood events; dynamic perception of population distribution and movement trajectories; quantitative assessment and classification of flood risks; and precise delivery and categorized distribution of early warning information. It fundamentally improves the timeliness, accuracy, targeting, and coverage of flood early warnings, providing reliable technical support for watershed flood control and disaster reduction. Breaking through the limitations of traditional monitoring, this method achieves early perception, rapid identification, accurate assessment, and precise delivery of flood warnings, significantly improving the timeliness, accuracy, and targeting of early warnings. It can be widely applied to flood control and disaster reduction early warning work in various watersheds.

[0182] This application also provides a flood early warning system based on multi-source social perception and swarm intelligence, which includes a flood early warning method based on multi-source social perception and swarm intelligence.

[0183] The multi-source fusion dataset construction module collects and preprocesses social media data and mobile LBS data to build a spatiotemporally aligned multi-source fusion dataset.

[0184] The standardized flood event information package determination module generates standardized flood event information packages based on the multi-source fusion dataset using NLP and spatial clustering algorithms.

[0185] The risk level classification module inputs the standardized flood event information package into the SWAT hydrological model to simulate floods, construct a three-dimensional risk indicator system, and classify risk levels.

[0186] The early warning push module performs tiered and categorized early warnings based on the risk level, user location, and push channel.

[0187] In practical applications, this embodiment takes a small to medium-sized watershed as the research object, with a watershed area of ​​approximately 500 km². 2 The area covers various underlying surface types, including urban areas, farmland, and mountains. There are currently three traditional hydrological monitoring stations and two meteorological stations, which have problems such as many monitoring blind spots and delayed early warnings. This application proposes a flood early warning method based on multi-source social perception and swarm intelligence to achieve refined early warning.

[0188] Experimental equipment: server (CPU: Intel Xeon E5-2690, memory: 64GB, hard disk: 1TB), ArcGIS 10.8 software, Python 3.9 environment (with scikit-learn, jieba, requests and other dependent libraries installed), hydrological monitoring data terminal.

[0189] Data sources: API interface data from social media platforms (Weibo, Douyin), anonymized mobile LBS data from local telecom operators, watershed DEM topographic data (30m resolution), soil type data, land use data, and measured data from hydrological stations.

[0190] Collect multi-source data and perform preprocessing to obtain accurate and unified dual-source data with a spatial reference.

[0191] 1. Collect social media data. Register Weibo and Douyin developer accounts, complete entity qualification verification, create applications to obtain core access credentials; study the platform API interface documentation, clarify interface endpoints, request parameters, and call quotas, construct a structured request parameter set, and use the Python requests library to initiate interface requests to obtain data such as text, publication time, and user location tags containing keywords such as "flood," "waterlogging," "water level," and "dam breach" within the research basin; at the same time, build a compliant crawler, limit the data usage and crawling scope, modify core parameters such as keywords and target links, crawl supplementary data not obtained through API interfaces, and save it in CSV format, including fields such as user nickname, publication time, post content, and location tags.

[0192] 2. Collect mobile LBS data. Sign cooperation agreements with local telecom operators, specifying data types (anonymous identifiers, latitude and longitude coordinates, timestamps, signal interaction status), spatiotemporal resolution (second-level timestamps, 10-meter-level coordinate accuracy), and coverage (the entire research basin). The operators will anonymize and desensitize the original user data in accordance with relevant personal information protection clauses and telecom industry data security standards, removing personally identifiable information to generate a standardized dataset containing core basic positioning data and population spatiotemporal distribution-derived data.

[0193] 3. Data Cleaning. Social media data is deduplicated (by post ID) and invalid content (posts without flood keywords) is removed. Core flood-related information is manually extracted from posts. Mobile LBS data is deduplicated for outliers, removing locations whose latitude and longitude exceed the study basin area or whose signals drift (coordinate fluctuations exceeding 500m / minute). Coordinate calibration is then completed.

[0194] 4. Unify the spatiotemporal benchmark. Let the social media dataset be... ,in For the total amount of social media data, each data sample Carry timestamp and space tags Mobile LBS data set is N_L represents the total amount of LBS data, and each sample Carry timestamp and precise geographic coordinates The WGS-84 coordinate system is used.

[0195] 5. Unify Time. Convert the timestamps of social media data to a Unix timestamp format consistent with LBS data, defining a time conversion function as follows. The converted social media data timestamps are Control conversion error .

[0196] 6. Time alignment. A dynamic time window method is used, defining the width of the time alignment window. It adaptively adjusts based on the data acquisition frequency; for any LBS data sample Constructing a temporal neighborhood Traverse social media data samples and filter those that meet the criteria. The samples constitute candidate spatiotemporal alignment pairs A time synchronization error correction model is introduced, and the time reference deviation is solved using the least squares method. The optimal value.

[0197] .

[0198] Where M is the initial number of candidate alignment pairs. ′、 These are the social media and LBS timestamps for the k-th pair of candidate data, respectively; the corrected social media timestamp is... The final time alignment criterion is Sample pairs that meet the conditions are included in the spatiotemporal alignment candidate set P.

[0199] 7. Spatial alignment. A geocoding mapping function is used. Convert social media text-based spatial tags into latitude and longitude rectangular areas in the WGS-84 coordinate system. Controlling mapping errors to ensure the area of ​​the region (Approximately 111km) 2 ); Calculate LBS coordinates With the region Spatial membership is determined using a distance-weighted model.

[0200] .

[0201] In the formula, The shortest distance from a point to a rectangular region , Set the distance attenuation coefficient (value 500m); set the membership threshold. , Spatial alignment is determined to be complete at the specified time and included in the final aligned dataset. ; For dataset Spatial error correction is performed, and the corrected LBS coordinates are as follows. .

[0202] in , This enables precise and unified spatial benchmarks for dual-source data.

[0203] Natural Language Processing (NLP) is used to process dual-source data, extract flood event elements, and construct a standardized flood event information package.

[0204] 1. Keyword extraction and optimization. Text collection. Improved TF-IDF weight formula: Domain-specific terms 1.2, non-domain-specific terms 1.0; (Merge duplicate phrases if the edit distance is ≤2).

[0205] 2. Sentiment-Risk Mapping. Mean Sentiment Tendency Credibility Risk Model Risk levels: [1,2) low risk, [2,3.5) medium risk, ≥3.5 high risk.

[0206] 3. Flood event element extraction and correlation analysis. Event element set. (Time, Location, Disaster Severity, Rescue Needs); Element Matching Degree , The elements are valid and constitute Event correlation model: , ; It is a strong correlation.

[0207] 4. Spatial clustering of mobile LBS data. The DBSCAN spatial clustering algorithm is used. Clustering region Profile coefficient evaluation formula: The average distance within the cluster. This embodiment represents the closest average distance between clusters. (Good clustering results); Spatial correlation model: , ; This is a highly correlated clustering region.

[0208] 5. Construct a standardized flood event information package. It adopts a three-level unique coding system: "administrative division code-date-time sequence number"; required fields (event identifier, occurrence time, occurrence location, hazard level, population density); optional fields (rescue needs, infrastructure impact); a unified JSON format and WGS-84 spatiotemporal reference; and the addition of source tracing identifiers. A four-dimensional structured information domain is constructed and encapsulated into a standardized information package.

[0209] Risk level thresholds are obtained through model simulation.

[0210] 1. The SWAT model was selected as the hydrological simulation tool. Based on ArcGIS 10.8 software, the watershed was divided into 12 sub-watersheds and 86 hydrological response units (HRUs) according to DEM topographic data. The underlying surface data such as topography, soil, and land use were integrated to simulate the complete hydrological process of "rainfall-interception-infiltration-runoff generation-confluence" in the watershed. The core quantitative formula of the runoff generation process was used.

[0211] .

[0212] 2. Parameter Correction. Import the standardized flood event information package into the ArcGIS platform, and match the corresponding sub-basin and HRU information through spatial overlay analysis; prioritize the use of first-level validity information, compare the measured and simulated values ​​of parameters, and quantify the adjustment basis through the deviation rate δ.

[0213] .

[0214] Set the δ threshold to 10%, when At the same time, the model parameters (such as the maximum retention capacity S of the watershed, infiltration coefficient, etc.) are adjusted iteratively based on the information package. After adjustment, the effectiveness level of the information package is updated, forming a closed-loop optimization. The spatial distribution changes of parameters can be viewed in real time through ArcGIS map to adapt to the spatial heterogeneity requirements. In this embodiment, the model parameter deviation rate is finally controlled within 8%, and the simulation accuracy meets the early warning requirements.

[0215] 3. Construct a risk indicator system. Based on the "ontology-impact-loss" chain, and combined with the data mentioned above, construct a three-level system, with the level classification aligned with emergency response.

[0216] The comprehensive flood risk assessment (R) is quantified using the following formula: ( (Determined through a combined weighting method). Taking the comprehensive risk index R (range 0-1) as the core, combined with the model inundation simulation results, information package personnel and asset data, and referring to the "Flood Disaster Risk Assessment Specification" (SL 483-2010), the risk level threshold was determined after demonstration by three experts in the field of hydrology and water resources.

[0217] Establish a mechanism for real-time collection and regular updating of multi-source data, and provide tiered early warnings and push notifications.

[0218] 1. Extract core related elements: Based on the risk assessment report, extract 5 types of core elements, namely basic risk information, risk level information, risk impact information, handling guidance information, and warning period information.

[0219] 2. Adapt and adjust channels. SMS messages should be strictly limited to 100 characters, prioritizing the retention of the warning title, risk level, core handling guidelines, and warning period; emails or announcements should be supplemented with complete quantitative reasoning, flooding simulation results, expert opinions, and other detailed content to ensure information integrity and traceability.

[0220] 3. Establish an early warning push mechanism and adopt a tiered push processing approach. Establish a multi-source data real-time collection and regular update mechanism to continuously collect social media data, mobile LBS data, and actual flood measurement data. Combine this with the early warning push effectiveness (reach rate, response rate) to dynamically adjust risk level thresholds, push channel weights, and update frequency, thereby continuously optimizing the tiered early warning push mechanism and improving early warning effectiveness.

[0221] The outstanding beneficial effects of this application are further verified by the technical features of the above embodiments, as detailed below.

[0222] 1. Wider flood sensing coverage and significantly improved real-time performance: By incorporating the technical features of multi-source data from social media and mobile LBS, the limitations of sparse traditional monitoring stations are overcome, achieving full coverage of the entire basin without blind spots; data is updated in seconds, providing early warnings 5 ​​to 30 minutes earlier than traditional stations, significantly improving early detection capabilities and effectively solving the problem of delayed warnings.

[0223] 2. More accurate and complete flood event identification: Improved technical features of TF-IDF, BERT entity recognition, and sentiment-risk mapping can automatically extract four core elements: time, location, severity, and demand. Verified by examples, the accuracy of flood event identification is ≥90%, and the consistency between the hazard level classification and the actual disaster situation is ≥85%, solving the problems of inaccurate flood event identification and missing elements.

[0224] 3. Precise perception of population dynamics and more scientific risk assessment: Combining the technical features of spatial clustering and flood event-population association models, high-risk cluster areas can be accurately located with a population exposure error of ≤15%, providing a reliable basis for personnel evacuation and rescue, and solving the problems of unknown population dynamics and insufficient targeted rescue.

[0225] 4. High accuracy of spatiotemporal fusion of multi-source data: Through the technical features of unified spatiotemporal benchmarks, dynamic time windows and spatial membership judgment, the time alignment error is ≤1s and the spatial matching error is ≤500m, providing a high-quality unified data foundation for subsequent model simulation and risk assessment, and solving the problem of insufficient fusion of multi-source data.

[0226] 5. Significantly improved flood simulation accuracy: By embedding socially-aware information packages into SWAT and iteratively correcting parameters, the model parameter deviation rate is controlled within 8%, and the simulation errors of inundation range and depth are reduced by 10% to 25% compared with traditional methods, thus improving the reliability of flood simulation.

[0227] 6. Risk assessment is quantifiable, reproducible, and verifiable: The technical features of a three-dimensional indicator system, combined weights, and clear thresholds, with a comprehensive risk index R∈[0,1], and fully quantifiable and unambiguous level classification, can be directly linked to emergency response activation standards, solving the problems of vague and unverifiable risk assessment.

[0228] 7. Highly targeted early warnings and significantly improved public response efficiency: The technical features of tiered and categorized push notifications and differentiated releases based on location / channel / risk ensure that the reach rate in high-risk areas is close to 100%, the overload of invalid information is reduced by more than 60%, and the speed of public evacuation response is increased by more than 30%, solving the problems of "one-size-fits-all" early warnings and low response efficiency.

[0229] 8. The early warning system is scalable, iterative, and highly adaptable to engineering applications: It features standardized, modular, and interface-based design throughout the entire process, and can be directly adapted to different river basins, operators, and early warning platforms. It can be implemented by those skilled in the art without creative effort and has extremely strong engineering application value.

[0230] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 4 As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database is used for flood early warning methods based on multi-source social sensing and swarm intelligence. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements flood early warning based on multi-source social sensing and swarm intelligence.

[0231] Those skilled in the art will understand that Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0232] In one exemplary embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0233] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0234] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0235] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).

[0236] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0237] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0238] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A flood early warning method based on multi-source social perception and swarm intelligence, characterized in that, include: Collect and preprocess social media data and mobile LBS data to construct a spatiotemporally aligned multi-source fusion dataset; Based on the multi-source fusion dataset, standardized flood event information packages are generated using NLP and spatial clustering algorithms; The standardized flood event information package is input into the SWAT hydrological model to simulate floods, construct a three-dimensional risk index system, and classify risk levels. Based on the aforementioned risk information, the user's real-time location, and the push channel, a tiered and categorized early warning system will be implemented.

2. The flood early warning method based on multi-source social perception and swarm intelligence according to claim 1, characterized in that, The process of collecting and preprocessing social media data and mobile LBS data to construct a spatiotemporally aligned multi-source fusion dataset includes: The social media data is obtained through API interfaces or web crawlers; the mobile LBS data includes: location coordinate data, timestamp information, movement trajectory data, and context information; Convert the timestamps of the social media data to Unix timestamp format; By using a dynamic time window method combined with time error, the converted social media data is synchronized and aligned with the time of mobile LBS data collection. By using geocoding mapping and spatial membership, the transformed social media data is spatially synchronized and aligned with mobile LBS data collection. Based on time-synchronized and spatially synchronized social media data and mobile LBS data, a spatiotemporally aligned multi-source fusion dataset is constructed.

3. The flood early warning method based on multi-source social perception and swarm intelligence according to claim 1, characterized in that, The process of generating standardized flood event information packages based on the multi-source fusion dataset using NLP and spatial clustering algorithms includes: Based on the improved TF-IDF algorithm in NLP, the matching coefficient of the core dictionary in the flood domain is introduced to optimize the weight of text keywords in the social media data and determine the optimized keywords. Based on the BERT entity recognition model in NLP, flood event elements are extracted according to the optimized keywords; The mobile phone LBS data is processed based on the spatial clustering algorithm to obtain the population gathering areas within the watershed; Based on the flood event elements and the densely populated areas within the basin, a spatial correlation model between flood events and densely populated areas is constructed to generate a standardized flood event information package; the flood event elements include time, location, disaster severity, and rescue needs.

4. The flood early warning method based on multi-source social perception and swarm intelligence according to claim 3, characterized in that, The spatial association model for: ; in, For spatial matching degree weight, As a weight for the risk level, =0.6, =0.4; For flood events Location of occurrence With densely populated areas within the watershed Spatial matching degree, The number of users in a densely populated area. This represents the total number of users within the watershed. The danger level of the flood event.

5. The flood early warning method based on multi-source social perception and swarm intelligence according to claim 1, characterized in that, The process of inputting the standardized flood event information package into the SWAT hydrological model for flood simulation, constructing a three-dimensional risk index system, and classifying risk levels specifically includes: The standardized flood event information package is input into the SWAT hydrological model, the model parameters and boundary conditions are corrected, the complete flood process within the basin is simulated, hydrological elements are output, and a three-dimensional risk indicator system is constructed. The complete flood process within the basin includes: rainfall, interception, infiltration, runoff generation, and confluence. The hydrological elements include: inundation range, inundation depth, and duration. The three-dimensional risk indicator system includes: hydrological inundation characteristics, risk of casualties, and risk of economic loss. Based on the aforementioned three-dimensional risk indicator system, the comprehensive risk index is calculated using the combined weighting method. Based on the comprehensive risk index, combined with industry standards and expert evaluation, risk levels are classified; these risk levels include low risk, medium risk, and high risk; the comprehensive risk index R is: ; Wherein, α is the hydrological inundation characteristic weight, β is the casualty risk weight, and γ is the economic loss risk weight. α+β+γ=1. α, β, and γ are determined by a combination weighting method combining the AHP analytic hierarchy process and the entropy weighting method. As a criterion for hydrological inundation characteristics, For the risk of casualties, This is a criterion for economic loss risk.

6. The flood early warning method based on multi-source social perception and swarm intelligence according to claim 5, characterized in that, The risk level is determined based on the comprehensive risk index, combined with industry standards and expert evaluation, specifically including: When the comprehensive risk index R ∈ [0, 0.3) and simultaneously meets the first condition in the industry standard, the current risk level is determined to be low risk level after expert evaluation; the first condition includes the simulated inundation depth < 0.5m and the relative inundation range. <10%, Population Risk Index <0.2 and estimated total economic loss <5%; When the comprehensive risk index R ∈ [0.3, 0.7) and simultaneously meets the second condition in the industry standard, the current risk level is determined to be medium risk level after expert evaluation; the second condition includes the simulated inundation depth of 0.5–1.5 m and the relative inundation range. Population risk index and estimated total economic losses It accounts for 5% to 20%; When the comprehensive risk index R ∈ [0.7, 1] and simultaneously meets the third condition in the industry standard, the current risk level is determined to be high risk level after expert evaluation; the third condition includes the simulated inundation depth ≥ 1.5m and the relative inundation range ≥30%, Population Risk Index ≥0.6 and estimated total economic loss The proportion is ≥20%.

7. The flood early warning method based on multi-source social perception and swarm intelligence according to claim 1, characterized in that, The step of implementing tiered and categorized early warnings based on the risk information, the user's real-time location, and the push channel specifically includes: Based on the risk information and the user's real-time location, differentiated early warning information is generated; the risk information includes: basic risk information, risk level information, and risk impact information; The content length of the push channels is adjusted according to the differentiated early warning information. At the same time, a hierarchical push mechanism is established based on the comprehensive risk index, which includes high-risk early warning and medium- and low-risk routine coverage, and hierarchical and classified early warning is implemented.

8. A flood early warning system based on multi-source social perception and swarm intelligence, characterized in that, The flood early warning method based on multi-source social perception and swarm intelligence as described in any one of claims 1-7 includes: The multi-source fusion dataset construction module collects and preprocesses social media data and mobile LBS data to construct a spatiotemporally aligned multi-source fusion dataset. The standardized flood event information package determination module generates standardized flood event information packages based on the multi-source fusion dataset using NLP and spatial clustering algorithms; The risk level classification module inputs the standardized flood event information package into the SWAT hydrological model to simulate floods, constructs a three-dimensional risk indicator system, and classifies risk levels. The early warning push module performs tiered and categorized early warnings based on the risk level, user location, and push channel.

9. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the flood early warning method based on multi-source social perception and swarm intelligence as described in any one of claims 1-7.

10. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the flood early warning method based on multi-source social perception and swarm intelligence as described in any one of claims 1-7.