Forest disaster risk assessment method based on machine learning
By constructing a database of forest disaster evolution trajectories and machine learning algorithms, combined with recommendation systems and natural language processing, the problem of insufficient utilization of time series features in traditional forest disaster risk assessment has been solved, achieving efficient and accurate risk identification and information dissemination.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-19
- Publication Date
- 2026-03-27
AI Technical Summary
Traditional forest disaster risk assessment methods cannot effectively utilize time series characteristics, make it difficult to accurately depict disaster development trends, ignore the coupling characteristics of environmental factors, resulting in biased risk identification results, failure to provide differentiated information, and affecting the timeliness of disaster prevention and mitigation work.
By constructing a database of forest disaster evolution trajectories, machine learning algorithms are used to analyze the similarity between environmental characteristics and historical disaster trajectories, disaster type inference and clustering are performed, and targeted risk information is generated by combining recommendation systems and natural language processing, and the assessment results are dynamically adjusted.
It improves the accuracy and interpretability of disaster risk identification, enhances the accuracy and response efficiency of risk assessment, and meets the needs of different management roles.
Smart Images

Figure CN121745686A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of forest disaster risk assessment technology, and specifically to a forest disaster risk assessment method based on machine learning. Background Technology
[0002] In forest resource management and ecological protection activities, disaster risk assessment is a crucial step in ensuring forest safety and enhancing disaster response capabilities. With climate warming, frequent extreme weather events, and changes in forest structure, forest disasters exhibit greater uncertainty and complexity, with their evolution paths, impact ranges, and triggering factors all displaying significant dynamic characteristics. Traditional forest disaster risk assessment methods largely rely on human experience, static indicators, or environmental information from single time segments, making it difficult to accurately depict disaster development trends and potential risks, and significantly limiting their predictive power and identification accuracy.
[0003] In existing technologies, the time-series characteristics of disasters are generally ignored in risk assessment. Historical disaster events often contain patterns of disaster evolution, such as the ignition and spread of fires, and the local outbreak and widespread spread of pests. Single static monitoring is insufficient to reveal the dynamic trends of these disasters. Furthermore, due to the highly coupled nature of forest environmental factors, such as temperature, humidity, wind field, vegetation type, slope, and moisture conditions, which all influence disaster triggering, traditional methods treat environmental characteristics in isolation, failing to effectively correlate them with disaster trajectories, leading to biases in risk identification results. With the increasing specialization of forest area regulatory bodies, different managers have significantly different focuses in risk assessment. For example, disaster command departments focus more on disaster intensity and spread direction, while forest farm managers focus more on local risk points and resource losses; ecological protection departments focus more on the extent of damage to ecological structures. However, existing risk assessment results are generally presented in a uniform format, making it difficult to provide differentiated risk information for different roles, resulting in low efficiency in understanding and high cost of use in practical applications. Furthermore, traditional risk assessment reports often present a large number of quantitative indicators, environmental data, and disaster descriptions in a stacked manner, lacking a hierarchical distinction between high-risk and low-risk characteristics, and also lacking the extraction and optimization of core risk information, making it difficult to identify key risk points in a timely manner. In scenarios where disasters evolve rapidly, if the information presentation method cannot support the rapid interpretation of core content, it will directly affect the timeliness of disaster prevention and mitigation work. Summary of the Invention
[0004] The purpose of this invention is to provide a forest disaster risk assessment method based on machine learning, thereby solving the problems existing in the prior art.
[0005] To achieve the above objectives, the present invention provides the following technical solution: a forest disaster risk assessment method based on machine learning, comprising: S1, constructing a forest disaster evolution trajectory database, modeling historical disaster events according to time series, analyzing the similarity between current forest environmental characteristics and historical disaster trajectories, inferring disaster types and key concerns, clustering disaster characteristics to form a risk category set; S2, based on the risk category set, initially splitting the quantitative indicators and descriptive content in the assessment information, wherein if the quantitative indicators exceed a preset threshold, they are marked as high-risk priority parts, otherwise they are classified as low-risk priority parts, to determine the boundary division of risk content; S3, using a recommendation system to match the priority content after boundary division with the disaster risk category set, and obtaining the risk information subset with the highest matching degree by comparing disaster concern preferences and content relevance; S4, extracting core quantitative indicators and text descriptions from the risk information subset with the highest matching degree, using natural language processing to perform semantic analysis and summary generation on these extracted contents, judging whether they meet the depth requirements of each disaster risk category set, if not, outputting a quantitative structured version, otherwise adjusting the summary structure to obtain a targeted expression form.
[0006] Preferably, step S1 includes: obtaining historical event sequences from a pre-established disaster trajectory database; arranging temperature change data, humidity change data, and wind speed change data in the historical event sequences in chronological order to generate trajectory data containing the evolution process of environmental parameters; using a dynamic time warping algorithm to calculate the distance between the current forest area environmental characteristics and the trajectory data; the dynamic time warping algorithm constructs a distance matrix to match time series of different lengths to obtain a matching degree value reflecting the degree of similarity; inferring possible disaster types based on the matching degree value; if the matching degree value exceeds a preset threshold, it is judged as a high-risk disaster type, and the corresponding coordinates of key environmental monitoring areas are obtained; extracting disaster feature parameters including occurrence time, duration, and impact range for disaster types; grouping and clustering by calculating the Euclidean distance between disaster feature parameters to form disaster category combinations containing different risk levels; determining the disaster development trend direction by the risk level distribution in the disaster category combinations, and generating corresponding risk warning level identifiers.
[0007] Preferably, step S2 includes obtaining assessment information from a preset set of risk categories, performing preliminary splitting of quantitative indicators and descriptive content in the assessment information, marking the quantitative indicators as high-risk priority parts if they exceed a preset threshold, and marking them as low-risk priority parts otherwise; for the high-risk priority parts and low-risk priority parts, the marking results are fused using a vegetation density monitoring method, and the boundary division is determined by comparing the correspondence between vegetation coverage and risk markings, thus determining the boundary division between high-risk content and low-risk content.
[0008] Preferably, step S3 includes obtaining high-priority and low-priority content after boundary division from a preset disaster risk category set, calculating initial matching values through content relevance comparison, and obtaining a preliminary matching matrix; for the preliminary matching matrix, adjusting the matching weights by integrating disaster concern preferences, and determining that if the preference exceeds a preset threshold, the corresponding weight is increased to determine a weighted matching matrix; obtaining the weighted matching matrix, and obtaining a content relevance ranking list by comparing the matching values within the matrix with relevance indicators; extracting the part with the highest matching value from the content relevance ranking list to generate a disaster risk matching sequence, and determining that if the sequence covers a risk category, it is marked as a priority sequence to determine a priority risk sequence; for the priority risk sequence, obtaining the risk information subset with the highest matching degree by integrating the matching degree within the sequence.
[0009] Preferably, step S4 includes extracting core quantitative indicators and text descriptions from the risk information subset with the highest matching degree, obtaining a quantitative set after indicator weight fusion and a text set after description completeness assessment; performing semantic association parsing on the quantitative set and text set using natural language processing to obtain the semantic feature set required for generating summary content; generating summary content based on the semantic feature set, and determining the optimized summary structure with targeted expression if the summary content meets the preset depth requirements and category coverage verification; if the summary content does not meet the preset depth requirements, outputting a quantitative structured version and determining the targeted expression form under risk category matching.
[0010] Preferably, the method further includes S5: obtaining the latest risk concern preferences based on the targeted expression form and combined with the dynamically changing historical feedback data of forest area user roles; if the preferences change, recalculating the matching degree; otherwise, maintaining the original matching degree, and determining the final risk information distribution version. Specifically, this includes obtaining the latest risk concern preferences from the historical feedback data of forest area user roles and integrating the preferences with dynamically changing forest area environmental monitoring data to obtain an updated preference set; for the updated preference set, if the preferences change, recalculating the matching degree, using the K-means algorithm to cluster the preference features in the set, where the input is the preference feature vector and the output is the cluster center, and obtaining the adjusted matching degree value by calculating the distance between the cluster center and the preset benchmark; if the preferences do not change, maintaining the original matching degree corresponding to the updated preference set, integrating the targeted expression form through the matching degree value to obtain an optimized risk information structure; and determining the final risk information distribution version based on the optimized risk information structure.
[0011] Preferably, it also includes S6: through the final risk information distribution version, reorganizing the overall structure of the assessment information, placing high-risk priority content at the front end and highlighting it, to obtain the reorganized complete risk assessment information package. Specifically, this includes obtaining the integrated forest area alarm data obtained through data monitoring through the final risk information distribution version, reorganizing the overall structure of the assessment information, placing high-risk priority content at the front end and highlighting it, to obtain the reorganized information.
[0012] Preferably, step S6 further includes determining, based on the reorganization information, whether to fuse the preference fusion data obtained through preference fusion if the environment changes, and thus determine the final complete risk assessment information package.
[0013] Preferably, it also includes S7, which involves performing risk information tiered distribution based on the recombined complete risk assessment information package and the disaster type requirements of the user role. If the path selection is insufficient to meet the requirements, it is rematched and adjusted to obtain accurate tiered results. Specifically, this includes performing risk information tiered distribution based on the disaster type requirements of the user role through the recombined complete risk assessment information package to obtain preliminary tiered results; and determining whether the path selection is insufficient to meet the requirements based on the preliminary tiered results, and determining the path adjustment requirements.
[0014] Preferably, step S7 further includes obtaining alarm integration data and re-matching and adjusting it if there is a need for path adjustment, to obtain adjusted classification information; and using the adjusted classification information, integrating environmental change data and high-risk priority content to determine the classification result for accurate push.
[0015] As can be seen from the above technical solution, the present invention has the following beneficial effects: This machine learning-based forest disaster risk assessment method constructs a forest disaster evolution trajectory database and employs a trajectory similarity learning algorithm to achieve dynamic correlation analysis between historical disaster events and the current forest environment. Based on evolutionary patterns, it can predict potential disaster types and their development trends in advance, effectively overcoming the problem of insufficient predictive ability caused by the inability of traditional methods to utilize time series features. By using disaster sensitivity clustering to structurally express risk features, the method achieves more accurate disaster feature classification and improves the refinement of risk identification. Through hierarchical decomposition of quantitative indicators and descriptive content, this invention can highlight high-risk features, suppress irrelevant information, and improve the readability of key risk points. Furthermore, this invention introduces a recommendation system and natural language processing technology to match, filter, and semantically refine disaster risk information. It can automatically generate differentiated risk information subsets based on the risk focus of different management roles, significantly improving the targeting and effectiveness of information delivery. Combined with a dynamic update mechanism, this invention can adaptively adjust risk focus preferences based on the historical usage behavior of management personnel, making the risk assessment results more consistent with actual application scenarios. Ultimately, by reorganizing and distributing risk information in a hierarchical manner, this invention can present disaster risk status with higher information density and clearer structure, thereby improving the speed of risk identification, reducing the burden of understanding caused by information overload, and significantly enhancing the accuracy, interpretability, and response efficiency of disaster risk assessment in forest areas. Attached Figure Description
[0016] Figure 1 This is a flowchart of the forest disaster risk assessment method based on machine learning of the present invention. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] like Figure 1As shown, this invention provides a technical solution: a forest disaster risk assessment method based on machine learning, comprising: S1, constructing a forest disaster evolution trajectory database, modeling historical disaster events according to time series, analyzing the similarity between current forest environmental characteristics and historical disaster trajectories, inferring disaster types and key concerns, clustering disaster characteristics to form a risk category set; S2, based on the risk category set, initially splitting the quantitative indicators and descriptive content in the assessment information, wherein if the quantitative indicators exceed a preset threshold, they are marked as high-risk priority parts, otherwise they are classified as low-risk priority parts, to determine the boundary division of risk content; S3, using a recommendation system to match the priority content after boundary division with the disaster risk category set, and obtaining the risk information subset with the highest matching degree by comparing disaster concern preferences and content relevance; S4, extracting core quantitative indicators and text descriptions from the risk information subset with the highest matching degree. The process involves: S1) Using natural language processing to semantically analyze and summarize the extracted content, determining whether it meets the depth requirements of each disaster risk category set. If not, a quantitative and structured version is output; otherwise, the summary structure is adjusted to obtain a targeted expression. S5) Based on the targeted expression, the latest risk concern preferences are obtained by combining dynamically changing historical feedback data of forest area users. If preferences change, the matching degree is recalculated; otherwise, the original matching degree is maintained to determine the final risk information distribution version. S6) Using the final risk information distribution version, the overall structure of the assessment information is reorganized, with high-risk priority content placed at the front end and highlighted to obtain a reorganized complete risk assessment information package. S7) Based on the reorganized complete risk assessment information package, risk information is distributed hierarchically according to the disaster type needs of user roles. If the path selection is insufficient to meet the needs, it is readjusted to obtain a precise hierarchical push result.
[0019] In the above implementation, by constructing a disaster evolution trajectory database and employing time series modeling, the dynamic evolution trend of forest disasters can be expressed as a function of climate, vegetation, topography, and human factors. Based on path similarity calculation (such as DTW dynamic time warping or a sequence similarity model based on encoder structure), the commonalities and differences between current forest disasters and historical disasters can be identified, thereby inferring potential disaster types and characteristics. Further, unsupervised clustering techniques can be used to cluster disaster features, extracting representative risk categories and providing a structured basis for subsequent matching steps. During the risk content decomposition process, thresholds can be used to prioritize quantitative indicators, enabling the assessment information to obtain a preliminary structured expression and providing a controllable input dimension for the recommendation system calculation. The recommendation system uses collaborative filtering or vector retrieval to match priority content with the risk category set, and achieves optimal risk subset selection by measuring the similarity between the attention preference vector and the content expression vector. The natural language processing module performs semantic analysis on the matching results. Through keyword extraction, entity recognition, dependency structure analysis, and generative model summarization, it achieves semantic compression and reorganization of risky content, and determines whether the content meets the presentation standards based on the depth requirements of the risk category set. With the dynamic participation of user role feedback, the system can achieve real-time personalized adjustment of risky content by updating preference vectors, ensuring that the final distribution version meets the actual needs of users. Finally, the system restructures the assessment information based on the generated distribution version, prioritizing and highlighting it according to risk priority, giving high-risk content a more prominent display position. The tiered distribution module uses multi-path calculation or tiered recommendation algorithms to push content according to the needs of different user roles, and ensures the accuracy and completeness of the push results through path deficiency detection and re-matching mechanisms.
[0020] In the above implementation, the construction of a disaster evolution trajectory database accurately reflects the patterns of disaster occurrence over time, improving the reliability of risk prediction. Prioritization using quantitative thresholds reduces data processing complexity, making subsequent recommendation matching more efficient. Introducing a recommendation system for content relevance analysis significantly improves the accuracy of risk assessment information filtering, making the output more targeted. Using natural language processing technology to structure and optimize the expression of risk information enhances clarity and comprehensibility, helping users quickly identify key risk points. A dynamic adjustment mechanism based on user role feedback enables the risk assessment model to be adaptive, continuously adapting to the changing needs of different user types. Reconstructing the assessment content and highlighting it based on priority enhances the readability of the overall risk assessment package, prioritizing high-risk information and improving the practical guiding effectiveness of early warnings. A tiered distribution mechanism ensures that different user roles receive content with risk levels more aligned with their responsibilities and needs, thereby improving information delivery efficiency and disaster response capabilities.
[0021] S1 includes obtaining historical event sequences from a pre-established disaster trajectory database, arranging temperature change data, humidity change data, and wind speed change data in the historical event sequences in chronological order to generate trajectory data containing the evolution process of environmental parameters; using a dynamic time warping algorithm to calculate the distance between the current forest area environmental characteristics and the trajectory data, the dynamic time warping algorithm constructs a distance matrix to match time series of different lengths to obtain a matching degree value reflecting the degree of similarity; inferring possible disaster types based on the matching degree value, if the matching degree value exceeds a preset threshold, it is judged as a high-risk disaster type, and obtaining the coordinates of the corresponding key environmental monitoring area; extracting disaster feature parameters including occurrence time, duration, and impact range for disaster types, and grouping and clustering by calculating the Euclidean distance between disaster feature parameters to form disaster category combinations containing different risk levels; determining the disaster development trend direction by the risk level distribution in the disaster category combinations, and generating corresponding risk warning level identifiers.
[0022] In this embodiment, firstly, in a pre-established disaster trajectory database, each historical event sequence corresponds to a forest disaster process or a long-term monitoring process that has already occurred. The sequence records temperature, humidity, and wind speed monitoring values for the forest area hourly or at fixed time intervals. Simultaneously, it saves the corresponding disaster type, occurrence time, duration, impact range, and the coordinates and risk level information of key environmental monitoring areas at the time of the disaster. In this step, temperature change data refers to a continuous temperature value sequence at the same monitoring point; humidity change data refers to a continuous humidity value sequence at the same monitoring point; wind speed change data refers to a continuous wind speed value sequence at the same monitoring point; occurrence time refers to the specific date and time when the disaster began to be recorded; duration refers to the duration from the start to the end of the disaster; and impact range refers to the actual forest area affected by the disaster or the area affected by the disaster. The number of grids and the coordinates of key environmental monitoring areas refer to the coordinate positions of the areas within the aforementioned impact range in a pre-established unified spatial coordinate system. When implementing this method, historical event sequences related to the target forest area are selected from the disaster trajectory database. Each event sequence is sorted from earliest to latest according to its recording time, and the corresponding temperature change data, humidity change data, and wind speed change data are arranged sequentially in the same time order, forming an environmental parameter evolution trajectory data arranged in ascending order of time, so that each time location corresponds to a set of temperature, humidity, and wind speed values. Simultaneously, from the current online monitoring data of the forest area, temperature, humidity, and wind speed monitoring values for the most recent period are extracted at the same time intervals as the historical data, and arranged from earliest to latest according to the monitoring time, constituting the current forest area environmental characteristic sequence. This sequence is structurally consistent with the historical trajectory data, ensuring comparability of each time location during subsequent matching calculations.Then, a dynamic time warping algorithm is used to compare the current forest area environmental feature sequence with each historical trajectory data. Specifically, a distance matrix is first established, where the row order corresponds to the time position of the current forest area environmental feature sequence, and the column order corresponds to the time position of a specific historical trajectory data point. For any cell in the matrix, the difference between the current temperature value at the corresponding time point in the row and the historical temperature value at the corresponding time point in the column is calculated. Then, the difference between the current humidity value and the historical humidity value, as well as the difference between the current wind speed value and the historical wind speed value, are calculated. These three differences are then compared... The absolute values or squares of the differences are taken separately and then combined into a single difference value according to a preset method. This value serves as the local difference measure for that cell. Then, starting from the top-left cell of the matrix, the calculation is performed cell by cell in a left-to-right, top-to-bottom order. In each cell, the local difference value is added to the minimum cumulative difference value among the three adjacent cells to its left, top, and top-left, to obtain the cumulative difference value for the current cell. This process is repeated until the bottom-right cell of the matrix is calculated. The cumulative difference value of the bottom-right cell is the distance between the current forest area environmental feature sequence and the historical trajectory data. The smaller the distance value, the closer the current environmental evolution process is to the historical trajectory. In the program, this distance value is then converted into a matching degree value through a monotonically decreasing mapping rule. The larger the matching degree value, the higher the similarity. The matching degree value is stored as a dimensionless value and used for subsequent threshold judgment. The process for determining the preset threshold is as follows: During the system deployment phase, several historical event sequences confirmed to have occurred for the target disaster type are selected from the database. The environmental evolution process of these sequences before the disaster is kept consistent with the structure of the current forest area environmental characteristic sequence, and their matching degree distribution is calculated. At the same time, several monitoring sequences that did not occur under the same climatic background are selected, and their matching degree distribution is calculated. The matching degree values of the two types are sorted and their overlapping ranges are compared. The system designers provide specific percentage values for acceptable false alarm and missed alarm ratios. A fixed value is selected from the matching degree values that meet the upper limit requirements of the false alarm and missed alarm ratios as the preset threshold and written into the system configuration. It is not changed during the method's operation or is recalculated uniformly once within a preset period. When the matching degree value of the current forest area environmental characteristics and a certain historical trajectory data is greater than or equal to the preset threshold, the system judges the disaster type corresponding to the historical trajectory as the high-risk disaster type at the current moment, and directly reads the coordinates of the key environmental monitoring areas stored in the database for the historical trajectory as the spatial location that needs to be focused on monitoring.Next, for all historical event sequences identified as high-risk disaster types, the system sequentially reads the occurrence time, duration, and impact range of each event from the database. These three values are combined into a set of disaster characteristic parameters. To facilitate distance calculation, the occurrence time is converted to a continuous number line using calendar time, the duration is converted to a uniform time unit, and the impact range is converted to a uniform area unit or by the number of grid cells. Necessary linear scaling can be performed within the same dimension to ensure the three values are within comparable orders of magnitude. Then, Euclidean distance is used for grouping and clustering. Specifically, between any two high-risk disaster events, the distances are calculated separately... Calculate the differences in their occurrence time, duration, and impact range. Square each difference, sum them, and then take the square root of the sum to obtain the Euclidean distance between the two events. List the Euclidean distances between all pairs of events in a distance list. Set a fixed upper limit for grouping distances. When the Euclidean distance between an event and any existing event in a group is less than or equal to the upper limit of the group distance, the event is assigned to that group. When the Euclidean distance between the event and all existing events in a group is greater than the upper limit of the group distance, a new group is created for the event. This process is repeated for all events until all events are assigned to specific groups. Each group constitutes a disaster category combination. For each disaster category combination, the system assigns a risk level to the disaster category combination by statistically analyzing the proportions of high-risk, medium-risk, and low-risk events within the group, based on the risk level information recorded in the database of historical events within the group. When the proportion of high-risk events in the total number of events in the group reaches or exceeds the specific percentage value given in the system configuration, the category combination is marked as high-risk. When the proportion of medium-risk events reaches or exceeds the configured value but the proportion of high-risk events does not meet the high-risk condition, the category combination is marked as medium-risk. In all other cases, it is marked as low-risk. Subsequently, the system arranges all disaster category combinations in chronological order of occurrence or in chronological order of the spatial location of the affected area, and counts the changes in the number of each risk level category combination in different time slices or different spatial partitions. When the number of high-risk level category combinations increases over time or expands to a larger area spatially, the disaster development trend is determined to be upward or spreading. When the number of high-risk level category combinations remains stable and is mainly concentrated in a fixed spatial area, the disaster development trend is determined to be stable. When the number of high-risk and medium-risk level category combinations decreases continuously and the spatial range shrinks, the disaster development trend is determined to be weakening.Finally, the system generates a risk warning level identifier based on the highest risk level and the direction of disaster development trend. Several risk warning levels are pre-set during the system initialization phase, such as low, medium, and high levels, each with corresponding numerical and color identifiers. During actual operation, when there is a high-risk level category combination and the development trend is upward or spreading, the highest-level risk warning level identifier is directly output. When there is no high-risk level category combination but there is a medium-risk level category combination and the development trend is stable, a medium-level risk warning level identifier is output. When there is only a low-risk level category combination or the development trend is weakening, the lowest-level risk warning level identifier is output. This risk warning level identifier is stored and displayed together with the warning results as the decision-making basis for forest area managers.
[0023] S2 includes obtaining assessment information from a preset set of risk categories, performing preliminary splitting of quantitative indicators and descriptive content in the assessment information, marking the quantitative indicators as high-risk priority parts if they exceed a preset threshold, and marking them as low-risk priority parts otherwise; for the high-risk priority parts and low-risk priority parts, the labeling results are fused using a vegetation density monitoring method, and the boundary division is determined by comparing the correspondence between vegetation coverage and risk labels, thus determining the boundary division between high-risk content and low-risk content.
[0024] In this implementation, firstly, one or more risk categories corresponding to the current forest disaster type are read from a preset risk category set. Each risk category has a pre-stored list of quantitative indicator names, the high-risk value range and low-risk value range for each quantitative indicator, as well as high-risk and low-risk intervals for vegetation coverage. The risk category set is obtained and stored in a fixed manner based on historical disaster samples during the system deployment phase. When the system performs the assessment, it first generates assessment information. The assessment information uses a monitoring unit in the forest area as the basic unit. Each monitoring unit corresponds to one assessment record. Each assessment record contains several quantitative indicators and at least one descriptive paragraph. The quantitative indicators are monitoring items with specific values, including the average temperature, average humidity, average wind speed, vegetation coverage calculated using a unified algorithm within the same coordinate range, and the number of historical disasters related to the monitoring unit during the assessment period. The descriptive paragraph consists of written descriptions by monitoring or management personnel regarding the accumulation of combustibles on the ground, the intensity of human activities, and recent forest understory management practices. After reading an evaluation record, the system first categorizes all numerical fields in the record into the quantitative indicator set and all textual fields into the descriptive content set. At the same time, it records the monitoring unit number to which each quantitative indicator belongs and its field position in the evaluation record, thus completing the initial separation of quantitative indicators and descriptive content in the evaluation information. Subsequently, the system performs threshold judgment individually for each quantitative indicator in the quantitative indicator set. In this embodiment, the preset threshold is a fixed value determined individually for each quantitative indicator. The specific determination method is as follows: During the system deployment phase, all known high-risk and low-risk samples in the historical disaster data are statistically analyzed. For a certain quantitative indicator, the values of the indicator for all high-risk samples are sorted from smallest to largest, and the values of the indicator for all low-risk samples are sorted from smallest to largest. Then, the number of misclassified samples is calculated one by one at the possible boundary positions. For each candidate boundary value, the number of low-risk samples above the value and the number of high-risk samples below the value are counted. These two numbers are added together to obtain the total number of errors for the candidate boundary value. Finally, the candidate boundary value with the smallest total number of errors is selected as the preset threshold corresponding to the quantitative indicator, and the threshold is written into the risk category set in the form of a specific value. It is not changed during system operation.During the operation of the method, the system reads the value of each quantitative indicator of a monitoring unit and directly compares it with the preset threshold recorded in the risk category set for that indicator. When the indicator is a positive risk indicator, such as temperature, wind speed, or the number of historical disasters, if the value is greater than or equal to the preset threshold, the indicator is marked as a high-risk priority part; if the value is less than the preset threshold, the indicator is marked as a low-risk priority part. When the indicator is a negative risk indicator, such as humidity, if the value is less than or equal to the preset threshold, the indicator is marked as a high-risk priority part; if the value is greater than the preset threshold, the indicator is marked as a low-risk priority part. The system stores the marking result along with the original value of the indicator in the structured data of the monitoring unit, thereby completing the marking of the high-risk priority part and the low-risk priority part of all quantitative indicators. Subsequently, the system integrates the above-mentioned labeling results based on the vegetation density monitoring method. The specific implementation of the vegetation density monitoring method is as follows: taking the monitoring unit obtained by dividing the forest area as the smallest spatial unit, for each monitoring unit, the system counts the area of the surface covered by vegetation and the total area of the monitoring unit within a fixed time window, divides the vegetation coverage area value by the total area value to obtain a vegetation coverage rate value between 0 and 1, and stores the value in the vegetation coverage rate field of the corresponding evaluation record. In the vegetation density monitoring method, the vegetation coverage threshold used for boundary delineation is determined through historical samples during the deployment phase. Specifically, a set of monitoring unit samples that have experienced forest disasters and a set of monitoring unit samples that have not experienced disasters in the same region and season are selected. The vegetation coverage value is calculated for each sample, and the quantitative indicators in each sample are marked as high-risk or low-risk based on the aforementioned preset threshold. The ratio of the number of high-risk priority markers to the total number of quantitative indicators in each sample is counted. All samples are sorted by vegetation coverage value from smallest to largest. For each candidate vegetation coverage boundary position, samples at and above the position are uniformly regarded as high-risk content, and samples below the position are uniformly regarded as low-risk content. The number of samples that are actually high-risk but are classified as low-risk content and the number of samples that are actually low-risk but are classified as high-risk content are counted. These two numbers are added together to obtain the total number of errors corresponding to the candidate vegetation coverage boundary position. Finally, the vegetation coverage value with the smallest total number of errors is selected as the vegetation coverage threshold and stored in the risk category set, which is recorded as the boundary delineation threshold.In actual operation, for each monitoring unit, the high-risk priority and low-risk priority markers of each quantitative indicator within the unit and the vegetation coverage value are obtained. First, the number of quantitative indicators marked as high-risk priority in the monitoring unit is counted, and then divided by the total number of quantitative indicators in the monitoring unit to obtain a high-risk marker ratio value. Then, the vegetation coverage threshold and a high-risk marker ratio threshold of the corresponding record in the risk category set are read. The high-risk marker ratio threshold is determined in the deployment phase through a statistical method similar to the preset threshold, and is a fixed value between 0 and 1. When determining the boundary, if the vegetation coverage value of a monitoring unit is less than or equal to the vegetation coverage threshold, and the high-risk marker ratio value of the monitoring unit is greater than or equal to the high-risk marker ratio threshold, the system marks the overall assessment information corresponding to the monitoring unit as high-risk content. If the vegetation coverage value of a monitoring unit is greater than the vegetation coverage threshold, and the high-risk marker ratio value of the monitoring unit is less than the high-risk marker ratio threshold, the system marks the overall assessment information corresponding to the monitoring unit as low-risk content. For monitoring units that do not meet the above two conditions, the system considers them as monitoring units near the boundary. A comprehensive deviation value is calculated based on the difference between the deviation value and the vegetation coverage threshold and the difference between the deviation value and the high-risk marker ratio threshold. When the comprehensive deviation value is closer to the high-risk end, the monitoring unit is marked as high-risk content; when it is closer to the low-risk end, it is marked as low-risk content. In this way, the system provides clear high-risk or low-risk content markings at each monitoring unit level, thereby completing the boundary division of high-risk and low-risk content in space, and storing the division result in the form of a marker field in the corresponding assessment record.
[0025] S3 includes: obtaining high-priority and low-priority content after boundary division from a preset set of disaster risk categories; calculating initial matching values by comparing content relevance to obtain a preliminary matching matrix; adjusting matching weights based on disaster concern preferences for the preliminary matching matrix; increasing corresponding weights if the preference exceeds a preset threshold to determine a weighted matching matrix; obtaining a content relevance ranking list by comparing matching values within the matrix with relevance indicators; extracting the highest matching values from the content relevance ranking list to generate a disaster risk matching sequence; marking a sequence as a priority sequence if it covers a risk category to determine a priority risk sequence; and integrating the priority risk sequences by matching degree within the sequences to obtain the subset of risk information with the highest matching degree.
[0026] In this implementation, the system first reads all disaster risk categories from a preset disaster risk category set. Each disaster risk category has a unique number, category name, key environmental feature description, typical triggering condition description, and risk level description. Simultaneously, the system reads content that has already been marked with high and low priorities from the data structure output from the aforementioned boundary delineation step. High-priority content consists of quantitative indicators marked as high-risk and their associated descriptive text, while low-priority content consists of quantitative indicators marked as low-risk and their associated descriptive text. The system assigns a content number to each content record and establishes an internal content index table, storing the content number in a one-to-one correspondence with the corresponding monitoring unit, indicator name, and descriptive text location. Subsequently, the system splits each content record into several content segments. The splitting rule is as follows: when a piece of content contains multiple quantitative indicators and multiple descriptive text segments, they are combined according to indicator category and text segment. Each combination constitutes an independent content segment. Each content segment records its content number, priority marker, high-risk or low-risk marker, the name and value of the included quantitative indicator, and a list of included keywords. The system establishes a corresponding feature record for each disaster risk category in the disaster risk category set. The feature record includes risk characteristic words such as high temperature, low humidity, high wind, sparse vegetation, and frequent human activity corresponding to that category in historical statistics, as well as descriptions of the high and low status of environmental indicators related to these characteristics. The system defines an initial matching value as a non-negative score to quantify the correlation between a single content fragment and a single disaster risk category. The calculation process for the initial matching value is as follows: For any content fragment and any disaster risk category, the system reads the keyword lists of both, counts the number of risk characteristic words commonly contained in both, and multiplies this number by the keyword weight score set during the system deployment phase. Then, it checks whether the high and low status of the quantitative indicators contained in the content fragment matches the high and low status description in the feature record of the disaster risk category. For each matching item, a quantitative status score is added. The quantitative status score is given as a fixed value by the system designers during the deployment phase. Finally, the keyword weight score and the quantitative status score are added to obtain a definite initial matching value, which is stored in association with the content fragment number and the disaster risk category number. The system performs the above matching calculations on all content fragments and all disaster risk categories. The resulting initial matching values are then filled into a two-dimensional data table, with each cell storing the corresponding initial matching value. This two-dimensional data table constitutes the preliminary matching matrix. Next, the system performs preference fusion and weight adjustment on the preliminary matching matrix.The disaster attention preference parameter is a attention score assigned to each disaster risk category. This score is stored in the system as a decimal between 0 and 1. The determination process is as follows: the system counts the number of clicks, views, and manually marked as important by users for each disaster risk category over a historical usage period. These three counts are converted into a total score according to the proportions set during deployment. The total scores for all categories are then normalized to between 0 and 1 to obtain the attention score for each category. The preference threshold is a fixed value between 0 and 1 used to distinguish between high-attention and general-attention categories. The preference threshold is determined as follows: during deployment, the system sorts all disaster risk categories by attention scores from highest to lowest. System designers specify a percentage value based on management needs, for example, requiring the top 30% of categories to be considered important. The system then finds the attention score at that percentage position in the sorting results and writes this score as the preference threshold into the configuration file. The weighting factor is a fixed value greater than 1 used to amplify the matching value corresponding to high-attention categories. This weighting factor is directly set by the system designers according to the desired amplification level. The system performs the following steps for each disaster risk category in the initial matching matrix: It reads the attention score for that category and compares it with a preference threshold. When the attention score is greater than or equal to the preference threshold, the system multiplies all initial matching values in that column by a weighting factor to obtain a new weighted matching value. When the attention score is less than the preference threshold, the system keeps the initial matching value in that column unchanged or makes a slight adjustment according to a smaller adjustment coefficient set in the configuration. After adjusting all columns, the values in all cells of the matrix are the weighted matching values, and this matrix is the weighted matching matrix. Subsequently, the system generates a content relevance ranking list based on the weighted matching matrix. The relevance index consists of two numerical thresholds used to classify three levels: high relevance, medium relevance, and low relevance. These thresholds are designated as the high relevance threshold and the medium relevance threshold. The determination process for these thresholds is as follows: The system collects a large number of weighted matching value samples generated during the deployment and testing phase. These samples are sorted from largest to smallest. System designers provide two percentages based on the proportions of highly relevant and moderately relevant samples, according to their priorities. The system finds the matching values at corresponding positions in the sorted results, using the matching values at higher percentage positions as the medium relevance threshold and the matching values at lower percentage positions as the high relevance threshold, ensuring that the high relevance threshold is greater than the medium relevance threshold. The system iterates through each cell of the weighted matching matrix, sequentially extracting the weighted matching values and comparing them with the high and medium relevance thresholds. When a weighted matching value is greater than or equal to the high relevance threshold, the record is marked as a highly relevant record; when the weighted matching value is less than the high relevance threshold but greater than or equal to the medium relevance threshold, the record is marked as a moderately relevant record; and when the weighted matching value is less than the medium relevance threshold, the record is marked as a low relevance record.For each record, the system simultaneously saves the content fragment number, disaster risk category number, weighted matching value, and relevance level label. Then, the system sorts all records from largest to smallest weighted matching value. When weighted matching values are the same, records with high relevance are prioritized, followed by records with medium relevance, and finally records with low relevance. The sorted records are then output sequentially to form a content relevance ranking list. When generating a disaster risk matching sequence, the system first reads records sequentially from the ranking list, determining the number of records to extract according to the maximum number of records parameter set in the configuration file. The maximum number of records parameter is a fixed integer, set during the deployment phase based on system processing capacity and business needs. The system arranges the content fragment numbers and disaster risk category numbers from the first few records in record order to form a disaster risk matching sequence and counts the set of disaster risk categories involved in this sequence. The covered risk category parameter is the set of disaster risk categories that should be covered. This set is specified by management personnel during the deployment phase based on the actual situation of the forest area and legal requirements. It can be all disaster risk categories or disaster risk categories with a risk level of at least medium. The system compares the actual set of disaster risk categories appearing in the disaster risk matching sequence with the coverage risk category parameters. When the disaster risk matching sequence contains all categories in the coverage risk category parameters, the system marks the sequence as a priority sequence and records the mapping relationship between its corresponding content segment number and disaster risk category number as a priority risk sequence. When the disaster risk matching sequence does not cover all target categories, the system increases the number of extracted records, continues to read records from the sorted list and add them to the matching sequence, repeating the coverage check until the matching sequence contains all categories in the coverage risk category parameters or the sorted list is completely read. If the coverage condition is still not met after reading, the system considers the set of categories appearing in the current matching sequence as the actual coverable set and marks it as a priority sequence to ensure that the algorithm still outputs a definite result with limited data. Finally, the system generates the risk information subset with the highest matching degree based on the priority risk sequence. To this end, the system establishes a summary record for each content segment number appearing in the priority risk sequence, traverses each matching record in the priority risk sequence, reads the content segment number and the corresponding weighted matching value, and accumulates the weighted matching values of the same content segment under different disaster risk categories to obtain the comprehensive matching degree score of the content segment. The overall matching degree is a defined non-negative number; the larger the value, the more important the content segment is to the overall disaster risk assessment. The overall matching degree threshold is used to filter the most important content segments. Its determination process is as follows: during the system deployment phase, based on the experience results of historical assessment tasks, the overall matching degree samples are sorted from largest to smallest, and the overall matching degree that is located at a specified percentage position in the sort is selected as the overall matching degree threshold. This percentage is given by the system designers in the configuration file.The system sorts all content fragments by their overall matching score from highest to lowest, prioritizing fragments with a matching score greater than or equal to a threshold as candidate risk information. When the number of candidates exceeds the maximum subset size parameter set in the configuration, fragments are truncated sequentially from highest to lowest matching score until the maximum subset size parameter is reached. The maximum subset size parameter is a fixed integer set during the deployment phase. When the number of content fragments with a matching score greater than or equal to the threshold is insufficient, the system supplements fragments with higher matching scores according to the sorting order until the minimum subset size parameter is reached. The minimum subset size parameter is also a fixed integer set during the deployment phase. Finally, the system back-links the selected content fragments with their corresponding original evaluation information to obtain a set of structured and numbered risk information records. This set of records constitutes the risk information subset with the highest matching score.
[0027] S4 includes extracting core quantitative indicators and text descriptions from the risk information subset with the highest matching degree, obtaining a quantitative set after indicator weight fusion and a text set after description completeness assessment; performing semantic association parsing on the quantitative set and text set using natural language processing to obtain the semantic feature set required for generating summary content; generating summary content based on the semantic feature set, and determining the targeted expression optimization summary structure if the summary content meets the preset depth requirements and category coverage verification; if the summary content does not meet the preset depth requirements, outputting a quantitative structured version and determining the targeted expression form under risk category matching.
[0028] In this embodiment, the system first receives the risk information subset with the highest matching degree output from step S3. Each risk information record in this subset contains several quantitative indicator fields and several text description fields. During the system deployment phase, a fixed indicator importance weight has been pre-set for each quantitative indicator. Each weight is a value between 0 and 1, which is determined as follows: The system selects samples that have experienced actual disasters from the historical disaster sample database as positive samples and samples that have not experienced disasters but participated in the assessment as negative samples. The system counts the number of successful warnings when the values of each quantitative indicator are abnormal in the positive samples and the number of false alarms when the values of each indicator are abnormal in the negative samples. The net contribution count is obtained by subtracting the number of false alarms from the number of successful warnings. The net contribution count of all indicators is then scaled proportionally to between 0 and 1. The scaled values are used as the weights of each indicator and written into the configuration file, where they remain unchanged during system operation. For each record in the risk information subset, the system sequentially reads the names and values of all quantitative indicators in that record. Based on the name, it finds the corresponding weight value in the configuration file and filters the core quantitative indicators according to the weight lower limit parameter set during the deployment phase. The weight lower limit parameter is a fixed value between 0 and 1. By gradually increasing the lower limit on historical samples and observing the changes in the early warning accuracy and the number of indicators involved in the calculation, the lower limit value corresponding to ensuring that the early warning accuracy is not lower than the preset ratio and the number of indicators involved in the calculation does not exceed the preset upper limit is written as the weight lower limit into the configuration file. The system identifies each quantitative indicator with a weight greater than or equal to the lower limit of its weight as a core quantitative indicator. It then performs standardization on the original values of these core quantitative indicators. The minimum and maximum values required for standardization are also statistically determined during the deployment phase. Specifically, the minimum and maximum values for each quantitative indicator in the historical samples are statistically determined and fixed in the configuration file. During runtime, the original values in the current record are linearly converted to standardized values between 0 and 1 according to this range. Subsequently, the system multiplies the standardized value of each core quantitative indicator by its corresponding weight to obtain a weighted value. The system then sequentially combines the indicator's name, original value, weight, standardized value, and weighted value into a structured record. All structured records corresponding to the core quantitative indicators in the same risk information record are grouped into a quantitative set, which is the quantitative set after indicator weight fusion.Meanwhile, the system extracts all text description fields from the same risk information record, strictly segments each field according to punctuation, treats each complete sentence as a text fragment, and evaluates the completeness of each text fragment. This completeness evaluation requires four types of parameters: a set of description elements, an element weight table, a minimum word count threshold, and a completeness threshold. The set of description elements is a fixed list, including at least five types of elements: time, location, disaster type, impact range, and response measures. This list is determined and configured by the system designers after analyzing a large amount of historical warning text. The element weight table represents the weight scores for the five categories of elements mentioned above. Each element corresponds to a fixed score greater than 0. This score is determined by statistically analyzing whether texts containing this element in historical warning texts are more likely to be adopted by users. Elements with higher adoption rates have higher weights. The minimum word count threshold is a positive integer. By statistically analyzing the word count distribution of historical texts, the minimum word count that can cover most effective texts is selected as the threshold. The completeness threshold is a positive number. By calculating the text retention rate and information loss rate under different thresholds in historical texts, the score that simultaneously meets the preset requirements for the text retention rate and information loss rate is selected as the fixed threshold. When evaluating a single text fragment, the system sequentially checks whether it contains date or time words to meet the time element, whether it contains place names or coordinate descriptions to meet the location element, whether it contains specific disaster type names to meet the disaster type element, whether it contains the affected area range or object names to meet the impact range element, and whether it contains specific actions that have been taken or are recommended to meet the response measures element. For each element detected, the corresponding weight is added to the fragment's completeness score. At the same time, the total number of words in the fragment is counted. When the total number of words is less than the minimum word count threshold, a pre-set penalty score is deducted from the completeness score, and finally, the descriptive completeness score of the fragment is obtained. The system records the descriptive completeness scores of all text fragments for the same risk information, sorts them from largest to smallest, and selects fragments with scores greater than or equal to the completeness threshold to form a text set, discarding the remaining fragments, thus obtaining the text set after descriptive completeness evaluation.Subsequently, the system performs a semantic association parsing process using the quantitative set and text set as input. During the deployment phase, a set of semantic feature types and their triggering conditions are predefined. This set includes eight types: risk category features, risk level features, time range features, spatial location features, key quantitative indicator group features, risk cause features, development trend features, and recommended measures features. Each type corresponds to a fixed list of trigger words and numerical conditions. For example, the triggering conditions for risk category features include the appearance of a disaster name in the text and a certain type of indicator in the quantitative set consistently being in a high-risk range. The triggering conditions for risk level features include the weighted values of several core indicators in the quantitative set reaching a preset level threshold range. The triggering conditions for time range features include the appearance of explicit start and end time phrases in the text. The trigger word lists and level threshold ranges corresponding to each semantic feature type are obtained during the deployment phase by analyzing historical disaster records and early warning text statistics and written into the configuration file. The specific values of the level threshold ranges are determined by dividing the historical samples marked with risk levels into weighted value ranges, ensuring that the number of samples within each level range reaches a preset proportion. The system performs word segmentation, part-of-speech tagging, and syntactic analysis on each text segment in the text set, identifying time words, location words, disaster type words, trend words, and measure words. These words are then matched one by one with the trigger word list in the semantic feature type set. Simultaneously, the system obtains the weighted values of each core quantitative indicator from the quantitative set, assigns these weighted values to specific risk level labels based on level threshold ranges, and associates the corresponding indicator names mentioned in the text with these labels. Based on this, the system generates a semantic feature record for each risk information record. This record contains specific risk category identifiers, risk level identifiers, time range identifiers, spatial location identifiers, key quantitative indicator group identifiers, risk cause identifiers, development trend identifiers, and recommended measure identifiers, thus forming the semantic feature set required for generating the summary content. When generating the summary, the system reads the semantic feature records one by one in descending order of risk level in the semantic feature set. It uses a fixed sentence template to combine the time range, spatial location, risk category name, risk level description, specific values of several core quantitative indicators and their high and low status, risk cause description and development trend description into a complete sentence. Then, it combines sentences of the same category or adjacent levels into a paragraph in logical order, so that the generated summary conforms to the conventional format of early warning reports in terms of structure.After generation, the system checks the summary content according to the preset depth requirements. The depth requirements consist of two parts: the semantic element quantity requirement and the core indicator quantity requirement. The semantic element quantity requirement stipulates that the number of semantic feature types actually appearing in the summary content must not be less than the specified minimum number of types. This minimum number of types is set to an integer of not less than 5 in the configuration to ensure that the summary contains sufficiently comprehensive information. The core indicator quantity requirement stipulates that the number of core quantitative indicators explicitly mentioned in the summary content in text form must not be less than the preset minimum number of core indicators. This value is set by the system designers to an integer of not less than 3 according to the decision-making needs. The system calculates the number of semantic feature types used and core quantitative indicators mentioned in the summary text. If both statistical results are above a preset threshold, the system determines that the depth requirement is met. Simultaneously, the system performs category coverage verification. The parameter used for category coverage verification is a risk category coverage list, which consists of all risk category names appearing in the subset of risk information with the highest matching degree. The system searches the summary text, checking whether each risk category name in the list appears at least once in the summary and is accompanied by a corresponding risk level and quantitative indicator description. When all risk category names are covered and the depth requirement is met, the system marks this version of the summary content as a targeted, optimized summary structure. This structure is stored in a fixed format and used for subsequent output. If the system detects that any aspect of the summary content does not meet the depth requirement or that at least one risk category name fails the coverage verification, the system determines that the generated summary content is insufficient in terms of information depth or category coverage. In this case, the system does not output the text summary but instead directly generates a quantitative structured version based on the semantic feature set. The steps for generating a quantitative structured version are as follows: The system groups semantic feature records according to risk category names, sorts them from high to low risk levels within each group, and sorts them from large to small weighted values of core quantitative indicators within the same level. The risk category name, time range, spatial location, name and value of the main core quantitative indicators, corresponding risk level and development trend description in each semantic feature record are concatenated into a structured text entry in a fixed field order. The entries are output one by one to form a structured list, which is the targeted expression form under risk category matching.
[0029] S5 involves obtaining the latest risk concern preferences from historical feedback data of forest area users and integrating these preferences with dynamically changing forest area environmental monitoring data to obtain an updated preference set. For the updated preference set, if preferences have changed, the matching degree is recalculated. The K-means algorithm is used to cluster the preference features in the set, where the input is the preference feature vector and the output is the cluster center. The adjusted matching degree value is obtained by calculating the distance between the cluster center and a preset benchmark. If preferences have not changed, the original matching degree corresponding to the updated preference set is maintained. The matching degree value is used to integrate targeted expression forms to obtain an optimized risk information structure. Based on the optimized risk information structure, the final risk information distribution version is determined.
[0030] In this implementation, the latest risk concern preferences are first obtained from the historical feedback data of forest area user roles. This historical feedback data includes the number of views, average reading time, number of confirmation responses to warning results, number of negative responses to false alarms, and number of times different presentation formats were selected, all recorded by user role and disaster risk category. During the deployment phase, the system sets a fixed statistical period, such as daily or weekly, and uses the feedback data collected within this period as the data basis for the current period. When executing this step, the system calculates the total number of views, total reading time, total number of confirmation responses, and total number of negative responses for each user role and each disaster risk category within that period. The total reading time is then divided by the number of views to obtain the average reading time. When the number of views is 0, the average reading time is directly recorded as 0. Then, the system, based on the deployment phase... The set preference weight coefficient combination converts the above statistical results into a preference score. The preference weight coefficient combination is a set of fixed non-negative values, which are used to multiply by the number of views, average reading time, number of confirmation feedbacks, and number of negative feedbacks, respectively. The confirmation feedback coefficient and the number of views coefficient are generally set to be large, while the negative feedback coefficient is set to a negative value or used for deduction. The system selects a set of coefficient combinations that maximizes the correct warning ratio and keeps the false alarm ratio below the preset upper limit through multiple combination experiments on historical data as a fixed configuration. For each user role's preference score in each disaster risk category, the system then performs normalization processing, scaling the preference scores in all categories to between 0 and 1 according to a relative ratio. The scaled preference scores of each category are arranged in a fixed order according to the disaster risk categories to form the preference feature vector of that user role. The preference feature vectors of all user roles form the original preference set. The dynamically changing forest area environmental monitoring data includes the actual number of alarms for each disaster risk category in the current period, the verified number of actual disasters, the number of times participation in response decisions, and the spatial coverage rate in the monitoring of the entire area. The system also calculates the actual occurrence frequency, warning trigger rate, and response participation rate of each category in the current period, and combines these values into environmental feature entries in a fixed order. To prevent scale differences from affecting subsequent fusion, the system scales each environmental feature entry to between 0 and 1 according to the category.Subsequently, the system fuses the preference feature vector with environmental feature entries. The fusion parameters are the preference weight coefficient and the environmental weight coefficient, both of which are fixed values between 0 and 1, and their sum is 1. During the deployment phase, the system tests the overall early warning accuracy under different weight combinations on historical data, and selects the combination with the highest accuracy and the actual disaster underreporting rate not exceeding the preset upper limit as the fixed value. During fusion, for each user role and each disaster risk category, the system multiplies the normalized preference score by the preference weight coefficient, combines the corresponding environmental feature values into an environmental impact score in a certain way, multiplies it by the environmental weight coefficient, and then adds the two results to obtain a comprehensive preference score. This comprehensive preference score reflects the comprehensive degree of user subjective concern and objective risk status. The system arranges each comprehensive preference score into a new preference feature vector according to the disaster risk category. The preference feature vectors of all user roles constitute the updated preference set. Next, the system needs to determine whether preferences have changed. To do this, at the end of each period, the updated preference set is saved as the current period's preference set. When this step is executed in the next period, the system reads the preference set from the previous period's storage. For each user role, the system calculates the difference between the current period's preference feature vector and the previous period's preference feature vector. The difference is calculated by subtracting the comprehensive preference score from each position in the vector, taking the absolute value of the difference, and then summing all the absolute differences to obtain a preference difference value. The larger this value, the greater the preference change. The threshold for judging preference changes is a preset preference change threshold, which is determined during the deployment phase by analyzing historical data from multiple periods: the system first marks which periods' preferences are considered "significantly changed" and which are considered "basically stable" in terms of business logic. Then, it calculates the preference difference values between all period pairs, observes the distribution range of the two types of difference values, and selects a specific value that best distinguishes the two types of situations and whose misjudgment ratio does not exceed a preset percentage as the preference change threshold and writes it into the configuration. During runtime, when the preference difference value for any user role is greater than or equal to the preference change threshold, the system determines that the preference has changed and the matching degree needs to be recalculated. If the preference difference values for all user roles are less than the preference change threshold, the system determines that the preference has not changed and the matching degree will not be recalculated. For cases where the preference has been determined to have changed, the system uses the K-means algorithm to cluster the updated preference set to obtain cluster centers representing different preference patterns and adjusts the matching degree value accordingly.The main parameters required for the K-means algorithm include the number of clusters, the maximum number of iterations, and the convergence threshold. The number of clusters is an integer greater than 1. During the deployment phase, the system tries different numbers of clusters on historical preference data, calculating the sum of intra-cluster differences and the difference in inter-cluster center distances for each number of clusters. When the number of clusters is too small, the intra-cluster differences are large; when the number of clusters is too large, the inter-cluster differences decrease and the model complexity increases. The system selects the number of clusters with the optimal comprehensive index of intra-cluster and inter-cluster differences as a fixed value. The maximum number of iterations is a positive integer used to ensure that the algorithm will stop in extreme cases, and is generally set directly based on experience. The convergence threshold is a non-negative integer, indicating that the clustering is considered convergent if the number of feature vectors whose cluster labels change in the two rounds of clustering does not exceed this value. In the actual clustering process, the system first selects several preference feature vectors as initial cluster centers from the updated preference set. The selection strategy can be to extract them from the set at uniform intervals or according to the principle of maximizing the distance between them. The extraction process is fixed in the deployment phase to ensure that the same rules are followed in each clustering. After initialization, the system enters the iteration phase. In each iteration, the system processes each preference feature vector in turn, calculating the distance between the feature vector and all cluster centers. The distance is calculated by subtracting the comprehensive preference score of each position in the feature vector from the score of the corresponding position in the cluster center, taking the absolute value, and finally summing the absolute differences of all positions to obtain a distance value. The system assigns the feature vector to the cluster corresponding to the cluster center with the smallest distance value, completing one cluster assignment after traversing all feature vectors. Subsequently, for all preference feature vectors within each cluster, the system calculates the arithmetic mean of all vector scores at each position, combines the average scores to form a new cluster center vector, and replaces the previous center with the new center. Then, it enters the next iteration, repeating the cluster assignment and center update process until the clustering results meet the convergence condition: the number of feature vectors whose cluster labels have changed between two iterations is less than or equal to the convergence threshold, or the number of iterations reaches the maximum iteration number parameter. At this point, the cluster center vector corresponding to each cluster is output. A preset benchmark is a set of benchmark preference feature vectors used to measure whether the preference pattern conforms to the overall management strategy. This set of benchmark vectors is selected by experts during the system deployment phase, combining policy requirements and historical high-quality assessment results: the system first selects a batch of historical periods deemed by experts to have reasonable focus and appropriate risk allocation. The preference feature vectors within these periods are averaged line by line according to the disaster risk category position, forming several benchmark feature vectors representing different management strategy orientations, and stored with fixed numbers.After clustering, the system calculates the distance between each cluster center vector and all preset baseline feature vectors. The calculation method is the same as the feature vector distance calculation mentioned above, which is to sum the absolute values of the differences in scores at each location, resulting in a set of distance values. The system selects the smallest distance value from these values as the baseline distance for that cluster center and records the corresponding baseline feature vector number. The smaller the baseline distance, the closer the preference pattern represented by the cluster center is to the ideal pattern represented by the baseline feature vector. The system adjusts the matching degree based on the baseline distance. The adjustment parameters include two fixed values: an amplification coefficient and a reduction coefficient. The amplification coefficient is a value greater than 1, and the reduction coefficient is a value between 0 and 1. These two values are determined during the deployment phase by simulating the changes in the overall evaluation results under different coefficient combinations. Specifically, multiple alternative coefficient combinations are applied to historical data, and the actual disaster identification rate and false alarm rate are compared under each combination. The combination that improves the identification rate without increasing the false alarm rate is selected as the final coefficient. In actual adjustment, the system first segments the baseline distance, sorts all cluster centers by baseline distance from smallest to largest, selects the top percentage of cluster centers as patterns close to the baseline, and multiplies the associated disaster risk category matching degree of the preference feature vectors corresponding to these clusters by an amplification factor. Cluster centers with lower baseline distances are classified as patterns deviating from the baseline, and the associated matching degree of the preference feature vectors corresponding to these clusters is multiplied by a reduction factor to obtain the adjusted matching degree value. This adjustment process is executed separately for each user role and each disaster risk category to ensure that the change in matching degree is consistent with the overall deviation of the preference pattern. For cases where preferences have not changed in the aforementioned judgment, the system skips the K-means clustering and matching degree adjustment steps, directly associating the updated preference set with the matching degree calculated in the previous period, keeping the matching degree value unchanged for each user role and each disaster risk category. Regardless of whether the matching degree is obtained through clustering adjustment or kept at its original value, the system will integrate the targeted expression forms generated in the aforementioned steps based on the final matching degree value. Specifically, the expression forms are grouped according to disaster risk categories, and within each category, they are sorted from largest to smallest matching degree. Priority is given to selecting expression items that rank higher and can cover more core quantitative indicators and key information elements. The expression items in different categories are combined into a hierarchical structure according to risk level and matching degree order to form an optimized risk information structure. This structure clearly stipulates which information content should be displayed first and in what order under different disaster risk categories and different user roles.Finally, based on the optimized risk information structure, the system assembles information items of each category and level into a complete risk information distribution version. In this distribution version, a concise summary, detailed content, and core quantitative indicator description are given for each type of disaster risk. The display order and display position are determined according to the matching degree and risk level. The system is output or pushed through a fixed format, thus forming the final risk information distribution version. This provides disaster risk information that matches the interests and actual risk conditions of users in different forest areas.
[0031] S6 includes obtaining integrated forest area warning data obtained through data monitoring through the final risk information distribution version, reorganizing the overall structure of the assessment information, placing high-risk priority content at the front end and highlighting it to obtain reorganized information; for the reorganized information, if the environment changes, it is determined that the preference fusion data obtained through preference fusion will be integrated to determine the final complete risk assessment information package.
[0032] In this implementation, the final risk information distribution version generated in the aforementioned steps is first received. Each piece of risk information in this distribution version includes at least a disaster risk category identifier, a corresponding forest area spatial location identifier, a risk level identifier, a matching degree value, and a marker indicating whether it belongs to high-risk priority content. During this step, the system extracts all single-point alarm records from the existing data monitoring process in the forest area within a fixed alarm statistics period. This statistics period parameter is uniformly set to a fixed time length of no less than one day and no more than one week during the deployment phase. The specific value is determined by analyzing the time interval of risk assessment updates in historical scheduling decisions. The system also includes monitored temperature exceedance alarms, humidity exceedance alarms, wind speed exceedance alarms, and alarms triggered by the comprehensive disaster model. Disaster early warning records are grouped according to the principle of monitoring unit location, disaster type, and time belonging to the same statistical period. Alarm records belonging to the same monitoring unit and with the same disaster type within the same statistical period are integrated to obtain forest area alarm integrated data. Each record in the integrated data includes at least the monitoring unit location, disaster type, total number of alarm triggers in the period, the highest alarm level in the period, the time of the first alarm occurrence, the last alarm end time, and the corresponding risk category identifier. During integration, the system first sorts each single-point alarm record by alarm time, then counts the number of alarms belonging to the same group and records the maximum alarm level. The earliest alarm time is used as the start time, and the latest cancellation time is used as the end time, thus forming a structurally complete integrated record. Subsequently, the system uses risk category identifiers and spatial location identifiers as connection keys to match each risk information in the final risk information distribution version with one or more records in the forest area alarm integration data. When multiple monitoring unit locations correspond to the same risk information entry, the system sums the total number of alarm triggers in these integration records, takes the maximum value of the highest alarm level, and counts the number of monitoring units involved in the alarm. This yields three sets of values for the risk information in the current period: the total number of alarms, the highest alarm level, and the size of the affected area. These values are then written into the runtime additional data of the risk information entry using a unified field. If a risk information entry does not match any record in the integration data, its total number of alarms, the highest alarm level, and the number of monitoring units involved are all set to 0.To restructure the overall evaluation information, the system defines a risk intensity score for each piece of risk information during the deployment phase. This score is calculated during runtime based on the total number of alarms, the highest alarm level, and the size of the affected area. The calculation process is as follows: The system first statistically analyzes historical data to determine the range of total alarm counts, the distribution of the highest alarm level, and the typical affected area for each disaster risk category during actual disasters. It then determines three parameters: the standardized upper limit for alarm counts, the standardized upper limit for alarm levels, and the standardized upper limit for affected area. Finally, it divides the total number of alarms in the current period by the standardized upper limit for alarm counts, the highest alarm level in the current period by the standardized upper limit for alarm levels, and the number of monitoring units involved in the current period by the standardized upper limit for affected area. If the ratio exceeds 1, it is truncated to 1, resulting in three standardized values between 0 and 1. During the deployment phase, fixed weight coefficients are assigned to these three standardized values, corresponding to the alarm frequency weight, alarm level weight, and scope weight, respectively. These three coefficients are determined through multiple trials on historical samples. Specifically, multiple sets of different weight combinations are taken, and the risk intensity score is calculated according to each weight. Then, the actual disaster identification rate and false alarm rate under each weight are compared, and the coefficient combination that results in the highest disaster identification rate and a false alarm rate not exceeding a predetermined upper limit is selected as the fixed value. During operation, the system multiplies the three standardized values by their corresponding weights, and then adds the three products together to obtain a risk intensity score between 0 and 1, which is then written into the structure of the risk information entry. The determination of high-risk priority content relies on a preset high-risk priority threshold, which is determined during the deployment phase as follows: The system extracts verified high-risk and non-high-risk scenario samples from historical data, calculates the corresponding risk intensity score for each sample, sorts the risk intensity scores of high-risk samples from largest to smallest, and sorts the risk intensity scores of non-high-risk samples from largest to smallest. Then, at several candidate thresholds, the system counts the number of samples classified as high-risk but actually non-high-risk and the number of samples classified as non-high-risk but actually high-risk when the threshold is set to that candidate value. The two types of misclassifications are added together to obtain the total number of misclassifications. Finally, the risk intensity score with the smallest total number of misclassifications is selected as the high-risk priority threshold and written into the configuration file, remaining unchanged during runtime. In this step, the system reads the risk intensity score of each risk information entry. When the risk intensity score is greater than or equal to the high-risk priority threshold, the entry is marked as high-risk priority content; otherwise, it is marked as low-risk priority content.After marking, the system restructures the overall evaluation information. Specifically, all risk information items are compiled into a list, first divided into two partitions based on whether they are high-risk priority items. Within the high-risk priority partition, items are arranged from highest to lowest risk intensity score. When risk intensity scores are the same or similar, they are then arranged from highest to lowest matching degree. Within the low-risk priority partition, items are arranged from highest to lowest matching degree, ultimately generating a sequence from high-risk priority to low-risk priority items. In the restructured data structure, a highlighting level field is added to each risk information item. The value of this field is calculated by combining the risk intensity score and risk level. During deployment, several risk intensity grading thresholds are preset to divide the risk intensity score range from 0 to 1. The risk intensity score is divided into several consecutive intervals, such as low, medium, and high, or more. The specific threshold is determined by statistically analyzing the proportion of severe disasters occurring in different risk intensity score ranges in historical samples, so that the proportion of severe disasters corresponding to high score ranges is significantly higher than that of low score ranges. During operation, when an item's risk intensity score falls into the high-level interval and the risk level is marked as high risk, its highlighting level is set to the highest. When the risk intensity score falls into the medium-level interval or the risk level is lower than high risk, its highlighting level is set to the middle value. When the risk intensity score falls into the low-level interval and the risk level is low, its highlighting level is set to the lowest. This ensures that high-risk priority content is clearly placed at the front and highlighted in the reorganized information. This sorted and marked list of risk information constitutes the reorganized information. Subsequently, the system needs to decide whether to introduce preferred fusion data for further fusion based on environmental changes. Therefore, at the end of each alarm statistical cycle, the system copies and saves the current forest area alarm integration data as the environmental baseline data for the previous cycle. When performing this step in the current cycle, for each disaster risk category, the system calculates the total number of alarms, the highest alarm level, and the number of monitoring units involved in the current cycle's integration data. It also extracts the corresponding values for the same disaster risk category from the environmental baseline data of the previous cycle. The system calculates the difference between the current cycle and the previous cycle item by item, takes the absolute values, and adds them together to obtain the environmental change intensity value for that disaster risk category. The larger this value, the greater the environmental risk intensity for that category. The more obvious the change, the more significant the environmental change. The judgment of whether the environmental change is significant depends on the preset environmental change threshold. This threshold is determined during the deployment phase using historical periodic data. Specifically, the environmental change intensity values are calculated for periodic pairs that were identified by experts as having "significant environmental changes" and periodic pairs that were identified as having "basically stable environments." The distribution ranges of the two types of values are compared, and a value that can separate the two types of periods as much as possible is selected as the environmental change threshold. The false judgment rate is constrained to not exceed a preset percentage. Finally, a fixed threshold is obtained and written into the configuration file. During operation, when the environmental change intensity value of any disaster risk category is greater than or equal to the environmental change threshold, the system determines that the environment has changed; otherwise, it determines that the environment has not changed.When an environmental change is determined to have occurred, the system reads the comprehensive preference score for each user role for each disaster risk category from the preference fusion data obtained in the previous steps. This comprehensive preference score, formed in the previous step through weighted fusion of historical feedback and environmental data, is a fixed value between 0 and 1 and is not recalculated in this step. Then, the system reads the disaster risk category identifier and corresponding comprehensive preference score for each risk information entry in the reconstructed information. This comprehensive preference score is combined with the risk intensity score of the entry according to the display weight coefficient set in the deployment phase to form a display priority score. The coefficient combination includes two parameters: risk intensity display weight and preference display weight. The sum of these two parameters equals 1. The specific values are determined by simulating different combinations on historical data and comparing user satisfaction indicators and disaster identification rates, and are fixed during the deployment phase. During runtime, the system multiplies the risk intensity score by the risk intensity display weight and the comprehensive preference score by the preference display weight, then adds them together to obtain the display priority score. This score is written into the structure of the corresponding entry. Then, from the perspective of each user role, the risk information entries visible to that user are sorted a second time according to their display priority scores from highest to lowest. When display priority scores are the same, the order continues to be determined by risk intensity display weight. The risk intensity score and matching degree value are used as the sorting criteria, and the highlighting level is fine-tuned according to the display priority score. For example, when the display priority score of an item is significantly higher than that of similar items, its highlighting level is increased by one level; when the display priority score of an item is significantly lower than that of similar items, its highlighting level is decreased by one level. This generates reconstructed information after fusion preference. Finally, the system uses the reconstructed information after fusion preference as the core, and combines all risk information items arranged in order, the corresponding disaster risk category identifier, risk level identifier, risk intensity score, comprehensive preference score, and display priority score. The highlighted level and other fields are organized into a structured dataset, which is combined with the basic description information, time range information and necessary statistical summary information of this assessment task in the same data structure. This structured dataset serves as the final complete risk assessment information package, which is used for hierarchical distribution and display to different user roles in subsequent steps. When the environmental change is determined to be unchanged, the system skips the above preference fusion step and directly uses the recombined information obtained from the first recombination as the core, combined with the basic description information and statistical summary information to generate a complete risk assessment information package, thereby maintaining the coherence and stability of the risk information structure under stable environmental conditions.
[0033] S7 includes distributing risk information in a tiered manner based on the disaster type requirements of user roles through a reorganized complete risk assessment information package, obtaining preliminary tiering results; determining whether the path selection is insufficient to meet the requirements based on the preliminary tiering results, and identifying path adjustment needs; if path adjustment needs exist, acquiring alarm integration data for rematching and adjustment to obtain adjusted tiering information; and using the adjusted tiering information, integrating environmental change data and high-risk priority content to determine the precise tiering results for push notifications.
[0034] In this implementation, the reconstructed complete risk assessment information package generated in the previous step is first read. Each risk information record in this package contains at least the following fields: disaster risk category identifier, risk level identifier, matching degree value, risk intensity score, high-risk priority marker, visibility marker of the user role, and current display order. During the system deployment phase, for each user role, disaster type requirement parameters are pre-configured. These parameters include four categories: the set of disaster risk categories the user role must pay attention to, the minimum number of information entries required at different risk levels, whether high-risk summary information is needed, and whether detailed quantitative information is needed. Simultaneously, distribution path parameters are configured for each user role, with a unique path marker configured for each available path. Each user role is assigned an urgency level in integer form, with urgency levels ranging from highest to lowest. For example, the highest urgency level is set to level 1, the next highest urgency level to level 2, and the lowest urgency level to level 3. For each user role, a minimum number of paths threshold and a maximum urgency requirement threshold are specified for each risk level. The minimum number of paths threshold is an integer not less than 1, used to constrain how many paths the system must select for that risk level. The maximum urgency requirement threshold is an integer not greater than the highest urgency level, used to constrain that at least one path must have an urgency level no higher than that level. These distribution path parameters and thresholds are determined by statistically analyzing the on-time delivery rate of information under different path combinations in historical information distribution tasks and the system load, and are written into the configuration file during the deployment phase. When the system performs risk information tiered distribution, it first breaks down the complete risk assessment information package according to user roles. Specifically, for each user role, it filters out all risk information records whose visibility is marked as visible to that role, forming a candidate information set for that user role. Then, within the candidate information set, based on the disaster risk category set configured in the disaster type requirement parameters, it deletes records not in the set, retaining only records belonging to the target disaster risk category. Next, the system stratifies the retained records according to risk level identifiers, classifying the highest risk level, medium risk level, and lower risk level into different levels. Within each level, it sorts the records by matching degree value from largest to smallest. When matching degree values are the same, it sorts them by risk intensity score from largest to smallest, obtaining the sorted list for each risk level level for that user role. Subsequently, based on the minimum number of information entries required for that user role in that risk level level, the system extracts records from the top of the sorted list that do not fall below the minimum number of information entries, and uses these records as the distribution candidate set for that risk level level.After obtaining the distribution candidate set for each level, the system selects a distribution path for each level based on the distribution path parameters. The specific process is as follows: The system reads the minimum path quantity threshold and the maximum urgency requirement threshold corresponding to the user role in the risk level level. It then filters out paths from the entire path list whose urgency level is less than or equal to the maximum urgency requirement threshold. These paths are sorted in ascending order of urgency level and descending order of historical delivery success rate. The historical delivery success rate was calculated and saved during the deployment phase based on historical task records using the method of "successful delivery count divided by total number of sends". If the number of paths meeting the urgency requirement is not less than the minimum path quantity threshold, the system selects the same number of paths as the minimum path quantity threshold from the head of the sorting table as the distribution path list for that risk level level. If the number of paths meeting the urgency requirement is less than the minimum path quantity threshold, the system selects all paths that meet the urgency requirement and marks the level as having insufficient paths during the path sufficiency check. At this point, the system generates a preliminary classification result for each risk level level for each user role. Each risk information record in the preliminary classification result is associated with its respective risk level level and the distribution path list for that level. Next, the system performs a path sufficiency assessment on the preliminary tiered results to determine if the selected paths are insufficient to meet the requirements. To this end, during the deployment phase, path sufficiency thresholds are pre-set for each user role and each risk level. These thresholds consist of two parts: a path quantity sufficiency threshold and a path urgency sufficiency threshold. The path quantity sufficiency threshold is an integer not less than the minimum path quantity threshold, while the path urgency sufficiency threshold is an integer not greater than the highest urgency level. These two thresholds are determined by analyzing the relationship between the proportion of information successfully delivered within a specified time in historical distribution tasks and the selected path quantity and urgency level. For example, different combinations of path quantity and urgency are enumerated in historical data to calculate the corresponding on-time delivery ratio, and paths with an on-time delivery ratio not less than the preset threshold are selected. The minimum number of paths and the highest urgency level required to achieve the target ratio are used as the path sufficiency threshold. During system operation, for each user role and each risk level layer, the number of paths actually selected at that level is counted and compared with the path sufficiency threshold. If the number of paths is less than the threshold, the path adjustment requirement for that level is directly marked as existing. If the number of paths is not less than the threshold, the minimum urgency level is found in the path list for that level. If the minimum urgency level is greater than the path urgency sufficiency threshold, it is considered that the level lacks paths with a sufficiently high urgency level, and the path adjustment requirement is also marked as existing. The system marks the path adjustment requirement for that level as non-existent if and only if the number of paths is not less than the path sufficiency threshold and the minimum urgency level is not greater than the path urgency sufficiency threshold.The system aggregates the marking results of all user roles and risk level layers. When at least one path adjustment requirement is marked as existing, it is confirmed that the current preliminary results of the classification are insufficient to fully meet the requirements in terms of path, and the path adjustment process needs to be initiated. If the path adjustment requirements of all levels are marked as non-existent, the path adjustment is skipped directly, and the subsequent precise push calculation is initiated.
[0035] When path adjustment is required, the system accesses the alarm integration data already constructed in the preceding steps. This data records fields such as the total number of alarm triggers for each disaster risk category within the current alarm statistics period, the highest alarm level for this period, and the number of monitoring units involved. For each user role and risk level tier marked as having path adjustment requirements, the system first groups the candidate set at that tier according to disaster risk category. For each disaster risk category, it reads the corresponding total number of alarms, the highest alarm level, and the number of monitoring units involved from the alarm integration data. Then, it calculates the path urgency score based on the risk urgency weighting coefficient set during the deployment phase. The system includes three non-negative parameters: alarm frequency weight, alarm level weight, and affected area weight. These three parameters were determined by repeatedly testing different combinations on historical data, and were fixed values that ensured the highest overall path resource utilization while prioritizing the addition of paths for high-frequency and high-level disasters. During operation, the system divides the total number of alarms in the current period by the alarm frequency upper limit parameter obtained from historical statistics to obtain a standardized alarm frequency value, divides the highest alarm level by the alarm level upper limit parameter to obtain a standardized alarm level value, and divides the number of monitoring units involved by the affected area upper limit parameter to obtain a standardized affected area value. If any of these values exceeds 1, it is truncated to 1. Then, each value is multiplied by its corresponding weight coefficient and summed. The summation result is then used to calculate the affected area. The path urgency score is used as the path urgency score for this disaster risk category. Next, the system searches the path pool for unassigned or shareable paths in this risk level layer, sorting them by urgency level from lowest to highest and historical delivery success rate from highest to lowest. Starting with the disaster risk category with the highest path urgency score, the system iterates through each category. When the path urgency score for a category is higher than a preset path urgency threshold and the current number of paths in that level is less than a path quantity sufficiency threshold, or there are no paths in that level with a sufficiently high urgency level, the system selects one unused path from the head of the path pool sorting table and adds it to the path list for that level. The path urgency threshold is determined during the deployment phase based on historical data. To ensure that paths are added only to categories with significantly increased risk and avoid wasting resources, the system calculates the path urgency scores for each category before a severe disaster occurs, selecting a score that covers most severe disaster scenarios without frequent triggering as a threshold. After adding paths to each category sequentially, the system recalculates the number of paths and the minimum urgency level for that risk level. When both indicators simultaneously meet the path sufficiency threshold, the system stops adding paths to that level and completes path rematching and adjustment. If the level still does not meet the path sufficiency threshold after the path pool is exhausted, the system retains the current adjustment result, marks the path insufficiency as true, and provides a notification in subsequent results. After the above adjustments are completed, the system writes the updated path list back to each risk information record, forming adjusted hierarchical information, which already includes the rematched distribution path configuration.Finally, the system generates precise push notifications based on the adjusted hierarchical information, integrated environmental change data, and high-risk priority content. The environmental change data, calculated in the preceding steps, includes the environmental change intensity value for each disaster risk category and a Boolean flag indicating whether the environment has changed significantly. The high-risk priority content, determined in the preceding steps by comparing the risk intensity score with the high-risk threshold, provides a flag field for each risk information record indicating whether it is high-risk priority content. During the deployment phase, the system pre-sets three types of weight coefficients for the final push ranking: hierarchical weight coefficient, environmental weight coefficient, and high-risk weight coefficient. The sum of these three is 1. The specific values are determined by simulating different coefficient combinations on historical data, using two indicators: the proportion of timely delivery of real disaster information to target user roles and the proportion of redundancy of non-critical information. The combination with the highest proportion of timely delivery of real disaster information and a redundancy ratio below the set upper limit is selected as the fixed value. During runtime, the system calculates the final push priority score for each record in the adjusted classification information. Specifically: based on the risk level layer to which the record belongs, the risk level is mapped to a fixed base level score, which is multiplied by the classification weight coefficient to obtain the classification weight score; if the environmental change of the disaster risk category to which the record belongs is marked as significant, the corresponding environmental change intensity value is standardized to between 0 and 1 and then multiplied by the environmental weight coefficient to obtain the environmental change score; if the environmental change is marked as significant, the environmental change score is set to 0; if the high-risk priority of the record is marked as true, the high-risk weight coefficient is used as the high-risk priority score; if it is marked as false, the high-risk priority score is set to 0; the three scores are added together to obtain the final push priority score for the record. Within each user role, the system sorts all visible risk information records from highest to lowest according to their final push priority score. During the sorting process, it maintains a structure of first grouping by risk level and then sorting by priority within each group. This ensures that records with high risk levels, significant environmental changes, and high-risk priority markers consistently appear at the top of the ranking. Records still marked with insufficient paths are marked with a "path needs optimization" message in the results. After sorting, the system packages the sorting results, associated path list, final push priority score, and relevant marker fields for each user role to form a precise push ranking result for that user role.
[0036] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A forest disaster risk assessment method based on machine learning, characterized by, Comprise: S1, by constructing forest disaster evolution trajectory database, historical disaster events are modeled according to time sequence, the similarity between current forest environment characteristics and historical disaster trajectory is analyzed, the disaster type and focus are inferred, the disaster characteristics are clustered, and a risk category set is formed; S2, according to the risk category set, the quantitative indicators and descriptive contents in the evaluation information are preliminarily split, wherein, if the quantitative indicators exceed the preset threshold, they are marked as high risk priority part, otherwise they are marked as low risk priority part, to determine the boundary division of risk content; S3, the recommended system is used to match and calculate the priority content after boundary division and the disaster risk category set, the matching degree of the highest risk information subset is obtained by comparing the disaster focus preference and the content correlation degree; S4, the core quantitative indicators and text description are extracted from the risk information subset with the highest matching degree, the extracted content is analyzed and summarized by natural language processing, and it is judged whether it meets the depth requirement of each disaster risk category set, if not, the quantitative structured version is output, otherwise the summary structure is adjusted to obtain the targeted expression form. 2.The forest disaster risk assessment method based on machine learning according to claim 1, wherein: The S1 comprises: obtain the historical event sequence from the pre-established disaster trajectory database, arrange the temperature change data, humidity change data and wind speed change data in the historical event sequence in time order, and generate trajectory data containing environmental parameter evolution process; The dynamic time warping algorithm is used to calculate the distance value between the current forest environment characteristics and the trajectory data, the dynamic time warping algorithm matches time series of different lengths by constructing distance matrix, and the matching degree value reflecting the similarity degree is obtained; According to the matching degree value, the possible disaster type is inferred, if the matching degree value exceeds the preset threshold, it is judged as a high risk disaster type, and the corresponding environmental monitoring key area coordinates are obtained; For disaster type, disaster characteristic parameters including occurrence time, duration period and influence range are extracted, grouping clustering is carried out by calculating the Euclidean distance between disaster characteristic parameters, and disaster category combination containing different risk levels is formed; The risk warning level mark is generated by determining the disaster development trend direction through the risk level distribution in the disaster category combination. 3.The forest disaster risk assessment method based on machine learning according to claim 1, characterized in that: The S2 comprises: obtain the evaluation information from the preset risk category set, preliminarily split the quantitative indicators and descriptive contents in the evaluation information, if the quantitative indicators exceed the preset threshold, mark them as high risk priority part, otherwise mark them as low risk priority part; For high risk priority part and low risk priority part, the marked results are fused by using vegetation density monitoring method, the boundary division is judged by comparing the corresponding relationship between vegetation coverage and risk marking, and the boundary division of high risk content and low risk content is determined. 4.The forest disaster risk assessment method based on machine learning according to claim 1, wherein: The S3 comprises: obtain the high priority content and low priority content after boundary division from the preset disaster risk category set, calculate the initial matching value by content correlation degree comparison, and obtain the preliminary matching matrix; For the preliminary matching matrix, the matching weight is adjusted by fusing disaster focus preference, it is judged whether the preference degree exceeds the preset threshold, the corresponding weight is improved, and the weighted matching matrix is determined; Obtaining a weighted matching matrix, comparing the matching value in the matrix with the correlation index to obtain a content correlation ranking list; Extracting the highest matching value part from the content correlation ranking list to generate a disaster risk matching sequence, and marking it as a priority sequence if the sequence covers the risk categories to determine the priority risk sequence; For the priority risk sequence, the matching degree is integrated to obtain the highest risk information subset. 5.The forest disaster risk assessment method based on machine learning according to claim 1, wherein: The S4 comprises: From the highest matching risk information subset, extract the core quantitative indicators and text descriptions, obtain the quantitative set after fusing the index weight and the text set after evaluating the description completeness; For the quantitative set and the text set, use natural language processing to perform semantic association analysis to obtain a semantic feature set required for generating summary content; According to the semantic feature set, generate summary content, and determine the summary structure after targeted expression optimization if the summary content meets the preset depth requirement and category coverage verification; If the summary content does not meet the preset depth requirement, output the quantitative structured version to determine the targeted expression form under risk category matching. 6.The forest disaster risk assessment method based on machine learning according to claim 1, wherein, It also includes S5, obtaining the latest risk attention preference from the dynamic changing forest user role historical feedback data according to the targeted expression form, recalculating the matching degree if the preference changes, otherwise keeping the original matching degree, and determining the final risk information distribution version, which specifically includes: Obtain the latest risk attention preference from the forest user role historical feedback data, and fuse the preference with the dynamically changing forest environment monitoring data to obtain an updated preference set; For the updated preference set, if the preference changes, recalculate the matching degree, use the K-means algorithm to cluster the preference features in the set, where the input is the preference feature vector and the output is the cluster center, and obtain the adjusted matching degree value by calculating the distance between the cluster center and the preset reference; If the preference does not change, keep the original matching degree corresponding to the updated preference set, integrate the targeted expression form through the matching degree value to obtain the optimized risk information structure; According to the optimized risk information structure, determine the final risk information distribution version. 7.The forest disaster risk assessment method based on machine learning according to claim 6, wherein, It also includes S6, reorganize the overall structure of the evaluation information through the final risk information distribution version, and place the high-risk priority content in the front end and highlight it to obtain the reorganized complete risk assessment information package, which specifically includes: Through the final risk information distribution version, obtain the forest warning integration data obtained through data monitoring, reorganize the overall structure of the evaluation information, place the high-risk priority content in the front end and highlight it to obtain the reorganized information.
8. The forest disaster risk assessment method based on machine learning according to claim 7, characterized in that: The S6 also includes: For the reorganized information, if the environment changes, fuse the preference fusion data obtained by the preference fusion to determine the final complete risk assessment information package. 9.The forest disaster risk assessment method based on machine learning according to claim 7, wherein, It also includes S7, based on the reorganized complete risk assessment information package, perform risk information hierarchical distribution according to the disaster type demand of the user role, if the path selection is insufficient to meet the demand, then re-match and adjust to obtain the hierarchical result of accurate push, which specifically includes: Through the recombined complete risk assessment information package, risk information is graded and distributed according to the disaster type demand of the user role, and a graded preliminary result is obtained; For the graded preliminary result, it is judged whether the path selection is insufficient to meet the demand, and the path adjustment demand is determined.
10. The forest disaster risk assessment method based on machine learning according to claim 9, characterized in that: The S7 further comprises: If the path adjustment demand exists, alarm integration data is obtained for re-matching adjustment, and the adjusted graded information is obtained; Through the adjusted graded information, the environment change data and the high-risk priority content are fused to determine the graded result of accurate push.