Forest fire danger monitoring method based on multi-modal feature encoding and attention fusion
Patent Information
- Application Number
- CN202610700663.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-20
- Publication Date
- 2026-08-18
AI Technical Summary
人工巡护和瞭望方式受巡查频次、视距条件及人员经验限制,难以实现对大范围林区的连续、实时监测,且存在主观性强、响应滞后等问题;基于单一信息源(如气象数据或视频图像)的火险监测方法通常只能反映森林环境的某一侧面特征,难以全面刻画复杂环境条件下森林火险的综合演化状态,对早期潜在火险及多因素耦合诱发的异常识别能力有限,易出现漏报或误报
本申请实施例提供一种基于多模态特征编码与注意力融合的森林火险监测方法,包括以下步骤:步骤S1、采集多模态数据;所述多模态数据包括热成像数据、可见光图像数据、振动数据、环境声音数据、气象传感数据、植被状态参数和林区环境工况参数;步骤S2、对所述多模态数据进行预处理,建立统一的时间戳,确保不同模态的数据在时间上对齐;步骤S3、分别建立与各个模态数据对应的编码器,将预处理后的多模态数据映射为一个高维特征向量,提取森林环境状态与火险演化高级特征;步骤S4、将各个模态的特征向量在特征层面进行多模态融合,对模型进行训练,使模型能够在不同气象条件与林区环境状态下自动学习,选择需重点关注的模态信息,得到训练好的模型;步骤S5、采用训练好的模型,基于学到的“正常”模式,计算实时数据的重构误差,超过设定的阈值时报警,得到调试好的模型;步骤S6、采用调试好的模型识别高风险火险类型,预测火险由当前状态向失控状态发展的剩余时间,实现提前干预与防控。
Smart Images

Figure CN122598341A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of fire risk early warning technology, and in particular to a forest fire risk monitoring method based on multimodal feature encoding and attention fusion. Background Technology
[0002] Forest ecosystems, as vital natural resources and ecological barriers, are characterized by their wide distribution, complex environmental types, and significant impact from both natural and human factors. Forest areas exhibit significant topographic relief, diverse vegetation types, and variable climate conditions. Typical fire hazard factors include understory combustibles, vegetation moisture content, surface temperature, meteorological conditions, and human activities. Under prolonged periods of high temperatures, drought, or strong winds, forest areas are prone to problems such as combustible accumulation, localized overheating, abnormal smoke, and potential fire hazards. If these issues are not identified and addressed promptly, they can rapidly escalate into forest fires, causing ecological damage, casualties, and substantial economic losses, severely impacting ecological security and social stability.
[0003] Currently, forest fire risk monitoring and control mainly rely on manual patrols, fixed-point lookouts, or monitoring methods based on a single information source. Manual patrols and lookouts are limited by patrol frequency, line-of-sight conditions, and personnel experience, making it difficult to achieve continuous, real-time monitoring of large-scale forest areas. They also suffer from strong subjectivity and delayed response. Fire risk monitoring methods based on a single information source (such as meteorological data or video images) can usually only reflect one aspect of the forest environment and cannot comprehensively depict the integrated evolution of forest fire risk under complex environmental conditions. They also have limited ability to identify early potential fire risks and anomalies induced by multiple coupled factors, making them prone to missed or false alarms.
[0004] In addition, forest fires are characterized by their suddenness and relatively low probability of occurrence. The number of tagged fire samples that can be obtained in practice is limited, which greatly restricts the training and generalization capabilities of supervised learning-based fire risk identification and early warning methods. This makes it difficult to adapt to the actual needs of forest fire risk monitoring in different regions, seasons and complex weather conditions.
[0005] Therefore, there is an urgent need for an intelligent monitoring and early warning method that can make full use of multi-source, multi-modal data generated during forest environmental monitoring, and achieve comprehensive perception of forest fire risk status, timely identification of abnormal fire risks, and prediction of fire risk development trends without the need for a large number of manually labeled samples, so as to improve the scientific nature, accuracy, and overall control efficiency of forest fire prevention work. Summary of the Invention
[0006] This application provides a forest fire risk monitoring method based on multimodal feature encoding and attention fusion, which realizes integrated early warning of fire risk anomaly identification, risk type determination and development trend assessment.
[0007] To address the aforementioned technical problems, in a first aspect, this application provides a forest fire risk monitoring method based on multimodal feature encoding and attention fusion, comprising the following steps: Step S1, collecting multimodal data; the multimodal data includes visible light image data, thermal imaging data, vibration data, temperature data, operation logs, environmental sound data, meteorological sensor data, vegetation status parameters, and forest area environmental condition parameters from power support equipment (such as water pumps, generators, etc.), monitoring and sensing equipment (such as infrared cameras), and communication equipment; Step S2, preprocessing the multimodal data, establishing a unified timestamp to ensure that data from different modalities are aligned in time; Step S3, establishing encoders corresponding to each modal data, mapping the preprocessed multimodal data into a high-dimensional feature vector, and extracting the forest environmental state and fire risk evolution high-dimensional feature vectors. Step S4: Perform multimodal fusion of the feature vectors of each modality at the feature level. Pre-train the model with a large amount of data in the source domain (such as a laboratory simulation environment or a mature forest area with abundant historical data) to learn common fault features so that the model can automatically learn under different meteorological conditions and forest environment conditions. Select the modal information that needs to be focused on to obtain a trained model. Step S5: Transfer the pre-trained model to the target domain (new forest area). Based on the learned "normal" mode, fine-tune the top layer of the model with only a small amount of new data to make it quickly adapt to the subtle differences in the new environment. Calculate the reconstruction error of real-time data. When it exceeds the set threshold, an alarm is triggered to obtain a debugged model. Step S6: Use the debugged model to identify high-risk fire hazard types and predict the remaining time for the fire hazard to develop from the current state to an out-of-control state to achieve early intervention and prevention.
[0008] In some exemplary embodiments, step S2 involves preprocessing the multimodal data, including: performing noise reduction processing on visible light image data, thermal imaging data, sound data, and vibration data respectively; and resampling meteorological sensor data and continuous environmental monitoring data.
[0009] In some exemplary embodiments, denoising the image data includes: performing bilateral filtering denoising processing on visible light image data and thermal imaging data, and denoising the location in the image. The pixel value after bilateral filtering The following formula is used to calculate:
[0010]
[0011] in, As the normalization factor, It revolves around pixels neighborhood window, It is the location of neighboring pixels. and They are pixels and The intensity value, It is a spatial domain weight function. , It is a range weight function. , and These are the standard deviations of the spatial domain and the range, respectively.
[0012] In some exemplary embodiments, noise reduction processing of the sound data includes: step S201, processing the collected ambient sound data... Perform Empirical Mode Decomposition (EMD) to remove spurious components and perform noise reduction; Step S202: Determine the acquired original noise signal. For all extreme points, the upper envelope curve of the data is formed by cubic spline interpolation over all maximum and minimum points. and lower envelope curve and take and average As shown in the following formula:
[0013] Step S203, let ,like Simultaneously satisfying both of the IMF's conditions: (1) Throughout the entire time course, the number of times the zero point is crossed is equal to or at most differs from the number of extreme points by 1. (2) If the upper envelope defined by the local maxima and the lower envelope defined by the local minima are locally symmetric about the time axis, then the mean of the upper envelope defined by the local maxima and the lower envelope defined by the local minima is 0. For the first-order IMF, if the conditions are not met, then... See as new , Given the mean of its upper and lower envelopes, we have ,like If not satisfied, repeat the process. Next, we obtain the following formula:
[0014] like and If the standard deviation SD is within the predetermined range, then the process is stopped. Original load signal The first-tier IMF, denoted as The standard deviation The calculation formula is:
[0015] In the formula, For load signal Total time length; Step S204, let ,Will See as new Repeat steps S201 to S203 to obtain the second-order IMF, denoted as Then, by analogy, other IMFs of various orders are obtained, denoted as follows: ... ,until The function is a monotonic function until it can no longer be divided into IMFs; at this point, it is denoted as... The signal after noise reduction of environmental sound signals in the forest area.
[0016] In some exemplary embodiments, noise reduction of the acquired vibration data includes: performing low-pass filtering and high-pass filtering on the acquired vibration data. During low-pass filtering, the cutoff frequency is set to 2-3 times the highest fault characteristic frequency of the original signal (e.g., the fault frequency of a water pump bearing is about 130Hz, so the cutoff frequency is set to 400Hz) to filter high-frequency noise while retaining fault characteristics and eliminating the influence of high-frequency noise such as electromagnetic interference. After low-pass filtering, high-pass filtering is then performed on the vibration data, with the cutoff frequency set to be lower than the lowest fault characteristic frequency (e.g., the rotation frequency is 25Hz, so the cutoff frequency is set to 10Hz) to filter low-frequency background vibration interference.
[0017] In some exemplary embodiments, resampling meteorological sensor data and continuous environmental monitoring data includes: determining a resampling factor; designing a filter; and resampling the data based on the resampling factor and the filter; wherein, when determining the resampling factor, it is assumed that the original sampling rate is fs. old The target sampling rate is fs new The resampling factor is: factor=fs new / fs old =L / M Where L and M are coprime integers, and L and M are chosen such that L / M is equal to or very close to the factor; when designing the filter, a low-pass filter is designed to simultaneously perform the tasks of anti-mirror and anti-aliasing, and the cutoff frequency of the filter is: fc = min(fs old , fs new ) / 2.0×0.9 Based on the resampling factor, the designed filter is used to perform filtering operations on the original signal to obtain the filtered signal x. filter For the filtered signal, one sample is extracted for every M samples to obtain the resampled signal.
[0018] In some exemplary embodiments, encoders corresponding to each modality of data are established, including: using a convolutional neural network to analyze the median-filtered visible light image and thermal imaging image, sliding a fixed-size convolutional kernel (such as 3×3 or 5×5), performing convolution operations on local pixel regions of the image, calculating the local correlation between pixels, and outputting feature vectors. For example, for infrared images of forest fire prevention equipment, a 3×3 convolutional kernel can capture the grayscale variation relationship between a pixel and its 8 neighbors, identifying features such as the edge and texture of the temperature gradient; for denoised vibration data, a 1D neural network is used for analysis, with convolutional layers using short convolutional kernels (length 3 / 5 / 7) to slide and cover the local temporal window, calculating the correlation of amplitude changes between adjacent sampling points, and capturing features such as "amplitude abrupt change (fault impact), local fluctuation (oscillation trend), and boundary of stable segment"; for denoised sound data, STFT / CWT is used to convert the noise sequence into a time-frequency image, and 2D CNN is used to extract the texture distribution features of the time-frequency image (e.g., uniform texture = Gaussian noise, discrete bright spots = impulse noise), outputting a time-frequency feature vector; for text data, the frequency of each word in the text is counted, generating a vocabulary list containing all core words, and for a single text, the number of times each word appears in the vocabulary list is counted, generating a vector.
[0019] In some exemplary embodiments, multimodal fusion is performed on the feature vectors of each modality at the feature level, including: each modal data first extracts high-level features through its corresponding encoder, and then performs fusion at the feature level using an attention mechanism. Specifically, the modal feature dimensions of image, vibration, sound, and text are unified to the same target dimension to solve the problem of inconsistent dimensions among different data modalities. Using the data modality with the largest modal feature dimension as the standard, zero values are padded to the end of the feature vectors of other data, keeping the original feature information unchanged and only expanding the dimension; this operation is simple and computationally insignificant. Weights are assigned according to the variance of the modal features (the larger the variance, the richer the information, and the higher the weight), and the five types of modal features are weighted and summed to perform feature fusion. The calculation method is as follows:
[0020]
[0021] in, (1) The eigenvectors after summation, , , , , These are the aligned image, thermal imaging, vibration, sound, and text modal feature vectors, respectively. (2) , , , , The fusion weights are respectively for image, thermal imaging, vibration, sound, and text modal features, satisfying... ; (3) , , , , These are the variances of the feature vectors for image, thermal imaging, vibration, sound, and text modalities, respectively. , , , The calculation method is similar.
[0022] In some exemplary embodiments, a trained model is used to calculate the reconstruction error of real-time data based on the learned "normal" pattern. An alarm is triggered when the error exceeds a set threshold, resulting in a calibrated model. This includes: using the trained model to calculate the reconstruction error of real-time data or its deviation from the "non-fire risk state cluster" in the joint embedding space based on the learned "normal forest environment state" pattern. An alarm is triggered when the error exceeds a set threshold. The threshold is repeatedly adjusted using existing historical fire risk or fire cases until the model output results are consistent with the actual fire risk situation, resulting in a calibrated model.
[0023] Secondly, this application also provides a forest fire risk monitoring system based on multimodal feature coding and attention fusion. The system uses the forest fire risk monitoring method based on multimodal feature coding and attention fusion as described in the above embodiments to monitor fire risk. The system includes: a multimodal data acquisition module, a preprocessing module, a feature extraction module, a model training module, a model debugging module, and a technology application module connected in sequence. The system includes a multimodal data acquisition module for collecting multimodal data, including visible light image data, thermal imaging data, vibration data, temperature data, operation logs, environmental sound data, meteorological sensor data, vegetation status parameters, and forest area environmental condition parameters from power supply equipment (such as water pumps and generators), monitoring and sensing equipment (such as infrared cameras), and communication equipment. A preprocessing module preprocesses the multimodal data, establishing a unified timestamp to ensure temporal alignment of data from different modalities. A feature extraction module establishes encoders corresponding to each modal data, mapping the preprocessed multimodal data into a high-dimensional feature vector to extract advanced features of forest environmental status and fire risk evolution. A model training module trains the feature vectors of each modality at the feature layer. Multimodal fusion is performed on the model. In the source domain (such as a laboratory simulation environment or a mature forest area with abundant historical data), a large amount of data is used to pre-train the model, allowing it to learn common fault characteristics. This enables the model to automatically learn under different meteorological conditions and forest environment states. The model selects the modal information that needs to be focused on, resulting in a well-trained model. The model debugging module is used to transfer the pre-trained model to the target domain (new forest area). Based on the learned "normal" mode, only a small amount of new data is used to fine-tune the top layer of the model, allowing it to quickly adapt to the subtle differences in the new environment. The reconstruction error of real-time data is calculated, and an alarm is triggered when it exceeds a set threshold, resulting in a well-tuned model. The technology application module is used to use the well-tuned model to identify high-risk fire hazard types and predict the remaining time for the fire hazard to develop from the current state to an out-of-control state, so as to achieve early intervention and prevention.
[0024] In some exemplary embodiments, the thermal imaging data is infrared temperature image data collected during forest area monitoring, with a single frame resolution of 320×240 pixels and a corresponding temperature range of 20℃~120℃. The data is stored in the form of a continuous frame sequence with a sampling time interval of 0.2 s, used to characterize the temperature distribution of the forest surface, vegetation, and local areas. The visible light image data is visible light image data collected during forest monitoring, with a single frame resolution of 1920×1080 pixels and a sampling frame rate of 15 frames / second. The image content includes forest vegetation distribution, surface conditions, smoke characteristics, and surrounding environmental information. The vibration data is vibration acceleration signals collected from key parts of the equipment during operation, with a vibration signal sampling frequency of 12.8 kHz and a single-channel vibration amplitude range of -10g~10g, stored in the form of a continuous time series. The environmental sound data is environmental sound signal data collected during forest area monitoring, with a sound signal sampling frequency of 16kHz and a quantization accuracy of 16. The data consists of a bit sample, with a single sound sample length of 5 seconds, used to reflect the acoustic characteristics related to environmental noise, abnormal sounds, and potential fire hazards in forest areas; meteorological sensor data is continuous environmental time-series data collected from forest monitoring points, including parameters such as temperature, humidity, wind speed, and wind direction, with a sampling frequency of 1Hz to 10Hz, stored in continuous time-series format; vegetation status parameters are indicators such as vegetation moisture content and combustible load obtained through remote sensing or ground sensing equipment, used to reflect the state of forest combustibles and fire hazard sensitivity; forest area environmental operating condition parameters are environmental and management parameters collected synchronously in forest areas, including precipitation and intensity of human activities, with a sampling period of 1s to 10s.
[0025] The technical solution provided in this application has at least the following advantages: This application provides a forest fire risk monitoring method based on multimodal feature encoding and attention fusion, comprising the following steps: Step S1, collecting multimodal data; the multimodal data includes thermal imaging data, visible light image data, vibration data, environmental sound data, meteorological sensor data, vegetation state parameters, and forest area environmental condition parameters; Step S2, preprocessing the multimodal data, establishing a unified timestamp to ensure that data from different modalities are aligned in time; Step S3, establishing encoders corresponding to each modal data, mapping the preprocessed multimodal data into a high-dimensional feature vector, and extracting forest ring data. Advanced features of environmental conditions and fire risk evolution; Step S4: Multimodal fusion of feature vectors of each modality at the feature level to train the model, enabling the model to automatically learn under different meteorological conditions and forest environment conditions, select the modal information that needs to be focused on, and obtain a trained model; Step S5: Using the trained model, based on the learned "normal" mode, calculate the reconstruction error of real-time data, and alarm when it exceeds the set threshold, to obtain a debugged model; Step S6: Using the debugged model to identify high-risk fire risk types, predict the remaining time for the fire risk to develop from the current state to an out-of-control state, and realize early intervention and prevention.
[0026] The forest fire risk monitoring method based on multimodal feature encoding and attention fusion provided in this application constructs a unified multimodal representation mechanism for complex forest environmental conditions. Addressing the characteristics of forest areas, such as their wide spatial range, complex environmental elements, and frequent changes in meteorological conditions, this application encodes and fuses features from thermal imaging, visible light images, vibration, environmental sound, meteorological monitoring data, vegetation state parameters, and forest environmental parameters under a unified time reference. This forms a unified feature representation capable of simultaneously characterizing abnormal temperature rises, smoke changes, environmental acoustic anomalies, and changes in fire risk triggers, significantly improving the completeness and stability of forest fire risk perception under complex natural environmental conditions.
[0027] On the other hand, this application proposes a multimodal fire risk fusion method based on adaptive allocation of attention weights. Unlike the existing technology that simply splices or weights multimodal features, this application introduces an attention mechanism at the feature layer, enabling the model to dynamically adjust the contribution weights of each modality according to different meteorological conditions, environmental states, and fire risk development stages. Thus, even when multiple fire risk factors overlap or environmental features interfere with each other, it can still accurately focus on the modal information that is most discriminative for current fire risk identification.
[0028] Secondly, this application proposes a fire risk anomaly detection mechanism based on normal forest environment patterns without prior knowledge. This application uses normal forest environment and non-fire risk state data as learning objects to establish a normal state feature space. Fire risk anomaly identification and early warning are achieved by real-time monitoring of the reconstruction error or deviation of data within this space. This eliminates the need for explicit fire risk thresholds or a large number of known fire samples, thus solving the problem of insufficient generalization ability of traditional supervised models due to the scarcity of forest fire samples and their high randomness.
[0029] Finally, this application achieves integrated early warning decision-making for fire risk anomaly identification, risk type determination, and development trend assessment. Within a unified model framework, this application realizes a continuous analysis process from abnormal fire risk detection and fire risk type and level determination to fire risk development trend assessment, transforming forest fire prevention from a reactive response to a proactive early warning and intervention, significantly reducing the risk of major forest fires, and improving the intelligence level and overall prevention and control effectiveness of forest fire monitoring and management. Attached Figure Description
[0030] One or more embodiments are illustrated by way of example with reference to the accompanying drawings. These illustrations do not constitute a limitation on the embodiments, and unless otherwise stated, the figures in the drawings are not to be limited by scale.
[0031] Figure 1 This is a flowchart illustrating a forest fire risk monitoring method based on multimodal feature encoding and attention fusion, provided as an embodiment of this application.
[0032] Figure 2 This is a schematic diagram of the structure of a forest fire risk monitoring system based on multimodal feature encoding and attention fusion, provided as an embodiment of this application.
[0033] Figure 3 This is a flowchart of EMD (Enhanced Motion Discharge) of collected forest environment sound signals provided in one embodiment of this application.
[0034] Figure 4 The spectrum diagrams of the signal before and after resampling are provided for an embodiment of this application.
[0035] Figure 5 This is a comparison image of a thermal image of a collection of deposits before and after noise reduction, provided as an embodiment of this application. Detailed Implementation
[0036] As can be seen from the background technology, existing methods for monitoring and controlling forest fire risk are difficult to achieve continuous and real-time monitoring of large-scale forest areas, and have problems such as strong subjectivity and delayed response. Moreover, they are difficult to fully depict the comprehensive evolution of forest fire risk under complex environmental conditions, and have limited ability to identify early potential fire risks and anomalies induced by the coupling of multiple factors, which easily leads to missed or false alarms.
[0037] To address the aforementioned technical problems, this application provides a forest fire risk monitoring method based on multimodal feature encoding and attention fusion, comprising the following steps: collecting multimodal data; preprocessing the multimodal data, establishing a unified timestamp to ensure temporal alignment of data from different modalities; establishing encoders corresponding to each modal data, mapping the preprocessed multimodal data into a high-dimensional feature vector, and extracting high-level features of forest environmental status and fire risk evolution; fusing the feature vectors of each modality at the feature level to train the model, enabling the model to automatically learn under different meteorological conditions and forest environmental states, selecting the modal information requiring key attention, and obtaining a trained model; using the trained model, based on the learned "normal" mode, calculating the reconstruction error of real-time data, and issuing an alarm when the error exceeds a set threshold, thus obtaining a debugged model; using the debugged model to identify high-risk fire risk types, predicting the remaining time for the fire risk to develop from its current state to an out-of-control state, and achieving early intervention and prevention. The forest fire risk monitoring method based on multimodal feature encoding and attention fusion provided in this application can achieve integrated early warning of fire risk anomaly identification, risk type determination, and development trend assessment.
[0038] The embodiments of this application will now be described in detail with reference to the accompanying drawings. However, those skilled in the art will understand that many technical details have been provided in the embodiments of this application to facilitate a better understanding of the application. However, the technical solutions claimed in this application can be implemented even without these technical details and various variations and modifications based on the following embodiments.
[0039] refer to Figure 1 This application provides a forest fire risk monitoring method based on multimodal feature encoding and attention fusion, comprising the following steps: Step S1: Collect multimodal data; the multimodal data includes visible light image data, thermal imaging data, vibration data, temperature data, operation logs, environmental sound data, meteorological sensor data, vegetation status parameters, and forest area environmental condition parameters of power support equipment (such as water pumps, generators, etc.), monitoring and sensing equipment (such as infrared cameras), and communication equipment. Step S2: Preprocess the multimodal data and establish a unified timestamp to ensure that the data of different modalities are aligned in time. Step S3: Establish encoders corresponding to each modal data, map the preprocessed multimodal data into a high-dimensional feature vector, and extract high-level features of forest environmental status and fire risk evolution. Step S4: Perform multimodal fusion of the feature vectors of each modality at the feature level, and pre-train the model with a large amount of data in the source domain (such as laboratory simulation environment, mature forest area with rich historical data) to learn the common fault features so that the model can automatically learn under different meteorological conditions and forest area environmental conditions, select the modal information that needs to be focused on, and obtain the trained model. Step S5: Transfer the pre-trained model to the target domain (new forest area). Based on the learned "normal" pattern, fine-tune the top layer of the model with only a small amount of new data to enable it to quickly adapt to the subtle differences in the new environment. Calculate the reconstruction error of the real-time data. If the error exceeds the set threshold, an alarm is triggered, and the debugged model is obtained. Step S6: Use the debugged model to identify high-risk fire hazard types, predict the remaining time for the fire hazard to develop from its current state to an out-of-control state, and achieve early intervention and prevention.
[0040] This application presents an intelligent forest fire risk monitoring method based on multimodal feature encoding and attention fusion, aiming to address the problems of reliance on a single information source, low information utilization, and insufficient accuracy in fire risk identification and early warning in existing forest fire risk monitoring processes. This method unifies the modeling of collected thermal infrared remote sensing data, visible light image data, vibration data, environmental sound data, meteorological sensor data, vegetation moisture content-related parameters, and forest area environmental condition parameters. After noise reduction, resampling, and time-series alignment, data encoders for different modalities are constructed, mapping multi-source heterogeneous data to a unified high-dimensional feature space to extract deep features of forest environmental status and fire risk evolution. Subsequently, an attention fusion mechanism is introduced at the feature level, enabling the model to adaptively allocate the feature weights of each modality according to different meteorological conditions, surface features, and fire risk triggers, thereby fully exploring the complementary relationships between multimodal information. Based on pre-trained normal forest environment and non-fire risk state patterns, real-time detection and early warning of forest fire risk anomalies are achieved using reconstruction errors or feature deviations. Furthermore, potential fire risk types and development trends are identified, enabling fire risk level assessment and early intervention. This method does not rely on a large amount of manually labeled data, and can significantly improve the accuracy and reliability of forest fire risk monitoring and early warning, reduce the cost of fire prevention patrols and management, and has high engineering application and promotion value.
[0041] In some embodiments, step S2 involves preprocessing the multimodal data, including: performing noise reduction processing on visible light image data, thermal imaging data, sound data, and vibration data respectively; and resampling meteorological sensor data and continuous environmental monitoring data.
[0042] In some embodiments, denoising the visible light image data and thermal imaging data includes: performing bilateral filtering denoising on the visible light image data and thermal imaging data, and determining the location in the image. The pixel value after bilateral filtering The following formula is used to calculate:
[0043]
[0044] in, As the normalization factor, It revolves around pixels neighborhood window, It is the location of neighboring pixels. and They are pixels and The intensity value, It is a spatial domain weight function. , It is a range weight function. , and These are the standard deviations of the spatial domain and the range, respectively. In some embodiments, noise reduction processing is performed on the audio data, and the specific flowchart is as follows: Figure 3 As shown; the noise reduction process includes: step S201, processing the collected ambient sound data. Perform Empirical Mode Decomposition (EMD) to remove spurious components and perform noise reduction; Step S202: Determine the acquired original noise signal. For all extreme points, the upper envelope curve of the data is formed by cubic spline interpolation over all maximum and minimum points. and lower envelope curve and take and average As shown in the following formula:
[0045] Step S203, let ,like Simultaneously satisfying both of the IMF's conditions: (1) Throughout the entire time course, the number of times the zero point is crossed is equal to or at most differs from the number of extreme points by 1. (2) If the upper envelope defined by the local maxima and the lower envelope defined by the local minima are locally symmetric about the time axis, then the mean of the upper envelope defined by the local maxima and the lower envelope defined by the local minima is 0. For the first-order IMF, if the conditions are not met, then... See as new , Given the mean of its upper and lower envelopes, we have ,like If not satisfied, repeat the process. Next, we obtain the following formula:
[0046] like and If the standard deviation SD is within the predetermined range, then the process is stopped. Original load signal The first-tier IMF, denoted as The standard deviation The calculation formula is:
[0047] In the formula, For load signal Total time length; Step S204, let ,Will See as new Repeat steps S201 to S203 to obtain the second-order IMF, denoted as Then, by analogy, other IMFs of various orders are obtained, denoted as follows: ... ,until The function is a monotonic function until it can no longer be divided into IMFs; at this point, it is denoted as... The signal after noise reduction of environmental sound signals in the forest area.
[0048] The method is illustrated below with specific embodiments. For noise signals... Its expression is ,in: Sampling time, ∈[0,2], It is a sine function. Pi It is Gaussian white noise with a mean of 0 and a variance of 1.
[0049] (1) Traverse the original noise signal Find all local maxima and local minima; (2) Fit the local maxima and local minima with cubic spline curves respectively to obtain the upper envelope of the signal. and lower envelope ; (3) Calculate the mean of the upper and lower envelopes to obtain a mean envelope. ; (4) Subtract the mean envelope from the original signal to obtain the preliminary components. ; (5) Repeat steps (1)-(4) above for multiple iterations until... If the two criteria of the IMF are met, then This is the first IMF component, denoted as IMF1; (6) Using the original signal Subtract IMF1 to obtain the remaining signal ; (7) For the remaining signal Repeat steps (1)-(6) above to sequentially extract IMF2, IMF3, ..., IMFn until the remaining signal is extracted. The IMF's criteria can no longer be met, at this point... That is, the residual term; (8) The decomposition ends, and the superposition of all IMF components and residual terms is the original signal.
[0050] Original signal after decomposition There are the following three signals , , and a residual signal :
[0051]
[0052]
[0053]
[0054] Remove residual signals back, , , The original signal can be obtained by adding them together. Noise-reduced signal .
[0055] In some embodiments, noise reduction of the acquired vibration data includes: performing low-pass filtering and high-pass filtering on the acquired vibration data. During low-pass filtering, the cutoff frequency is set to 2-3 times the highest fault characteristic frequency of the original signal (e.g., the fault frequency of a water pump bearing is about 130Hz, so the cutoff frequency is set to 400Hz) to filter high-frequency noise while retaining fault characteristics and eliminating the influence of high-frequency noise such as electromagnetic interference. The vibration data after low-pass filtering is then subjected to high-pass filtering, with the cutoff frequency set to be lower than the lowest fault characteristic frequency (e.g., the rotation frequency is 25Hz, so the cutoff frequency is set to 10Hz) to filter low-frequency background vibration interference.
[0056] In some embodiments, resampling meteorological sensor data and continuous environmental monitoring data includes: determining a resampling factor; designing a filter; and resampling the data based on the resampling factor and the filter; wherein, when determining the resampling factor, it is assumed that the original sampling rate is fs. old The target sampling rate is fs new The resampling factor is: factor=fs new / fs old =L / M Where L and M are coprime integers, and L and M are chosen such that L / M is equal to or very close to the factor; when designing the filter, a low-pass filter is designed to simultaneously perform the tasks of anti-mirror and anti-aliasing, and the cutoff frequency of the filter is: fc = min(fs old , fs new ) / 2.0×0.9 Based on the resampling factor, the designed filter is used to perform filtering operations on the original signal to obtain the filtered signal x. filter For the filtered signal, one sample is extracted for every M samples to obtain the resampled signal.
[0057] The method is illustrated below with specific embodiments. Assume that the collected environmental monitoring time series signal... Its expression is Where t is the sampling time, t∈[0,2], It is a sine function. Pi It is Gaussian white noise with a mean of 0 and a variance of 1.
[0058] (1) Calculate the resampling factor: the original sampling frequency fs of the signal old The Hz frequency is 2048 Hz, and the 2-second data contains 4096 points, which is a large amount of data. The target sampling rate during resampling is fs. new The frequency was set to 1024Hz, reducing the data volume by half while ensuring no loss of environmental change characteristics, thus improving analysis efficiency. The resampling factor was [value missing]. factor=fs new / fs old =1 / 2 (2) Filter design: Calculate the cutoff frequency of the filter, design the filter, and filter the original signal to obtain the filtered signal. .
[0059] (3) Resampling: for Data is extracted according to the rule of "taking 1 point for every 2 points" (that is, keeping the 1st, 3rd, 5th...4095th points, or the 2nd, 4th, 6th...4096th points, the two methods have the same effect).
[0060] The original signal underwent a Fast Fourier Transform (FFT), with peak frequencies concentrated at 25Hz, 50Hz, 105.3Hz, and 129.7Hz. The resampled signal was then subjected to an FFT. Figure 4 The graph shows a spectrum comparison, revealing that the peak frequency of the resampled signal is identical to that of the original signal. After resampling, the signal data size is reduced by half, while all core frequency components are fully preserved without distortion.
[0061] In some embodiments, encoders corresponding to each modality of data are established, including: using convolutional neural networks to analyze the median-filtered visible light image and thermal imaging image, sliding a fixed-size convolution kernel (such as 3×3 or 5×5), performing convolution operations on local pixel regions of the image, calculating the local correlation between pixels, and outputting feature vectors. For example, for infrared images of forest fire prevention equipment, a 3×3 convolutional kernel can capture the grayscale variation relationship between a pixel and its 8 neighbors, identifying features such as the edge and texture of the temperature gradient; for denoised vibration data, a 1D neural network is used for analysis, with short convolutional kernels (length 3 / 5 / 7) sliding to cover the local temporal window, calculating the correlation of amplitude changes between adjacent sampling points, and capturing features such as "amplitude abrupt change (fault impact), local fluctuation (oscillation trend), and boundary of stable segment"; for denoised sound data, STFT / CWT is used to convert the noise sequence into a time-frequency image, and 2D CNN is used to extract the texture distribution features of the time-frequency image (e.g., uniform texture = Gaussian noise, discrete bright spots = impulse noise), outputting a time-frequency feature vector; for text data, the frequency of each word in the text is counted, generating a vocabulary list containing all core words, and for a single text, the number of times each word appears in the vocabulary list is counted, generating a vector.
[0062] In some embodiments, the feature vectors of each modality are fused at the feature level, including: each modal data first extracts high-level features through its corresponding encoder, and then performs fusion at the feature level using an attention mechanism. Specifically, the modal feature dimensions of image, vibration, sound, and text are unified to the same target dimension to solve the problem of inconsistent dimensions among different data modalities. Using the modal dimension with the largest modal feature dimension as the standard, zero values are padded to the end of the feature vectors of other data, keeping the original feature information unchanged and only expanding the dimension; this operation is simple and computationally insignificant. Weights are assigned according to the variance of the modal features (the larger the variance, the richer the information, and the higher the weight), and the five types of modal features are weighted and summed to perform feature fusion. The calculation method is as follows:
[0063]
[0064] in, (1) The eigenvectors after summation, , , , , These are the aligned image, thermal imaging, vibration, sound, and text modal feature vectors, respectively. (2) , , , , The fusion weights are respectively for image, thermal imaging, vibration, sound, and text modal features, satisfying... ; (3) , , , , These are the variances of the feature vectors for image, thermal imaging, vibration, sound, and text modalities, respectively. , , , The calculation method is similar.
[0065] In some embodiments, a trained model is used to calculate the reconstruction error of real-time data based on the learned "normal" pattern. An alarm is triggered when the error exceeds a set threshold, resulting in a calibrated model. This includes: using the trained model to calculate the reconstruction error of real-time data or its deviation from the "non-fire risk state cluster" in the joint embedding space based on the learned "normal forest environment state" pattern. An alarm is triggered when the error exceeds a set threshold. The threshold is repeatedly calibrated using existing historical fire risk or fire cases until the model output results are consistent with the actual fire risk situation, resulting in a calibrated model.
[0066] See Figure 2This application also provides a forest fire risk monitoring system based on multimodal feature coding and attention fusion. The system uses the forest fire risk monitoring method based on multimodal feature coding and attention fusion as described in the above embodiments for fire risk monitoring. The system includes: a multimodal data acquisition module, a preprocessing module, a feature extraction module, a model training module, a model debugging module, and a technology application module connected in sequence. The multimodal data acquisition module is used to acquire multimodal data, including thermal imaging data, visible light image data, vibration data, environmental sound data, meteorological sensor data, vegetation state parameters, and forest area environmental condition parameters. The preprocessing module is used to preprocess the multimodal data, establish a unified timestamp, and ensure that data from different modalities are aligned in time. The feature extraction module... The module is used to establish encoders corresponding to each modality of data, mapping the preprocessed multimodal data into a high-dimensional feature vector, and extracting high-level features of forest environmental status and fire risk evolution. The model training module is used to perform multimodal fusion of the feature vectors of each modality at the feature level to train the model, enabling the model to automatically learn under different meteorological conditions and forest environmental conditions, select the modal information that needs to be focused on, and obtain a trained model. The model debugging module is used to use the trained model, based on the learned "normal" mode, to calculate the reconstruction error of real-time data, and to alarm when it exceeds the set threshold, thus obtaining a debugged model. The technology application module is used to use the debugged model to identify high-risk fire risk types, predict the remaining time for the fire risk to develop from the current state to an out-of-control state, and realize early intervention and prevention.
[0067] In some embodiments, the thermal imaging data is infrared temperature image data collected during forest area monitoring. The resolution of a single frame is 320×240 pixels, and the corresponding temperature range is 20 ℃~120 ℃. The data is stored in the form of a continuous frame sequence with a sampling time interval of 0.2 s, and is used to characterize the temperature distribution of the forest surface, vegetation and local areas.
[0068] The visible light image data is the visible light image data collected during the forest monitoring process. The resolution of a single frame image is 1920×1080 pixels, the sampling frame rate is 15 frames / second, and the image content includes the distribution of forest vegetation, surface conditions, smoke characteristics and surrounding environmental information.
[0069] The vibration data consists of vibration acceleration signals collected from key parts of the equipment during operation. The vibration signal sampling frequency is 12.8 kHz, and the single-channel vibration amplitude range is -10g to 10g. The data is stored in the form of a continuous time series.
[0070] The environmental sound data refers to the environmental sound signal data collected during the monitoring of forest areas. The sound signal sampling frequency is 16kHz, the quantization accuracy is 16 bits, and the length of a single sound sample is 5s. It is used to reflect the acoustic characteristics of environmental noise, abnormal sounds, and potential fire hazards in forest areas.
[0071] The meteorological sensor data is continuous environmental time series data collected from forest area monitoring points, including parameters such as temperature, humidity, wind speed, and wind direction. The sampling frequency is 1Hz to 10Hz, and it is stored in the form of continuous time series.
[0072] The vegetation status parameters are indicators such as vegetation moisture content and combustible load obtained through remote sensing or ground sensing equipment, which are used to reflect the state of forest combustibles and the degree of fire risk sensitivity.
[0073] The environmental parameters of the forest area are environmental and management parameters collected synchronously in the forest area, including precipitation, intensity of human activities, etc., with a sampling period of 1s to 10s.
[0074] Figure 5 The images show thermal images of forest areas. The left thermal image is of a forest area under normal fire-free conditions, with a relatively high temperature center but no obvious abnormally high temperature areas. The right thermal image is of a local high temperature hazard area and a potential fire hazard area, showing that there is a significant abnormal temperature rise in some areas, indicating the possibility of a fire.
[0075] Visible light image data consists of images collected during forest monitoring. Each frame has a resolution of 1920×1080 pixels and a sampling frame rate of 15 frames per second. The images include forest vegetation distribution, surface conditions, smoke characteristics, and surrounding environmental information. Vibration data consists of vibration acceleration signals collected from key parts of the equipment during operation. The vibration signal sampling frequency is 12.8 kHz, and the single-channel vibration amplitude range is -10g to 10g, stored in continuous time series format. Environmental sound data consists of environmental sound signals collected during forest area monitoring. The sound signal sampling frequency is 16 kHz, and the quantization precision is 16. The data consists of a bit sample, with a single sound sample length of 5 seconds, used to reflect the acoustic characteristics related to environmental noise, abnormal sounds, and potential fire hazards in forest areas; meteorological sensor data is continuous environmental time-series data collected from forest monitoring points, including parameters such as temperature, humidity, wind speed, and wind direction, with a sampling frequency of 1Hz to 10Hz, stored in continuous time-series format; vegetation status parameters are indicators such as vegetation moisture content and combustible load obtained through remote sensing or ground sensing equipment, used to reflect the state of forest combustibles and fire hazard sensitivity; forest area environmental operating condition parameters are environmental and management parameters collected synchronously in forest areas, including precipitation and intensity of human activities, with a sampling period of 1s to 10s.
[0076] In summary, the key innovation of this application lies in: (1) This application proposes a multimodal fire risk fusion method based on adaptive allocation of attention weights. Unlike the existing technology that simply splices or weights multimodal features, this application introduces an attention mechanism at the feature layer, enabling the model to dynamically adjust the contribution weights of each modality according to different meteorological conditions, environmental states and fire risk development stages. Thus, even when multiple fire risk factors are superimposed or environmental features interfere with each other, it can still accurately focus on the modal information that is most discriminative for current fire risk identification.
[0077] (2) This application proposes a fire risk anomaly detection mechanism based on normal forest environment patterns without prior knowledge. This application uses normal forest environment and non-fire risk state data as learning objects to establish a normal state feature space. By monitoring the reconstruction error or deviation of data in this space in real time, fire risk anomaly identification and early warning are achieved. It does not rely on a clear fire risk threshold or a large number of known fire samples, and solves the problem that the scarcity of forest fire samples and the strong randomness of their occurrence lead to insufficient generalization ability of traditional supervised models.
[0078] (3) This application also constructs a multimodal unified representation mechanism for complex forest environmental conditions. In view of the characteristics of forest areas such as wide spatial range, complex environmental elements and frequent changes in meteorological conditions, this application encodes and fuses thermal imaging, visible light images, environmental sound, meteorological monitoring data, vegetation status parameters and forest area environmental conditions parameters under a unified time reference to form a unified feature representation that can simultaneously characterize abnormal temperature rise, smoke changes, environmental acoustic anomalies and changes in fire risk factors, significantly improving the integrity and stability of forest fire risk perception under complex natural environmental conditions.
[0079] Research and development results diagram Figure 4 and Figure 5 As shown. Figure 4 and Figure 5 The images show a spectral comparison chart and a thermal image of the forest area. Figure 4 As can be seen, the peak frequency of the signal after resampling is consistent with the original signal. After resampling, the signal data volume is reduced by half, and all core frequency components are completely preserved without distortion. Figure 5 The thermal image on the left shows a forest area under normal fire-free conditions. The temperature distribution center is relatively high, but there are no obvious abnormally high temperature areas. The thermal image on the right shows a local high temperature hazard area and a potential fire hazard area. It can be seen that there is a significant abnormal temperature rise in some areas, which may indicate the possibility of a fire.
[0080] This application proposes a forest fire risk monitoring method based on multimodal feature coding and attention fusion, achieving integrated early warning of fire risk anomaly identification, risk type determination, and development trend assessment. Within a unified model framework, this application realizes a continuous analysis process from anomaly fire risk detection, fire risk type and level determination to fire risk development trend assessment, transforming forest fire prevention from a reactive response to a proactive early warning and intervention, significantly reducing the risk of major forest fires, and improving the intelligence level and overall prevention and control effectiveness of forest fire monitoring and management.
[0081] Based on the above technical solutions, this application provides a forest fire risk monitoring method based on multimodal feature encoding and attention fusion, including the following steps: Step S1, collecting multimodal data; the multimodal data includes visible light image data, thermal imaging data, vibration data, temperature data, operation logs, environmental sound data, meteorological sensor data, vegetation status parameters, and forest area environmental condition parameters of power support equipment (such as water pumps, generators, etc.), monitoring and sensing equipment (such as infrared cameras), and communication equipment; Step S2, preprocessing the multimodal data, establishing a unified timestamp to ensure that the data of different modalities are aligned in time; Step S3, establishing encoders corresponding to each modal data, mapping the preprocessed multimodal data into a high-dimensional feature vector, and extracting high-level features of forest environmental status and fire risk evolution; Step S4: Perform multimodal fusion of the feature vectors of each modality at the feature level. Pre-train the model with a large amount of data in the source domain (such as a laboratory simulation environment or a mature forest area with abundant historical data) to learn common fault characteristics so that the model can automatically learn under different meteorological conditions and forest environment conditions. Select the modal information that needs to be focused on to obtain a trained model. Step S5: Transfer the pre-trained model to the target domain (new forest area). Based on the learned "normal" mode, fine-tune the top layer of the model with only a small amount of new data to make it quickly adapt to the subtle differences in the new environment. Calculate the reconstruction error of real-time data and alarm when it exceeds the set threshold to obtain a debugged model. Step S6: Use the debugged model to identify high-risk fire hazard types and predict the remaining time for the fire hazard to develop from the current state to an out-of-control state to achieve early intervention and prevention.
[0082] Those skilled in the art will understand that the above-described embodiments are specific examples of implementing this application, and in practical applications, various changes in form and detail may be made without departing from the spirit and scope of this application. Any person skilled in the art can make their own modifications and alterations without departing from the spirit and scope of this application; therefore, the scope of protection of this application should be determined by the scope defined in the claims.
Claims
1. A forest fire risk monitoring method based on multimodal feature encoding and attention fusion, characterized in that, Includes the following steps: Step S1: Collect multimodal data; the multimodal data includes visible light image data, thermal imaging data, vibration data, temperature data, operation logs, environmental sound data, meteorological sensor data, vegetation status parameters, and forest area environmental condition parameters of power support equipment, monitoring and sensing equipment, and communication equipment. Step S2: Preprocess the multimodal data and establish a unified timestamp to ensure that the data of different modalities are aligned in time. Step S3: Establish encoders corresponding to each modal data, map the preprocessed multimodal data into a high-dimensional feature vector, and extract high-level features of forest environmental status and fire risk evolution. Step S4: Perform multimodal fusion of the feature vectors of each modality at the feature level, pre-train the model with a large amount of data in the source domain, so that it learns the common fault features and can automatically learn under different meteorological conditions and forest environment conditions, select the modal information that needs to be focused on, and obtain the trained model. Step S5: Transfer the pre-trained model to the target domain. Based on the learned "normal" pattern, fine-tune the top layer of the model with only a small amount of new data to enable it to quickly adapt to the subtle differences in the new environment. Calculate the reconstruction error of the real-time data. If the error exceeds the set threshold, an alarm is triggered, and the debugged model is obtained. Step S6: Use the debugged model to identify high-risk fire hazard types, predict the remaining time for the fire hazard to develop from its current state to an out-of-control state, and achieve early intervention and prevention.
2. The forest fire risk monitoring method based on multimodal feature encoding and attention fusion according to claim 1, characterized in that, In step S2, the multimodal data is preprocessed, including: Noise reduction processing is performed on visible light image data, thermal imaging data, sound data, and vibration data respectively; Meteorological sensor data and continuous environmental monitoring data are resampled.
3. The forest fire risk monitoring method based on multimodal feature encoding and attention fusion according to claim 2, characterized in that, Noise reduction for visible light image data and thermal imaging data, including: Bilateral filtering noise reduction is performed on visible light image data and thermal imaging data, and the location in the image... The pixel value after bilateral filtering The following formula is used to calculate: in, As the normalization factor, It revolves around pixels neighborhood window, It is the location of neighboring pixels. and They are pixels and The intensity value, It is a spatial domain weight function. , It is a range weight function. , and These are the standard deviations of the spatial domain and the range, respectively.
4. The forest fire risk monitoring method based on multimodal feature encoding and attention fusion according to claim 2, characterized in that, Noise reduction processing of audio data includes: Step S201: Process the collected environmental sound data Empirical mode decomposition is performed to remove spurious components, followed by noise reduction. Step S202: Determine the acquired raw noise signal For all extreme points, the upper envelope curve of the data is formed by cubic spline interpolation over all maximum and minimum points. and lower envelope curve and take and average As shown in the following formula: Step S203, let ,like Simultaneously satisfying both of the IMF's conditions: (1) Throughout the entire time course, the number of times the zero point is crossed is equal to or at most differs from the number of extreme points by 1. (2) If the upper envelope defined by the local maxima and the lower envelope defined by the local minima are locally symmetric about the time axis, then the mean of the upper envelope defined by the local maxima and the lower envelope defined by the local minima is 0. For the first-order IMF, if the conditions are not met, then... See as new , Given the mean of its upper and lower envelopes, we have ,like If not satisfied, repeat the process. Next, we obtain the following formula: like and If the standard deviation SD is within the predetermined range, then the process is stopped. Original load signal The first-tier IMF, denoted as The standard deviation The calculation formula is: In the formula, For load signal Total time length; Step S204, let ,Will See as new Repeat steps S201 to S203 to obtain the second-order IMF, denoted as Then, by analogy, other IMFs of various orders are obtained, denoted as follows: ... ,until The function is a monotonic function until it can no longer be divided into IMFs; at this point, it is denoted as... The signal after noise reduction of environmental sound signals in the forest area.
5. The forest fire risk monitoring method based on multimodal feature encoding and attention fusion according to claim 2, characterized in that, The collected vibration data is denoised, including: The collected vibration data is low-pass filtered, and the cutoff frequency is set to 2 to 3 times the highest fault characteristic frequency of the original signal to filter high-frequency noise while retaining fault characteristics and eliminating the influence of electromagnetic interference and high-frequency noise. The vibration data after low-pass filtering is then subjected to high-pass filtering, with the cutoff frequency set below the lowest fault characteristic frequency to filter out low-frequency background vibration interference.
6. The forest fire risk monitoring method based on multimodal feature encoding and attention fusion according to claim 2, characterized in that, Resampling of meteorological sensor data and continuous environmental monitoring data includes: Determine the resampling factor; Design filters; Data is resampled based on resampling factors and filters; In determining the resampling factor, it is assumed that the original sampling rate is fs. old The target sampling rate is fs new The resampling factor is: factor=fs new / fs old =L / M Where L and M are coprime integers, and L and M are chosen such that L / M is equal to or very close to the factor; when designing the filter, a low-pass filter is designed to simultaneously perform the tasks of anti-mirror and anti-aliasing, and the cutoff frequency of the filter is: fc = min(fs old , fs new ) / 2.0×0.9 Based on the resampling factor, the designed filter is used to perform filtering operations on the original signal to obtain the filtered signal x. filter For the filtered signal, one sample is extracted for every M samples to obtain the resampled signal.
7. The forest fire risk monitoring method based on multimodal feature encoding and attention fusion according to claim 1, characterized in that, In step S3, encoders corresponding to each modal data are established, including: Convolutional neural networks are used to analyze the median-filtered visible light images and thermal images. A fixed-size convolution kernel is slid to perform convolution operations on local pixel regions of the image, calculate the local correlation between pixels, and output feature vectors. The denoised vibration data were analyzed using a 1D neural network with short convolutional kernels in the convolutional layers. The local temporal window was covered by sliding to calculate the correlation of amplitude changes between adjacent sampling points and capture the features of "amplitude abrupt change, local fluctuation, and boundary of stable segment". For the denoised audio data, the noise sequence is converted into a time-frequency image using STFT / CWT, and the texture distribution features of the time-frequency image are extracted using 2D CNN to output the time-frequency feature vector. For text data, the frequency of each word in the text is counted, and a vocabulary list containing all core words is generated. For a single text, the number of times each word appears in the vocabulary list is counted, and a vector is generated.
8. The forest fire risk monitoring method based on multimodal feature encoding and attention fusion according to claim 1, characterized in that, In step S4, the feature vectors of each modality are fused at the feature level, including: This method unifies the modal feature dimensions of images, vibrations, sounds, and texts into the same target dimension, resolving the issue of inconsistent dimensions among different data modalities. It uses the modal dimension with the largest modal feature dimension as the standard, and fills the tail of the feature vectors of other data with 0 values, keeping the original feature information unchanged and only expanding the dimension. This method is simple to operate and requires very little computation. Weights are assigned based on the variance of the modal features, and the five types of modal features are weighted and summed to perform feature fusion. The calculation method is as follows: in, (1) The eigenvectors after summation, , , , , These are the aligned image, thermal imaging, vibration, sound, and text modal feature vectors, respectively. (2) , , , , The fusion weights are respectively for image, thermal imaging, vibration, sound, and text modal features, satisfying... ; (3) , , , , These are the variances of the feature vectors for image, thermal imaging, vibration, sound, and text modalities, respectively. , , , The calculation method is similar.
9. The forest fire risk monitoring method based on multimodal feature encoding and attention fusion according to claim 1, characterized in that, In step S5, the trained model is used to calculate the reconstruction error of real-time data based on the learned "normal" pattern. An alarm is triggered when the error exceeds a set threshold, resulting in a calibrated model, including: Using the trained model, based on the learned "normal forest environment state" pattern, the reconstruction error of real-time data or its deviation from the "non-fire risk state cluster" in the joint embedding space is calculated, and an alarm is triggered if the error exceeds the set threshold. The threshold was repeatedly adjusted using existing historical fire risk or fire case studies until the model output results were consistent with the actual fire risk situation.
10. A forest fire risk monitoring system based on multimodal feature encoding and attention fusion, wherein the forest fire risk monitoring method based on multimodal feature encoding and attention fusion as described in any one of claims 1 to 9 is used for fire risk monitoring, characterized in that, The system includes: a multimodal data acquisition module, a preprocessing module, a feature extraction module, a model training module, a model debugging module, and a technology application module, connected in sequence; among them, The multimodal data acquisition module is used to acquire multimodal data; the multimodal data includes visible light image data, thermal imaging data, vibration data, temperature data, operation logs, environmental sound data, meteorological sensor data, vegetation status parameters, and forest area environmental condition parameters of power support equipment, monitoring and sensing equipment, and communication equipment. The preprocessing module is used to preprocess the multimodal data, establish a unified timestamp, and ensure that the data of different modalities are aligned in time. The feature extraction module is used to establish encoders corresponding to each modal data, map the preprocessed multimodal data into a high-dimensional feature vector, and extract high-level features of forest environmental status and fire risk evolution. The model training module is used to perform multimodal fusion of the feature vectors of each modality at the feature level, pre-train the model with a large amount of data in the source domain, so that it learns the common fault features and can automatically learn under different meteorological conditions and forest environment conditions, select the modal information that needs to be focused on, and obtain the trained model. The model debugging module is used to transfer the pre-trained model to the target domain. Based on the learned "normal" pattern, it fine-tunes the top layer of the model with only a small amount of new data, so that it can quickly adapt to the subtle differences in the new environment. It calculates the reconstruction error of the real-time data, and alarms when it exceeds the set threshold, thus obtaining the debugged model. The technology application module is used to identify high-risk fire hazard types using a debugged model, predict the remaining time for a fire hazard to develop from its current state to an out-of-control state, and achieve early intervention and prevention.