Multimodal fire burn detection method, device, equipment and medium
By setting a memory buffer in the semantic segmentation model, the dependence on single modal data and 'catastrophic forgetting' in fire track detection is solved, and efficient integration of multimodal data and the accuracy and timeliness of detection results are achieved.
Patent Information
- Application Number
- CN202411451427.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-17
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-10-17
AI Technical Summary
The existing fire-spot detection technology relies on single-modal remote sensing data, and there are factors such as seasonal changes in vegetation, cloud cover and light conditions that affect the accuracy and timeliness of detection. The model is prone to ‘catastrophic forgetting’ when learning new tasks.
Setting memory buffers in the semantic segmentation model enables the model to continuously learn and update, adaptively process multimodal data, and alleviate the problem of ‘catastrophic forgetting’.
Through the integration of multimodal data and the use of memory buffers, the cross-modal generalization ability and applicability of fire trace detection are improved, ensuring the accuracy and timeliness of detection.
Smart Images

Figure CN119251682B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of fire mark detection, and in particular relates to a multi-modal fire mark detection method, device, equipment and medium. Background Art
[0002] At present, the detection of burnt areas mainly relies on remote sensing imaging technology, which analyzes images obtained by satellites, drones and other equipment. Traditional burnt area detection technology mainly uses single-modality remote sensing data to extract specific spectral indices, texture information and multi-spectral band information, and distinguishes between burned and unburned areas based on appropriate thresholds or using machine learning methods such as random forests. However, these single-modality methods face multiple challenges, such as seasonal changes in vegetation, cloud cover and changes in lighting conditions, which may affect the accuracy and timeliness of detection.
[0003] Due to the limitations of imaging conditions, such as weather, time, terrain and other factors, it is impossible to obtain image data of a specific modality every time a detection is performed, which affects the timeliness and accuracy of the burn detection results. In order to overcome the limitations of specific modality data, it is necessary to design a multimodal burn detection model that can process image data of different modalities, such as optical images, infrared images, radar images, etc., reduce the dependence of the burn detection model on single modality data, and improve the cross-modal generalization ability and applicability of the model.
[0004] However, cross-modal detection models face another challenging problem, that is, the model may not be able to obtain data from all modalities at once during the training process. In real-world problems, data from different modalities may need to be acquired gradually, which requires the detection model to have the ability to continuously learn new modal knowledge, to continuously learn new modal data, and to update and optimize the detection model in a timely manner. However, general machine learning models face the problem of "catastrophic forgetting" when learning new tasks, that is, when learning new knowledge, they will forget the old knowledge they have learned before, resulting in poor performance of the model on old tasks. Summary of the invention
[0005] 1. Technical issues to be resolved
[0006] In order to solve at least one of the above-mentioned technical problems arising in the detection of burnt areas in the prior art, the embodiments of the present invention provide a multimodal burnt area detection method, apparatus, equipment and medium. By setting a memory buffer in the semantic segmentation model, the model can be continuously learned, adaptively updated and optimized according to new modal data, and the catastrophic forgetting problem can be alleviated.
[0007] (II) Technical solution
[0008] In view of the above problems, embodiments of the present invention provide a multi-modal fire burn detection method, device, equipment and medium.
[0009] According to a first aspect of the present invention, a multimodal burnt area detection method is provided, comprising: acquiring target burnt area remote sensing observation data; performing feature extraction on the target burnt area remote sensing observation data to obtain target multidimensional feature data; and inputting the target multidimensional feature data into a pre-trained target semantic segmentation model to perform burnt area detection to obtain burnt area detection data corresponding to the target burnt area remote sensing observation data, wherein the target burnt area remote sensing observation data includes data before and after a fire in the same area; the target semantic segmentation model is capable of processing multiple modalities of burnt area remote sensing observation data; and the target semantic segmentation model includes a memory buffer to enable the model to continuously learn, and adaptively update and optimize the model according to new modal data to alleviate the problem of catastrophic forgetting.
[0010] In some exemplary embodiments, pre-training a target semantic segmentation model includes: obtaining multiple groups of task data grouped by modality, each group of task data serving as an independent training task; constructing an initial semantic segmentation model for burn site detection, the initial semantic segmentation model including a memory buffer; training the initial semantic segmentation model using a first group of task data, calculating the segmentation loss using a first loss function, updating the parameters in the initial semantic segmentation model according to the segmentation loss, and updating the memory buffer with data from the current task; and training subsequent task data in sequence, updating the parameters of the semantic segmentation model trained with the last task data, and updating the memory buffer with data from the current task, until the training of all task data is completed to obtain the target semantic segmentation model, wherein the first group of task data is any one of the multiple groups of task data.
[0011] In some exemplary embodiments, obtaining multiple groups of task data grouped by modality includes: obtaining multimodal burnt area remote sensing observation data and burnt area mask data corresponding to the burnt area remote sensing observation data; performing feature extraction on the multimodal burnt area remote sensing observation data and the burnt area mask data to obtain multidimensional feature data; and grouping the multidimensional feature data and the mask data by modality to obtain task data grouped by modality, wherein multidimensional feature data of different modalities have the same feature dimension; the multidimensional feature data contains information before and after the fire; and the task data is divided into training samples and verification samples according to a preset ratio.
[0012] In some exemplary embodiments, the initial semantic segmentation model uses a deep learning semantic segmentation model that combines ResNet and U-Net, wherein ResNet serves as the encoder of U-Net, and each layer in the encoder is spliced with the corresponding layer in the decoder, so that the deep learning semantic segmentation model can learn both shallow and deep features at the same time.
[0013] In some exemplary embodiments, the memory buffer is used to store the training samples used by the semantic segmentation model in the previous task training process and the output results of the semantic segmentation model for the training samples obtained after the previous task training, wherein, if the amount of data stored in the memory buffer is less than the size of the memory buffer, the data being processed is directly stored in the memory buffer; if the amount of data stored in the memory buffer is equal to the size of the memory buffer, it is randomly determined according to the reservoir sampling method whether to replace the task data stored in the memory buffer with the data being processed.
[0014] In some exemplary embodiments, training subsequent task data in sequence, updating the parameters of the semantic segmentation model trained with the last task data, and updating the memory buffer with the data of the current task include: calculating the segmentation loss using a first loss function based on the semantic segmentation model trained with the last task data and the current task data to obtain the segmentation loss; calculating the mean square loss as the distillation loss using a second loss function based on the task data stored in the memory buffer to reduce the model's forgetting of old task knowledge during the current task training process; constructing a third loss function using the weighted sum of the segmentation loss and the distillation loss to calculate the total loss; and updating the parameters of the semantic segmentation model based on the total loss.
[0015] In some exemplary embodiments, the segmentation loss function includes a focal loss function.
[0016] The second aspect of the present invention provides a multimodal burnt area detection device, comprising the following modules: an acquisition module, used to acquire target burnt area remote sensing observation data; a feature extraction module, used to perform feature extraction on the target burnt area remote sensing observation data to obtain target multidimensional feature data; and a detection module, used to input the target multidimensional feature data into a pre-trained target semantic segmentation model for burnt area detection, and obtain burnt area detection data corresponding to the target burnt area remote sensing observation data, wherein the target burnt area remote sensing observation data includes data before and after the fire in the same area; the target semantic segmentation model is capable of processing multimodal burnt area remote sensing observation data; and the target semantic segmentation model includes a memory buffer so that the model can continue to learn, and adaptively update and optimize the model according to new modal data to alleviate the problem of catastrophic forgetting.
[0017] A third aspect of the present invention provides an electronic device, comprising: one or more processors and a memory, the memory being used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors execute the above method.
[0018] The fourth aspect of the present invention also provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, causes the processor to execute the above method.
[0019] (III) Beneficial effects
[0020] It can be seen from the above technical solutions that the multi-modal fire burn detection method, device, equipment and medium provided by the embodiments of the present invention have at least one of the following beneficial effects:
[0021] (1) By setting a memory buffer in the semantic segmentation model, the model can continue to learn, adaptively update and optimize the model according to new modal data, and alleviate the catastrophic forgetting problem.
[0022] (2) The present invention can integrate multimodal data and maintain efficient detection performance even when data availability is limited.
[0023] (3) By dividing the training tasks into different modalities and using a memory buffer, the model can retain the memory of old knowledge and learn the knowledge of new tasks, thereby improving the cross-modal generalization ability and applicability of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The above contents and other objects, features and advantages of the present invention will become more apparent through the following description of the embodiments of the present invention with reference to the accompanying drawings, in which:
[0025] Figure 1 A schematic diagram of a process of obtaining a target semantic segmentation model by pre-training according to an embodiment of the present invention is shown;
[0026] Figure 2 A schematic diagram of a process of acquiring multiple groups of task data grouped according to modalities according to an embodiment of the present invention is shown;
[0027] Figure 3 A schematic diagram of a process of sequentially training subsequent task data, updating parameters of a semantic segmentation model obtained by training the previous task data, and updating a memory buffer with data of the current task according to an embodiment of the present invention is shown;
[0028] Figure 4 A schematic diagram of a process flow of a multi-modal fire scar detection method according to an embodiment of the present invention is shown;
[0029] Figure 5 A schematic structural block diagram of a multi-modal fire scar detection device 800 according to an embodiment of the present invention is shown; and
[0030] Figure 6 A block diagram of an electronic device for a multi-modal burn mark detection method according to an embodiment of the present invention is schematically shown.
[0031] Illustration Description:
[0032] 600-electronic device; 601-processor; 602-read only memory (ROM); 603-random access memory (RAM); 604-bus; 605-input / output (I / O) interface; 606-input part; 607-output part; 608-storage part; 609-communication part; 610-drive; 611-removable medium. DETAILED DESCRIPTION
[0033] In order to make the purpose, technical scheme and advantages of the present invention more clearly understood, the present invention is further described in detail below in combination with specific embodiments and with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0034] Figure 1 A flowchart of a method for pre-training a target semantic segmentation model according to an embodiment of the present invention is schematically shown.
[0035] like Figure 1 As shown, a method for pre-training a target semantic segmentation model according to an embodiment of the present invention includes steps S110-S140.
[0036] In step S110, multiple groups of task data grouped according to modality are obtained, and each group of task data serves as an independent training task.
[0037] In some exemplary embodiments, step S110 includes steps S111-S113, see Figure 2 .
[0038] In step S111, multimodal burnt area remote sensing observation data and burnt area mask data corresponding to the burnt area remote sensing observation data are obtained, wherein the remote sensing observation data includes data before and after the fire in the same area.
[0039] For example, different satellite observation data of multiple fires are obtained as multimodal burnt area remote sensing observation data. The burnt area remote sensing observation data of the same fire includes pre-fire and post-fire data in the same area. The National Burned Area Composite (NBAC) burned area polygon data is obtained as burnt area mask data.
[0040] In step S112, feature extraction is performed on the multimodal burnt area remote sensing observation data and the burnt area mask data to obtain multidimensional feature data, wherein the multidimensional feature data of different modes have the same feature dimension; and the multidimensional feature data contains information before and after the fire.
[0041] For example, the pre-fire and post-fire remote sensing observation data of the same fire are cropped to the same spatial range and resampled to a spatial resolution of 20m. Feature extraction is performed on remote sensing observation data of different modes of burnt areas so that the remote sensing observation data of different modes of burnt areas have the same feature dimension, specifically including: for the first satellite data, extracting 3D features of red, near infrared, and short-wave infrared; for the second satellite data, obtain VV polarization data and VH polarization data, calculate the Normalized Difference Backscatter Index (NDBI), extract VV, VH, and NDBI as the 3D features of the second satellite data; for the third satellite data, obtain HH polarization data and HV polarization data, calculate NDBI, and extract HH, HV, and NDBI as the 3D features of the third satellite data.
[0042] For the second satellite, NDBI 第二卫星 The calculation formula is as follows:
[0043] (1)
[0044] For the third satellite NDBI 第三卫星 The calculation formula is as follows:
[0045] (2)
[0046] The data features before and after the fire are combined to generate 6-dimensional feature data to ensure that each data contains information before and after the fire, which is conducive to the model learning more relevant knowledge about the burnt area and the common features of the burnt areas between different modes, and obtaining more accurate detection results. Finally, all multi-modal remote sensing observation data of the burnt areas are standardized.
[0047] In step S113, the multidimensional feature data and the mask data are grouped according to the modality to obtain task data grouped according to the modality, wherein the task data is divided into training samples and verification samples according to a preset ratio.
[0048] For example, the remote sensing observation data of the burnt area of each mode after preprocessing and the corresponding burnt area mask data are used as independent training tasks, specifically including: dividing the optical data of the first satellite and the corresponding burnt area mask data into the first task, the SAR c-band data of the second satellite and the corresponding burnt area mask data into the second task, and the SAR l-band data of the second satellite and the corresponding burnt area mask data into the third task; 80% of all the remote sensing observation data and the corresponding burnt area mask data of each task are divided into training samples, and 20% are divided into verification samples.
[0049] In step S120, an initial semantic segmentation model is constructed for fire scar detection, and the initial semantic segmentation model includes a memory buffer.
[0050] In an embodiment of the present invention, a deep learning semantic segmentation model combining ResNet and U-Net is used, that is, ResNet is used as the encoder of U-Net, and each layer in the encoder is spliced with the corresponding layer in the decoder, which is beneficial for the deep learning semantic segmentation model to learn shallow and deep features at the same time; the deep learning semantic segmentation model requires that the input data have the same feature dimension, and the output is two-category probability map of "burned area" and "non-burned area" for each pixel, and the output has the same spatial resolution as the input.
[0051] In an embodiment of the present invention, a memory buffer with a fixed size of M is constructed, where M is 200. The size of the memory buffer determines the amount of data stored in the entire continuous learning task, which is used to prevent catastrophic forgetting of the model; the task data stored in the memory buffer includes the training samples used by the deep learning semantic segmentation model in the previous task training process and the output results of the deep learning semantic segmentation model after the previous task training for the training samples.
[0052] In step S130, the initial semantic segmentation model is trained using the first set of task data, the segmentation loss is calculated using the first loss function, the parameters in the initial semantic segmentation model are updated according to the segmentation loss, and the memory buffer is updated with the data of the current task, wherein the first set of task data is any set of multiple sets of task data.
[0053] In the embodiment of the present invention, when the deep learning semantic segmentation model starts training the first task, the segmentation loss of the current task is calculated. Since the model only needs to learn the knowledge of the current task, the model parameters are updated only according to the segmentation loss. In order to alleviate the problem of class imbalance between burned areas and non-burned areas in the semantic segmentation of burnt areas, and make the model pay more attention to difficult-to-train samples, Focal Loss is used as the segmentation loss, and the first loss function formula is as follows:
[0054] (3)
[0055] in, is the segmentation loss, For the mask, It is the output result of the deep learning semantic segmentation model during the training process of the current task.
[0056] Process the current task data in sequence. If the amount of data stored in the memory buffer is less than M, the data being processed will be directly stored in the memory buffer. If the amount of data stored in the memory buffer is equal to M, the reservoir sampling method is used to randomly decide whether to replace the task data stored in the memory buffer with the data being processed. This method can reduce the usage of computing memory while ensuring the uniformity and accuracy of the sampling data.
[0057] In step S140, subsequent task data are trained in sequence, the parameters of the semantic segmentation model obtained by the previous task data training are updated, and the memory buffer is updated with the data of the current task until the training of all task data is completed to obtain the target semantic segmentation model.
[0058] In some exemplary embodiments, step S140 includes steps S141 - S144 .
[0059] In step S141, based on the semantic segmentation model obtained by training the previous task data and the current task data, the segmentation loss is calculated using the first loss function to obtain the segmentation loss.
[0060] In step S142, based on the task data stored in the memory buffer, the second loss function is used to calculate the mean square loss as the distillation loss to reduce the model's forgetting of old task knowledge during the current task training process.
[0061] In the embodiment of the present invention, when the deep learning semantic segmentation model starts to train the second and third tasks, in addition to calculating the segmentation loss, it is also necessary to use the task data stored in the memory buffer to calculate the mean square loss as the distillation loss, so as to reduce the model's forgetting of old task knowledge during the current task training process. The second loss function calculation formula is as follows:
[0062] (4)
[0063] in, is the mean square loss, Outputs the results of the deep learning semantic segmentation model for the previously trained samples of the task stored in the memory buffer. It is the output result of the deep learning semantic segmentation model in the current task training process for the old task training sample data.
[0064] In step S143, a third loss function is constructed using the weighted sum of the segmentation loss and the distillation loss, and the total loss is calculated.
[0065] During the current task training process, the deep learning semantic segmentation model calculates the total loss based on the weighted sum of the segmentation loss and the distillation loss, and updates the deep learning semantic segmentation model parameters based on the total loss. The third loss function is used to calculate the total loss using the following formula:
[0066] (5)
[0067] in, is the total loss, is the segmentation loss, namely Focal Loss, is the mean square loss, i.e. the distillation loss.
[0068] In step S144, the parameters of the semantic segmentation model are updated based on the total loss.
[0069] Process the current task data in sequence. If the amount of data stored in the memory buffer is less than M, the data being processed will be directly stored in the memory buffer. If the amount of data stored in the memory buffer is equal to M, it is randomly decided according to the reservoir sampling method whether to replace the task data stored in the memory buffer with the data being processed.
[0070] In the embodiment of the present invention, when it comes to updating the memory buffer, a reservoir sampling method is used to store new task data, and an integer is randomly generated between 0 and the total number of all processed task data. If the integer is less than the fixed size M of the buffer, the integer is used as an index to replace the task data at the index position in the memory buffer with the data being processed; if the integer is greater than or equal to the fixed size M of the buffer, the data being processed is not stored in the memory buffer. The reservoir sampling method can ensure that each data has an equal probability of being stored in the memory buffer in a limited memory space.
[0071] Figure 4 The following is a schematic flow chart of a multi-modal burn mark detection method according to an embodiment of the present invention.
[0072] like Figure 4As shown, a multi-modal fire scar detection method according to an embodiment of the present invention uses Figure 1 The trained target semantic segmentation model is used for detection, including steps S210-S230.
[0073] In step S210, remote sensing observation data of the target fire-burned area is obtained, wherein the remote sensing observation data of the target fire-burned area includes data before and after the fire in the same area.
[0074] In step S220, feature extraction is performed on the target fire site remote sensing observation data to obtain target multi-dimensional feature data.
[0075] In step S230, the target multi-dimensional feature data is input into a pre-trained target semantic segmentation model for burnt area detection to obtain burnt area detection data corresponding to the target burnt area remote sensing observation data, wherein the target semantic segmentation model is capable of processing multiple modalities of burnt area remote sensing observation data; and the target semantic segmentation model includes a memory buffer to enable the model to continuously learn, and adaptively update and optimize the model according to new modal data to alleviate the problem of catastrophic forgetting.
[0076] For example, when the deep learning semantic segmentation model is trained in all tasks, the model is used to detect burnt areas on new data of the trained modality, and the F1 coefficient is calculated for the detection results. The F1 coefficient for the first satellite data is 0.893, the F1 for the second satellite data is 0.911, and the F1 for the third satellite data is 0.847.
[0077] Figure 5 The structure block diagram of a multi-modal fire mark detection device 800 according to an embodiment of the present invention is schematically shown.
[0078] like Figure 5 As shown, a multimodal burnt scar detection device 800 according to an embodiment of the present invention includes an acquisition module 810 , a feature extraction module 820 and a detection module 830 .
[0079] The acquisition module 810 is used to acquire remote sensing observation data of the target fire area.
[0080] The feature extraction module 820 is used to extract features from the target fire-scarred area remote sensing observation data to obtain target multi-dimensional feature data. The target fire-scarred area remote sensing observation data includes data before and after the fire in the same area.
[0081] The detection module 830 is used to input the target multi-dimensional feature data into a pre-trained target semantic segmentation model to perform burnt area detection, and obtain burnt area detection data corresponding to the target burnt area remote sensing observation data, wherein the target semantic segmentation model is capable of processing multimodal burnt area remote sensing observation data; and the target semantic segmentation model includes a memory buffer to enable the model to continuously learn, adaptively update and optimize the model according to new modal data, so as to alleviate the problem of catastrophic forgetting.
[0082] In some specific embodiments, any multiple modules among the acquisition module 810, the feature extraction module 820 and the detection module 830 can be combined into one module for implementation, or any one of the modules can be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module.
[0083] In some specific embodiments, at least one of the acquisition module 810, the feature extraction module 820, and the detection module 830 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or may be implemented by hardware or firmware in any other reasonable manner of integrating or packaging the circuit, or in any one of the three implementation methods of software, hardware, and firmware, or in any appropriate combination of any of them. Alternatively, at least one of the acquisition module 810, the feature extraction module 820, and the detection module 830 may be at least partially implemented as a computer program module, and when the computer program module is run, the corresponding function may be executed.
[0084] Figure 6 A block diagram of an electronic device for a multi-modal burn mark detection method according to an embodiment of the present invention is schematically shown.
[0085] like Figure 6 As shown, the electronic device 600 according to an embodiment of the present invention includes a processor 601, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 602 or the program loaded from the storage part 608 to the random access memory (RAM) 603. The processor 601 may include, for example, a general-purpose microprocessor (such as a CPU), an instruction set processor and / or a related chipset and / or a special-purpose microprocessor (for example, an application-specific integrated circuit (ASIC)), etc. The processor 601 may also include an on-board memory for caching purposes. The processor 601 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.
[0086] In RAM603, various programs and data required for the operation of electronic device 600 are stored. Processor 601, ROM602 and RAM603 are connected to each other via bus 604. Processor 601 performs various operations of the method flow according to an embodiment of the present invention by executing the program in ROM602 and / or RAM603. It should be noted that the program can also be stored in one or more memories other than ROM602 and RAM603. Processor 601 can also perform various operations of the method flow according to an embodiment of the present invention by executing the program stored in one or more memories.
[0087] In some specific embodiments, the electronic device 600 may further include an input / output (I / O) interface 605, which is also connected to the bus 604. The electronic device 600 may further include one or more of the following components connected to the I / O interface 605: an input portion 606 including a keyboard, a mouse, etc.; an output portion 607 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage portion 608 including a hard disk, etc.; and a communication portion 609 including a network interface card such as a LAN card, a modem, etc. The communication portion 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as needed, so that a computer program read therefrom is installed into the storage portion 608 as needed.
[0088] An embodiment of the present invention shows a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiment; or may exist independently without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiment of the present invention is implemented.
[0089] In some specific embodiments, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, may include but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, an apparatus or a device. For example, according to an embodiment of the present invention, the computer-readable storage medium may include the ROM602 and / or RAM603 described above and / or one or more memories other than ROM602 and RAM603.
[0090] An embodiment of the present invention shows a computer program product, which includes a computer program, and the computer program contains program code for executing the method shown in the flowchart. When the computer program product is run in a computer system, the program code is used to enable the computer system to implement the method provided by the embodiment of the present invention.
[0091] The computer program executes the above functions defined in the system / device of the embodiment of the present invention when it is executed by the processor 601. According to the embodiment of the present invention, the system, device, module, unit, etc. described above can be implemented by a computer program module.
[0092] In some specific embodiments, the computer program may rely on tangible storage media such as optical storage devices, magnetic storage devices, etc. In another embodiment, the computer program may also be transmitted and distributed in the form of signals on a network medium, and downloaded and installed through the communication part 609, and / or installed from a removable medium 611. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0093] In some specific embodiments, the computer program can be downloaded and installed from the network through the communication part 609, and / or installed from the removable medium 611. When the computer program is executed by the processor 601, the above functions defined in the system of the embodiment of the present invention are performed. According to the embodiment of the present invention, the system, device, apparatus, module, unit, etc. described above can be implemented by a computer program module.
[0094] In some specific embodiments, the program code for executing the computer program provided by the embodiment of the present invention can be written in any combination of one or more programming languages, specifically, these computing programs can be implemented using high-level process and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, such as Java, C++, python, "C" language or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on the remote computing device, or entirely on the remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., using an Internet service provider to connect through the Internet).
[0095] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present invention. In this regard, each box in the flow chart or block diagram can represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of the boxes in the block diagram or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0096] It will be appreciated by those skilled in the art that the features described in the various embodiments of the present invention may be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features described in the various embodiments of the present invention may be combined and / or combined in various ways. All of these combinations and / or combinations fall within the scope of the present invention.
[0097] The embodiments of the present invention are described above. However, these embodiments are only for the purpose of illustration, and are not intended to limit the scope of the present invention. Although each embodiment is described above, it does not mean that the measures in each embodiment cannot be used in combination advantageously. Without departing from the scope of the present invention, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present invention.
Claims
1. A multi-modal fire burn detection method, characterized in that: include: Obtain remote sensing observation data of target fire-scarred areas; Extracting features from the target fire scar remote sensing observation data to obtain target multi-dimensional feature data; as well as Inputting the target multidimensional feature data into a pre-trained target semantic segmentation model to perform burnt area detection, and obtaining burnt area detection data corresponding to the target burnt area remote sensing observation data, wherein the target burnt area remote sensing observation data includes data before and after the fire in the same area; The target semantic segmentation model can process multiple modalities of remote sensing observation data of burnt areas, and each modality of remote sensing observation data of burnt areas and corresponding mask data of burnt areas are used as independent training tasks; and The target semantic segmentation model includes a memory buffer so that the model can continue to learn and adaptively update and optimize the model according to new modality data; The memory buffer is used to store the training samples used by the semantic segmentation model in the previous task training process and the output results of the semantic segmentation model obtained after the previous task training for the training samples; The adaptive updating and optimization of the model according to the new modal data includes: Based on the semantic segmentation model obtained by training the previous task data and the current task data, the segmentation loss is calculated using the first loss function to obtain the segmentation loss; Based on the task data stored in the memory buffer, the second loss function is used to calculate the mean square loss as the distillation loss to reduce the model's forgetting of old task knowledge during the current task training process; Constructing a third loss function using a weighted sum of the segmentation loss and the distillation loss, and calculating a total loss; and Update the parameters of the semantic segmentation model based on the total loss.
2. The method according to claim 1, characterized in that The target semantic segmentation model obtained by pre-training includes: Obtain multiple groups of task data grouped by modality, with each group of task data serving as an independent training task; Constructing an initial semantic segmentation model for fire scar detection, wherein the initial semantic segmentation model includes a memory buffer; Using the first set of task data to train the initial semantic segmentation model, using a first loss function to calculate the segmentation loss, updating the parameters in the initial semantic segmentation model according to the segmentation loss, and updating the memory buffer with the data of the current task; and Train the subsequent task data in sequence, update the parameters of the semantic segmentation model obtained by the previous task data training, and update the memory buffer with the data of the current task until the training of all task data is completed to obtain the target semantic segmentation model. The first group of task data is any group among the multiple groups of task data.
3. The method according to claim 2, characterized in that The obtaining of multiple groups of task data grouped according to the modality includes: Acquire multi-modal burnt area remote sensing observation data and burnt area mask data corresponding to the burnt area remote sensing observation data; Performing feature extraction on the multi-modal burnt area remote sensing observation data and the burnt area mask data to obtain multi-dimensional feature data; and The multi-dimensional feature data and the mask data are grouped according to the modality to obtain task data grouped according to the modality, Among them, the multidimensional feature data of different modalities have the same feature dimension; The multi-dimensional feature data includes information before and after the fire; and The task data is divided into training samples and verification samples according to a preset ratio.
4. The method according to claim 2, characterized in that: The initial semantic segmentation model uses a deep learning semantic segmentation model that combines ResNet and U-Net, wherein ResNet is used as the encoder of U-Net, and each layer in the encoder is spliced with the corresponding layer in the decoder, so that the deep learning semantic segmentation model can learn shallow and deep features at the same time.
5. The method according to claim 3, characterized in that: If the amount of data stored in the memory buffer is less than the size of the memory buffer, the task data being processed is directly stored in the memory buffer; If the amount of data stored in the memory buffer is equal to the size of the memory buffer, a random decision is made based on the reservoir sampling method as to whether the data being processed will replace the task data stored in the memory buffer.
6. The method according to claim 2, characterized in that The segmentation loss function includes a focal loss function.
7. A multi-modal fire burn detection device, characterized in that: The device comprises the following modules: An acquisition module is used to obtain remote sensing observation data of the target fire area; A feature extraction module is used to extract features from the target fire scar remote sensing observation data to obtain target multi-dimensional feature data; as well as The detection module is used to input the target multi-dimensional feature data into a pre-trained target semantic segmentation model to perform fire scar detection, and obtain fire scar detection data corresponding to the target fire scar remote sensing observation data. The target fire-scarred area remote sensing observation data includes data before and after the fire in the same area; The target semantic segmentation model can process multi-modal burn scar remote sensing observation data, with each modality of burn scar remote sensing observation data and corresponding burn scar mask data as independent training tasks; and The target semantic segmentation model includes a memory buffer so that the model can continue to learn and adaptively update and optimize the model according to new modality data; The memory buffer is used to store the training samples used by the semantic segmentation model in the previous task training process and the output results of the semantic segmentation model obtained after the previous task training for the training samples; The adaptive updating and optimization of the model according to the new modal data includes: Based on the semantic segmentation model obtained by training the previous task data and the current task data, the segmentation loss is calculated using the first loss function to obtain the segmentation loss; Based on the task data stored in the memory buffer, the second loss function is used to calculate the mean square loss as the distillation loss to reduce the model's forgetting of old task knowledge during the current task training process; Constructing a third loss function using a weighted sum of the segmentation loss and the distillation loss, and calculating a total loss; and Update the parameters of the semantic segmentation model based on the total loss.
8. An electronic device, wherein: include: one or more processors; as well as a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors execute the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, causes the processor to execute the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Target identification method, system and device based on multi-modal continuous learning and medium
CN118196645A
Training method of social event classification model based on time sequence
CN118536051A