Earthquake survivor identification method and device based on knowledge anti-forgetting

By introducing elastic weight constraints and feature multiplexing losses into the earthquake survivor identification model, the problem of 'knowledge catastrophic forgetting' of survivor identification algorithm in earthquake disaster scenarios is solved, and the accuracy of identification is improved.

CN114926856BActive Publication Date: 2025-06-06INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210474859.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-29
Publication Date
2025-06-06
Estimated Expiration
2042-04-29

AI Technical Summary

Technical Problem

In the prior art, survivor identification algorithms in earthquake disaster scenarios have the problem of ‘knowledge catastrophic forgetting’, resulting in inaccurate identification results.

Method used

The earthquake survivor identification model based on multiple training tasks and loss functions is adopted. The parameters updates between adjacent training tasks are constrained by elastic weights and feature multiplexing losses, and the features of the old task are reused when new tasks are learned.

Benefits of technology

It effectively improves the accuracy of earthquake survivor identification and reduces the phenomenon of knowledge forgetting in new scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114926856B_ABST
    Figure CN114926856B_ABST
Patent Text Reader

Abstract

The present application provides a method and device for earthquake survivor identification based on knowledge anti-forgetting, the method comprising: inputting audio information and visual information of a target video into an earthquake survivor identification model, and obtaining an earthquake survivor identification result of the target video output by the earthquake survivor identification model; the earthquake survivor identification model is obtained by training a historical model based on multiple training tasks and loss functions, the loss function is determined based on elastic weight constraint loss, feature reuse loss and classification loss, the elastic weight constraint loss is used to constrain parameter updates between two adjacent training tasks, and the feature reuse loss is used to reuse the data of trained training tasks when training based on training tasks, thereby improving the accuracy of earthquake survivor identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a method and device for identifying earthquake survivors based on knowledge anti-forgetting. Background Art

[0002] In an earthquake disaster scenario, how to accurately and quickly confirm whether there are survivors in the honeycomb-like holes of collapsed buildings is a basic rescue task.

[0003] In the prior art, searching for human targets through artificial intelligence algorithms based on video or audio data of earthquake scenes can improve the speed of disaster relief. However, the dynamic changes in earthquake disaster scenes make the survivors and the surrounding environment change frequently and highly uncertain, resulting in a serious "catastrophic forgetting of knowledge" problem in artificial intelligence algorithms. Specifically, in the process of survivor identification based on ideal scenarios, researchers often focus on how to improve the model's expressive ability in new unknown scenarios, while ignoring the model's ability to remember the learned knowledge, which will cause a serious "catastrophic forgetting" problem in artificial intelligence algorithms, leading to inaccurate recognition results. Summary of the invention

[0004] The present application provides a method and device for earthquake survivor identification based on knowledge anti-forgetting, so as to solve the defect of inaccurate identification results in the prior art and improve the accuracy of earthquake survivor identification.

[0005] The present application provides an earthquake survivor identification method based on knowledge anti-forgetting, comprising:

[0006] Inputting the audio information and visual information of the target video into an earthquake survivor recognition model, and obtaining an earthquake survivor recognition result of the target video output by the earthquake survivor recognition model;

[0007] The earthquake survivor identification model is obtained by training a historical model based on multiple training tasks and loss functions. The loss function is determined based on elastic weight constraint loss, feature reuse loss and classification loss. The elastic weight constraint loss is used to constrain parameter updates between two adjacent training tasks, and the feature reuse loss is used to reuse the data of trained training tasks when training based on training tasks.

[0008] According to a method for earthquake survivor identification based on knowledge anti-forgetting provided by the present application, the audio information and visual information of the target video are input into an earthquake survivor identification model, and an earthquake survivor identification result of the target video output by the earthquake survivor identification model is obtained, including:

[0009] Inputting the audio information and visual information of the target video into the earthquake survivor recognition model, extracting audio features based on the audio information in the target video, and extracting visual features based on the visual information in the target video;

[0010] Based on the audio features and the visual features, feature fusion is performed to obtain cross-modal audio features and cross-modal visual features;

[0011] An audio classification probability is obtained based on the cross-modal audio feature, a visual classification probability is obtained based on the cross-modal visual feature, and an earthquake survivor recognition result of the target video output by the earthquake survivor recognition model is obtained based on the audio classification probability and the visual classification probability.

[0012] According to a method for identifying earthquake survivors based on knowledge anti-forgetting provided by the present application, the method further includes:

[0013] Dividing the total training samples into training samples corresponding to each training task in the multiple training tasks;

[0014] Inputting the training sample corresponding to the i-th training task among the multiple training tasks into the historical model for training, where i is a positive integer;

[0015] When i is greater than 1, supervised learning is performed on the historical model based on the memory module samples, and the elastic weight constraint loss is determined based on the parameters of the historical model after training the training samples corresponding to the i-th training task and the parameters of the historical model after training the training samples corresponding to the i-1-th training task;

[0016] randomly sampling the training samples corresponding to the i-th training task to obtain a sampling result, and updating the memory module sample based on the sampling result;

[0017] Determining the feature reuse loss based on the memory module sample and the label corresponding to the memory module sample;

[0018] Determining whether the loss function converges based on the feature reuse loss, the elastic weight constraint loss, and the classification loss;

[0019] If the loss function has not converged, performing an increment operation on i to perform training based on the training samples corresponding to the next training task until the loss function converges;

[0020] When the loss function converges, the parameters of the historical model are saved to obtain the earthquake survivor identification model.

[0021] According to a method for identifying earthquake survivors based on knowledge anti-forgetting provided by the present application, the elastic weight constraint loss is determined based on the parameters of the historical model after training the training samples corresponding to the i-th training task and the parameters of the historical model after training the training samples corresponding to the i-1-th training task, including:

[0022] Determine the elastic weight constraint loss using an elastic weight constraint loss calculation formula based on the parameters of the historical model after training the training sample corresponding to the i-th training task and the parameters of the historical model after training the training sample corresponding to the i-1-th training task;

[0023] The elastic weight constraint loss calculation formula is as follows:

[0024]

[0025] in, represents the elastic weight constraint loss, θ z represents the zth parameter of the historical model after the training sample corresponding to the i-th training task is trained, represents the zth parameter of the historical model after training the training sample corresponding to the i-1th training task, λ is a hyperparameter representing the importance of the i-th training task, α z is a parameter indicating the importance of the zth parameter of the model, and Φ represents the number of parameters of the historical model.

[0026] According to a method for identifying earthquake survivors based on knowledge anti-forgetting provided by the present application, the memory module samples include: audio memory module samples, two-dimensional visual memory module samples and three-dimensional visual memory module samples;

[0027] The determining the feature reuse loss based on the memory module sample and the label corresponding to the memory module sample includes:

[0028] Based on the audio memory module samples, the labels corresponding to the audio memory module samples, the two-dimensional visual memory module samples, the labels corresponding to the two-dimensional visual memory module samples, the three-dimensional visual memory module samples, and the labels corresponding to the three-dimensional visual memory module samples, the feature reuse loss is determined using a feature reuse loss calculation formula;

[0029] The feature reuse loss calculation formula is as follows:

[0030]

[0031] in, represents the feature reuse loss, M a represents the audio memory module sample, Ya Indicates the label corresponding to the audio memory module sample, M 2d represents the two-dimensional visual memory module sample, Y 2d represents the label corresponding to the sample of the two-dimensional visual memory module, M 3d represents the three-dimensional visual memory module sample, Y 3d represents the label corresponding to the sample of the three-dimensional visual memory module, and CE represents the cross entropy loss.

[0032] The present application provides an earthquake survivor identification device based on knowledge anti-forgetting, comprising: an identification module, used to input audio information and visual information of a target video into an earthquake survivor identification model, and obtain an earthquake survivor identification result of the target video output by the earthquake survivor identification model;

[0033] The earthquake survivor identification model is obtained by training a historical model based on multiple training tasks and loss functions. The loss function is determined based on elastic weight constraint loss, feature reuse loss and classification loss. The elastic weight constraint loss is used to constrain parameter updates between two adjacent training tasks, and the feature reuse loss is used to reuse the data of trained training tasks when training based on training tasks.

[0034] According to an earthquake survivor identification device based on knowledge anti-forgetting provided by the present application, the identification module is specifically used for:

[0035] Inputting the audio information and visual information of the target video into the earthquake survivor recognition model, extracting audio features based on the audio information in the target video, and extracting visual features based on the visual information in the target video;

[0036] Based on the audio features and the visual features, feature fusion is performed to obtain cross-modal audio features and cross-modal visual features;

[0037] An audio classification probability is obtained based on the cross-modal audio feature, a visual classification probability is obtained based on the cross-modal visual feature, and an earthquake survivor recognition result of the target video output by the earthquake survivor recognition model is obtained based on the audio classification probability and the visual classification probability.

[0038] The present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method for identifying earthquake survivors based on knowledge anti-forgetting as described above is implemented.

[0039] The present application also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method for identifying earthquake survivors based on knowledge anti-forgetting as described in any one of the above is implemented.

[0040] The present application also provides a computer program product, including a computer program, which, when executed by a processor, implements any of the above-mentioned knowledge-based anti-forgetting earthquake survivor identification methods.

[0041] The earthquake survivor identification method and device based on knowledge anti-forgetting provided in the present application address the "catastrophic forgetting" problem of artificial intelligence algorithms in dynamic environments, and propose an earthquake survivor identification model based on elastic weight constraints of audiovisual models and feature reuse. The elastic weight constraints of the audiovisual models constrain the parameter updates between two adjacent training tasks so that the model does not stray far from the parameters of the old task when learning new task data, thereby achieving the purpose of knowledge anti-forgetting. The anti-forgetting mechanism of feature reuse reuses the data of the trained training tasks when training the training tasks, and achieves the purpose of reviewing old knowledge by reusing the features of the old tasks when learning the new tasks, thereby improving the accuracy of earthquake survivor identification. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions in the present application or the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0043] Figure 1 It is a flowchart of the earthquake survivor identification method based on knowledge anti-forgetting provided by the present application;

[0044] Figure 2 It is a flow chart of inputting audio information and visual information of a target video into an earthquake survivor recognition model to obtain an earthquake survivor recognition result of the target video output by the earthquake survivor recognition model, as provided by the present application;

[0045] Figure 3 It is a flowchart of training an earthquake survivor identification model provided by the present application;

[0046] Figure 4 It is a schematic diagram of the network structure provided by this application;

[0047] Figure 5 It is a schematic diagram of the structure of the earthquake survivor identification device based on knowledge anti-forgetting provided by the present application;

[0048] Figure 6 It is a structural schematic diagram of the electronic device provided by this application. DETAILED DESCRIPTION

[0049] In order to make the purpose, technical solutions and advantages of this application clearer, the technical solutions in this application will be clearly and completely described below in conjunction with the drawings in this application. Obviously, the described embodiments are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0050] Figure 1 is a flow chart of an earthquake survivor identification method based on knowledge anti-forgetting provided in an embodiment of the present application, such as Figure 1 As shown, the earthquake survivor identification method based on knowledge anti-forgetting includes step 100.

[0051] Step 100: input the audio information and visual information of the target video into an earthquake survivor recognition model to obtain an earthquake survivor recognition result of the target video output by the earthquake survivor recognition model.

[0052] Alternatively, the survivor identification task can be defined as:

[0053]

[0054] Among them, f() is the classification module, a() represents the audio feature mapping module, v() represents the visual feature mapping module, and the feature mapping module maps the backbone features into and g() is a feature fusion module. and Perform cross-modal fusion to obtain cross-modal audio features and cross-modal visual features is the prediction result, indicating whether the video data contains human targets.

[0055] It can be understood that the earthquake survivor identification model includes a pre-training model, an audio feature mapping module, a visual feature mapping module, a feature fusion module and a classification module.

[0056] The pre-trained model is used to extract backbone features of audio information and visual information, including audio backbone features, visual two-dimensional backbone features, and visual three-dimensional backbone features.

[0057] The audio feature mapping module is used to map the audio backbone features into audio features, and to map the visual two-dimensional backbone features and the visual three-dimensional backbone features into visual features.

[0058] The visual feature mapping module is used to perform homomodal and cross-modal feature fusion on audio features and visual features to obtain cross-modal audio features and cross-modal visual features.

[0059] The classification module is used to perform classification based on cross-modal audio features and cross-modal visual features to obtain earthquake survivor identification results of the target video.

[0060] In some embodiments, step 100 includes step 200, step 201, and step 202.

[0061] Step 200: input the audio information and visual information of the target video into the earthquake survivor recognition model, extract audio features based on the audio information in the target video, and extract visual features based on the visual information in the target video.

[0062] Optionally, the audio information and visual information of the target video are input into an earthquake survivor recognition model, the backbone features are extracted using a pre-trained model of the earthquake survivor recognition model, and then the backbone features are mapped into audio features and visual features using an audio feature mapping module and a visual feature mapping module.

[0063] Step 201: Based on the audio features and the visual features, feature fusion is performed to obtain cross-modal audio features and cross-modal visual features.

[0064] Optionally, a feature fusion module is used to perform homomodal and cross-modal feature fusion on audio features and visual features to obtain cross-modal audio features and cross-modal visual features.

[0065] Step 202: obtain an audio classification probability based on the cross-modal audio feature, obtain a visual classification probability based on the cross-modal visual feature, and obtain an earthquake survivor recognition result of the target video output by the earthquake survivor recognition model based on the audio classification probability and the visual classification probability.

[0066] Optionally, cross-modal audio features and cross-modal visual features are used for classification to obtain earthquake survivor identification results of the target video.

[0067] To achieve video classification, we use a shared fully connected layer The time series segment features and Project it into the category space and use the sigmoid function (σ()) to get the classification probability:

[0068]

[0069]

[0070] in, is the audio classification probability, is the visual classification probability. It should be noted that i is a positive integer less than or equal to N, and N is the number of video clips included in the target video. Each video clip corresponds to a cross-modal audio feature and a cross-modal audio feature, that is, each video clip corresponds to an audio classification probability and a visual classification probability.

[0071] Finally, the single-mode prediction results of each segment of the video are integrated to classify the video, that is, the prediction results of all modal categories of all video segments in the target video are summed up and comprehensively considered. The final classification probability of the target video is:

[0072]

[0073] in, is the classification probability of the target video, we can As the earthquake survivor identification result of the target video, it can also be based on Get the earthquake survivor identification results of the target video.

[0074] Due to the existence of scene differences, improving the model's recognition ability in new scenes often weakens the model's memory of old knowledge, that is, the survivor recognition model can only process the current scene data. However, due to the ever-changing characteristics of real disaster scenes and the real-time requirements of disaster relief scenes, it is unrealistic to retain a recognition model for each scene and it contains a lot of redundant information.

[0075] When faced with a new scenario, the knowledge contained in the new data can be learned by updating the trained model. At the same time, the time cost of updating the trained model is often lower than the time required to retrain a model, which meets the real-time requirements of disaster relief scenarios.

[0076] In an embodiment of the present application, the earthquake survivor identification model is obtained by training a historical model based on multiple training tasks and loss functions. The loss function is determined based on elastic weight constraint loss, feature reuse loss and classification loss. The elastic weight constraint loss is used to constrain parameter updates between two adjacent training tasks, and the feature reuse loss is used to reuse the data of trained training tasks when training is performed based on training tasks.

[0077] The earthquake survivor identification method based on knowledge anti-forgetting provided in the embodiment of the present application aims at the "catastrophic forgetting" problem of artificial intelligence algorithms in dynamic environments, and proposes an earthquake survivor identification model based on elastic weight constraints of visual and auditory models and feature reuse. The elastic weight constraints of the visual and auditory models constrain the parameter updates between two adjacent training tasks so that the model does not deviate from the parameters of the old task when learning new task data, thereby achieving the purpose of knowledge anti-forgetting. The anti-forgetting mechanism of feature reuse reuses the data of the trained training tasks when training the training tasks, and achieves the purpose of reviewing old knowledge by reusing the features of the old tasks when learning the new tasks, thereby improving the accuracy of earthquake survivor identification.

[0078] Figure 3 1 is a flow chart of training an earthquake survivor identification model provided in an embodiment of the present application. Figure 3 As shown, the training of the earthquake survivor identification model includes: step 300, step 301, step 302, step 303, step 304, step 305, step 306 and step 307.

[0079] Step 300: Divide the total training samples into training samples corresponding to each training task in the multiple training tasks.

[0080] Optionally, the training samples corresponding to each training task may include at least one video and a label corresponding to the video.

[0081] Step 301: input the training sample corresponding to the i-th training task among the multiple training tasks into the historical model for training, where i is a positive integer.

[0082] It can be understood that the history model learns the training tasks one by one, the i-th training task can be any one of multiple training tasks, and the initial value of i is 1.

[0083] Step 302: When i is greater than 1, supervised learning is performed on the historical model based on the memory module samples, and the elastic weight constraint loss is determined based on the parameters of the historical model after training the training samples corresponding to the i-th training task and the parameters of the historical model after training the training samples corresponding to the i-1-th training task.

[0084] In some embodiments, step 302 includes:

[0085] Determine the elastic weight constraint loss using an elastic weight constraint loss calculation formula based on the parameters of the historical model after training the training sample corresponding to the i-th training task and the parameters of the historical model after training the training sample corresponding to the i-1-th training task;

[0086] The elastic weight constraint loss calculation formula is as follows:

[0087]

[0088] in, represents the elastic weight constraint loss, θ z represents the zth parameter of the historical model after the training sample corresponding to the i-th training task is trained, represents the zth parameter of the historical model after training the training sample corresponding to the i-1th training task, λ is a hyperparameter representing the importance of the i-th training task, α z is a parameter indicating the importance of the zth parameter of the history model, and Φ indicates the number of parameters of the history model.

[0089] α z It can be measured by second-order derivative information, and the specific calculation formula is as follows:

[0090]

[0091] Among them, D o represents the dataset of i-1 training tasks, represents the second-order derivative information, Represents the first-order derivative information.

[0092] Because the amount of calculation required to calculate the second-order derivative is very large, in some embodiments, the square of the first-order derivative can also be used for approximation. The specific calculation formula is as follows:

[0093]

[0094] In some embodiments, Figure 4 As shown in Figure 1, the earthquake survivor recognition model includes a backbone network module, a feature mapping module, a cross-modal mixed attention module, and a classification module. Since the visual and auditory backbone networks are all based on pre-trained models, the pre-trained models do not have the phenomenon of "catastrophic forgetting" of knowledge. Therefore, Constraints are only used to constrain other network modules.

[0095] Therefore, the final elastic weight constraint of this paper is as follows:

[0096]

[0097] in, is the loss of the feature mapping module, is the loss of the cross-modal mixed attention module, is the loss of the classification module, w 1 is the importance weight of the feature mapping module, w 2is the importance weight of the cross-modal mixed attention module, w 3 is the importance weight of the classification module. The specific implementation is as follows:

[0098]

[0099] Where a is the number of parameters of the audio feature mapping module, is the zth parameter of the audio feature mapping module after training the training sample corresponding to the i-th training task, is the zth parameter of the audio feature mapping module after training the training sample corresponding to the i-1th training task, is a parameter representing the importance of the zth parameter of the audio feature mapping module.

[0100] v is the number of parameters of the visual feature mapping module, is the zth parameter of the visual feature mapping module after training the training sample corresponding to the i-th training task, is the zth parameter of the visual feature mapping module after training the training sample corresponding to the i-1th training task, is a parameter representing the importance of the zth parameter of the visual feature mapping module.

[0101]

[0102] g is the number of parameters of the cross-modal mixed attention module, is the zth parameter of the cross-modal mixed attention module after training the training sample corresponding to the i-th training task, is the zth parameter of the cross-modal mixed attention module after training the training sample corresponding to the i-1th training task, is a parameter representing the importance of the zth parameter of the cross-modal mixed attention module.

[0103]

[0104] f is the number of parameters of the classification module, is the zth parameter of the classification module after training the training sample corresponding to the i-th training task, is the zth parameter of the classification module after training the training sample corresponding to the i-1th training task, is a parameter representing the importance of the zth parameter of the classification module.

[0105] Step 302: Obtain a sampling result by randomly sampling the training samples corresponding to the i-th training task, and update the memory module sample based on the sampling result.

[0106] Optionally, when the historical model is trained based on the training samples corresponding to the first training task, there is no old data, that is, when training the training samples D corresponding to the first task 1 When the first task training is completed, the memory module is switched from D 1 Random sampling B 1 Samples update memory module sample M, that is, M = {sample (D 1 ,B 1 )}. Sample represents a random sampling operation.

[0107] When the model trains the vth (v ≥ 2) task D v When Figure 4 As shown, the algorithm performs supervised learning on samples of the memory module:

[0108]

[0109] Among them, Y V,old represents the label of M. When the training of the vth task is completed, the memory module is v Random sampling B v The memory module is updated by using samples. The updating process is as follows:

[0110] M=M∪{sample(D v ,B v )}.

[0111] Step 304: Determine the feature reuse loss based on the memory module sample and the label corresponding to the memory module sample.

[0112] In some embodiments, the memory module samples include: audio memory module samples, two-dimensional visual memory module samples, and three-dimensional visual memory module samples;

[0113] Step 304 includes:

[0114] Based on the audio memory module samples, the labels corresponding to the audio memory module samples, the two-dimensional visual memory module samples, the labels corresponding to the two-dimensional visual memory module samples, the three-dimensional visual memory module samples, and the labels corresponding to the three-dimensional visual memory module samples, the feature reuse loss is determined using a feature reuse loss calculation formula;

[0115] The feature reuse loss calculation formula is as follows:

[0116]

[0117] in, represents the feature reuse loss, M arepresents the audio memory module sample, Y a Indicates the label corresponding to the audio memory module sample, M 2d represents the two-dimensional visual memory module sample, Y 2d represents the label corresponding to the sample of the two-dimensional visual memory module, M 3d represents the three-dimensional visual memory module sample, Y 3d represents the label corresponding to the sample of the three-dimensional visual memory module, and CE represents the cross entropy loss.

[0118] It is understandable that in order to reuse cross-modal features, this embodiment constructs three memory modules based on audio, two-dimensional vision and three-dimensional vision backbone features. The corresponding losses obtained from these three memory modules are added together to obtain the feature reuse loss.

[0119] Step 305: Determine whether the loss function converges based on the feature reuse loss, the elastic weight constraint loss and the classification loss.

[0120] Optionally, each training task includes a training video sample and a label corresponding to the training video sample. V And the classification probability based on the training video sample video The video classification loss can be obtained, and the classification loss calculation formula is as follows:

[0121]

[0122] Optionally, the loss function is:

[0123]

[0124] in, is the loss function, is the feature reuse loss, is the elastic weight constraint loss, is the classification loss.

[0125] When the value of the loss function is less than the target value, the loss function is considered to have converged.

[0126] Step 306: When the loss function has not converged, perform an increment operation on i to perform training based on the training samples corresponding to the next training task until the loss function converges.

[0127] Optionally, when the loss function has not converged, it means that the training of the historical model has not been completed and the historical model needs to continue to be trained. After adding one to i, repeat steps 301 to 305 to perform a new round of training on the model, and determine again whether the loss function has converged until the model training is completed.

[0128] Step 307: When the loss function converges, save the parameters of the historical model to obtain the earthquake survivor identification model.

[0129] When the loss function does not converge, it means that the training of the historical model is not completed. The parameters of the historical model are saved to obtain the earthquake survivor identification model.

[0130] Table 1 shows the experimental results of the present application and other methods on various data sets. It can be seen that the method of the present application has a significant effect on the accuracy of earthquake survivor identification.

[0131] Table 1 Experimental results of this application and other methods on various datasets

[0132] method Task1 Task2 Task3 Task4 Task5 Avg M-LSTM 70.37 78.97 77.29 79.49 84.97 78.22 M-GRU 70.13 82.71 83.05 86.29 85.90 81.62 UMP 73.71 88.05 87.48 88.21 81.84 82.86 This application 73.71 86.51 87.53 86.67 87.76 84.44

[0133] The earthquake survivor identification method based on knowledge anti-forgetting provided in the embodiment of the present application aims at the "catastrophic forgetting" problem of artificial intelligence algorithms in dynamic environments, and proposes an earthquake survivor identification model based on elastic weight constraints of visual and auditory models and feature reuse. The elastic weight constraints of the visual and auditory models constrain the parameter updates between two adjacent training tasks so that the model does not deviate from the parameters of the old task when learning new task data, thereby achieving the purpose of knowledge anti-forgetting. The anti-forgetting mechanism of feature reuse reuses the data of the trained training tasks when training the training tasks, and achieves the purpose of reviewing old knowledge by reusing the features of the old tasks when learning the new tasks, thereby improving the accuracy of earthquake survivor identification.

[0134] The earthquake survivor identification device based on knowledge anti-forgetting provided by the present application is described below. The earthquake survivor identification device based on knowledge anti-forgetting described below and the earthquake survivor identification method based on knowledge anti-forgetting described above can be referenced to each other.

[0135] like Figure 5 As shown, the earthquake survivor identification device 500 based on knowledge anti-forgetting includes: an identification module 510.

[0136] The recognition module 510 is used to input the audio information and visual information of the target video into the earthquake survivor recognition model to obtain the earthquake survivor recognition result of the target video output by the earthquake survivor recognition model;

[0137] The earthquake survivor identification model is obtained by training a historical model based on multiple training tasks and loss functions. The loss function is determined based on elastic weight constraint loss, feature reuse loss and classification loss. The elastic weight constraint loss is used to constrain parameter updates between two adjacent training tasks, and the feature reuse loss is used to reuse the data of trained training tasks when training based on training tasks.

[0138] Optionally, the identification module 510 is specifically configured to:

[0139] Inputting the audio information and visual information of the target video into the earthquake survivor recognition model, extracting audio features based on the audio information in the target video, and extracting visual features based on the visual information in the target video;

[0140] Based on the audio features and the visual features, feature fusion is performed to obtain cross-modal audio features and cross-modal visual features;

[0141] An audio classification probability is obtained based on the cross-modal audio feature, a visual classification probability is obtained based on the cross-modal visual feature, and an earthquake survivor recognition result of the target video output by the earthquake survivor recognition model is obtained based on the audio classification probability and the visual classification probability.

[0142] Optionally, the earthquake survivor identification device 500 based on knowledge anti-forgetting further includes a training module.

[0143] The training module is used to:

[0144] Dividing the total training samples into training samples corresponding to each training task in the multiple training tasks;

[0145] Inputting the training sample corresponding to the i-th training task among the multiple training tasks into the historical model for training, where i is a positive integer;

[0146] When i is greater than 1, supervised learning is performed on the historical model based on the memory module samples, and the elastic weight constraint loss is determined based on the parameters of the historical model after training the training samples corresponding to the i-th training task and the parameters of the historical model after training the training samples corresponding to the i-1-th training task;

[0147] randomly sampling the training samples corresponding to the i-th training task to obtain a sampling result, and updating the memory module sample based on the sampling result;

[0148] Determining the feature reuse loss based on the memory module sample and the label corresponding to the memory module sample;

[0149] Determining whether the loss function converges based on the feature reuse loss, the elastic weight constraint loss, and the classification loss;

[0150] If the loss function has not converged, performing an increment operation on i to perform training based on the training samples corresponding to the next training task until the loss function converges;

[0151] When the loss function converges, the parameters of the historical model are saved to obtain the earthquake survivor identification model.

[0152] Optionally, determining the elastic weight constraint loss based on the parameters of the historical model after the training samples corresponding to the i-th training task are trained and the parameters of the historical model after the training samples corresponding to the i-1-th training task are trained includes:

[0153] Determine the elastic weight constraint loss using an elastic weight constraint loss calculation formula based on the parameters of the historical model after training the training sample corresponding to the i-th training task and the parameters of the historical model after training the training sample corresponding to the i-1-th training task;

[0154] The elastic weight constraint loss calculation formula is as follows:

[0155]

[0156] in, represents the elastic weight constraint loss, θ z represents the zth parameter of the historical model after the training sample corresponding to the i-th training task is trained, represents the zth parameter of the historical model after training the training sample corresponding to the i-1th training task, λ is a hyperparameter representing the importance of the i-th training task, α z is a parameter indicating the importance of the zth parameter of the model, and Φ represents the number of parameters of the historical model.

[0157] Optionally, the memory module samples include: audio memory module samples, two-dimensional visual memory module samples and three-dimensional visual memory module samples;

[0158] The determining the feature reuse loss based on the memory module sample and the label corresponding to the memory module sample includes:

[0159] Based on the audio memory module samples, the labels corresponding to the audio memory module samples, the two-dimensional visual memory module samples, the labels corresponding to the two-dimensional visual memory module samples, the three-dimensional visual memory module samples, and the labels corresponding to the three-dimensional visual memory module samples, the feature reuse loss is determined using a feature reuse loss calculation formula;

[0160] The feature reuse loss calculation formula is as follows:

[0161]

[0162] in, represents the feature reuse loss, M a represents the audio memory module sample, Y a Indicates the label corresponding to the audio memory module sample, M 2d represents the two-dimensional visual memory module sample, Y 2d represents the label corresponding to the sample of the two-dimensional visual memory module, M 3d represents the three-dimensional visual memory module sample, Y 3d represents the label corresponding to the sample of the three-dimensional visual memory module, and CE represents the cross entropy loss.

[0163] It should be noted here that the above-mentioned device provided in the embodiment of the present application can implement all the method steps implemented in the above-mentioned method embodiment, and can achieve the same technical effect. The parts and beneficial effects of this embodiment that are the same as those in the method embodiment will not be described in detail here.

[0164] Figure 6 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 6 As shown, the electronic device may include: a processor 610, a communication interface 620, a memory 630 and a communication bus 640, wherein the processor 610, the communication interface 620 and the memory 630 communicate with each other through the communication bus 640. The processor 610 may call the logic instructions in the memory 630 to execute the earthquake survivor identification method based on knowledge anti-forgetting, and the method includes:

[0165] Inputting the audio information and visual information of the target video into an earthquake survivor recognition model, and obtaining an earthquake survivor recognition result of the target video output by the earthquake survivor recognition model;

[0166] The earthquake survivor identification model is obtained by training a historical model based on multiple training tasks and loss functions. The loss function is determined based on elastic weight constraint loss, feature reuse loss and classification loss. The elastic weight constraint loss is used to constrain parameter updates between two adjacent training tasks, and the feature reuse loss is used to reuse the data of trained training tasks when training based on training tasks.

[0167] In addition, the logic instructions in the above-mentioned memory 630 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art, and the computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk.

[0168] On the other hand, the present application also provides a computer program product, the computer program product includes a computer program, the computer program can be stored on a non-transitory computer-readable storage medium, when the computer program is executed by a processor, the computer can execute the earthquake survivor identification method based on knowledge anti-forgetting provided by the above methods, the method includes:

[0169] Inputting the audio information and visual information of the target video into an earthquake survivor recognition model, and obtaining an earthquake survivor recognition result of the target video output by the earthquake survivor recognition model;

[0170] The earthquake survivor identification model is obtained by training a historical model based on multiple training tasks and loss functions. The loss function is determined based on elastic weight constraint loss, feature reuse loss and classification loss. The elastic weight constraint loss is used to constrain parameter updates between two adjacent training tasks, and the feature reuse loss is used to reuse the data of trained training tasks when training based on training tasks.

[0171] On the other hand, the present application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which is implemented when the computer program is executed by a processor to perform the earthquake survivor identification method based on knowledge anti-forgetting provided by the above methods, the method comprising:

[0172] Inputting the audio information and visual information of the target video into an earthquake survivor recognition model, and obtaining an earthquake survivor recognition result of the target video output by the earthquake survivor recognition model;

[0173] The earthquake survivor identification model is obtained by training a historical model based on multiple training tasks and loss functions. The loss function is determined based on elastic weight constraint loss, feature reuse loss and classification loss. The elastic weight constraint loss is used to constrain parameter updates between two adjacent training tasks, and the feature reuse loss is used to reuse the data of trained training tasks when training based on training tasks.

[0174] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0175] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0176] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for identifying earthquake survivors based on knowledge-resistant forgetting. It is characterized in that include: Inputting the audio information and visual information of the target video into an earthquake survivor recognition model, and obtaining an earthquake survivor recognition result of the target video output by the earthquake survivor recognition model; The earthquake survivor identification model is obtained by training the historical model based on multiple training tasks and loss functions, wherein the loss function is determined based on elastic weight constraint loss, feature reuse loss and classification loss, wherein the elastic weight constraint loss is used to constrain parameter updates between two adjacent training tasks, and the feature reuse loss is used to reuse the data of the trained training tasks when training is performed based on the training tasks; Based on the parameters of the historical model after the training samples corresponding to the i-th training task and the parameters of the historical model after the training samples corresponding to the i-1-th training task, the elastic weight constraint loss is determined using an elastic weight constraint loss calculation formula; The elastic weight constraint loss calculation formula is as follows: ; in, represents the elastic weight constraint loss, represents the number of training samples corresponding to the i-th training task in the history model after training parameters, represents the number of training samples corresponding to the i-1th training task in the history model after training. parameters, is a hyperparameter representing the importance of the i-th training task, is a parameter indicating the importance of the zth parameter of the model, represents the number of parameters of the history model; The memory module samples include: audio memory module samples, two-dimensional visual memory module samples and three-dimensional visual memory module samples; Based on the prediction probability corresponding to the audio memory module sample, the label corresponding to the audio memory module sample, the prediction probability corresponding to the two-dimensional visual memory module sample, the label corresponding to the two-dimensional visual memory module sample, the prediction probability corresponding to the three-dimensional visual memory module sample, and the label corresponding to the three-dimensional visual memory module sample, the feature reuse loss is determined using a feature reuse loss calculation formula; The feature reuse loss calculation formula is as follows: ; in, represents the feature reuse loss, represents the prediction probability corresponding to the audio memory module sample, Indicates the label corresponding to the audio memory module sample, represents the predicted probability corresponding to the sample of the two-dimensional visual memory module, represents the label corresponding to the sample of the two-dimensional visual memory module, represents the predicted probability corresponding to the sample of the three-dimensional visual memory module, represents the label corresponding to the sample of the three-dimensional visual memory module, represents the cross entropy loss.

2. The earthquake survivor identification method based on knowledge anti-forgetting according to claim 1, It is characterized in that The step of inputting the audio information and visual information of the target video into the earthquake survivor recognition model to obtain the earthquake survivor recognition result of the target video output by the earthquake survivor recognition model comprises: Inputting the audio information and visual information of the target video into the earthquake survivor recognition model, extracting audio features based on the audio information in the target video, and extracting visual features based on the visual information in the target video; Based on the audio features and the visual features, feature fusion is performed to obtain cross-modal audio features and cross-modal visual features; An audio classification probability is obtained based on the cross-modal audio feature, a visual classification probability is obtained based on the cross-modal visual feature, and an earthquake survivor recognition result of the target video output by the earthquake survivor recognition model is obtained based on the audio classification probability and the visual classification probability.

3. The earthquake survivor identification method based on knowledge anti-forgetting according to claim 1, It is characterized in that Also includes: Dividing the total training samples into training samples corresponding to each training task in the multiple training tasks; Inputting the training sample corresponding to the i-th training task among the multiple training tasks into the historical model for training, where i is a positive integer; When i is greater than 1, supervised learning is performed on the historical model based on the memory module samples, and the elastic weight constraint loss is determined based on the parameters of the historical model after training the training samples corresponding to the i-th training task and the parameters of the historical model after training the training samples corresponding to the i-1-th training task; randomly sampling the training samples corresponding to the i-th training task to obtain a sampling result, and updating the memory module sample based on the sampling result; Determining the feature reuse loss based on the memory module sample and the label corresponding to the memory module sample; Determining whether the loss function converges based on the feature reuse loss, the elastic weight constraint loss, and the classification loss; If the loss function has not converged, performing an increment operation on i to perform training based on the training samples corresponding to the next training task until the loss function converges; When the loss function converges, the parameters of the historical model are saved to obtain the earthquake survivor identification model.

4. An earthquake survivor identification device based on knowledge anti-forgetting, It is characterized in that include: A recognition module, used for inputting audio information and visual information of a target video into an earthquake survivor recognition model, and obtaining an earthquake survivor recognition result of the target video output by the earthquake survivor recognition model; The earthquake survivor identification model is obtained by training the historical model based on multiple training tasks and loss functions, wherein the loss function is determined based on elastic weight constraint loss, feature reuse loss and classification loss, wherein the elastic weight constraint loss is used to constrain parameter updates between two adjacent training tasks, and the feature reuse loss is used to reuse the data of the trained training tasks when training is performed based on the training tasks; Based on the parameters of the historical model after the training samples corresponding to the i-th training task and the parameters of the historical model after the training samples corresponding to the i-1-th training task, the elastic weight constraint loss is determined using an elastic weight constraint loss calculation formula; The elastic weight constraint loss calculation formula is as follows: ; in, represents the elastic weight constraint loss, represents the number of training samples corresponding to the i-th training task in the history model after training parameters, represents the number of training samples corresponding to the i-1th training task in the history model after training. parameters, is a hyperparameter representing the importance of the i-th training task, is a parameter indicating the importance of the zth parameter of the model, represents the number of parameters of the history model; The memory module samples include: audio memory module samples, two-dimensional visual memory module samples and three-dimensional visual memory module samples; Based on the prediction probability corresponding to the audio memory module sample, the label corresponding to the audio memory module sample, the prediction probability corresponding to the two-dimensional visual memory module sample, the label corresponding to the two-dimensional visual memory module sample, the prediction probability corresponding to the three-dimensional visual memory module sample, and the label corresponding to the three-dimensional visual memory module sample, the feature reuse loss is determined using a feature reuse loss calculation formula; The feature reuse loss calculation formula is as follows: ; in, represents the feature reuse loss, represents the prediction probability corresponding to the audio memory module sample, Indicates the label corresponding to the audio memory module sample, represents the predicted probability corresponding to the sample of the two-dimensional visual memory module, represents the label corresponding to the sample of the two-dimensional visual memory module, represents the predicted probability corresponding to the sample of the three-dimensional visual memory module, represents the label corresponding to the sample of the three-dimensional visual memory module, represents the cross entropy loss.

5. The earthquake survivor identification device based on knowledge anti-forgetting according to claim 4, It is characterized in that The identification module is specifically used for: Inputting the audio information and visual information of the target video into the earthquake survivor recognition model, extracting audio features based on the audio information in the target video, and extracting visual features based on the visual information in the target video; Based on the audio features and the visual features, feature fusion is performed to obtain cross-modal audio features and cross-modal visual features; An audio classification probability is obtained based on the cross-modal audio feature, a visual classification probability is obtained based on the cross-modal visual feature, and an earthquake survivor recognition result of the target video output by the earthquake survivor recognition model is obtained based on the audio classification probability and the visual classification probability.

6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, It is characterized in that When the processor executes the program, the earthquake survivor identification method based on knowledge anti-forgetting as described in any one of claims 1 to 3 is implemented.

7. A non-transitory computer-readable storage medium having a computer program stored thereon, It is characterized in that When the computer program is executed by a processor, the earthquake survivor identification method based on knowledge anti-forgetting as claimed in any one of claims 1 to 3 is implemented.

8. A computer program product comprising a computer program, It is characterized in that When the computer program is executed by a processor, the earthquake survivor identification method based on knowledge anti-forgetting as claimed in any one of claims 1 to 3 is implemented.