Multi-mode-based behavior early warning method and device, electronic equipment and storage medium

Through the application of multimodal data fusion and behavioral intention model, the shortcomings of traditional single modal behavior monitoring methods in identifying complex behavioral intentions are solved, and higher recognition accuracy and environmental adaptability are achieved, and the safety management effect is enhanced.

CN119989189APending Publication Date: 2025-05-13GREE ELECTRIC APPLIANCE INC OF ZHUHAI +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411882749.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-19
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Traditional single-modal behavior monitoring methods are difficult to accurately and comprehensively identify the complex behavioral intentions of people, especially under changing environmental conditions.

Method used

A multimodal-based behavioral warning method is adopted, by obtaining image data, voice data and sensor data, these data are fused to generate fused feature data, and input them into the trained behavioral intention model to identify abnormal behavioral intentions and issue early warnings.

Benefits of technology

It improves the accuracy and stability of the identification of personnel's behavioral intentions, can adapt to different environmental conditions, such as lighting changes and noise interference, timely identify abnormal behavioral intentions and issue early warnings, enhancing the effectiveness of safety management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989189A_ABST
    Figure CN119989189A_ABST
Patent Text Reader

Abstract

The invention provides a behavior early warning method and device based on multiple modes, electronic equipment and a storage medium, and the method comprises the steps: obtaining multi-mode target monitoring data, and fusing a plurality of pieces of target monitoring data to obtain target fusion feature data; inputting the target fusion feature data into a preset behavior intention model to obtain a target behavior intention category; the behavior intention model is obtained by training fusion feature data marked with behavior intention categories; and when the target behavior intention category is an abnormal behavior intention category, performing early warning on behaviors aiming at the multiple pieces of target monitoring data. By fusing monitoring data of multiple modes, the accuracy of identifying the behavior intention of the personnel can be improved. And the method can adapt to different environmental conditions, such as illumination change and noise interference, and the stability and reliability of behavior monitoring are improved. In addition, the workload of manual monitoring can be reduced, the management efficiency is improved, and the management cost is also reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of behavior warning, and in particular to a multimodal behavior warning method, a multimodal behavior warning method device, an electronic device and a computer-readable storage medium. Background Art

[0002] With the continuous development of science and technology, the demand for monitoring and management of human behavior is increasing. In various public places such as shopping malls, schools, libraries, stations, etc., it is very important to ensure safety and order.

[0003] Traditional single-modality behavior monitoring methods have limitations and are difficult to accurately and comprehensively identify people's complex behavioral intentions. Summary of the invention

[0004] In view of the above problems, a multimodal behavior warning method, a multimodal behavior warning method device, an electronic device and a computer-readable storage medium are proposed to overcome the above problems or at least partially solve the above problems, including:

[0005] A multimodal behavior early warning method, the method comprising:

[0006] Acquire multi-modal target monitoring data, and fuse multiple target monitoring data to obtain target fusion feature data;

[0007] Inputting the target fusion feature data into a preset behavior intention model to obtain a target behavior intention category; the behavior intention model is trained based on the fusion feature data labeled with the behavior intention category;

[0008] When the target behavior intention category is an abnormal behavior intention category, an early warning is issued for the behaviors targeted by the multiple target monitoring data.

[0009] Optionally, the issuing of early warning for the behaviors targeted by the multiple target monitoring data includes:

[0010] Determine the target warning method according to the target behavior intention category;

[0011] The target early warning method is used to issue an early warning for the behavior targeted by the multiple target monitoring data.

[0012] Optionally, the method further comprises:

[0013] Acquire multimodal monitoring data for training, and fuse multiple monitoring data for training to obtain fused feature data for training;

[0014] Inputting the training monitoring data of different modalities into the corresponding prediction model to obtain the first behavior intention category corresponding to the training monitoring data of each modality;

[0015] determining a second behavioral intention category based on the plurality of first behavioral intention categories;

[0016] The training fusion feature data is labeled with a second behavior intention category, and a preset model is trained based on the labeled training fusion feature data to obtain the behavior intention model.

[0017] Optionally, determining the second behavior intention category according to the plurality of first behavior intention categories includes:

[0018] The multiple first behavior intention categories are merged to obtain the second behavior intention category.

[0019] Optionally, the fusing of a plurality of monitoring data for training to obtain fused feature data for training includes:

[0020] Preprocess each training monitoring data;

[0021] Extract features from the preprocessed multiple training monitoring data to obtain feature data corresponding to each training monitoring data;

[0022] The feature data corresponding to each training monitoring data are fused to obtain the training fused feature data.

[0023] Optionally, the preprocessing includes at least one of the following:

[0024] Noise reduction, enhancement, calibration, normalization, filtering.

[0025] Optionally, the fusion method for fusing the feature data corresponding to each training monitoring data includes any one of the following:

[0026] Weighted fusion, feature mapping, and sequential concatenation.

[0027] Optionally, a fusion method for fusing multiple first behavior intention categories includes any one of the following:

[0028] Voting method, weighted average method, neural network fusion method.

[0029] Optionally, the target monitoring data includes the following data:

[0030] Image data, voice data, sensor data.

[0031] The embodiment of the present invention further provides a multi-modal behavior warning device, the device comprising:

[0032] A fusion module is used to obtain multi-modal target monitoring data and fuse multiple target monitoring data to obtain target fusion feature data;

[0033] A prediction model is used to input the target fusion feature data into a preset behavior intention model to obtain a target behavior intention category; the behavior intention model is trained based on the fusion feature data labeled with the behavior intention category;

[0034] The early warning module is used to issue an early warning for the behavior targeted by the multiple target monitoring data when the target behavior intention category is an abnormal behavior intention category.

[0035] Optionally, the early warning module is used to determine a target early warning method according to the target behavior intention category; and use the target early warning method to issue an early warning for the behavior targeted by the multiple target monitoring data.

[0036] Optionally, the device further comprises:

[0037] A training module is used to obtain multi-modal training monitoring data and fuse multiple training monitoring data to obtain training fused feature data; input the training monitoring data of different modalities into the corresponding prediction model to obtain the first behavior intention category corresponding to the training monitoring data of each modality; determine the second behavior intention category based on multiple first behavior intention categories; label the training fused feature data with the second behavior intention category, and train the preset model based on the labeled training fused feature data to obtain the behavior intention model.

[0038] Optionally, the training module is used to fuse multiple first behavior intention categories to obtain the second behavior intention category.

[0039] Optionally, the training module is used to preprocess each training monitoring data; extract features from the preprocessed multiple training monitoring data to obtain feature data corresponding to each training monitoring data; and fuse the feature data corresponding to each training monitoring data to obtain the training fused feature data.

[0040] Optionally, the preprocessing includes at least one of the following: noise reduction, enhancement, calibration, normalization, and filtering.

[0041] Optionally, a fusion method for fusing feature data corresponding to each training monitoring data includes any one of the following: weighted fusion, feature mapping, and sequential concatenation.

[0042] Optionally, a fusion method for fusing multiple first behavior intention categories includes any one of the following: a voting method, a weighted average method, and a neural network fusion method.

[0043] Optionally, the target monitoring data includes the following data: image data, voice data, and sensor data.

[0044] An embodiment of the present invention further provides an electronic device, comprising a processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein the computer program implements the above-mentioned multimodal-based behavior warning method when executed by the processor.

[0045] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the multimodal-based behavior warning method as described above is implemented.

[0046] The embodiments of the present invention have the following advantages:

[0047] In the embodiment of the present invention, multi-modal target monitoring data is obtained, and multiple target monitoring data are fused to obtain target fusion feature data; the target fusion feature data is input into a preset behavior intention model to obtain a target behavior intention category; the behavior intention model is trained based on the fusion feature data labeled with the behavior intention category; when the target behavior intention category is an abnormal behavior intention category, an early warning is issued for the behavior targeted by the multiple target monitoring data. By fusing monitoring data of multiple modes, the advantages of different data sources can be comprehensively utilized to make up for the shortcomings of a single mode, thereby improving the accuracy of identifying the behavior intention of personnel.

[0048] It can also adapt to different environmental conditions, such as lighting changes, noise interference, etc., which improves the stability and reliability of behavior monitoring.

[0049] In addition, the implementation of the present invention can also timely identify abnormal behavior intentions and issue warnings, initiate corresponding treatment measures, and effectively enhance the effect of security management and prevent the occurrence of potential security risks and bad behaviors. And the fully automatic method can reduce the workload of manual monitoring, improve management efficiency, and also reduce management costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the technical solution of the present invention, the accompanying drawings required for use in the description of the present invention will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative labor.

[0051] Figure 1 is a flowchart of a multi-modal behavior warning method according to an embodiment of the present invention;

[0052] Figure 2is a flowchart of another multi-modal behavior warning method according to an embodiment of the present invention;

[0053] Figure 3 is a flowchart of steps for training a behavior intention model according to an embodiment of the present invention;

[0054] Figure 4 is a flowchart of another step of training a behavior intention model according to an embodiment of the present invention;

[0055] Figure 5 is a flowchart of another step of training a behavior intention model according to an embodiment of the present invention;

[0056] Figure 6 is a flow chart of a training and application of an embodiment of the present invention;

[0057] Figure 7 It is a structural diagram of a multi-modal behavior warning device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0058] In order to make the above-mentioned purposes, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0059] In actual applications, relying solely on camera image monitoring may not be able to accurately determine behavioral intentions due to factors such as viewing angle limitations and lighting conditions; relying solely on voice monitoring may be disturbed by environmental noise.

[0060] To this end, an embodiment of the present invention provides a multimodal-based behavior warning method, which can integrate data from multiple modalities and comprehensively utilize the advantages of different data sources such as voice, images, and sensors to predict and warn behavioral intentions, thereby improving the accuracy and reliability of behavioral intention recognition.

[0061] For details, please refer to Figure 1 , showing a step flow chart of a multimodal behavior warning method according to an embodiment of the present invention.

[0062] like Figure 1 As shown, the multimodal behavior warning method may include the following steps:

[0063] Step 101: Acquire multi-modal target monitoring data, and fuse multiple target monitoring data to obtain target fusion feature data.

[0064] In an embodiment of the present invention, target monitoring data of different modes may be acquired first; illustratively, the target monitoring data may include the following data:

[0065] Image data, voice data, sensor data.

[0066] For example, high-definition cameras can be installed at different key locations, including but not limited to entrances, exits, corridors, doors of important rooms, public activity areas, etc. Ensure that the angle and field of view of the camera can cover as many human activity ranges as possible, while taking into account the influence of lighting conditions and obstructions.

[0067] Then, the camera parameters such as resolution, frame rate, etc. can be set to obtain high-quality image data.

[0068] In practical applications, cameras with different resolutions can be selected according to actual needs. For example, high-resolution cameras can be used in areas where clearer details are required, while lower-resolution cameras can be used in some areas where image quality requirements are not high to save resources.

[0069] In addition, microphones can be set up in areas with large traffic and frequent communication, such as conference rooms, rest areas, public service areas, etc. Specifically, you can choose a directional microphone or an omnidirectional microphone, depending on the needs of the specific scene.

[0070] The microphone's sensitivity and gain can then be adjusted to ensure that voice data is captured clearly while avoiding excessive ambient noise interference.

[0071] In actual applications, tests can be carried out, and then the sensitivity and gain of the microphone can be adjusted according to the actual test results to improve the collection effect of the microphone.

[0072] In some feasible embodiments, appropriate sensors may be selected and deployed in the environment to be monitored according to different behavioral intention monitoring requirements.

[0073] For example, to monitor the movement trajectory of people, motion sensors, proximity sensors, infrared sensors, etc. can be used; to monitor environmental parameters, temperature sensors, humidity sensors, light sensors, etc. can be used. These sensors can monitor the environment and the people in the environment and generate corresponding sensor data.

[0074] Exemplarily, motion sensors, proximity sensors, infrared sensors, and the like can generate sensor data of a person's motion trajectory; temperature sensors can generate sensor data of the temperature value of the environment; humidity sensors can generate sensor data of the humidity value of the environment; light sensors can generate sensor data of the light intensity of the environment, etc., and the embodiments of the present invention are not limited to this.

[0075] After obtaining target monitoring data of different modalities, the features of the multiple target monitoring data may be fused to obtain target fusion feature data.

[0076] In some feasible embodiments, target monitoring data of different modes may be monitoring data obtained by monitoring the same scene, or may be monitoring data obtained by monitoring the same area, and the embodiments of the present invention are not limited to this.

[0077] Step 102: Input the target fused feature data into a preset behavior intention model to obtain the target behavior intention category; the behavior intention model is trained based on the fused feature data labeled with the behavior intention category.

[0078] After obtaining the target fusion features, the target fusion features can be input into a preset behavior intention model.

[0079] The preset behavior intention model is based on the target fusion feature and can integrate monitoring data of different modalities to predict the target behavior intention category corresponding to multiple target monitoring data. The target behavior intention category may refer to the behavior intention category of people in the scene targeted by the target monitoring data; illustratively, the behavior intention category may be a normal behavior intention category (for example: normal walking), a suspicious behavior intention category (for example: wandering) or an abnormal behavior intention category (for example: graffiti, littering, stealing, etc.), etc., and the embodiments of the present invention are not limited to this.

[0080] In some feasible embodiments, different fused feature data may be labeled with behavioral intention categories in advance; then, a preset model may be trained based on the fused feature data labeled with behavioral intention categories, thereby obtaining a behavioral intention model that can analyze the fused feature data obtained by multimodal fusion and output the corresponding behavioral intention category.

[0081] Step 103: When the target behavior intention category is an abnormal behavior intention category, an early warning is issued for the behaviors targeted by the multiple target monitoring data.

[0082] In the embodiment of the present invention, if the target behavior intention category is a normal behavior intention category, new target monitoring data may continue to be acquired, and step 101 to step 103 may be re-executed.

[0083] On the contrary, if the target behavior intention category is an abnormal intention category, for example, when the target behavior intention category is a suspicious behavior intention category or an abnormal behavior intention category, an early warning can be issued for the behavior occurring in the scenarios targeted by multiple target monitoring data.

[0084] Exemplarily, the early warning may include playing a warning sound, lighting a warning light, etc., and may also include pushing information to surrounding on-duty personnel so that the enforcement personnel can go and handle it. The embodiment of the present invention is not limited to this.

[0085] In the embodiment of the present invention, multi-modal target monitoring data is obtained, and multiple target monitoring data are fused to obtain target fusion feature data; the target fusion feature data is input into a preset behavior intention model to obtain a target behavior intention category; the behavior intention model is trained based on the fusion feature data labeled with the behavior intention category; when the target behavior intention category is an abnormal behavior intention category, an early warning is issued for the behavior targeted by the multiple target monitoring data. By fusing monitoring data of multiple modes, the advantages of different data sources can be comprehensively utilized to make up for the shortcomings of a single mode, thereby improving the accuracy of identifying the behavior intention of personnel.

[0086] It can also adapt to different environmental conditions, such as lighting changes, noise interference, etc., which improves the stability and reliability of behavior monitoring.

[0087] In addition, the implementation of the present invention can also timely identify abnormal behavior intentions and issue warnings, initiate corresponding treatment measures, and effectively enhance the effect of security management and prevent the occurrence of potential security risks and bad behaviors. And the fully automatic method can reduce the workload of manual monitoring, improve management efficiency, and also reduce management costs.

[0088] Reference Figure 2 , shows a flowchart of another multimodal behavior warning method according to an embodiment of the present invention, which may include the following steps:

[0089] Step 201: Acquire multi-modal target monitoring data, and fuse multiple target monitoring data to obtain target fusion feature data.

[0090] In practical applications, target monitoring data of different modalities may be obtained first; after obtaining the target monitoring data of different modalities, the features of the multiple target monitoring data may be fused to obtain target fusion feature data.

[0091] Step 202: Input the target fusion feature data into a preset behavior intention model to obtain the target behavior intention category.

[0092] After obtaining the target fusion features, the target fusion features can be input into a preset behavior intention model.

[0093] The preset behavior intention model is based on the target fusion feature and can integrate monitoring data of different modalities to predict the target behavior intention category corresponding to multiple target monitoring data. The target behavior intention category may refer to the behavior intention category of the person in the scene targeted by the target monitoring data.

[0094] Step 203: When the target behavior intention category is an abnormal behavior intention category, determine the target early warning method according to the target behavior intention category.

[0095] In the embodiment of the present invention, if the target behavior intention category is a normal behavior intention category, new target monitoring data may continue to be acquired, and step 201 to step 203 may be re-executed.

[0096] On the contrary, if the target behavior intention category is an abnormal intention category, for example, when the target behavior intention category is a suspicious behavior intention category or an abnormal behavior intention category, an early warning can be issued for the behavior occurring in the scenarios targeted by multiple target monitoring data.

[0097] Specifically, it can be determined first whether the target behavior intention category is a suspicious behavior intention category or an abnormal behavior intention category.

[0098] Then, the target warning method corresponding to the target behavior intention category can be determined; illustratively, different warning methods can be set in advance for different abnormal intention categories; for example: for minor suspicious behavior intention categories, SMS notification can be used for warning; for serious abnormal behavior intention categories, sound and light alarm can be used for warning.

[0099] Furthermore, different treatment measures and emergency plans can be implemented for specific behaviors. For example, for peeping behavior, warnings and expulsions can be implemented; for graffiti behavior, cleaning measures can be implemented; for dangerous behavior, alarms can be implemented, and the embodiments of the present invention do not limit this.

[0100] Step 204: Use a target warning method to issue a warning for the behaviors targeted by the multiple target monitoring data.

[0101] After determining the target warning method, warnings can be issued for behaviors occurring in scenes or areas targeted by multiple target monitoring data.

[0102] In the embodiment of the present invention, multi-modal target monitoring data is obtained, and multiple target monitoring data are fused to obtain target fusion feature data; the target fusion feature data is input into a preset behavior intention model to obtain a target behavior intention category; when the target behavior intention category is an abnormal behavior intention category, a target early warning method is determined according to the target behavior intention category; and the target early warning method is used to warn the behavior targeted by multiple target monitoring data. By fusing monitoring data of multiple modes, the advantages of different data sources can be comprehensively utilized to make up for the shortcomings of a single mode, thereby improving the accuracy of identifying the behavior intention of personnel.

[0103] It can also adapt to different environmental conditions, such as lighting changes, noise interference, etc., which improves the stability and reliability of behavior monitoring.

[0104] In addition, the implementation of the present invention can also timely identify abnormal behavior intentions and issue warnings, initiate corresponding treatment measures, and effectively enhance the effect of security management and prevent the occurrence of potential security risks and bad behaviors. And the fully automatic method can reduce the workload of manual monitoring, improve management efficiency, and also reduce management costs.

[0105] In one embodiment of the present invention, a method for training a behavior intention model is also provided; specifically, refer to Figure 3 , showing a flowchart of the steps of training a behavior intention model in an embodiment of the present invention.

[0106] like Figure 3 As shown, the steps of the behavioral intention model may include:

[0107] Step 301: Acquire multimodal monitoring data for training, and fuse multiple monitoring data for training to obtain fused feature data for training.

[0108] In practical applications, multimodal monitoring data for training may be obtained first; the monitoring data for training may refer to monitoring data such as image data, voice data, sensor data, etc. collected in advance for the scenes and areas to be monitored.

[0109] Exemplarily, there may be multiple groups of multimodal training monitoring data; for each group of multimodal training monitoring data, multiple training monitoring data in the group may be fused to obtain training fused feature data.

[0110] After fusing each set of multimodal monitoring data for training, multiple fused feature data for training can be obtained.

[0111] Step 302: Input the training monitoring data of different modalities into the corresponding prediction model to obtain the first behavior intention category corresponding to the training monitoring data of each modality.

[0112] On the other hand, the training monitoring data of different modalities may also be input into the corresponding prediction model to obtain the first behavior intention category corresponding to the training monitoring data of each modality.

[0113] Exemplarily, different model structures and parameters can be used to perform special training for different modalities, thereby obtaining prediction models corresponding to different modalities. The prediction model can predict the training monitoring data and output the behavior intention category corresponding to the behavior monitored by the training monitoring data.

[0114] For example, for the image modality, a convolutional neural network (CNN) can be used for training; for the speech modality, a recurrent neural network (RNN) can be used for training; for the sensor modality, a support vector machine (SVM) can be used for training.

[0115] After obtaining the prediction model corresponding to each modality, the training monitoring data can be input into the corresponding prediction model.

[0116] After obtaining the training monitoring data, each prediction model can analyze the data and output the behavior intention category corresponding to the behavior monitored by the training monitoring data, that is, the first behavior intention.

[0117] Step 303: Determine a second behavior intention category based on multiple first behavior intention categories.

[0118] After obtaining the first behavior intention categories corresponding to each of a group of multimodal training monitoring data, the second behavior intention category corresponding to the group of training monitoring data can be determined based on the multiple first behavior intention categories.

[0119] Step 304: annotate the training fusion feature data with the second behavior intention category, and train the preset model based on the annotated training fusion feature data to obtain a behavior intention model.

[0120] Then, the training fused feature data obtained by fusing the group of multimodal training monitoring data may be labeled with the second behavior intention category.

[0121] Next, the preset model can be trained based on the labeled training fusion feature data to obtain a behavior intention model.

[0122] For example, a convolutional neural network (CNN) can be selected as the machine learning algorithm for the behavior intention model. CNN is a deep learning model that can automatically learn features in images and has strong object recognition and behavior intention modeling capabilities. During the training process, CNN continuously adjusts its weights and biases so that its output results are close to the actual labeled labels, thereby ultimately obtaining a behavior intention model.

[0123] In an embodiment of the present invention, multimodal training monitoring data is obtained, and multiple training monitoring data are fused to obtain training fused feature data; the training monitoring data of different modalities are input into the corresponding prediction model to obtain the first behavior intention category corresponding to the training monitoring data of each modality; the second behavior intention category is determined based on the multiple first behavior intention categories; the training fused feature data is labeled with the second behavior intention category, and the preset model is trained based on the labeled training fused feature data to obtain a behavior intention model. The behavior intention model constructed by the embodiment of the present invention can comprehensively utilize the advantages of different data sources and make up for the shortcomings of a single modality by fusing monitoring data of multiple modalities, thereby improving the accuracy of identifying personnel behavior intentions.

[0124] It can also adapt to different environmental conditions, such as lighting changes, noise interference, etc., which improves the stability and reliability of behavior monitoring.

[0125] In addition, the implementation of the present invention can also timely identify abnormal behavior intentions and issue warnings, initiate corresponding treatment measures, and effectively enhance the effect of security management and prevent the occurrence of potential security risks and bad behaviors. And the fully automatic method can reduce the workload of manual monitoring, improve management efficiency, and also reduce management costs.

[0126] Reference Figure 4 , shows a flowchart of the steps of another training behavior intention model according to an embodiment of the present invention.

[0127] Step 401: Acquire multimodal monitoring data for training, and fuse multiple monitoring data for training to obtain fused feature data for training.

[0128] In practical applications, multimodal monitoring data for training may be obtained first; the monitoring data for training may refer to monitoring data such as image data, voice data, sensor data, etc. collected in advance for the scenes and areas to be monitored.

[0129] For each set of multimodal training monitoring data, multiple training monitoring data in the set may be fused to obtain training fused feature data. After fusion of each set of multimodal training monitoring data, multiple training fused feature data may be obtained.

[0130] Step 402: Input the training monitoring data of different modalities into the corresponding prediction model to obtain the first behavior intention category corresponding to the training monitoring data of each modality.

[0131] On the other hand, the training monitoring data of different modalities may also be input into the corresponding prediction model to obtain the first behavior intention category corresponding to the training monitoring data of each modality.

[0132] Exemplarily, different model structures and parameters can be used to perform special training for different modalities, thereby obtaining prediction models corresponding to different modalities. The prediction model can predict the training monitoring data and output the behavior intention category corresponding to the behavior monitored by the training monitoring data.

[0133] Specifically, for the image modality, a convolutional neural network (CNN) can be used for training; for the speech modality, a recurrent neural network (RNN) can be used for training; for the sensor modality, a support vector machine (SVM) can be used for training.

[0134] After obtaining the prediction model corresponding to each modality, the training monitoring data can be input into the corresponding prediction model.

[0135] After obtaining the training monitoring data, each prediction model can analyze the data and output the behavior intention category corresponding to the behavior monitored by the training monitoring data, that is, the first behavior intention.

[0136] Step 403: merge multiple first behavior intention categories to obtain a second behavior intention category.

[0137] After obtaining the first behavior intention categories corresponding to each of a group of multimodal training monitoring data, the multiple first behavior intention categories can be fused to obtain the second behavior intention category corresponding to the group of training monitoring data.

[0138] In an embodiment of the present invention, a fusion method for fusing multiple first behavior intention categories includes any one of the following:

[0139] Voting method, weighted average method, neural network fusion method.

[0140] For example, in the voting method, independent prediction models are trained for each modality, and these prediction models are used to predict the training monitoring data to obtain the first behavior intention category corresponding to each modality. For example, for the image modality, the prediction model predicts "normal behavior intention category" or "suspicious behavior intention category (voyeurism)", etc.; the voice modality and sensor modality also have their own first behavior intention category.

[0141] Then, the results of each modality are voted and counted. For example, if there are three modalities (image, voice, sensor), for a group of training monitoring data, the result of the image modality is "normal behavior intention category", the result of the voice modality is "suspicious behavior intention category", and the sensor modality prediction is "normal behavior intention category"; then in the voting method, the "normal behavior intention category" gets two votes, and the "suspicious behavior intention category" gets one vote, and finally the second behavior intention category of the group is the normal behavior intention category.

[0142] If a tie occurs, the final result may be determined according to a preset rule, such as giving priority to the result of a certain mode, or randomly selecting a result, etc., which is not limited in this embodiment of the present invention.

[0143] Taking the weighted average method as an example, we can first train the prediction models of different modalities separately and obtain their first behavioral intention categories.

[0144] But on this basis, a weight needs to be assigned to each modality. The weight can be determined based on the reliability of the modality, the accuracy on the training data, expert experience, or through some optimization algorithm.

[0145] Assume that the image modality weight is w1, the speech modality weight is w2, and the sensor modality weight is w3 (w1+w2+w3=1).

[0146] For the second behavior intention category of each group, if it is a classification problem, for example, the behavior intention category is divided into two categories: "normal behavior intention category" (represented by 0) and "abnormal behavior intention category" (represented by 1), the result of image modality prediction is p1, the result of voice modality prediction is p2, and the result of sensor modality prediction is p3, then the weighted average fusion prediction result p = w1*p1+w2*p2+w3*p3. If p is greater than a certain threshold (such as 0.5), it is determined to be an "abnormal behavior intention category", otherwise it is a "normal behavior intention category".

[0147] For multi-classification problems, the principle is similar, but the calculation process is more complicated. It is necessary to perform weighted summation according to the prediction probability corresponding to each category, and select the category with the largest probability as the final fusion result. The embodiment of the present invention does not limit this.

[0148] Taking neural network fusion as an example, we can first train the prediction models of different modalities separately and obtain their first behavioral intention categories for the training monitoring data, and then use these first behavioral intention categories as input features to build a new neural network.

[0149] For example, if the prediction model of the image modality outputs a feature vector of length n1 representing its predicted probability distribution for different behavioral intentions, the prediction model of the speech modality outputs a feature vector of length n2, and the prediction model of the sensor modality outputs a feature vector of length n3, then these vectors are concatenated to form a new input vector with a dimension of n1+n2+n3.

[0150] The constructed neural network can have multiple hidden layers, and through training on training data, it can learn how to fuse the prediction results of different modalities to obtain a more accurate final decision result.

[0151] For example, the output layer of the neural network can set the corresponding number of neurons according to the number of categories of behavioral intentions, convert the output into the probability distribution of each category through the softmax function, and then select the category with the highest probability as the final fusion prediction result. When training this neural network, a large amount of labeled data is usually required to adjust the weights and biases of the neural network to optimize its fusion performance.

[0152] Step 404: annotate the training fusion feature data with the second behavior intention category, and train the preset model based on the annotated training fusion feature data to obtain a behavior intention model.

[0153] Then, the training fused feature data obtained by fusing the group of multimodal training monitoring data may be labeled with the second behavior intention category.

[0154] Next, the preset model can be trained based on the labeled training fusion feature data to obtain a behavior intention model.

[0155] For example, a convolutional neural network (CNN) can be selected as the machine learning algorithm for the behavior intention model. CNN is a deep learning model that can automatically learn features in images and has strong object recognition and behavior intention modeling capabilities. During the training process, CNN continuously adjusts its weights and biases so that its output results are close to the actual labeled labels, thereby ultimately obtaining a behavior intention model.

[0156] In an embodiment of the present invention, multimodal training monitoring data is obtained, and multiple training monitoring data are fused to obtain training fused feature data; the training monitoring data of different modalities are input into the corresponding prediction model to obtain the first behavior intention category corresponding to the training monitoring data of each modality; multiple first behavior intention categories are fused to obtain a second behavior intention category; the training fused feature data is labeled with the second behavior intention category, and the preset model is trained based on the labeled training fused feature data to obtain a behavior intention model. The behavior intention model constructed by the embodiment of the present invention can comprehensively utilize the advantages of different data sources and make up for the shortcomings of a single modality by fusing monitoring data of multiple modalities, thereby improving the accuracy of identifying personnel behavior intentions.

[0157] It can also adapt to different environmental conditions, such as lighting changes, noise interference, etc., which improves the stability and reliability of behavior monitoring.

[0158] In addition, the implementation of the present invention can also timely identify abnormal behavior intentions and issue warnings, initiate corresponding treatment measures, and effectively enhance the effect of security management and prevent the occurrence of potential security risks and bad behaviors. And the fully automatic method can reduce the workload of manual monitoring, improve management efficiency, and also reduce management costs.

[0159] Reference Figure 5 , shows a flowchart of the steps of another training behavior intention model according to an embodiment of the present invention.

[0160] Step 501: Obtain multi-modal monitoring data for training.

[0161] In practical applications, multimodal monitoring data for training may be obtained first; the monitoring data for training may refer to monitoring data such as image data, voice data, sensor data, etc. collected in advance for the scenes and areas to be monitored.

[0162] Step 502: pre-process each training monitoring data.

[0163] For each group of multimodal training monitoring data, multiple training monitoring data in the group may be preprocessed.

[0164] In one embodiment of the present invention, the preprocessing includes at least one of the following:

[0165] Noise reduction, enhancement, calibration, normalization, filtering.

[0166] In some feasible embodiments, different preprocessing methods may be used to process training monitoring data of different modalities.

[0167] Specifically, Gaussian filtering can be used to remove noise from image data. Gaussian filtering is a linear smoothing filter that achieves the purpose of noise reduction by weighted averaging each pixel in the image with the pixels in its neighborhood. It can also improve the quality and clarity of the image by contrast enhancement and brightness adjustment.

[0168] For speech data, spectral subtraction can be used to remove noise from speech data. Spectral subtraction is a noise reduction method based on signal spectrum analysis. It estimates the power spectrum of speech data and the power spectrum of noise signal, and then subtracts the power spectrum of noise signal from the power spectrum of speech data to obtain the power spectrum of pure speech signal. Finally, the power spectrum of pure speech data is converted back to time domain through inverse Fourier transform to obtain the denoised speech data.

[0169] In some feasible embodiments, continuous voice data can also be segmented into individual voice segments for subsequent analysis and processing. In addition, a voice activity detection (VAD) algorithm or other voice segmentation algorithms can also be used to segment the voice data according to characteristics such as energy and frequency.

[0170] For sensor data, the sensor can be calibrated to ensure the accuracy of its measurement results. The sensor data can also be filtered using a low-pass filter to remove noise and interference. The sensor data can also be normalized to keep its value within a certain range for subsequent analysis and processing.

[0171] Step 503: extract features from the preprocessed plurality of training monitoring data to obtain feature data corresponding to each training monitoring data.

[0172] After preprocessing each training monitoring data, feature extraction may be performed from the preprocessed plurality of training monitoring data to obtain feature data corresponding to each training monitoring data.

[0173] For example, for image data, the Canny edge detection algorithm can be used to extract the contour features of a person. Specifically, the steps of the Canny edge detection algorithm include: first, performing Gaussian filtering on the image data to remove noise; then, calculating the gradient amplitude and direction of the image data; then, performing non-maximum suppression on the gradient amplitude to remove false edges; finally, performing double threshold processing on the gradient amplitude after non-maximum suppression to obtain the edge contour of the image data.

[0174] In addition, convolutional neural networks (CNNs) can be used to identify objects in image data and extract the features of the objects. Specifically, the structure of CNN includes input layer, convolution layer, pooling layer, fully connected layer and output layer. During the training process, CNN continuously adjusts its weights and biases to make its output as close to the actual label as possible. During the testing process, CNN can identify objects in the input image data and extract the features of the objects.

[0175] In some feasible embodiments, the distribution of different colors in the image data may also be analyzed to extract color distribution features, which may be extracted using methods such as color histograms and color moments.

[0176] By analyzing the differences between consecutive frames of image data, the motion features of people, such as motion speed, motion direction, motion trajectory, etc., can be extracted. Motion feature extraction can be performed using algorithms such as optical flow and background subtraction. Optical flow is a method that calculates the speed and direction of an object based on the motion information of pixels in an image sequence. Background subtraction is a method that extracts the motion features of foreground objects by differentiating the current frame image from the background image.

[0177] For voice data, the duration and number of words in the voice data can be calculated to obtain the speech speed feature. Speech recognition software or tools can be used to convert the voice data into text and then calculate the speech speed.

[0178] You can also analyze the frequency changes of speech data to extract intonation features. You can use speech signal processing software or tools to perform spectrum analysis on speech data to obtain intonation features.

[0179] You can also use the TF-IDF (Term Frequency-Inverse Document Frequency) algorithm to extract keywords from speech data. The TF-IDF algorithm is an algorithm used to calculate the importance of words in a text. It measures the importance of words by calculating the frequency of occurrence of words in the text (Term Frequency, TF) and the inverse document frequency (Inverse Document Frequency, IDF). In speech data, you can convert the speech data into text and then use the TF-IDF algorithm to extract keywords.

[0180] It is also possible to analyze the emotional information in the speech data and extract emotional features. You can use emotional analysis algorithms, such as sentiment analysis algorithms based on machine learning and sentiment analysis algorithms based on deep learning.

[0181] For sensor data, motion speed feature extraction can be performed; specifically, the motion speed of a person can be calculated based on the sensor data. The motion speed can be obtained by integrating the data of motion sensors, such as accelerometers, gyroscopes, etc.

[0182] For sensor data, acceleration feature extraction can also be performed; specifically, the acceleration information in the sensor data can be analyzed to extract acceleration features. The magnitude and direction of acceleration can be calculated using accelerometer data.

[0183] For sensor data, directional features can also be extracted; specifically, the direction of movement of a person can be determined based on the sensor data. The data of motion sensors, such as a compass, a gyroscope, etc., can be used in combination with map information to determine the direction of movement of a person.

[0184] For sensor data, distance feature extraction can also be performed; specifically, the sensor data is used to calculate the distance between a person and other objects or locations. Proximity sensors, infrared sensors, etc. can be used to measure the distance between a person and an object.

[0185] Step 504: Fuse the feature data corresponding to each training monitoring data to obtain training fused feature data.

[0186] After obtaining the feature data corresponding to each preprocessed training monitoring data, for a group of training monitoring data, the feature data corresponding to each training monitoring data in the group can be fused to obtain the training fused feature data corresponding to the group.

[0187] In one embodiment of the present invention, the fusion method for fusing the feature data corresponding to each training monitoring data includes any one of the following:

[0188] Weighted fusion, feature mapping, and sequential concatenation.

[0189] In some feasible embodiments, the feature vectors (i.e., feature data) extracted from different modalities can be directly spliced ​​together in a predetermined order to form a new feature vector with a higher dimension. For example, assuming that the dimension of the image feature vector is n1, the dimension of the speech feature vector is n2, and the dimension of the sensor feature vector is n3, the dimension of the high-dimensional feature vector obtained after splicing is n1+n2+n3.

[0190] In other feasible embodiments, corresponding weights may be assigned to features of different modalities, and the weights may be determined based on prior knowledge, expert experience, or through some data-driven methods (such as preliminary experimental comparisons on small-scale labeled data, etc.). Then, each element in each modality feature vector is multiplied by the corresponding weight, and then concatenated to form a high-dimensional feature vector.

[0191] In some other feasible embodiments, a specific mapping function (which may be a linear function, a nonlinear function, etc.) can be used to map the features of different modalities into a unified feature space, and then splice or further process them in this unified space to form the final high-dimensional feature vector.

[0192] Step 505: Input the training monitoring data of different modalities into the corresponding prediction model to obtain the first behavior intention category corresponding to the training monitoring data of each modality.

[0193] Next, the training monitoring data of different modalities may be input into the corresponding prediction model to obtain the first behavior intention category corresponding to the training monitoring data of each modality.

[0194] Exemplarily, different model structures and parameters can be used to perform special training for different modalities, thereby obtaining prediction models corresponding to different modalities. The prediction model can predict the training monitoring data and output the behavior intention category corresponding to the behavior monitored by the training monitoring data.

[0195] Specifically, for the image modality, a convolutional neural network (CNN) can be used for training; for the speech modality, a recurrent neural network (RNN) can be used for training; for the sensor modality, a support vector machine (SVM) can be used for training.

[0196] After obtaining the prediction model corresponding to each modality, the training monitoring data can be input into the corresponding prediction model.

[0197] After obtaining the training monitoring data, each prediction model can analyze the data and output the behavior intention category corresponding to the behavior monitored by the training monitoring data, that is, the first behavior intention.

[0198] Step 506: merge multiple first behavior intention categories to obtain a second behavior intention category.

[0199] After obtaining the first behavior intention categories corresponding to each of a group of multimodal training monitoring data, the multiple first behavior intention categories can be fused to obtain the second behavior intention category corresponding to the group of training monitoring data.

[0200] Step 507: label the training fusion feature data with the second behavior intention category, and train the preset model based on the labeled training fusion feature data to obtain a behavior intention model.

[0201] Then, the training fused feature data obtained by fusing the group of multimodal training monitoring data may be labeled with the second behavior intention category.

[0202] Next, the preset model can be trained based on the labeled training fusion feature data to obtain a behavior intention model.

[0203] For example, a convolutional neural network (CNN) can be selected as the machine learning algorithm for the behavior intention model. CNN is a deep learning model that can automatically learn features in images and has strong object recognition and behavior intention modeling capabilities. During the training process, CNN continuously adjusts its weights and biases so that its output results are close to the actual labeled labels, thereby obtaining a behavior intention model.

[0204] In an embodiment of the present invention, multimodal training monitoring data is obtained; each training monitoring data is preprocessed; feature extraction is performed from the preprocessed multiple training monitoring data to obtain feature data corresponding to each training monitoring data; the feature data corresponding to each training monitoring data is fused to obtain fused training feature data; the training monitoring data of different modalities are input into the corresponding prediction model to obtain the first behavior intention category corresponding to the training monitoring data of each modality; multiple first behavior intention categories are fused to obtain a second behavior intention category; the training fused feature data is labeled with the second behavior intention category, and the preset model is trained based on the labeled training fused feature data to obtain a behavior intention model. The behavior intention model constructed by the embodiment of the present invention can comprehensively utilize the advantages of different data sources and make up for the shortcomings of a single modality by fusing monitoring data of multiple modalities, thereby improving the accuracy of identifying personnel behavior intentions.

[0205] It can also adapt to different environmental conditions, such as lighting changes, noise interference, etc., which improves the stability and reliability of behavior monitoring.

[0206] In addition, the implementation of the present invention can also timely identify abnormal behavior intentions and issue warnings, initiate corresponding treatment measures, and effectively enhance the effect of security management and prevent the occurrence of potential security risks and bad behaviors. And the fully automatic method can reduce the workload of manual monitoring, improve management efficiency, and also reduce management costs.

[0207] For example, Figure 6 As shown, data collection may be performed first; specifically, it may include image data collection, voice data collection, and sensor data collection.

[0208] After data collection is completed, data preprocessing can be performed; specifically, it can include image data preprocessing, voice data preprocessing, and sensor data preprocessing.

[0209] Next, feature extraction may be performed; specifically, it may include image feature extraction of image data, speech feature extraction of speech data, and sensor feature extraction of sensor data.

[0210] After feature extraction is completed, early fusion of features and late fusion of prediction results can be performed.

[0211] Next, behavioral intention modeling can be performed; specifically, behavioral intention categories can be defined, machine learning algorithms can be selected, and model training and verification can be performed to obtain a behavioral intention model.

[0212] After obtaining the behavior intention model, the currently collected target monitoring data can be input into the behavior intention model; then the early warning signal is issued and the treatment measures are initiated.

[0213] It should be noted that, for the sake of simplicity, the method embodiments are described as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present invention are not limited by the order of the actions described, because according to the embodiments of the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.

[0214] Reference Figure 7 , shows a schematic diagram of the structure of a multi-modal behavior warning device according to an embodiment of the present invention, which may include the following modules:

[0215] The fusion module 701 is used to obtain multi-modal target monitoring data and fuse multiple target monitoring data to obtain target fusion feature data;

[0216] Prediction model 702, used to input the target fusion feature data into a preset behavior intention model to obtain the target behavior intention category; the behavior intention model is trained based on the fusion feature data labeled with the behavior intention category;

[0217] The early warning module 703 is used to issue an early warning for the behaviors targeted by the multiple target monitoring data when the target behavior intention category is an abnormal behavior intention category.

[0218] In an optional embodiment of the present invention, the early warning module 703 is used to determine a target early warning method according to the target behavior intention category; and use the target early warning method to issue an early warning for the behavior targeted by multiple target monitoring data.

[0219] In an optional embodiment of the present invention, the device further comprises:

[0220] The training module is used to obtain multi-modal training monitoring data and fuse multiple training monitoring data to obtain training fused feature data; input the training monitoring data of different modalities into the corresponding prediction model to obtain the first behavior intention category corresponding to the training monitoring data of each modality; determine the second behavior intention category based on multiple first behavior intention categories; label the training fused feature data with the second behavior intention category, and train the preset model based on the labeled training fused feature data to obtain a behavior intention model.

[0221] In an optional embodiment of the present invention, the training module is used to fuse multiple first behavior intention categories to obtain a second behavior intention category.

[0222] In an optional embodiment of the present invention, the training module is used to preprocess each training monitoring data; extract features from the preprocessed multiple training monitoring data to obtain feature data corresponding to each training monitoring data; and fuse the feature data corresponding to each training monitoring data to obtain fused feature data for training.

[0223] In an optional embodiment of the present invention, the preprocessing includes at least one of the following: noise reduction, enhancement, calibration, normalization, and filtering.

[0224] In an optional embodiment of the present invention, a fusion method for fusing feature data corresponding to each training monitoring data includes any one of the following: weighted fusion, feature mapping, and sequential splicing.

[0225] In an optional embodiment of the present invention, a fusion method for fusing multiple first behavior intention categories includes any one of the following: a voting method, a weighted average method, and a neural network fusion method.

[0226] In an optional embodiment of the present invention, the target monitoring data includes the following data: image data, voice data, and sensor data.

[0227] In the embodiment of the present invention, multi-modal target monitoring data is obtained, and multiple target monitoring data are fused to obtain target fusion feature data; the target fusion feature data is input into a preset behavior intention model to obtain a target behavior intention category; the behavior intention model is trained based on the fusion feature data labeled with the behavior intention category; when the target behavior intention category is an abnormal behavior intention category, an early warning is issued for the behavior targeted by the multiple target monitoring data. By fusing monitoring data of multiple modes, the advantages of different data sources can be comprehensively utilized to make up for the shortcomings of a single mode, thereby improving the accuracy of identifying the behavior intention of personnel.

[0228] It can also adapt to different environmental conditions, such as lighting changes, noise interference, etc., which improves the stability and reliability of behavior monitoring.

[0229] In addition, the implementation of the present invention can also timely identify abnormal behavior intentions and issue warnings, initiate corresponding treatment measures, and effectively enhance the effect of security management and prevent the occurrence of potential security risks and bad behaviors. And the fully automatic method can reduce the workload of manual monitoring, improve management efficiency, and also reduce management costs.

[0230] An embodiment of the present invention further provides an electronic device, including a processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein when the computer program is executed by the processor, the multimodal-based behavior warning method as described above is implemented.

[0231] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the multi-modal-based behavior warning method as described above is implemented.

[0232] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0233] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0234] Those skilled in the art will appreciate that the embodiments of the present invention may be provided as methods, devices, or computer program products. Therefore, the embodiments of the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the embodiments of the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0235] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0236] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0237] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0238] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.

[0239] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements.

[0240] The above provides a detailed introduction to a multimodal behavior warning method, a multimodal behavior warning method device, an electronic device and a computer-readable storage medium. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for general technical personnel in this field, according to the idea of ​​the present invention, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.

Claims

1. A multimodal behavior warning method, characterized in that: The method comprises: Acquire multi-modal target monitoring data, and fuse multiple target monitoring data to obtain target fusion feature data; Inputting the target fusion feature data into a preset behavior intention model to obtain a target behavior intention category; the behavior intention model is trained based on the fusion feature data labeled with the behavior intention category; When the target behavior intention category is an abnormal behavior intention category, an early warning is issued for the behaviors targeted by the multiple target monitoring data.

2. The method according to claim 1, characterized in that The warning for the behaviors targeted by the multiple target monitoring data includes: Determine the target warning method according to the target behavior intention category; The target early warning method is used to issue an early warning for the behavior targeted by the multiple target monitoring data.

3. The method according to claim 1, characterized in that The method further comprises: Acquire multimodal monitoring data for training, and fuse multiple monitoring data for training to obtain fused feature data for training; Inputting the training monitoring data of different modalities into the corresponding prediction model to obtain the first behavior intention category corresponding to the training monitoring data of each modality; determining a second behavioral intention category based on the plurality of first behavioral intention categories; The training fusion feature data is labeled with a second behavior intention category, and a preset model is trained based on the labeled training fusion feature data to obtain the behavior intention model.

4. The method according to claim 3, characterized in that: The determining the second behavior intention category according to the plurality of first behavior intention categories includes: The multiple first behavior intention categories are merged to obtain the second behavior intention category.

5. The method according to claim 3, characterized in that: The fusing of a plurality of monitoring data for training to obtain fused feature data for training includes: Preprocess each training monitoring data; Extract features from the preprocessed multiple training monitoring data to obtain feature data corresponding to each training monitoring data; The feature data corresponding to each training monitoring data are fused to obtain the training fused feature data.

6. The method according to claim 5, characterized in that The pretreatment includes at least one of the following: Noise reduction, enhancement, calibration, normalization, filtering.

7. The method according to claim 5, characterized in that The fusion method for fusing the feature data corresponding to each training monitoring data includes any of the following: Weighted fusion, feature mapping, and sequential concatenation.

8. The method according to claim 4, characterized in that The fusion method of fusion of multiple first behavior intention categories includes any of the following: Voting method, weighted average method, neural network fusion method.

9. The method according to claim 1, characterized in that: The target monitoring data includes the following data: Image data, voice data, sensor data.

10. A multi-modal behavior warning device, characterized in that: The device comprises: A fusion module is used to obtain multi-modal target monitoring data and fuse multiple target monitoring data to obtain target fusion feature data; A prediction model is used to input the target fusion feature data into a preset behavior intention model to obtain a target behavior intention category; the behavior intention model is trained based on the fusion feature data labeled with the behavior intention category; The early warning module is used to issue an early warning for the behavior targeted by the multiple target monitoring data when the target behavior intention category is an abnormal behavior intention category.

11. An electronic device, characterized in that: It comprises a processor, a memory and a computer program stored in the memory and capable of running on the processor, wherein when the computer program is executed by the processor, the multimodal-based behavior warning method as claimed in any one of claims 1 to 9 is implemented.

12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the multimodal-based behavior warning method as described in any one of claims 1 to 9 is implemented.