A dangerous behavior identification and early warning method based on multimodal analysis

By combining multimodal analysis with video, audio, and X-ray detection, the problems of high labor costs and low recognition accuracy in existing security systems are solved, and efficient and accurate dangerous behavior recognition and warning are achieved.

CN117437574BActive Publication Date: 2025-09-26ANHUI HEXIN TECH DEV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202311396026.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-26
Publication Date
2025-09-26
Estimated Expiration
2043-10-26

AI Technical Summary

Technical Problem

The existing security system has problems such as high labor costs, low recognition efficiency and low recognition accuracy.

Method used

Through multimodal analysis methods, combined with video data, audio data and X-ray detection, the target person's image, voice information, expression and behavior information are obtained, and comprehensive analysis is performed using image fusion technology and emotion recognition models to generate a warning report.

Benefits of technology

It improves the accuracy of identifying dangerous behaviors, reduces the degree of manual intervention, and improves recognition efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117437574B_ABST
    Figure CN117437574B_ABST
Patent Text Reader

Abstract

The present invention discloses a dangerous behavior identification and warning method based on multimodal analysis, comprising the following steps: obtaining video data and audio data of a target person in a preset area, and extracting multiple frames of image data from the video data; processing and analyzing the audio data, image data and video data respectively, and making a risk assessment of the target person; drawing a warning report based on the risk assessment, and sending the warning report to a preset terminal; the present invention captures the image of the target person through video, collects voice information and video behavior, detects facial expressions, performs emotion recognition on audio information, classifies and judges video behavior, conducts comprehensive analysis, predicts and warns of dangerous behaviors, and multi-angle analysis can reduce data deviations caused by accidental states, improve recognition accuracy, and at the same time reduce the degree of manual intervention, which can greatly improve recognition efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of security, and in particular to a dangerous behavior identification and early warning method based on multimodal analysis. Background Art

[0002] With the advancement of technology, security requirements are increasing. Conventional video surveillance and identity recognition methods are no longer sufficient for organizations or institutions with high security demands. Dangerous behaviors such as fighting, loitering, and stalking are rampant in some settings. Early warning and timely response to these behaviors are crucial, making behavioral recognition increasingly important.

[0003] Among the existing behavior recognition methods in the security field, most are performed by security personnel through video surveillance to observe the abnormal behavior of suspicious persons, or through real-time monitoring to perform behavior recognition, by simulating the movement profile of the person and searching the database to identify the behavior category. Although these methods are effective, they have high labor costs, low recognition efficiency, and low recognition accuracy. Therefore, a dangerous behavior recognition and early warning method based on multimodal analysis is proposed. Summary of the Invention

[0004] The technical problem to be solved by the present invention is how to solve the problems of high labor cost, low recognition efficiency and low recognition accuracy of existing safety protection systems, and provide a dangerous behavior recognition and early warning method based on multimodal analysis.

[0005] The present invention solves the above technical problems through the following technical solutions, which include the following steps:

[0006] S1. Obtain video data and audio data of a target person in a preset area, and extract multiple frames of image data from the video data;

[0007] S2. Process and analyze the audio data, image data, and video data to determine the risk of the target person;

[0008] S3. A warning report is generated based on the risk assessment and sent to a preset terminal.

[0009] Preferably, the video data and audio data in S1 are collected by a preset device, which includes an image acquisition device with a sound receiving function. The number of preset devices is m, where m is a positive integer greater than 3. S1 specifically includes:

[0010] Select image acquisition devices at different locations;

[0011] Determine the effectiveness of the image acquisition device selection based on the location of the image acquisition device and the location of the image acquisition device and the target person;

[0012] Synchronously obtain the image data and audio data of the target person;

[0013] Then, multiple frames of image data are read and acquired frame by frame, and the image fusion technology is used to acquire the first type of image data at the same moment.

[0014] Preferably, the process of judging the effectiveness of the image acquisition device selection is specifically as follows:

[0015] With the geometric center of the target person as the origin and the origin as the initial point, a spatial rectangular coordinate system is established, wherein the plane where the X-axis and the Y-axis are located is parallel to the horizontal plane, and the spatial rectangular coordinate system is established with the vertical direction as the preset dimension of the Z-axis as the unit;

[0016] Randomly select two image acquisition devices and create two straight lines connecting the origin to the center of the image acquisition device;

[0017] Read the coordinates A(x1, y1, z1) and B(x2, y2, z2) of the geometric centers of the two image acquisition devices respectively and calculate the angle α between the two lines. The specific calculation process is as follows:

[0018]

[0019] Then use trigonometric functions to calculate the angle α;

[0020] When the angle α is greater than or equal to the preset threshold C, a first device reselection instruction is generated.

[0021] Preferably, when the angle α is less than the preset threshold C1, the following judgment process is further performed:

[0022] Create a first type of area in the shooting direction of the target person, and evenly create multiple reference points in the first type of area;

[0023] Use any two image acquisition devices to simultaneously shoot the same person target;

[0024] Count the number of reference points that can be captured by the two image acquisition devices, and record them as M1 and M2 respectively;

[0025] Count the total number of reference points M3 in the first type of area;

[0026] Calculate the overlap E of the two image acquisition devices capturing the target person. The specific calculation process is as follows:

[0027]

[0028] When the overlap amount E is less than the preset threshold value C2, a second device reselection instruction is generated.

[0029] Preferably, the process of determining the effectiveness of the image acquisition device selection further includes:

[0030] According to the formula Calculate the distance D between multiple image acquisition devices and the target person respectively, remove the maximum and minimum values ​​of the distance D to calculate the average distance D;

[0031] Using the formula method Calculate the deviation degree F of the image acquisition device;

[0032] When the degree of deviation F>the preset threshold C3, a clearing instruction is generated.

[0033] Preferably, the S2 specifically includes:

[0034] Input the image data into the image recognition model to obtain the first type of emotional information of the target person;

[0035] Input the audio data into the emotion recognition model to obtain the second type of emotional information of the target person;

[0036] The first type of emotional information and the second type of emotional information both include excitement and calmness;

[0037] Comparing the first type of emotional information with the second type of emotional information, and formulating a warning alert report when the first type of information and the second type of information are the same;

[0038] Inputting video data into a video behavior recognition model to obtain behavior information of the target person, the behavior information including danger and safety;

[0039] When the target person's behavior information is dangerous, a warning report will be issued.

[0040] Preferably, before S3, the method further includes detecting the type of items carried by the target person by using an X-ray detection device to determine the danger level. The specific detection process is as follows:

[0041] Use X-ray detection equipment to irradiate the target person and obtain data on the target items carried by the target person;

[0042] Import X-ray comparison item data into the database;

[0043] Extracting the first type of features from the target item and the second type of features from the comparison item;

[0044] Record the number of first-type features H1 and the number of second-type features H2 respectively;

[0045] Compare the first type of features with the second type of features to obtain a number H3 of common features between the first type of features and the second type of features;

[0046] Determine the similarity Q between two items. The specific judgment process is:

[0047] Q=2H3 / (H1+H2)

[0048] Classify the target item data based on the comparison item data whose similarity degree Q ≥ preset threshold N;

[0049] Determine whether the target item is dangerous based on the dangerousness of the comparative item data classified by the target item data, and generate a monitoring frequency report;

[0050] When the target object is dangerous, the target person is monitored in real time to increase the frequency of obtaining video and audio data of the target person.

[0051] Preferably, when the approximation degree Q is less than the preset threshold N, the following processing is further performed:

[0052] First, the position of the X-ray detection equipment is changed multiple times to re-acquire the first type of features in the target object, and multiple approximation degrees Q are repeatedly calculated. Then, the target object data is classified according to the comparison object data with approximation degree Q ≥ preset threshold N; finally, the monitoring frequency report is verified based on the dangerousness of the target object.

[0053] Compared with the existing technology, the present invention has the following advantages: the dangerous behavior identification and warning method based on multimodal analysis captures the target person's image through video, collects voice information and video behavior, detects facial expressions, recognizes emotions in audio information, classifies and judges video behavior, and conducts comprehensive analysis to predict and alarm dangerous behaviors. Multi-angle analysis can reduce data deviations caused by accidental states, improve recognition accuracy, and at the same time reduce the degree of manual intervention, which can greatly improve recognition efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 It is the overall flow chart of the present invention;

[0055] Figure 2 It is a schematic diagram of establishing the spatial rectangular coordinate system in the present invention. DETAILED DESCRIPTION

[0056] The following is a detailed description of an embodiment of the present invention. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process. However, the protection scope of the present invention is not limited to the following embodiment.

[0057] like Figure 1 As shown, this embodiment provides a technical solution: a dangerous behavior identification and early warning method based on multimodal analysis, comprising the following steps:

[0058] S1. Obtain video data and audio data of a target person in a preset area, and extract multiple frames of image data from the video data;

[0059] S2. Process and analyze the audio data, image data, and video data to determine the risk of the target person;

[0060] S3. A warning report is generated based on the risk assessment and sent to a preset terminal.

[0061] This recognition and warning method captures the target person's image through video, collects voice information and video behavior, detects facial expressions, recognizes emotions in audio information, classifies and judges video behavior, and conducts comprehensive analysis to predict and alarm dangerous behaviors. Multi-angle analysis can reduce data deviations caused by accidental states, improve recognition accuracy, and reduce the degree of manual intervention, which can greatly improve recognition efficiency.

[0062] The video data and audio data in S1 are collected by a preset device, which includes an image acquisition device with a sound receiving function. The number of preset devices is m, where m is a positive integer greater than 3. S1 specifically includes:

[0063] Select image acquisition devices at different locations;

[0064] Determine the effectiveness of the image acquisition device selection based on the location of the image acquisition device and the location of the image acquisition device and the target person;

[0065] Synchronously obtain the image data and audio data of the target person;

[0066] Then, multiple frames of image data are read and acquired frame by frame, and the image fusion technology is used to acquire the first type of image data at the same moment.

[0067] It should be noted that the image acquisition device can be a camera, which first obtains the video stream through the camera, then reads the monitoring picture in the video stream frame by frame, and then obtains the audio stream through the camera's microphone.

[0068] Reference Figure 2 ,Furthermore, the judgment process of the effectiveness of the image acquisition device selection is as follows:

[0069] With the geometric center of the target person as the origin and the origin as the initial point, a spatial rectangular coordinate system is established, wherein the plane where the X-axis and the Y-axis are located is parallel to the horizontal plane, and the spatial rectangular coordinate system is established with the vertical direction as the preset dimension of the Z-axis as the unit;

[0070] Randomly select two image acquisition devices and create two straight lines connecting the origin to the center of the image acquisition device;

[0071] Read the coordinates A(x1, y1, z1) and B(x2, y2, z2) of the geometric centers of the two image acquisition devices respectively and calculate the angle α between the two lines. The specific calculation process is as follows:

[0072]

[0073] Then use trigonometric functions to calculate the angle α;

[0074] When the angle α is greater than or equal to the preset threshold C, a first device reselection instruction is generated.

[0075] It should be noted that the first reselect device instruction indicates that the shooting angles of the two image acquisition devices are too large, and the video overlap of the same target person taken by different image acquisition devices is low. It is necessary to narrow the angle between the image acquisition devices to improve the effectiveness of data acquisition.

[0076] Furthermore, when the angle α is less than the preset threshold C1, the following judgment process is performed:

[0077] Create a first type of area in the shooting direction of the target person, and evenly create multiple reference points in the first type of area;

[0078] Use any two image acquisition devices to simultaneously shoot the same person target;

[0079] Count the number of reference points that can be captured by the two image acquisition devices, and record them as M1 and M2 respectively;

[0080] Count the total number of reference points M3 in the first type of area;

[0081] Calculate the overlap E of the two image acquisition devices capturing the target person. The specific calculation process is as follows:

[0082]

[0083] When the overlap amount E is less than the preset threshold value C2, a second device reselection instruction is generated.

[0084] It should be noted that the second reselect device instruction indicates that the shooting reduction of the two image acquisition devices is too small to make up for the shortcomings of the image shooting at this angle and further improve the efficiency of data collection.

[0085] Furthermore, the process of determining the effectiveness of the image acquisition device selection also includes:

[0086] According to the formula Calculate the distance D between multiple image acquisition devices and the target person respectively, remove the maximum and minimum values ​​of distance D to calculate the average distance

[0087] Using the formula method Calculate the deviation degree F of the image acquisition device;

[0088] When the degree of deviation F>the preset threshold C3, a clearing instruction is generated.

[0089] It should be noted that the clear command indicates that the image acquisition device is far away from the target person, the clarity is low, and the effectiveness is poor, that is, the data information of the device is cleared, and the collected data is further denoised, thereby improving the efficiency and accuracy of data analysis.

[0090] Among them, S2 specifically includes:

[0091] Input the image data into the image recognition model to obtain the first type of emotional information of the target person;

[0092] Input the audio data into the emotion recognition model to obtain the second type of emotional information of the target person;

[0093] Both the first and second types of emotional information include excitement and calmness;

[0094] Comparing the first type of emotional information with the second type of emotional information, and formulating a warning alert report when the first type of information and the second type of information are the same;

[0095] Input the video data into the video behavior recognition model to obtain the target person's behavior information, including danger and safety;

[0096] When the target person's behavior information is dangerous, a warning report will be issued.

[0097] It should be noted that the image recognition model is the facial expression recognition calculation of the convolutional neural network CNN; the emotion recognition model is the deep fusion technology based on factor decomposition bilinear pooling; and the video behavior recognition model is the SlowFast technology for video behavior recognition.

[0098] In a further embodiment, before S3, the method further includes detecting the type of items carried by the target person by using an X-ray detection device to determine the danger level. The specific detection process is as follows:

[0099] Use X-ray detection equipment to irradiate the target person and obtain data on the target items carried by the target person;

[0100] Import X-ray comparison item data into the database;

[0101] Extracting the first type of features from the target item and the second type of features from the comparison item;

[0102] Record the number of first-type features H1 and the number of second-type features H2 respectively;

[0103] Compare the first type of features with the second type of features to obtain a number H3 of common features between the first type of features and the second type of features;

[0104] Determine the similarity Q between two items. The specific judgment process is:

[0105] Q=2H3 / (H1+H2)

[0106] Classify the target item data based on the comparison item data whose similarity degree Q ≥ preset threshold N;

[0107] Determine whether the target item is dangerous based on the dangerousness of the comparative item data classified by the target item data, and generate a monitoring frequency report;

[0108] When the target object is dangerous, the target person is monitored in real time to increase the frequency of obtaining video and audio data of the target person.

[0109] By using X-ray detection equipment to detect the items carried by the target person, we can further determine the dangerousness of the target person and increase the monitoring frequency according to the dangerousness, thereby further improving the reliability of this method.

[0110] Furthermore, when the approximation degree Q is less than the preset threshold N, the following processing is performed:

[0111] First, the position of the X-ray detection equipment is changed multiple times to re-acquire the first type of features in the target object, and multiple approximation degrees Q are repeatedly calculated. Then, the target object data is classified according to the comparison object data with approximation degree Q ≥ preset threshold N; finally, the monitoring frequency report is verified based on the dangerousness of the target object.

[0112] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of the present invention, "plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0113] In the description of this specification, the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.

[0114] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.

Claims

1. A dangerous behavior identification and early warning method based on multimodal analysis, characterized in that: The following steps are involved: S1. Obtain video data and audio data of a target person in a preset area, and extract multiple frames of image data from the video data; S2. Process and analyze the audio data, image data, and video data to determine the risk of the target person; S3. Draw up a warning report based on the risk assessment and send the warning report to a preset terminal; The video data and audio data in S1 are collected by a preset device, which includes an image acquisition device with a sound receiving function. The number of preset devices is m, where m is a positive integer greater than 3. S1 specifically includes: Select image acquisition devices at different locations; Determine the effectiveness of the image acquisition device selection based on the location of the image acquisition device and the location of the image acquisition device and the target person; Synchronously obtain the image data and audio data of the target person; Then read and obtain multiple frames of image data frame by frame, and use image fusion technology to obtain the first type of image data at the same time; The specific process of judging the effectiveness of the image acquisition device selection is as follows: With the geometric center of the target person as the origin and the origin as the initial point, a spatial rectangular coordinate system is established, wherein the plane where the X-axis and the Y-axis are located is parallel to the horizontal plane, and the spatial rectangular coordinate system is established with the vertical direction as the preset dimension of the Z-axis as the unit; Randomly select two image acquisition devices and create two straight lines connecting the origin to the center of the image acquisition device; Read the coordinates A(x1,y1,z1) and B(x2,y2,z2) of the geometric centers of the two image acquisition devices respectively And calculate the angle α between the two straight lines. The specific calculation process is: Then use trigonometric functions to calculate the angle α; When the angle α is greater than or equal to the preset threshold C, a first reselection device instruction is generated; When the angle α is less than the preset threshold C1, the following judgment process is also performed: Create a first type of area in the shooting direction of the target person, and evenly create multiple reference points in the first type of area; Use any two image acquisition devices to simultaneously shoot the same person target; Count the number of reference points that can be captured by the two image acquisition devices, and record them as M1 and M2 respectively; Count the total number of reference points M3 in the first type of area; Calculate the overlap E of the two image acquisition devices capturing the target person. The specific calculation process is as follows: When the overlap amount E is less than the preset threshold value C2, a second reselection device instruction is generated; The process of judging the validity of the image acquisition device selection further includes: According to the formula Calculate the distance D between multiple image acquisition devices and the target person respectively, remove the maximum and minimum values ​​of distance D to calculate the average distance Using the formula method Calculate the deviation degree F of the image acquisition device; When the degree of deviation F>the preset threshold C3, a clearing instruction is generated.

2. The method for identifying and warning dangerous behaviors based on multimodal analysis according to claim 1, characterized in that: The S2 specifically includes: Input the image data into the image recognition model to obtain the first type of emotional information of the target person; Input the audio data into the emotion recognition model to obtain the second type of emotional information of the target person; The first type of emotional information and the second type of emotional information both include excitement and calmness; Comparing the first type of emotional information with the second type of emotional information, and formulating a warning alert report when the first type of information and the second type of information are the same; Inputting video data into a video behavior recognition model to obtain behavior information of the target person, the behavior information including danger and safety; When the target person's behavior information is dangerous, a warning report will be issued.

3. The method for identifying and warning dangerous behaviors based on multimodal analysis according to claim 2, characterized in that: The above-mentioned step S3 also includes detecting the type of items carried by the target person through X-ray detection equipment to determine the danger. The specific detection process is as follows: Use X-ray detection equipment to irradiate the target person and obtain data on the target items carried by the target person; Import X-ray comparison item data into the database; Extracting the first type of features from the target item and the second type of features from the comparison item; Record the number of first-type features H1 and the number of second-type features H2 respectively; Compare the first type of features with the second type of features to obtain a number H3 of common features between the first type of features and the second type of features; Determine the similarity Q between two items. The specific judgment process is: Q=2H3 / (H1+H2) Classify the target item data based on the comparison item data whose similarity degree Q ≥ preset threshold N; Determine whether the target item is dangerous based on the dangerousness of the comparative item data classified by the target item data, and generate a monitoring frequency report; When the target object is dangerous, the target person is monitored in real time to increase the frequency of obtaining video and audio data of the target person.

4. The method for identifying and warning dangerous behaviors based on multimodal analysis according to claim 3 is characterized by: When the approximation degree Q is less than the preset threshold N, the following processing is also performed: First, the position of the X-ray detection equipment is changed multiple times to re-acquire the first type of features in the target object, and multiple approximation degrees Q are repeatedly calculated. Then, the target object data is classified according to the comparison object data with approximation degree Q ≥ preset threshold N; finally, the monitoring frequency report is verified based on the dangerousness of the target object.

Citation Information

Patent Citations

  • Intelligent recognition shooting method and system for multi-person scene and storage medium

    CN110572570A

  • Employee abnormal behavior detection method and device, and system

    CN112883932A

  • Emergency alarm system for school

    CN114566023A