Video monitoring event early warning method and system based on multi-modal data fusion

Through multimodal data fusion and event prediction model, the problems of inaccurate and untimely early warning in the video surveillance system are solved, and accurate early warning and resource optimization for different event types are achieved.

CN120375288AActive Publication Date: 2025-07-25WUHAN XINKAILI COMM ENG CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510493660.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-25
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

The existing video surveillance system has inaccurate and timely early warnings, lacks multimodal data fusion and active alarm information integration, and the warning strategy lacks targeted and flexible, and it is impossible to accurately determine the location and type of incidents.

Method used

By integrating video surveillance data, audio data, infrared data and active alarm information, combined with event prediction models, we determine the location and type of expected events, and generate targeted early warning strategies based on different event types.

Benefits of technology

It improves the accuracy and timeliness of early warnings, reduces false alarms and missed reports, improves resource utilization efficiency, and reduces monitoring costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375288A_ABST
    Figure CN120375288A_ABST
Patent Text Reader

Abstract

The invention provides a video monitoring event early warning method and system based on multi-modal data fusion, and relates to the field of city monitoring, and the method comprises the steps: determining an expected event occurrence place and an expected event type according to event state information corresponding to a current time period, the expected event type comprises a tumble event, a fight event, a theft event, a fire event and a theft event; determining actual monitoring equipment according to the expected event occurrence place and the expected event type, and obtaining real-time monitoring data of the next time period; when the expected event type is a fire event or a theft event, determining a first expected event occurrence probability according to the video monitoring data and the real-time monitoring data; and when the expected event type is a tumble event, a fight event or a stealing event, inputting real-time monitoring data to a preset event prediction model, and determining a second expected event occurrence probability. According to the invention, by fusing the multi-modal data and integrating the active alarm information, the accuracy of early warning is obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of urban surveillance, and in particular, to a video surveillance event early warning method and system based on multi-modal data fusion. Background Art

[0002] In the current video surveillance system, although it is possible to give a certain degree of early warning of security events through video surveillance data, there are still many problems in the accuracy and timeliness of the early warning. On the one hand, existing early warning systems often rely only on single video surveillance data and lack the fusion and utilization of other modal data, resulting in inaccurate early warning results. On the other hand, for different types of security events, existing early warning systems often adopt the same early warning strategy, lacking pertinence and flexibility, and it is difficult to meet the requirements in practical applications. In addition, existing video surveillance early warning systems also have the problem of insufficiently intelligent processing of active alarm information. When a user actively alarms, the system can often only simply record the alarm information, and cannot quickly and accurately determine the expected event occurrence location and expected event type based on the alarm information, thus unable to activate the early warning mechanism in a timely and effective manner.

[0003] Chinese Patent No. CN112257546A discloses an event early warning method, device, electronic device and storage medium. The event early warning method includes: obtaining a surveillance video captured by a surveillance device and identifying a current event corresponding to the surveillance video; querying a predicted event associated with the current event by using an event knowledge graph; wherein, the event knowledge graph stores the association relationship between events; if the predicted event is a target type event, then generating an early warning information corresponding to the predicted event, which can give an early warning before a malicious event occurs. However, it has disadvantages such as a single data source, lack of integration of active alarm information, lack of pertinence and flexibility in the early warning strategy, inability to update and adjust the early warning model in real time, and inaccurate prediction of the probability of event occurrence. At present, there is no technical solution that can solve the above technical problems, and there is no video surveillance event early warning method and system based on multi-modal data fusion. Summary of the Invention

[0004] The present invention provides a video surveillance event early warning method and system based on multi-modal data fusion, which solves the problems of inaccurate and untimely early warning in the existing video surveillance system, and the lack of pertinence and flexibility in the early warning strategy for different types of events.

[0005] In a first aspect, the present invention provides a video surveillance event early warning method based on multi-modal data fusion, including:

[0006] Determine the expected event occurrence location and the expected event type according to the event status information corresponding to the current time period. The event status information includes the active alarm information from any user and the video surveillance data determined by any passive monitoring device associated with the active alarm information. The expected event types include fall events, fighting events, pickpocketing events, fire events, and theft events;

[0007] Determine the actual monitoring device according to the expected event occurrence location and the expected event type, and obtain the real-time monitoring data for the next time period from the actual monitoring device;

[0008] In the case where the expected event type is a fire event or a theft event, determine the first expected event occurrence probability according to the video surveillance data and the real-time monitoring data, and generate a first target warning strategy according to the first expected event occurrence probability;

[0009] In the case where the expected event type is a fall event, a fighting event, or a pickpocketing event, input the real-time monitoring data into a preset event prediction model, determine the second expected event occurrence probability output by the preset event prediction model, and generate a second target warning strategy according to the second expected event occurrence probability.

[0010] According to the video surveillance event warning method based on multi-modal data fusion provided by the present invention, the determining the expected event occurrence location and the expected event type according to the event status information corresponding to the current time period includes:

[0011] Perform speech recognition on the active alarm information, obtain the alarm event location and the alarm event type, call all the passive monitoring devices within the area where the alarm event location is located, and obtain the video surveillance data corresponding to each passive monitoring device in the current time period. Different passive monitoring devices correspond to different event occurrence locations, and the passive monitoring device is a video surveillance device;

[0012] For any video surveillance data, obtain the current time period target analysis result corresponding to the video surveillance data. The current time period target analysis result includes the initial occurrence probabilities of different events. In the case where the initial occurrence probability of the event corresponding to any alarm event type is greater than the preset occurrence probability, determine the alarm event type as the expected event type;

[0013] Determine all the passive monitoring devices whose initial occurrence probabilities of the events corresponding to the alarm event type are greater than the preset occurrence probability as candidate monitoring devices, sort the initial occurrence probabilities corresponding to all the candidate monitoring devices in descending order, and determine the event occurrence location corresponding to the candidate monitoring device with the largest initial occurrence probability as the expected event occurrence location.

[0014] According to the video surveillance event early warning method based on multi-modal data fusion provided by the present invention, before obtaining the current period target analysis result corresponding to the video surveillance data, the method further includes:

[0015] For any passive monitoring device, obtain the historical monitoring data corresponding to the passive monitoring device in different historical periods. For the historical monitoring data of any historical period, input the historical monitoring data into a preset event classification model to obtain the historical period event analysis result corresponding to the historical period output by the preset event classification model;

[0016] Traverse all historical periods, obtain each historical period event analysis result corresponding to all historical periods. For any event type, determine the current period target analysis result according to each historical period event analysis result and the current period event analysis result;

[0017] The current period event analysis result is determined after inputting the current monitoring data of the current period into the preset event classification model; the preset event classification model is determined after being trained according to all sample monitoring data and the sample classification label corresponding to each sample monitoring data, and the sample classification label includes the occurrence probability of each event type corresponding to the sample monitoring data.

[0018] According to the video surveillance event early warning method based on multi-modal data fusion provided by the present invention, determining the current period target analysis result according to each historical period event analysis result and the current period event analysis result includes:

[0019] Sequentially determine the three periods before the current period as the first target period, the second target period, and the third target period, and determine the first analysis result corresponding to the first target period, the second analysis result corresponding to the second target period, and the third analysis result corresponding to the third target period from the historical period event analysis results;

[0020] For any event type, determine the target event result corresponding to the event type according to the first event result, the second event result, the third event result, and the current event result. Traverse all event types to determine the current period target analysis result;

[0021] Wherein, the first event result is the occurrence probability of the event corresponding to the event type in the first analysis result, the second event result is the occurrence probability of the event corresponding to the event type in the second analysis result, the third event result is the occurrence probability of the event corresponding to the event type in the third analysis result, and the current event result is the occurrence probability of the event corresponding to the event type in the current period event analysis result.

[0022] According to the video surveillance event early warning method based on multi-modal data fusion provided by the present invention, determining the target event result corresponding to the event type according to the first event result, the second event result, the third event result, and the current event result includes:

[0023] T = 0.1×R1 + 0.2×R2 + 0.3×R3 + 0.4×R4;

[0024] Wherein, T is the target event result, R1 is the first event result, R2 is the second event result, R3 is the third event result, and R4 is the current event result.

[0025] According to the video surveillance event early warning method based on multi-modal data fusion provided by the present invention, determining the actual monitoring device according to the expected event occurrence location and the expected event type, and obtaining the real-time monitoring data of the next time period from the actual monitoring device includes:

[0026] In the case where the expected event type is a fire event, determining the video surveillance device and the infrared sensing device associated with the expected event occurrence location as the actual monitoring device, and obtaining the video data and infrared data of the next time period from the actual monitoring device;

[0027] In the case where the expected event types are fall events, fight events, and pickpocketing events, determining the video surveillance device and the audio surveillance device associated with the expected event occurrence location as the actual monitoring device, and obtaining the video data and audio data of the next time period from the actual monitoring device;

[0028] In the case where the expected event type is a theft event, determining the video surveillance device associated with the expected event occurrence location as the actual monitoring device, and obtaining the video data of the next time period from the actual monitoring device.

[0029] According to the video surveillance event early warning method based on multi-modal data fusion provided by the present invention, determining the first expected event occurrence probability according to the video surveillance data and the real-time monitoring data, and generating the first target early warning strategy according to the first expected event occurrence probability, includes:

[0030] In the case where the expected event type is a fire event, determine the average brightness value and the average red saturation value of the video frames based on the video data, determine the highest temperature value based on the infrared data, and determine the first fire event occurrence probability based on the average brightness value of the video frames, the preset brightness weight, the average red saturation value, the preset saturation weight, the highest temperature value, and the preset temperature weight; determine the fire event occurrence probability based on the first fire event occurrence probability, the first fire weight, the second fire event occurrence probability corresponding to the video surveillance data, and the second fire weight. In the case where it is determined that the fire event occurrence probability is greater than the preset fire probability, generate a fire indication instruction, and the fire indication instruction is used to indicate going to the location where the expected event occurs to extinguish the fire;

[0031] In the case where the expected event type is a theft event, determine all the first items in the current time period based on the video surveillance data, determine all the second items in the next time period based on the real-time surveillance data, compare the first items and the second items, determine the missing items, and in the case where the missing items are on the preset surveillance list, determine that a theft event has occurred, capture the head portrait of the target thief using the video surveillance data and the real-time surveillance data, and send the head portrait of the target thief to a third-party early warning platform.

[0032] According to the video surveillance event early warning method based on multi-modal data fusion provided by the present invention, the preset event prediction model includes a fall event prediction model, a fight event prediction model, and a pickpocketing event prediction model;

[0033] In the case where the expected event type is a fall event, input the real-time surveillance data into the fall event prediction model, and determine the fall event occurrence probability output by the fall event prediction model;

[0034] In the case where the expected event type is a fight event, input the real-time surveillance data into the fight event prediction model, and determine the fight event occurrence probability output by the fight event prediction model;

[0035] In the case where the expected event type is a pickpocketing event, input the real-time surveillance data into the pickpocketing event prediction model, and determine the pickpocketing event occurrence probability output by the pickpocketing event prediction model.

[0036] According to the video surveillance event early warning method based on multi-modal data fusion provided by the present invention, generating a second target early warning strategy according to the second expected event occurrence probability includes:

[0037] In the case where the second expected event occurrence probability is less than or equal to the first preset occurrence probability, no early warning is issued;

[0038] When the probability of occurrence of the second expected event is greater than the first preset probability of occurrence and less than or equal to the second preset probability of occurrence, the video surveillance data and the real-time surveillance data are merged into suspected event surveillance data, and the suspected event surveillance data is sent to a third-party manual review platform to implement manual review of the suspected event surveillance data;

[0039] When the probability of occurrence of the second expected event is greater than the second preset probability of occurrence, the video surveillance data and the real-time surveillance data are used to capture the head portrait of the target event person, and the head portrait of the target event person is sent to a third-party early warning platform.

[0040] Second, a video surveillance event early warning system based on multi-modal data fusion is provided, including:

[0041] A determination unit, which is used to determine the expected event occurrence location and the expected event type according to the event status information corresponding to the current time period. The event status information includes the active alarm information from any user and the video surveillance data determined by any passive surveillance device associated with the active alarm information. The expected event types include fall events, fight events, pickpocketing events, fire events, and theft events;

[0042] An acquisition unit, which is used to determine the actual surveillance device according to the expected event occurrence location and the expected event type, and acquire the real-time surveillance data of the next time period from the actual surveillance device;

[0043] A first generation unit, which is used, when the expected event type is a fire event or a theft event, to determine the first expected event occurrence probability according to the video surveillance data and the real-time surveillance data, and generate a first target early warning strategy according to the first expected event occurrence probability;

[0044] A second generation unit, which is used, when the expected event type is a fall event, a fight event, or a pickpocketing event, to input the real-time surveillance data into a preset event prediction model, determine the second expected event occurrence probability output by the preset event prediction model, and generate a second target early warning strategy according to the second expected event occurrence probability.

[0045] The present invention comprehensively obtains various types of information related to events by fusing multi-modal data (including video surveillance data, audio data, and infrared data) and integrating users' active alarm information, thereby more accurately predicting potential security events, effectively avoiding the problem of inaccurate early warnings that may be caused by a single data source, and significantly improving the accuracy of early warnings; it can obtain and analyze surveillance data in real time, quickly generate corresponding early warning strategies according to different expected event types, ensure that early warnings can be issued in a timely manner at the initial stage of the event, and significantly improve the timeliness of early warnings;

[0046] For different types of expected events, the present invention adopts different early warning strategies. By using differentiated early warning strategy design, the early warning measures are more in line with the actual event requirements, improving the pertinence and effectiveness of early warnings. By introducing technical means such as speech recognition and machine learning, the automatic processing of active alarm information and the intelligent analysis of real-time surveillance data are realized, which not only reduces the burden of manual monitoring, but also improves the efficiency and accuracy of the system in processing massive data; by accurately predicting the probability and location of event occurrence, reasonably allocating surveillance resources and emergency forces, avoiding waste and duplicate investment of resources, which helps to improve the resource utilization efficiency and reduce the surveillance cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0048] Figure 1 It is a schematic flowchart of a video surveillance event early warning method based on multi-modal data fusion provided by the present invention;

[0049] Figure 2 It is a schematic structural diagram of a video surveillance event early warning system based on multi-modal data fusion provided by the present invention;

[0050] Figure 3 It is a schematic structural diagram of an electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention fall within the protection scope of the present invention.

[0052] To solve the problems of inaccurate and untimely early warnings and the lack of pertinence and flexibility in early warning strategies in existing video surveillance systems, the present invention proposes a method for early warning of urban video surveillance events based on multi-modal data fusion. By combining various types of data, including video surveillance data, audio data, and infrared data, and combining user-initiated alarm information and real-time data from passive monitoring devices, it can more accurately predict and early warn various types of security events, improving the accuracy and timeliness of early warnings. At the same time, it can also adopt different early warning strategies according to different types of security events, improving the pertinence and flexibility of early warnings. Figure 1 It is a schematic flowchart of the method for early warning of video surveillance events based on multi-modal data fusion provided by the present invention. The method for early warning of video surveillance events based on multi-modal data fusion includes:

[0053] Step 101: Determine the expected event occurrence location and the expected event type according to the event status information corresponding to the current time period. The event status information includes the user-initiated alarm information from any user and the video surveillance data determined by any passive monitoring device associated with the user-initiated alarm information. The expected event types include fall events, fighting events, pickpocketing events, fire events, and theft events.

[0054] Step 102: Determine the actual monitoring devices according to the expected event occurrence location and the expected event type, and obtain the real-time monitoring data for the next time period from the actual monitoring devices.

[0055] Step 103: In the case where the expected event type is a fire event or a theft event, determine the first expected event occurrence probability according to the video surveillance data and the real-time monitoring data, and generate a first target early warning strategy according to the first expected event occurrence probability.

[0056] Step 104: In the case where the expected event type is a fall event, a fighting event, or a pickpocketing event, input the real-time monitoring data into a preset event prediction model, determine the second expected event occurrence probability output by the preset event prediction model, and generate a second target early warning strategy according to the second expected event occurrence probability.

[0057] In step 101, step 101 mainly includes determining the occurrence location and type of the expected event based on the event status information of the current period, which involves the processing of active alarm information and the integration of passive monitoring device data. First, the present invention can convert the user's active alarm information into text or structured data through speech recognition technology, extract key elements in the alarm information, such as the approximate event location and event type, and then call the associated passive monitoring devices according to the approximate event location, such as corresponding cameras, to obtain the video monitoring data of these devices in the current period. According to the active alarm information and the passive monitoring device data, combined with preset rules or algorithms, determine the occurrence location and type of the expected event. Here, the occurrence location of the expected event is a more specific event occurrence location compared to the approximate event location, and the expected event type needs to determine whether the approximate event type is true according to the event status information corresponding to the current period. If it is true, it is determined as the expected event type.

[0058] In an alternative embodiment, assume that in a certain shopping mall, the user reports a possible pickpocketing incident through an alarm call and provides an approximate location. The system first converts the alarm information into text through speech recognition technology, and extracts that the event location is "the clothing area on the second floor of a certain shopping mall" and the event type is "pickpocketing". Then, the system calls all the cameras in the clothing area on the second floor of the shopping mall to obtain the video monitoring data of the current period. Through the analysis of the video data, the system finds that the personnel flow in this area is dense and there are characteristics of pickpocketing behavior. Combining the alarm information and the video data, the system finally determines that the expected event is a pickpocketing incident and confirms the location where the relevant camera where the pickpocketing behavior occurred is located, and determines it as the occurrence location of the expected event.

[0059] Optionally, determining the occurrence location of the expected event and the expected event type according to the event status information corresponding to the current period includes:

[0060] Performing speech recognition on the active alarm information to obtain the alarm event location and the alarm event type, calling all the passive monitoring devices in the area where the alarm event location is located, and obtaining the video monitoring data corresponding to each passive monitoring device in the current period, where different passive monitoring devices correspond to different event occurrence locations, and the passive monitoring devices are video monitoring devices;

[0061] For any video monitoring data, obtaining the current period target analysis result corresponding to the video monitoring data, where the current period target analysis result includes the initial occurrence probabilities of different events. In the case where the initial occurrence probability of the event corresponding to any alarm event type is greater than the preset occurrence probability, determining the alarm event type as the expected event type;

[0062] All passive monitoring devices with the initial occurrence probability of the event corresponding to the alarm event type greater than the preset occurrence probability are determined as candidate monitoring devices. The initial occurrence probabilities corresponding to all candidate monitoring devices are sorted in descending order, and the event occurrence location corresponding to the candidate monitoring device with the largest initial occurrence probability is determined as the expected event occurrence location.

[0063] Optionally, the present invention uses speech recognition technology to convert the user's active alarm information, such as telephone alarm or voice alarm, etc., into processable text or structured data, extracts the alarm event location and alarm event type from the converted data, and then calls all passive monitoring devices within the area where the location is located according to the extracted alarm event location. In the present invention, mainly video monitoring devices are involved, and each video monitoring data corresponding to all passive monitoring devices in the current period is obtained. In the present invention, all video monitoring data perform preliminary event behavior judgments all the time in different historical periods. The monitoring means are relatively simple, and the determined event types also have a large misjudgment rate. That is, the present invention first conducts a preliminary analysis on the video monitoring data provided by each passive monitoring device to obtain the target analysis result of the current period, and then directly obtains the corresponding target analysis result of the current period after determining that all passive monitoring devices need to be retrieved. The present invention determines the initial occurrence probabilities of different events based on the target analysis result of the current period. These probabilities are calculated based on the features in the monitoring screen (such as human behavior, environmental changes, etc.), or can be preliminarily determined according to model recognition.

[0064] Furthermore, compare the alarm event type with the initial occurrence probability of the event in the target analysis result of the current period. If the initial occurrence probability of the event corresponding to any alarm event type is greater than the preset occurrence probability threshold, then determine the alarm event type as the expected event type, and determine all passive monitoring devices corresponding to the alarm event type and with the initial occurrence probability greater than the preset threshold as candidate monitoring devices. That is, this embodiment is to find out the monitoring devices with the highest corresponding event occurrence probability and the highest event relevance in all the video monitoring data recorded by the current event behavior among all passive monitoring devices, so as to improve the accuracy and reliability of the data source for subsequent event early warning. Specifically, sort the candidate monitoring devices in descending order according to the initial occurrence probability, and select the event occurrence location corresponding to the candidate monitoring device with the largest initial occurrence probability as the expected event occurrence location, thus completing the determination of the candidate monitoring devices, as well as the determination of the expected event type and the expected event occurrence location.

[0065] By integrating the active alarm information and the data of passive monitoring devices, the present invention can comprehensively understand the event status, reduce false alarms and missed alarms. By quickly calling and analyzing the data of monitoring devices, it can quickly determine the expected event type and occurrence location, sort and select the monitoring device with the highest initial occurrence probability as the expected event occurrence location, improve the accuracy and efficiency of positioning, and only deeply analyze the monitoring devices related to the alarm event and with a relatively high initial occurrence probability, reducing unnecessary calculations and resource consumption.

[0066] Optionally, before obtaining the current period target analysis result corresponding to the video monitoring data, the method further includes:

[0067] For any passive monitoring device, obtain the historical monitoring data corresponding to the passive monitoring device in different historical periods. For the historical monitoring data of any historical period, input the historical monitoring data into a preset event classification model to obtain the historical period event analysis result corresponding to the historical period output by the preset event classification model;

[0068] Traverse all historical periods, obtain each historical period event analysis result corresponding to all historical periods. For any event type, determine the current period target analysis result according to each historical period event analysis result and the current period event analysis result;

[0069] The current period event analysis result is determined after inputting the current monitoring data of the current period into the preset event classification model; the preset event classification model is determined after training according to all sample monitoring data and the sample classification label corresponding to each sample monitoring data, and the sample classification label includes the occurrence probability of each event type corresponding to the sample monitoring data.

[0070] Optionally, before obtaining the current period target analysis result corresponding to the video monitoring data, the present invention constantly performs a preliminary analysis of historical monitoring data all the time. As an optional embodiment of the present invention, it can also be to obtain historical data in real time and analyze the historical video data to obtain the analysis result of the historical period. Specifically, for any passive monitoring device, collect its historical monitoring data in different historical periods. These historical data may include video monitoring records in the past few minutes. Input the historical monitoring data of each collected historical period into a pre-trained event classification model. The event classification model analyzes these historical monitoring data and outputs the historical period event analysis result corresponding to each historical period. Traverse all historical periods, collect and integrate the event analysis results of each historical period. These results constitute a comprehensive analysis of historical monitoring data and provide an important reference for determining the current period target analysis result.

[0071] Meanwhile, the present invention also inputs the current monitoring data of the current time period into a preset event classification model to obtain the event analysis result of the current time period. By combining the event analysis results of the historical time period and the current time period, the target analysis result of the current time period is comprehensively determined. This target analysis result includes the initial occurrence probabilities of different events and is the key basis for subsequent judgment of the expected event type and location. Specifically, the preset event classification model is trained based on all sample monitoring data and corresponding sample classification labels. The sample classification labels contain the occurrence probability of each event type corresponding to the sample monitoring data, providing rich training data for the model. In such an embodiment, the present invention can collect and store the historical monitoring data of passive monitoring devices using a database or a big data storage platform, construct an event classification model using machine learning or deep learning techniques, train the model with a large amount of sample monitoring data and corresponding sample classification labels so that it can accurately identify the event type in the video and predict the occurrence probability of the event, and then use statistical methods (such as weighted average, moving average, etc.) to combine historical data and current data to determine a more accurate target analysis result for the current time period. By combining the event analysis results of the historical time period and the current time period, the present invention can more comprehensively understand the laws and trends of event occurrence, which helps to more accurately predict the event type and its occurrence probability that may occur in the current time period and improve the accuracy of early warning.

[0072] Optionally, determining the target analysis result of the current time period according to the event analysis results of each historical time period and the event analysis result of the current time period includes:

[0073] Sequentially determining the three time periods before the current time period as the first target time period, the second target time period, and the third target time period, and determining the first analysis result corresponding to the first target time period, the second analysis result corresponding to the second target time period, and the third analysis result corresponding to the third target time period from the event analysis results of the historical time period;

[0074] For any event type, determine the target event result corresponding to the event type according to the first event result, the second event result, the third event result, and the current event result, and traverse all event types to determine the target analysis result of the current time period;

[0075] Wherein, the first event result is the occurrence probability of the event corresponding to the event type in the first analysis result, the second event result is the occurrence probability of the event corresponding to the event type in the second analysis result, the third event result is the occurrence probability of the event corresponding to the event type in the third analysis result, and the current event result is the occurrence probability of the event corresponding to the event type in the event analysis result of the current time period.

[0076] In an optional embodiment, the three time periods before the current time period are sequentially determined as the first target time period, the second target time period, and the third target time period. In other embodiments, more target time periods can also be set, such as the first target time period, the second target time period, the third target time period, and the fourth target time period, the first target time period, the second target time period, the third target time period, the fourth target time period, and the fifth target time period, etc. The selection of these time periods is based on temporal continuity to better capture the trends and patterns of event occurrences. From the historical time period event analysis results, the first analysis result corresponding to the first target time period, the second analysis result corresponding to the second target time period, and the third analysis result corresponding to the third target time period are respectively extracted. These analysis results include the occurrence probabilities of different event types within each time period. For any event type in the present invention, the following operations are performed: Based on the event occurrence probability in the first analysis result (the first event result), the event occurrence probability in the second analysis result (the second event result), the event occurrence probability in the third analysis result (the third event result), and the event occurrence probability in the current time period event analysis result (the current event result), the target event result corresponding to this event type is comprehensively determined. The target event result represents a comprehensive occurrence probability of this event type in the current time period, considering the combined effects of historical data and current data. Repeat the above steps for all possible event types, calculate the target event result corresponding to each event type. Finally, the target event results constitute the target analysis result for the current time period, providing an important basis for subsequent early warnings and decision-making.

[0077] Optionally, determining the target event result corresponding to the event type according to the first event result, the second event result, the third event result, and the current event result includes:

[0078] T = 0.1×R1 + 0.2×R2 + 0.3×R3 + 0.4×R4;

[0079] Wherein, T is the target event result, R1 is the first event result, R2 is the second event result, R3 is the third event result, and R4 is the current event result.

[0080] Optionally, the present invention adopts a weighted average algorithm to calculate the target event result by combining the event occurrence probabilities of historical time periods and the event occurrence probabilities of the current time period. By comprehensively considering the event occurrence probabilities of historical time periods and the current time period, it is possible to more accurately predict the event types that may occur in the current time period and their occurrence probabilities, which helps to reduce false alarms and missed alarms and improve the accuracy and reliability of early warnings.

[0081] In step 102, determining the actual monitoring device according to the expected event occurrence location and the expected event type, and obtaining the real-time monitoring data for the next time period from the actual monitoring device includes:

[0082] In the case where the expected event type is a fire event, determine the video surveillance device and the infrared sensing device associated with the location where the expected event occurs as the actual monitoring devices, and obtain the video data and infrared data for the next time period from the actual monitoring devices;

[0083] In the case where the expected event types are fall events, fight events, and pickpocketing events, determine the video surveillance device and the audio surveillance device associated with the location where the expected event occurs as the actual monitoring devices, and obtain the video data and audio data for the next time period from the actual monitoring devices;

[0084] In the case where the expected event type is a theft event, determine the video surveillance device associated with the location where the expected event occurs as the actual monitoring device, and obtain the video data for the next time period from the actual monitoring device.

[0085] Optionally, when the expected event type is a fire event, determine the video surveillance device and the infrared sensing device associated with the location where the expected event occurs as the actual monitoring devices, and obtain the video data and infrared data for the next time period from these actual monitoring devices. The video data is used to observe the spread of the fire, the evacuation of people, etc., and the infrared data is used to detect the fire source, high-temperature areas, etc.; when the expected event types are fall events, fight events, or pickpocketing events, determine the video surveillance device and the audio surveillance device associated with the location where the expected event occurs as the actual monitoring devices, and obtain the video data and audio data for the next time period from these actual monitoring devices. The video data is used to observe the specific process of the event, the behavior of people, etc., and the audio data is used to capture the on-site sounds, conversations, etc., providing more clues for event analysis; when the expected event type is a theft event, determine the video surveillance device associated with the location where the expected event occurs as the actual monitoring device, and obtain the video data for the next time period from these actual monitoring devices. The video data is used to observe the occurrence of the theft behavior, the characteristics of the target person, etc.

[0086] The present invention utilizes a Geographic Information System (GIS) or a monitoring device management system to associate the location where an expected event occurs with the associated monitoring devices. By means of device identification, location information, etc., it accurately identifies and controls the actual monitoring devices. It uses the API interface or data transmission protocol of video monitoring devices to obtain real-time monitoring data for the next time period from the actual monitoring devices. For infrared sensing devices and audio monitoring devices, corresponding data acquisition and transmission technologies are also adopted to ensure the real-time and accuracy of the data. The present invention selectively chooses actual monitoring devices according to the type of expected event, avoiding blind monitoring and resource waste. The obtained video data, infrared data, and audio data provide a rich source of information for event analysis. The acquisition and analysis of real-time monitoring data help to timely detect potential safety hazards and abnormal behaviors, providing strong support for taking preventive measures.

[0087] If the type of the expected event is a fire event or a theft event, then step 103 is executed. If the type of the expected event is a fall event, a fight event, or a pickpocketing event, then step 104 is executed.

[0088] In step 103, the determining the first probability of the occurrence of the expected event according to the video monitoring data and the real-time monitoring data, and generating a first target warning strategy according to the first probability of the occurrence of the expected event, includes:

[0089] In the case where the type of the expected event is a fire event, determine the average brightness value and the average red saturation value of the video frames according to the video data, determine the highest temperature value according to the infrared data, and determine the first probability of the occurrence of the fire event according to the average brightness value of the video frames, the preset brightness weight, the average red saturation value, the preset saturation weight, the highest temperature value, and the preset temperature weight; determine the probability of the occurrence of the fire event according to the first probability of the occurrence of the fire event, the first fire weight, the second probability of the occurrence of the fire event corresponding to the video monitoring data, and the second fire weight. In the case where it is determined that the probability of the occurrence of the fire event is greater than the preset fire probability, generate a fire indication instruction, and the fire indication instruction is used to indicate going to the location where the expected event occurs to extinguish the fire;

[0090] In the case where the type of the expected event is a theft event, determine all the first items in the current time period according to the video monitoring data, determine all the second items in the next time period according to the real-time monitoring data, compare the first items and the second items to determine the lost items. In the case where the lost items are in the preset monitoring list, determine that a theft event has occurred, capture the head portrait of the target thief using the video monitoring data and the real-time monitoring data, and send the head portrait of the target thief to a third-party warning platform.

[0091] Optionally, the present invention can calculate the average brightness value and the average red saturation value of video frames using video processing algorithms (such as OpenCV), extract the highest temperature value using infrared data processing technology, and combine the average brightness value, the preset brightness weight, the average red saturation value, the preset saturation weight, the highest temperature value, and the preset temperature weight to calculate the probability of the first fire event occurring. Further, by combining the probability of the first fire event occurring, the first fire weight, the probability of the second fire event occurring corresponding to the video surveillance data, and the second fire weight, the final probability of the fire event occurring is determined:

[0092] P final = W1×P1 + W2×P2

[0093] Wherein, P final is the probability of the fire event occurring, W1 is the first fire weight, W2 is the second fire weight, P1 is the probability of the first fire event occurring, and P2 is the probability of the second fire event occurring corresponding to the video surveillance data. Among them, the calculation formula for the probability P1 of the first fire event occurring is:

[0094] P1 = W L ×L avg + W S ×S avg + W T ×T max

[0095] Wherein, W L is the preset brightness weight, L avg is the average brightness value of the video frame, S avg is the average red saturation value, W S is the preset saturation weight, T max is the highest temperature value, and W T is the preset temperature weight.

[0096] In an optional embodiment, it is assumed that it is necessary to analyze the video surveillance data and the real-time surveillance data detected by a certain actual surveillance device in a certain shopping mall. The system calculates the probability of the first fire event occurring according to the preset weights and algorithms, and combines historical data to calculate the final probability of the fire event occurring. If the probability of the fire event occurring is greater than the preset fire preset probability, the system generates a fire indication instruction to notify the firefighters to go to the area to extinguish the fire.

[0097] Optionally, in the case where the expected event type is a theft event, the present invention uses an item recognition algorithm (such as a deep learning model) to compare and identify items in the current time period and the next time period, and preset a preset monitoring list for quickly determining whether the lost item is an important or sensitive item. Determine all first items in the current time period based on the video surveillance data, determine all second items in the next time period based on the real-time surveillance data, compare the first items and the second items to determine the lost item. If the lost item is in the preset monitoring list, it is determined that a theft event has occurred. Combine the video surveillance data and the real-time surveillance data with face recognition technology to capture the head portrait of the target thief, and send the head portrait of the target thief to a third-party warning platform through a preset API interface or message queue.

[0098] In an optional embodiment, assume that the monitoring system of a jewelry store detects that there are several pieces of jewelry on the counter in the current time period, but the real-time surveillance data in the next time period shows that some of the jewelry has disappeared. The system compares the items in the current time period and the next time period, determines that the lost jewelry belongs to the items in the preset monitoring list, captures the head portrait of the target thief, and sends it to a third-party warning platform for law enforcement officers to intervene in a timely manner.

[0099] By combining video data, infrared data, and historical data, the present invention can more accurately calculate the probability of a fire event occurring and identify theft events. When a fire event occurs, it can quickly generate a fire indication instruction to notify firefighters to go to extinguish the fire, reducing the losses caused by the fire; when a theft event occurs, it can quickly capture the head portrait of the target thief and send it to a third-party warning platform, which helps law enforcement officers to intervene and track in a timely manner; through real-time analysis and processing of surveillance data, it can timely discover potential safety hazards and abnormal behaviors, providing strong support for taking preventive measures, and the use of the preset monitoring list helps to quickly identify the loss of important or sensitive items, improving the pertinence and effectiveness of security prevention.

[0100] In step 104, the preset event prediction model includes a fall event prediction model, a fight event prediction model, and a pickpocketing event prediction model;

[0101] In the case where the expected event type is a fall event, input the real-time surveillance data into the fall event prediction model to determine the probability of a fall event output by the fall event prediction model;

[0102] In the case where the expected event type is a fight event, input the real-time surveillance data into the fight event prediction model to determine the probability of a fight event output by the fight event prediction model;

[0103] When the expected event type is a pickpocketing event, input the real-time monitoring data into the pickpocketing event prediction model to determine the probability of a pickpocketing event occurring output by the pickpocketing event prediction model.

[0104] Optionally, the present invention uses video processing technology (such as OpenCV) to preprocess the real-time monitoring video, extract key frames and features, process the audio data, extract sound features (such as volume, frequency, etc.), input the processed real-time monitoring data into the corresponding prediction model, and calculate and output the probability of an event occurring according to the input data: When the expected event type is a fall event, based on historical fall event data, extract relevant features, such as human body posture, movement trajectory, speed change, etc., use machine learning algorithms, such as random forest, support vector machine, deep learning, etc., to train the model, input the real-time monitoring data into the fall event prediction model to obtain the probability of a fall event occurring output by the fall event prediction model; When the expected event type is a fight event, analyze the historical fight event videos, extract features, such as the degree of personnel gathering, the intensity of actions, sound features, etc., train this model to identify fight behaviors, input the real-time monitoring data into the fight event prediction model to obtain the probability of a fight event occurring output by the fight event prediction model; When the expected event type is a pickpocketing event, collect pickpocketing event cases, extract features, such as the degree of personnel proximity, hand movements, item movements, etc., train the model to predict pickpocketing behaviors, input the real-time monitoring data into the pickpocketing event prediction model to obtain the probability of a pickpocketing event occurring output by the pickpocketing event prediction model. Those skilled in the art understand that the construction and application of the fall event prediction model, the fight event prediction model, and the pickpocketing event prediction model can refer to the relevant event behavior prediction models in the prior art. The content of the construction and application of these three models is not an improvement point of the present invention and does not affect the specific implementation of the present invention, so it will not be elaborated here.

[0105] In other embodiments, when the expected event type is a fall event, a fight event, or a pickpocketing event, input the real-time monitoring data into a preset event prediction model to determine the second expected event occurrence result output by the preset event prediction model, and generate a second target warning strategy according to the second expected event occurrence result. In such embodiments, the output is not the probability of the second expected event occurring, but the specific result, such as whether a fall occurs, whether there is a fight behavior, whether there is a pickpocketing event, etc. Further, the training set and test set data of the corresponding prediction model can be adaptively adjusted according to the above scheme.

[0106] Optionally, the generating a second target warning strategy according to the probability of the second expected event occurring includes:

[0107] In the case where the occurrence probability of the second expected event is less than or equal to the first preset occurrence probability, no warning is issued;

[0108] In the case where the occurrence probability of the second expected event is greater than the first preset occurrence probability and less than or equal to the second preset occurrence probability, the video monitoring data and the real-time monitoring data are combined into suspected event monitoring data, and the suspected event monitoring data is sent to a third-party manual review platform to implement manual review of the suspected event monitoring data;

[0109] In the case where the occurrence probability of the second expected event is greater than the second preset occurrence probability, the video monitoring data and the real-time monitoring data are used to capture the head portrait of the target event person, and the head portrait of the target event person is sent to a third-party warning platform.

[0110] Optionally, this embodiment describes the specific process of generating the second target warning strategy according to the occurrence probability of the second expected event. Specifically, different warning measures are taken according to different ranges of the occurrence probability of the second expected event, including not issuing a warning, sending the suspected event monitoring data to a third-party manual review platform, and sending the head portrait of the target event person to a third-party warning platform. When the occurrence probability of the second expected event is less than or equal to the first preset occurrence probability, the system considers that the possibility of the event occurring is relatively low, so no warning measures are taken; when the occurrence probability of the second expected event is greater than the first preset occurrence probability and less than or equal to the second preset occurrence probability, the system combines the video monitoring data and the real-time monitoring data into suspected event monitoring data, and sends the suspected event monitoring data to a third-party manual review platform for manual review to confirm whether the event actually occurs; when the occurrence probability of the second expected event is greater than the second preset occurrence probability, the system considers that the possibility of the event occurring is relatively high, and the system uses the video monitoring data and the real-time monitoring data to capture the head portrait of the target event person, and sends the head portrait of the target event person to a third-party warning platform to take warning measures in a timely manner.

[0111] By taking different warning measures according to different ranges of the occurrence probability of the second expected event, the present invention avoids unnecessary warnings and false alarms, improves the accuracy of warnings. When the possibility of the event occurring is relatively high, sending the head portrait of the target event person to a third-party warning platform in a timely manner helps to quickly take emergency response measures, reduce the losses and impacts caused by the event. By manually reviewing the suspected event monitoring data, it can be further confirmed whether the event actually occurs, improving the reliability and effectiveness of security prevention.

[0112] The present invention comprehensively obtains various types of information related to events by fusing multi-modal data (including video surveillance data, audio data, and infrared data) and integrating the active alarm information of users, thereby more accurately predicting potential security events, effectively avoiding the problem of inaccurate early warnings that may be caused by a single data source, and significantly improving the accuracy of early warnings; it can obtain and analyze surveillance data in real time, quickly generate corresponding early warning strategies according to different expected event types, ensure that early warnings can be issued in a timely manner at the initial stage of the event, and significantly improve the timeliness of early warnings;

[0113] For different types of expected events, the present invention adopts different early warning strategies. By using differential early warning strategy design, the early warning measures are made more in line with the actual event requirements, improving the pertinence and effectiveness of early warnings. By introducing technical means such as speech recognition and machine learning, the automatic processing of active alarm information and the intelligent analysis of real-time surveillance data are realized, which not only reduces the burden of manual monitoring, but also improves the efficiency and accuracy of the system in processing massive data; by accurately predicting the probability and location of event occurrence, the monitoring resources and emergency forces are reasonably allocated, avoiding waste and duplicate investment of resources, which helps to improve the resource utilization efficiency and reduce the monitoring cost.

[0114] Figure 2 It is a schematic structural diagram of a video surveillance event early warning system based on multi-modal data fusion provided by the present invention. The video surveillance event early warning system based on multi-modal data fusion includes a determination unit 1. The determination unit 1 is used to determine the expected event occurrence location and the expected event type according to the event status information corresponding to the current time period. The event status information includes the active alarm information from any user and the video surveillance data determined by any passive monitoring device associated with the active alarm information. The expected event types include fall events, fight events, pickpocket events, fire events, and theft events. The working principle of the determination unit 1 can refer to the foregoing step 101 and will not be elaborated here.

[0115] The video surveillance event early warning system based on multi-modal data fusion further includes an acquisition unit 2. The acquisition unit 2 is used to determine the actual monitoring devices according to the expected event occurrence location and the expected event type, and obtain the real-time monitoring data of the next time period from the actual monitoring devices. The working principle of the acquisition unit 2 can refer to the foregoing step 102 and will not be elaborated here.

[0116] The video surveillance event early warning system based on multi-modal data fusion further includes a first generating unit 3. The first generating unit 3 is configured to, when the expected event type is a fire event or a theft event, determine a first probability of the occurrence of the expected event according to the video surveillance data and the real-time monitoring data, and generate a first target early warning strategy according to the first probability of the occurrence of the expected event. The working principle of the first generating unit 3 can refer to the foregoing step 103 and will not be elaborated herein.

[0117] The video surveillance event early warning system based on multi-modal data fusion further includes a second generating unit 4. The second generating unit 4 is configured to, when the expected event type is a fall event, a fight event or a pickpocketing event, input the real-time monitoring data into a preset event prediction model, determine a second probability of the occurrence of the expected event output by the preset event prediction model, and generate a second target early warning strategy according to the second probability of the occurrence of the expected event. The working principle of the second generating unit 4 can refer to the foregoing step 104 and will not be elaborated herein.

[0118] By fusing multi-modal data (including video surveillance data, audio data, and infrared data) and integrating the user's active alarm information, the present invention obtains various types of information related to events more comprehensively, thereby predicting potential security events more accurately, effectively avoiding the problem of inaccurate early warning that may be caused by a single data source, and significantly improving the accuracy of early warning; it can obtain and analyze monitoring data in real time, quickly generate corresponding early warning strategies according to different expected event types, ensure that early warnings can be issued in a timely manner at the initial stage of the event, and significantly improve the timeliness of early warning;

[0119] For different types of expected events, the present invention adopts different early warning strategies. By using differentiated early warning strategy design, the early warning measures are made closer to the actual event requirements, improving the pertinence and effectiveness of early warning. By introducing technical means such as speech recognition and machine learning, the automatic processing of active alarm information and the intelligent analysis of real-time monitoring data are realized, which not only reduces the burden of manual monitoring, but also improves the efficiency and accuracy of the system in processing massive data; by accurately predicting the probability and location of event occurrence, the monitoring resources and emergency forces are reasonably allocated, avoiding waste and repeated investment of resources, which helps to improve the resource utilization efficiency and reduce the monitoring cost.

[0120] Figure 3 is a schematic structural diagram of the electronic device provided by the present invention. As Figure 3As shown in the figure, the electronic device may include: a processor 110, a communications interface 120, a memory 130, and a communication bus 140. Among them, the processor 110, the communications interface 120, and the memory 130 complete mutual communication through the communication bus 140. The processor 110 may call the logical instructions in the memory 130 to execute a video surveillance event early warning method based on multimodal data fusion. The method includes: determining the expected event occurrence location and the expected event type according to the event status information corresponding to the current time period, where the event status information includes the active alarm information from any user and the video surveillance data determined by any passive monitoring device associated with the active alarm information, and the expected event types include fall events, fight events, pickpocketing events, fire events, and theft events; determining the actual monitoring device according to the expected event occurrence location and the expected event type, and obtaining the real-time monitoring data of the next time period from the actual monitoring device; in the case where the expected event type is a fire event or a theft event, determining the first expected event occurrence probability according to the video surveillance data and the real-time monitoring data, and generating a first target early warning strategy according to the first expected event occurrence probability; in the case where the expected event type is a fall event, a fight event, or a pickpocketing event, inputting the real-time monitoring data into a preset event prediction model, determining the second expected event occurrence probability output by the preset event prediction model, and generating a second target early warning strategy according to the second expected event occurrence probability.

[0121] In addition, the logical instructions in the above-mentioned memory 130 can be implemented in the form of software functional units and, when sold or used as an independent product, can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0122] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute a video surveillance event early warning method based on multi-modal data fusion provided by the above-mentioned various methods. The method includes: determining the expected event occurrence location and the expected event type according to the event status information corresponding to the current time period, where the event status information includes the active alarm information from any user and the video surveillance data determined by any passive monitoring device associated with the active alarm information, and the expected event types include fall events, fight events, pickpocketing events, fire events, and theft events; determining the actual monitoring device according to the expected event occurrence location and the expected event type, and obtaining the real-time monitoring data of the next time period from the actual monitoring device; in the case where the expected event type is a fire event or a theft event, determining the first expected event occurrence probability according to the video surveillance data and the real-time monitoring data, and generating a first target early warning strategy according to the first expected event occurrence probability; in the case where the expected event type is a fall event, a fight event, or a pickpocketing event, inputting the real-time monitoring data into a preset event prediction model, determining the second expected event occurrence probability output by the preset event prediction model, and generating a second target early warning strategy according to the second expected event occurrence probability.

[0123] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the video surveillance event early warning method based on multi-modal data fusion provided by the above-mentioned various methods. The method includes: determining the expected event occurrence location and the expected event type according to the event status information corresponding to the current time period, where the event status information includes the active alarm information from any user and the video surveillance data determined by any passive monitoring device associated with the active alarm information, and the expected event types include fall events, fight events, pickpocketing events, fire events, and theft events; determining the actual monitoring device according to the expected event occurrence location and the expected event type, and obtaining the real-time monitoring data of the next time period from the actual monitoring device; in the case where the expected event type is a fire event or a theft event, determining the first expected event occurrence probability according to the video surveillance data and the real-time monitoring data, and generating a first target early warning strategy according to the first expected event occurrence probability; in the case where the expected event type is a fall event, a fight event, or a pickpocketing event, inputting the real-time monitoring data into a preset event prediction model, determining the second expected event occurrence probability output by the preset event prediction model, and generating a second target early warning strategy according to the second expected event occurrence probability.

[0124] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative work.

[0125] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0126] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A video surveillance event early warning method based on multi-modal data fusion, characterized in that Including: Determine the expected event occurrence location and the expected event type according to the event status information corresponding to the current time period. The event status information includes the active alarm information from any user and the video surveillance data determined by any passive monitoring device associated with the active alarm information. The expected event types include fall events, fight events, pickpocketing events, fire events, and theft events; Determine the actual monitoring device according to the expected event occurrence location and the expected event type, and obtain the real-time monitoring data of the next time period from the actual monitoring device; In the case where the expected event type is a fire event or a theft event, determine the first expected event occurrence probability according to the video surveillance data and the real-time monitoring data, and generate a first target warning strategy according to the first expected event occurrence probability; In the case where the expected event type is a fall event, a fight event, or a pickpocketing event, input the real-time monitoring data into a preset event prediction model, determine the second expected event occurrence probability output by the preset event prediction model, and generate a second target warning strategy according to the second expected event occurrence probability.

2. The method for video surveillance event early warning based on multi-modal data fusion according to claim 1, wherein The determining the expected event occurrence location and the expected event type according to the event status information corresponding to the current time period includes: Perform speech recognition on the active alarm information to obtain the alarm event location and the alarm event type, call all passive monitoring devices within the area where the alarm event location is located, and obtain the video surveillance data corresponding to each passive monitoring device in the current time period. Different passive monitoring devices correspond to different event occurrence locations, and the passive monitoring device is a video surveillance device; For any video surveillance data, obtain the current time period target analysis result corresponding to the video surveillance data. The current time period target analysis result includes the initial occurrence probabilities of different events. In the case where the initial occurrence probability of the event corresponding to any alarm event type is greater than the preset occurrence probability, determine the alarm event type as the expected event type; Determine all passive monitoring devices whose initial occurrence probabilities of the events corresponding to the alarm event type are greater than the preset occurrence probability as candidate monitoring devices, sort the initial occurrence probabilities corresponding to all candidate monitoring devices in descending order, and determine the event occurrence location corresponding to the candidate monitoring device with the largest initial occurrence probability as the expected event occurrence location.

3. The video surveillance event early warning method based on multi-modal data fusion according to claim 2, characterized in that, Before obtaining the current time period target analysis result corresponding to the video surveillance data, the method further includes: For any passive monitoring device, obtain the historical monitoring data corresponding to the passive monitoring device in different historical time periods. For the historical monitoring data of any historical time period, input the historical monitoring data into a preset event classification model to obtain the historical time period event analysis result corresponding to the historical time period output by the preset event classification model; Traverse all historical time periods, obtain each historical time period event analysis result corresponding to all historical time periods, and for any event type, determine the current time period target analysis result according to each historical time period event analysis result and the current time period event analysis result. The analysis result of the current time period event is determined by inputting the current monitoring data of the current time period into the preset event classification model; the preset event classification model is determined after training according to all sample monitoring data and the sample classification label corresponding to each sample monitoring data, and the sample classification label includes the occurrence probability of each event type corresponding to the sample monitoring data.

4. The method for warning of video surveillance events based on multi-modal data fusion according to claim 3, wherein Determining the target analysis result of the current time period according to the analysis result of each historical time period event and the analysis result of the current time period event includes: Sequentially determining the three time periods before the current time period as the first target time period, the second target time period, and the third target time period, and determining the first analysis result corresponding to the first target time period, the second analysis result corresponding to the second target time period, and the third analysis result corresponding to the third target time period from the analysis results of historical time period events; For any event type, determining the target event result corresponding to the event type according to the first event result, the second event result, the third event result, and the current event result, and traversing all event types to determine the target analysis result of the current time period; Wherein, the first event result is the occurrence probability of the event corresponding to the event type in the first analysis result, the second event result is the occurrence probability of the event corresponding to the event type in the second analysis result, the third event result is the occurrence probability of the event corresponding to the event type in the third analysis result, and the current event result is the occurrence probability of the event corresponding to the event type in the analysis result of the current time period event.

5. The method for video surveillance event early warning based on multi-modal data fusion according to claim 4, wherein, Determining the target event result corresponding to the event type according to the first event result, the second event result, the third event result, and the current event result includes: T = 0.1×R1 + 0.2×R2 + 0.3×R3 + 0.4×R4; Wherein, T is the target event result, R1 is the first event result, R2 is the second event result, R3 is the third event result, and R4 is the current event result.

6. The video surveillance event early warning method based on multi-modal data fusion according to claim 1, characterized in that, Determining the actual monitoring device according to the expected event occurrence location and the expected event type, and obtaining the real-time monitoring data of the next time period from the actual monitoring device includes: In the case where the expected event type is a fire event, determining the video monitoring device and the infrared sensing device associated with the expected event occurrence location as the actual monitoring device, and obtaining the video data and infrared data of the next time period from the actual monitoring device; In the case where the expected event types are fall event, fight event, and pickpocketing event, determining the video monitoring device and the audio monitoring device associated with the expected event occurrence location as the actual monitoring device, and obtaining the video data and audio data of the next time period from the actual monitoring device; In the case where the expected event type is a theft event, determining the video monitoring device associated with the expected event occurrence location as the actual monitoring device, and obtaining the video data of the next time period from the actual monitoring device.

7. The method for video surveillance event early warning based on multi-modal data fusion according to claim 6, wherein Determining a first expected event occurrence probability based on the video monitoring data and the real-time monitoring data, and generating a first target warning strategy according to the first expected event occurrence probability, including: When the expected event type is a fire event, determining an average brightness value and an average red saturation value of a video frame according to the video data, determining a maximum temperature value according to the infrared data, and determining a first fire event occurrence probability according to the average brightness value of the video frame, a preset brightness weight, the average red saturation value, a preset saturation weight, the maximum temperature value, and a preset temperature weight; determining a fire event occurrence probability according to the first fire event occurrence probability, a first fire weight, a second fire event occurrence probability corresponding to the video monitoring data, and a second fire weight, and generating a fire indication instruction when it is determined that the fire event occurrence probability is greater than a fire preset probability, where the fire indication instruction is used to indicate going to the location where the expected event occurs to extinguish the fire; When the expected event type is a theft event, determining all first items in the current time period according to the video monitoring data, determining all second items in the next time period according to the real-time monitoring data, comparing the first items and the second items, determining the lost items, determining that a theft event has occurred when the lost items are on a preset monitoring list, capturing a target thief's head image using the video monitoring data and the real-time monitoring data, and sending the target thief's head image to a third-party warning platform.

8. The method for video surveillance event early warning based on multi-modal data fusion according to claim 1, wherein The preset event prediction model includes a fall event prediction model, a fight event prediction model, and a pickpocketing event prediction model; When the expected event type is a fall event, inputting the real-time monitoring data into the fall event prediction model to determine the fall event occurrence probability output by the fall event prediction model; When the expected event type is a fight event, inputting the real-time monitoring data into the fight event prediction model to determine the fight event occurrence probability output by the fight event prediction model; When the expected event type is a pickpocketing event, inputting the real-time monitoring data into the pickpocketing event prediction model to determine the pickpocketing event occurrence probability output by the pickpocketing event prediction model.

9. The method for video surveillance event early warning based on multi-modal data fusion according to claim 1, wherein Generating a second target warning strategy according to the second expected event occurrence probability, including: When the second expected event occurrence probability is less than or equal to a first preset occurrence probability, no warning is issued; When the second expected event occurrence probability is greater than the first preset occurrence probability and less than or equal to a second preset occurrence probability, combining the video monitoring data and the real-time monitoring data into suspected event monitoring data, and sending the suspected event monitoring data to a third-party manual review platform to implement manual review of the suspected event monitoring data; When the second expected event occurrence probability is greater than the second preset occurrence probability, capturing a target event person's head image using the video monitoring data and the real-time monitoring data, and sending the target event person's head image to a third-party warning platform.

10. A video surveillance event early warning system based on multi-modal data fusion, characterized in that, Including: A determination unit, configured to determine an expected event occurrence location and an expected event type according to event status information corresponding to a current time period, where the event status information includes active alarm information from any user and video monitoring data determined by any passive monitoring device associated with the active alarm information, and the expected event types include a fall event, a fight event, a pickpocketing event, a fire event, and a theft event; An acquisition unit, configured to determine an actual monitoring device according to the expected event occurrence location and the expected event type, and acquire real-time monitoring data for a next time period from the actual monitoring device; A first generation unit, configured to, when the expected event type is a fire event or a theft event, determine a first expected event occurrence probability according to the video monitoring data and the real-time monitoring data, and generate a first target warning strategy according to the first expected event occurrence probability; A second generation unit, configured to, when the expected event type is a fall event, a fight event, or a pickpocketing event, input the real-time monitoring data into a preset event prediction model, determine a second expected event occurrence probability output by the preset event prediction model, and generate a second target warning strategy according to the second expected event occurrence probability.

Citation Information

Patent Citations

  • Event early warning method and device, electronic equipment and storage medium

    CN112257546A

  • Substation fire alarm method and system

    CN112242036A

  • Image processing method and device, equipment and medium

    CN115294492A

  • Transportation alarm method and system based on customs lock

    CN115587777A

  • Urban rail transit security and protection integrated monitoring method and security and protection integrated platform

    CN115802011A