Video monitoring event early warning method and system based on multi-modal data fusion
By combining multimodal data fusion and event prediction models with user alarm information, the accuracy and timeliness of early warnings in video surveillance systems have been improved, solving the problems of inaccurate and untimely early warnings in existing technologies, and improving targeting and resource utilization efficiency.
Patent Information
- Application Number
- CN202510493660.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2045-04-18
AI Technical Summary
Existing video surveillance systems suffer from inaccurate and untimely early warnings, lack multimodal data fusion and targeted early warning strategies, and are unable to effectively handle proactive alarm information.
By integrating video surveillance data, audio data, and infrared data, combined with user-initiated alarm information, and utilizing event prediction models and voice recognition technology, the location and type of expected events are determined, and differentiated early warning strategies are adopted to generate highly targeted early warning strategies.
It improved the accuracy and timeliness of early warnings, reduced false alarms and missed alarms, improved resource utilization efficiency, and reduced monitoring costs.
Smart Images

Figure CN120375288B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of urban monitoring, and in particular to a video monitoring event early warning method and system based on multi-modal data fusion. BACKGROUND
[0002] In the current video monitoring system, although the security event can be warned to a certain extent through the video monitoring data, the accuracy and timeliness of the warning still have many problems. On the one hand, the existing warning system often only relies on single video monitoring data, lacks the fusion and utilization of other modal data, resulting in inaccurate warning results; on the other hand, for different types of security events, the existing warning system often uses the same warning strategy, lacks pertinence and flexibility, and is difficult to meet the needs in actual application. In addition, the existing video monitoring warning system also has the problem that the processing of active alarm information is not intelligent enough. When the user actively alarms, the system can only simply record the alarm information, and cannot quickly and accurately determine the expected event occurrence location and expected event type according to the alarm information, so as to start the warning mechanism in time and effectively.
[0003] Chinese invention patent with publication number CN112257546A discloses an event early warning method, device, electronic equipment and storage medium, the event early warning method comprises: acquiring a monitoring picture shot by a monitoring device, and identifying a current event corresponding to the monitoring picture; querying a predicted event associated with the current event by using an event knowledge graph; wherein the event knowledge graph stores the association relationship between events; if the predicted event is a target type event, generate the warning information corresponding to the predicted event, which can warn before the occurrence of a malignant event. However, it has the disadvantages of single data source, lack of integration of active alarm information, lack of pertinence and flexibility of warning strategy, inability to update and adjust the warning model in real time, and inaccurate probability prediction of event occurrence, etc. At present, there is no technical solution that can solve the above technical problems, and there is no video monitoring event early warning method and system based on multi-modal data fusion. SUMMARY
[0004] The present application provides a video monitoring event early warning method and system based on multi-modal data fusion, which solves the problems of inaccurate and untimely warning in the existing video monitoring system, and lack of pertinence and flexibility of warning strategy for different types of events.
[0005] In a first aspect, the present application provides a video monitoring event early warning method based on multi-modal data fusion, comprising:
[0006] determine an expected event occurrence location and an expected event type according to event state information corresponding to a current period, the event state information including active alarm information from any user and video monitoring data determined by any passive monitoring device associated with the active alarm information, the expected event type including a fall event, a fight event, a pickpocketing event, a fire event, and a theft event;
[0007] determine an actual monitoring device according to the expected event occurrence location and the expected event type, and obtain real-time monitoring data of a next period from the actual monitoring device;
[0008] in a case where the expected event type is the fire event or the theft event, determine a first expected event occurrence probability according to the video monitoring data and the real-time monitoring data, and generate a first target early warning strategy according to the first expected event occurrence probability;
[0009] in a case where the expected event type is the fall event, the fight event, or the pickpocketing event, input the real-time monitoring data into a preset event prediction model, determine a second expected event occurrence probability output by the preset event prediction model, and generate a second target early warning strategy according to the second expected event occurrence probability.
[0010] The video monitoring event early warning method based on multi-modal data fusion provided by the present application comprises the following steps:
[0011] voice recognize the active alarm information, obtain an alarm event location and an alarm event type, call all passive monitoring devices in a region where the alarm event location is located, and obtain video monitoring data corresponding to a current period of each passive monitoring device, wherein different passive monitoring devices correspond to different event occurrence locations, and the passive monitoring device is a video monitoring device;
[0012] for any video monitoring data, obtain a current period target analysis result corresponding to the video monitoring data, the current period target analysis result including initial occurrence probabilities of different events, and in a case where an initial occurrence probability of an event corresponding to any alarm event type is greater than a preset occurrence probability, determine that the alarm event type is the expected event type;
[0013] determine all passive monitoring devices, for which an initial occurrence probability of an event corresponding to the alarm event type is greater than a preset occurrence probability, as candidate monitoring devices, sort initial occurrence probabilities of all candidate monitoring devices in descending order, and determine an event occurrence location corresponding to a candidate monitoring device with the largest initial occurrence probability as the expected event occurrence location.
[0014] The method for video monitoring event early warning based on multi-modal data fusion provided by the application comprises the following steps:
[0015] For any passive monitoring device, historical monitoring data corresponding to the passive monitoring device in different historical periods is acquired, and for the historical monitoring data in any historical period, the historical monitoring data is input into a preset event classification model to obtain a historical period event analysis result corresponding to the historical period output by the preset event classification model;
[0016] All historical period event analysis results corresponding to all historical periods are acquired by traversing all historical periods, and for any event type, the current period target analysis result is determined according to each historical period event analysis result and a current period event analysis result;
[0017] The current period event analysis result is determined by inputting current monitoring data in the current period into the preset event classification model, and the preset event classification model is determined by training according to all sample monitoring data and a sample classification label corresponding to each sample monitoring data, wherein the sample classification label comprises the occurrence probability of each event type corresponding to the sample monitoring data.
[0018] The method for video monitoring event early warning based on multi-modal data fusion provided by the application comprises the following steps:
[0019] Three periods before the current period are sequentially determined as a first target period, a second target period and a third target period, a first analysis result corresponding to the first target period, a second analysis result corresponding to the second target period and a third analysis result corresponding to the third target period are determined from the historical period event analysis result;
[0020] For any event type, a target event result corresponding to the event type is determined according to a first event result, a second event result, a third event result and a current event result, and all event types are traversed to determine the current period target analysis result;
[0021] The first event result is the event occurrence probability corresponding to the event type in the first analysis result, the second event result is the event occurrence probability corresponding to the event type in the second analysis result, the third event result is the event occurrence probability corresponding to the event type in the third analysis result, and the current event result is the event occurrence probability corresponding to the event type in the current period event analysis result.
[0022] According to the video monitoring event early warning method based on multi-modal data fusion provided by the application, the target event result corresponding to the event type is determined according to the first event result, the second event result, the third event result and the current event result, and the target event result comprises:
[0023] T=0.1*R1+0.2*R2+0.3*R3+0.4*R4;
[0024] Wherein, T is the target event result, R1 is the first event result, R2 is the second event result, R3 is the third event result, and R4 is the current event result.
[0025] According to the video monitoring event early warning method based on multi-modal data fusion provided by the application, the actual monitoring device is determined according to the expected event occurrence location and the expected event type, and real-time monitoring data of the next period is obtained from the actual monitoring device, and the real-time monitoring data comprises:
[0026] In the case that the expected event type is a fire event, the video monitoring device and the infrared sensing device associated with the expected event occurrence location are determined as the actual monitoring device, and video data and infrared data of the next period are obtained from the actual monitoring device;
[0027] In the case that the expected event type is a fall event, a fight event and a pickpocketing event, the video monitoring device and the audio monitoring device associated with the expected event occurrence location are determined as the actual monitoring device, and video data and audio data of the next period are obtained from the actual monitoring device;
[0028] In the case that the expected event type is a theft event, the video monitoring device associated with the expected event occurrence location is determined as the actual monitoring device, and video data of the next period is obtained from the actual monitoring device.
[0029] According to the video monitoring event early warning method based on multi-modal data fusion provided by the application, the first expected event occurrence probability is determined according to the video monitoring data and the real-time monitoring data, and the first target early warning strategy is generated according to the first expected event occurrence probability, and the first expected event occurrence probability comprises:
[0030] In a case where the expected event type is a fire event, a luminance average value and a red saturation average value of a video frame are determined according to the video data, a highest temperature value is determined according to the infrared data, a first fire event occurrence probability is determined according to the luminance average value of the video frame, a preset luminance weight, the red saturation average value, a preset saturation weight, the highest temperature value, and a preset temperature weight, a fire event occurrence probability is determined according to the first fire event occurrence probability, a first fire weight, a second fire event occurrence probability corresponding to the video monitoring data, and a second fire weight, and in a case where the fire event occurrence probability is greater than a preset fire probability, a fire indication instruction is generated, and the fire indication instruction is used to instruct to go to the expected event occurrence location to extinguish fire.
[0031] In a case where the expected event type is a theft event, all first articles in a current period are determined according to the video monitoring data, all second articles in a next period are determined according to the real-time monitoring data, a lost article is determined by comparing the first articles and the second articles, a theft event is determined in a case where the lost article exists in a preset monitoring list, a target theft personnel portrait is captured by using the video monitoring data and the real-time monitoring data, and the target theft personnel portrait is sent to a third-party early warning platform.
[0032] According to the video monitoring event early warning method based on multi-modal data fusion provided in the application, the preset event prediction model includes a fall event prediction model, a fight event prediction model, and a pickpocketing event prediction model.
[0033] In a case where the expected event type is a fall event, the real-time monitoring data is input to the fall event prediction model, and a fall event occurrence probability output by the fall event prediction model is determined.
[0034] In a case where the expected event type is a fight event, the real-time monitoring data is input to the fight event prediction model, and a fight event occurrence probability output by the fight event prediction model is determined.
[0035] In a case where the expected event type is a pickpocketing event, the real-time monitoring data is input to the pickpocketing event prediction model, and a pickpocketing event occurrence probability output by the pickpocketing event prediction model is determined.
[0036] According to the video monitoring event early warning method based on multi-modal data fusion provided in the application, the second target early warning strategy is generated according to the second expected event occurrence probability, and the generating the second target early warning strategy according to the second expected event occurrence probability includes:
[0037] In a case where the second expected event occurrence probability is less than or equal to a first preset occurrence probability, no early warning is performed.
[0038] In a case where the second expected event occurrence probability is greater than the first preset event occurrence probability and less than or equal to the second preset event occurrence probability, the video monitoring data and the real-time monitoring data are combined into suspected event monitoring data, and the suspected event monitoring data is sent to a third-party artificial review platform to realize artificial review of the suspected event monitoring data.
[0039] In a case where the second expected event occurrence probability is greater than the second preset event occurrence probability, a target event personnel portrait is captured by using the video monitoring data and the real-time monitoring data, and the target event personnel portrait is sent to a third-party early warning platform.
[0040] In a second aspect, a video monitoring event early warning system based on multi-modal data fusion is provided, and the system comprises:
[0041] A determination unit is configured to determine an expected event occurrence location and an expected event type according to event state information corresponding to a current time period, the event state information comprising active alarm information from any user and video monitoring data determined by any passive monitoring device associated with the active alarm information, and the expected event type comprising a falling event, a fighting event, a pickpocketing event, a fire event and a theft event.
[0042] An acquisition unit is configured to determine an actual monitoring device according to the expected event occurrence location and the expected event type, and acquire real-time monitoring data of a next time period from the actual monitoring device.
[0043] A first generation unit is configured to determine a first expected event occurrence probability according to the video monitoring data and the real-time monitoring data in a case where the expected event type is a fire event or a theft event, and generate a first target early warning strategy according to the first expected event occurrence probability.
[0044] A second generation unit is configured to input the real-time monitoring data into a preset event prediction model in a case where the expected event type is a falling event, a fighting event or a pickpocketing event, determine a second expected event occurrence probability output by the preset event prediction model, and generate a second target early warning strategy according to the second expected event occurrence probability.
[0045] The application more comprehensively acquires various types of information related to the event by fusing multi-modal data (including video monitoring data, audio data, infrared data) and integrating user's active alarm information, so as to more accurately predict potential safety events, effectively avoid the inaccurate early warning problem caused by a single data source, and significantly improve the accuracy of early warning; the monitoring data can be acquired and analyzed in real time, the corresponding early warning strategy is quickly generated according to different expected event types, the early warning can be timely sent in the early stage of the event, and the timeliness of early warning is significantly improved;
[0046] The application adopts different early warning strategies for different types of expected events, the differentiated early warning strategy design makes the early warning measures more close to the actual event demand, improves the pertinence and effectiveness of early warning, the automatic processing of active alarm information and the intelligent analysis of real-time monitoring data are realized by introducing voice recognition, machine learning and other technical means, not only the burden of manual monitoring is reduced, but also the efficiency and accuracy of the system processing massive data are improved; by accurately predicting the probability and location of the event, the monitoring resources and emergency forces are reasonably allocated, the waste and repeated investment of resources are avoided, which helps to improve the resource utilization efficiency and reduce the monitoring cost. BRIEF DESCRIPTION OF DRAWINGS
[0047] In order to more clearly illustrate the technical solutions in the present application or prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor.
[0048] Figure 1 is a flowchart of the video monitoring event early warning method based on multi-modal data fusion provided by the present application;
[0049] Figure 2 is a structural schematic diagram of the video monitoring event early warning system based on multi-modal data fusion provided by the present application;
[0050] Figure 3 is a structural schematic diagram of the electronic device provided by the present application. DETAILED DESCRIPTION
[0051] In order to make the purpose, technical scheme and advantages of the present application more clear, the technical scheme in the present application will be described clearly and completely in the following combined with the drawings in the present application. Obviously, the described embodiments are some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the present application.
[0052] In order to solve the problems of inaccurate and untimely prewarning, and lack of pertinence and flexibility of prewarning strategy in the existing video monitoring system, the present application provides a smart city video monitoring event prewarning method based on multi-modal data fusion, which can more accurately predict and prewarn various types of safety events by combining data of multiple modes including video monitoring data, audio data and infrared data, and combining user active alarm information and real-time data of passive monitoring equipment, improve the accuracy and timeliness of prewarning, and can also adopt different prewarning strategies according to different safety event types, improve the pertinence and flexibility of prewarning. Figure 1 It is a flowchart of the video monitoring event prewarning method based on multi-modal data fusion provided by the present application, which comprises:
[0053] Step 101, determining the expected event occurrence place and the expected event type according to the event state information corresponding to the current period, wherein the event state information includes active alarm information from any user, and video monitoring data determined by any passive monitoring equipment associated with the active alarm information, and the expected event type includes fall event, fighting event, pickpocketing event, fire event and theft event;
[0054] Step 102, determining the actual monitoring equipment according to the expected event occurrence place and the expected event type, and acquiring real-time monitoring data of the next period from the actual monitoring equipment;
[0055] Step 103, in the case that the expected event type is fire event or theft event, determining a first expected event occurrence probability according to the video monitoring data and the real-time monitoring data, and generating a first target prewarning strategy according to the first expected event occurrence probability;
[0056] Step 104, in the case that the expected event type is fall event, fighting event or pickpocketing event, inputting the real-time monitoring data into a preset event prediction model, determining a second expected event occurrence probability output by the preset event prediction model, and generating a second target prewarning strategy according to the second expected event occurrence probability.
[0057] In step 101, step 101 mainly includes determining the occurrence location and type of the expected event according to the event state information of the current period, which involves the processing of active alarm information and the integration of passive monitoring device data. First, the application can convert the user's active alarm information into text or structured data through voice recognition technology, extract the key elements in the alarm information, such as the approximate event location and event type, and then call the passive monitoring devices associated with the approximate event location, such as the corresponding camera, to obtain the video monitoring data of these devices in the current period. According to the active alarm information and passive monitoring device data, combined with the preset rules or algorithms, the occurrence location and type of the expected event are determined. The expected event occurrence location is a more specific event occurrence location compared to the approximate event location, and the expected event type needs to be determined according to the event state information of the current period to determine whether the approximate event type is true, and if it is true, it is determined as the expected event type.
[0058] In an optional embodiment, assuming that a user reports a possible pickpocketing event through an alarm phone in a mall and provides an approximate location, the system first converts the alarm information into text through voice recognition technology and extracts the event location as "the second floor clothing area of the mall" and the event type as "pickpocketing". Then, the system calls all cameras in the second floor clothing area of the mall to obtain the video monitoring data of the current period. Through analysis of the video data, the system finds that there is a characteristic of pickpocketing behavior in the area with dense personnel flow. Combined with the alarm information and video data, the system finally determines that the expected event is a pickpocketing event and confirms the location of the related camera where the pickpocketing behavior occurs as the expected event occurrence location.
[0059] Optionally, the determining of the expected event occurrence location and the expected event type according to the event state information of the current period comprises:
[0060] Voice recognition of the active alarm information to obtain the alarm event location and the alarm event type, calling all passive monitoring devices in the area where the alarm event location is located to obtain the video monitoring data of each passive monitoring device in the current period corresponding to the video monitoring data, wherein different passive monitoring devices correspond to different event occurrence locations, and the passive monitoring device is a video monitoring device;
[0061] For any video monitoring data, the current period target analysis result corresponding to the video monitoring data is obtained, the current period target analysis result includes the initial occurrence probability of different events, and in the case that the initial occurrence probability of any alarm event type corresponding event is greater than the preset occurrence probability, the alarm event type is determined as the expected event type;
[0062] Determine all passive monitoring devices corresponding to the event whose initial occurrence probability is greater than the preset occurrence probability as candidate monitoring devices, sort the initial occurrence probabilities of all candidate monitoring devices in descending order, and determine the event occurrence location corresponding to the candidate monitoring device with the greatest initial occurrence probability as the expected event occurrence location.
[0063] Optionally, the present application uses voice recognition technology to convert the active alarm information of the user, such as telephone alarm or voice alarm, into processable text or structured data, extracts the alarm event location and alarm event type from the converted data, and then calls all passive monitoring devices in the region where the location is located according to the extracted alarm event location, mainly video monitoring devices in the present application, and obtains each video monitoring data corresponding to all passive monitoring devices in the current period. In the present application, all video monitoring data performs preliminary event behavior judgment at different historical periods, the monitoring means is relatively simple, and the determined event type also has a large misjudgment, that is, the present application preliminarily analyzes the video monitoring data provided by each passive monitoring device to obtain the target analysis result of the current period, and then directly obtains the target analysis result of the corresponding current period after determining that all passive monitoring devices need to be called. The present application determines the initial occurrence probability of different events based on the target analysis result of the current period. These probabilities are calculated based on the features in the monitoring screen (such as personnel behavior, environmental change, etc.), or are preliminarily determined according to model recognition.
[0064] Further, compare the alarm event type with the initial occurrence probability of the event in the target analysis result of the current period. If the initial occurrence probability of the event corresponding to any alarm event type is greater than the preset occurrence probability threshold, determine that the alarm event type is the expected event type, and determine all passive monitoring devices corresponding to the alarm event type and having an initial occurrence probability greater than the preset threshold as candidate monitoring devices. That is, the present embodiment is to find out the monitoring device with the highest event occurrence probability and the highest event correlation in all video monitoring data recorded by the current event behavior among all passive monitoring devices, so as to improve the accuracy and reliability of the data source for subsequent event warning. Specifically, sort the candidate monitoring devices in descending order of initial occurrence probability, select the event occurrence location corresponding to the candidate monitoring device with the greatest initial occurrence probability as the expected event occurrence location, thereby completing the determination of the candidate monitoring device, the expected event type, and the expected event occurrence location.
[0065] The application can comprehensively understand the event state by fusing active alarm information and passive monitoring device data, reduce false positives and false negatives, quickly determine the expected event type and location by quickly calling and analyzing monitoring device data, sort and select the monitoring device with the highest initial occurrence probability as the expected event occurrence location, improve the accuracy and efficiency of positioning, and reduce unnecessary calculation and resource consumption.
[0066] Optionally, before obtaining the current period target analysis result corresponding to the video monitoring data, the method further comprises:
[0067] For any passive monitoring device, obtain historical monitoring data corresponding to the passive monitoring device in different historical periods, and for the historical monitoring data in any historical period, input the historical monitoring data into a preset event classification model to obtain a historical period event analysis result corresponding to the historical period output by the preset event classification model;
[0068] Iterate through all historical periods to obtain each historical period event analysis result corresponding to all historical periods, and for any event type, determine the current period target analysis result according to each historical period event analysis result and a current period event analysis result;
[0069] The current period event analysis result is determined by inputting the current monitoring data in the current period into the preset event classification model; the preset event classification model is determined by training according to all sample monitoring data and a sample classification label corresponding to each sample monitoring data, and the sample classification label includes the occurrence probability of each event type corresponding to the sample monitoring data.
[0070] Optionally, before obtaining the current period target analysis result corresponding to the video monitoring data, the application is not performing preliminary analysis of historical monitoring data at all times, but as an optional embodiment of the application, historical data can also be obtained in real time, and the analysis result of the historical period is obtained by analyzing the historical video data. Specifically, for any passive monitoring device, collect historical monitoring data thereof in different historical periods. These historical data can include video monitoring records in the past few minutes. Input each historical period historical monitoring data collected into a pre-trained event classification model, the event classification model analyzes the historical monitoring data, and outputs a historical period event analysis result corresponding to each historical period. Iterate through all historical periods to collect and integrate the event analysis result of each historical period. These results constitute a comprehensive analysis of the historical monitoring data, and provide an important reference for subsequent determination of the current period target analysis result.
[0071] Meanwhile, the current monitoring data of the current period is input into a preset event classification model to obtain an event analysis result of the current period, and the target analysis result of the current period is determined comprehensively by combining the event analysis result of the historical period and the event analysis result of the current period. The target analysis result includes the initial occurrence probability of different events and is a key basis for judging the expected event type and location. Specifically, the preset event classification model is trained based on all sample monitoring data and corresponding sample classification labels. The sample classification labels include the occurrence probability of each event type corresponding to the sample monitoring data, and provide rich training data for the model. In such an embodiment, the historical monitoring data of the passive monitoring device can be collected and stored by using a database or a big data storage platform, and the event classification model can be constructed by using machine learning or deep learning technology. The model is trained by using a large amount of sample monitoring data and corresponding sample classification labels, so that the model can accurately identify the event type in the video and predict the occurrence probability of the event. Then, a statistical method (such as weighted average, moving average, etc.) is used to combine the historical data and the current data to determine a more accurate target analysis result of the current period. By combining the event analysis result of the historical period and the event analysis result of the current period, the event occurrence rule and trend can be more comprehensively understood, which helps to more accurately predict the event type and its occurrence probability that may occur in the current period, and improves the accuracy of the early warning.
[0072] Optionally, the target analysis result of the current period is determined according to the event analysis result of each historical period and the event analysis result of the current period, and includes:
[0073] The three periods before the current period are sequentially determined as a first target period, a second target period and a third target period, and the first analysis result corresponding to the first target period, the second analysis result corresponding to the second target period and the third analysis result corresponding to the third target period are determined from the historical period event analysis result.
[0074] For any event type, the target event result corresponding to the event type is determined according to the first event result, the second event result, the third event result and the current event result, and all event types are traversed to determine the target analysis result of the current period.
[0075] The first event result is the event occurrence probability corresponding to the event type in the first analysis result, the second event result is the event occurrence probability corresponding to the event type in the second analysis result, the third event result is the event occurrence probability corresponding to the event type in the third analysis result, and the current event result is the event occurrence probability corresponding to the event type in the current period event analysis result.
[0076] In an optional embodiment, three time periods before the current time period are sequentially determined as the first target time period, the second target time period and the third target time period, and in other embodiments, more target time periods can also be set, such as the first target time period, the second target time period, the third target time period and the fourth target time period, the first target time period, the second target time period, the third target time period, the fourth target time period and the fifth target time period, and the like. The selection of these time periods is based on the continuity in time so as to better capture the trend and regularity of event occurrence. From the historical time period event analysis results, a first analysis result corresponding to the first target time period, a second analysis result corresponding to the second target time period and a third analysis result corresponding to the third target time period are extracted. These analysis results contain the occurrence probability of different event types in each time period. For any event type, the following operations are performed by the application: according to the event occurrence probability in the first analysis result (the first event result), the event occurrence probability in the second analysis result (the second event result), the event occurrence probability in the third analysis result (the third event result) and the event occurrence probability in the current time period event analysis result (the current event result), the target event result corresponding to the event type is comprehensively determined. The target event result represents a comprehensive occurrence probability of the event type in the current time period, considering the joint action of historical data and current data. The above steps are repeated for all possible event types to calculate the target event result corresponding to each event type. Finally, the target event result constitutes the target analysis result of the current time period, which provides an important basis for subsequent early warning and decision-making.
[0077] Optionally, the determination of the target event result corresponding to the event type according to the first event result, the second event result, the third event result and the current event result comprises:
[0078] T=0.1×R1+0.2×R2+0.3×R3+0.4×R4;
[0079] Wherein, T is the target event result, R1 is the first event result, R2 is the second event result, R3 is the third event result, and R4 is the current event result.
[0080] Optionally, the application adopts a weighted average algorithm to calculate the target event result by combining the event occurrence probability of the historical time period and the event occurrence probability of the current time period. By comprehensively considering the event occurrence probability of the historical time period and the current time period, the event type and its occurrence probability that may occur in the current time period can be more accurately predicted, which helps to reduce false positives and omissions, and improve the accuracy and reliability of early warning.
[0081] In step 102, the real-time monitoring data of the next time period is obtained from the actual monitoring device according to the expected event occurrence location and the expected event type, comprising:
[0082] In the case that the expected event type is a fire event, video monitoring devices and infrared sensing devices associated with the expected event occurrence location are determined as the actual monitoring devices, and video data and infrared data of the next time period are acquired from the actual monitoring devices;
[0083] In the case that the expected event type is a fall event, a fight event or a pickpocketing event, video monitoring devices and audio monitoring devices associated with the expected event occurrence location are determined as the actual monitoring devices, and video data and audio data of the next time period are acquired from the actual monitoring devices;
[0084] In the case that the expected event type is a theft event, video monitoring devices associated with the expected event occurrence location are determined as the actual monitoring devices, and video data of the next time period is acquired from the actual monitoring devices.
[0085] Optionally, when the expected event type is a fire event, video monitoring devices and infrared sensing devices associated with the expected event occurrence location are determined as the actual monitoring devices, and video data and infrared data of the next time period are acquired from the actual monitoring devices, the video data is used to observe fire spread, personnel evacuation, etc., and the infrared data is used to detect fire sources, high-temperature areas, etc.; when the expected event type is a fall event, a fight event or a pickpocketing event, video monitoring devices and audio monitoring devices associated with the expected event occurrence location are determined as the actual monitoring devices, and video data and audio data of the next time period are acquired from the actual monitoring devices, the video data is used to observe specific processes of the event, personnel behavior, etc., and the audio data is used to capture on-site sounds, conversations, etc., to provide more clues for event analysis; when the expected event type is a theft event, video monitoring devices associated with the expected event occurrence location are determined as the actual monitoring devices, and video data of the next time period is acquired from the actual monitoring devices. The video data is used to observe the occurrence of theft behavior, characteristics of target personnel, etc.
[0086] The application utilizes a geographic information system (GIS) or a monitoring device management system to associate an expected event occurrence location with an associated monitoring device, accurately identifies and controls the actual monitoring device through device identification, location information and other means, adopts an API interface or a data transmission protocol of a video monitoring device to obtain real-time monitoring data of the next period from the actual monitoring device, for infrared sensing devices and audio monitoring devices, corresponding data acquisition and transmission technologies are also adopted to ensure the real-time and accuracy of the data, the application selects the actual monitoring device according to the type of the expected event, avoids blind monitoring and resource waste, and the obtained video data, infrared data and audio data provide rich information sources for event analysis, and the acquisition and analysis of real-time monitoring data help to discover security risks and abnormal behaviors in time, and provide strong support for taking preventive measures.
[0087] If the expected event type is a fire event or a theft event, step 103 is performed, and if the expected event type is a fall event, a fight event or a pickpocketing event, step 104 is performed.
[0088] In step 103, the first expected event occurrence probability is determined according to the video monitoring data and the real-time monitoring data, and a first target early warning strategy is generated according to the first expected event occurrence probability, which comprises:
[0089] In the case that the expected event type is a fire event, the average brightness value and the average red saturation value of the video frame are determined according to the video data, the highest temperature value is determined according to the infrared data, the first fire event occurrence probability is determined according to the average brightness value of the video frame, the preset brightness weight, the average red saturation value, the preset saturation weight, the highest temperature value and the preset temperature weight, the fire event occurrence probability is determined according to the first fire event occurrence probability, the first fire weight, the second fire event occurrence probability corresponding to the video monitoring data and the second fire weight, and in the case that the fire event occurrence probability is greater than the preset fire probability, a fire indication instruction is generated, which is used to instruct to go to the expected event occurrence location to extinguish the fire.
[0090] In the case that the expected event type is a theft event, all first articles in the current period are determined according to the video monitoring data, all second articles in the next period are determined according to the real-time monitoring data, the lost articles are determined by comparing the first articles and the second articles, in the case that the lost articles exist in the preset monitoring list, it is determined that a theft event occurs, the head portrait of the target thief is captured by using the video monitoring data and the real-time monitoring data, and the head portrait of the target thief is sent to a third party early warning platform.
[0091] Optionally, the application can use a video processing algorithm (such as OpenCV) to calculate the average brightness and average red saturation of the video frame, use infrared data processing technology to extract the highest temperature value, combine the average brightness, the preset brightness weight, the average red saturation, the preset saturation weight, the highest temperature value, and the preset temperature weight to calculate the first fire event occurrence probability, and further combine the first fire event occurrence probability, the first fire weight, the second fire event occurrence probability corresponding to the video monitoring data, and the second fire weight to determine the final fire event occurrence probability:
[0092] P final =W1×P1+W2×P2
[0093] wherein P final is the fire event occurrence probability, W1 is the first fire weight, W2 is the second fire weight, P1 is the first fire event occurrence probability, and P2 is the second fire event occurrence probability corresponding to the video monitoring data, wherein the calculation formula of the first fire event occurrence probability P1 is:
[0094] P1=W L ×L avg +W S ×S avg +W T ×T max
[0095] wherein W L is the preset brightness weight, L avg is the average brightness of the video frame, S avg is the average red saturation, W S is the preset saturation weight, T max is the highest temperature value, and W T is the preset temperature weight.
[0096] In an optional embodiment, it is assumed that the video monitoring data and the real-time monitoring data detected by an actual monitoring device of a certain shopping mall need to be analyzed, the system calculates the first fire event occurrence probability according to the preset weight and algorithm, and combines the historical data to calculate the final fire event occurrence probability, if the fire event occurrence probability is greater than the preset fire preset probability, the system generates a fire indication instruction to notify the fire personnel to go to the area to extinguish the fire.
[0097] Optionally, in the case of the expected event type being a theft event, the application uses an item identification algorithm (such as a deep learning model) to compare and identify items in the current period and the next period, and presets a preset monitoring list for quickly determining whether the lost item is an important or sensitive item, determines all first items in the current period according to the video monitoring data, determines all second items in the next period according to the real-time monitoring data, compares the first items and the second items to determine the lost item, and if the lost item exists in the preset monitoring list, it is determined that a theft event has occurred, and the video monitoring data and real-time monitoring data are combined with face recognition technology to capture the target thief's portrait, which is sent to a third-party early warning platform through a preset API interface or message queue.
[0098] In an optional embodiment, assuming that a jewelry store's monitoring system detects that there are several pieces of jewelry on the counter in the current period, but the real-time monitoring data of the next period shows that some of the jewelry is missing, the system compares the items in the current period and the next period to determine that the lost jewelry belongs to the items in the preset monitoring list, the system captures the target thief's portrait and sends it to a third-party early warning platform so that law enforcement personnel can intervene in a timely manner.
[0099] The application can more accurately calculate the probability of a fire event occurring and identify a theft event by combining video data, infrared data, and historical data. In the event of a fire, it can quickly generate a fire indication instruction to notify firefighters to put out the fire, reducing the loss caused by the fire. In the event of a theft, it can quickly capture the target thief's portrait and send it to a third-party early warning platform, which helps law enforcement personnel intervene and track in a timely manner. Through real-time analysis and processing of monitoring data, security risks and abnormal behavior can be detected in a timely manner, providing strong support for taking preventive measures. The use of a preset monitoring list helps to quickly identify the loss of important or sensitive items, improving the relevance and effectiveness of security precautions.
[0100] In step 104, the preset event prediction model includes a fall event prediction model, a fight event prediction model, and a pickpocketing event prediction model;
[0101] In the case of the expected event type being a fall event, the real-time monitoring data is input into the fall event prediction model to determine the fall event occurrence probability output by the fall event prediction model;
[0102] In the case of the expected event type being a fight event, the real-time monitoring data is input into the fight event prediction model to determine the fight event occurrence probability output by the fight event prediction model;
[0103] In the case that the expected event type is a pickpocketing event, the real-time monitoring data is input to the pickpocketing event prediction model to determine a pickpocketing event occurrence probability output by the pickpocketing event prediction model.
[0104] Optionally, the present application uses video processing technology (such as OpenCV) to pre-process the real-time monitoring video, extract key frames and features, process audio data, extract sound features (such as volume, frequency, etc.), input the processed real-time monitoring data to the corresponding prediction model, and calculate and output the event occurrence probability according to the input data: in the case that the expected event type is a fall event, based on historical fall event data, relevant features such as human posture, motion trajectory, speed change, etc. are extracted, and a machine learning algorithm such as random forest, support vector machine, deep learning, etc. is used to train the model. The real-time monitoring data is input to the fall event prediction model to obtain a fall event occurrence probability output by the fall event prediction model; in the case that the expected event type is a fight event, historical fight event videos are analyzed to extract features such as the degree of personnel gathering, the degree of action intensity, sound features, etc. This model is trained to identify fighting behavior. The real-time monitoring data is input to the fight event prediction model to obtain a fight event occurrence probability output by the fight event prediction model; in the case that the expected event type is a pickpocketing event, pickpocketing event cases are collected, features such as personnel proximity, hand movements, and object movements are extracted, and the model is trained to predict pickpocketing behavior. The real-time monitoring data is input to the pickpocketing event prediction model to obtain a pickpocketing event occurrence probability output by the pickpocketing event prediction model. Those skilled in the art understand that the construction and application of the fall event prediction model, the fight event prediction model, and the pickpocketing event prediction model can refer to related event behavior prediction models in the prior art. The construction and application of the three models are not the improvement points of the present application and do not affect the specific implementation of the present application. Therefore, they will not be described here.
[0105] In other embodiments, in the case that the expected event type is a fall event, a fight event, or a pickpocketing event, the real-time monitoring data is input to a preset event prediction model to determine a second expected event occurrence result output by the preset event prediction model, and a second target warning strategy is generated according to the second expected event occurrence result. In such embodiments, the output is not a second expected event occurrence probability, but a specific result, such as whether to fall, whether to fight, whether to pickpocket, etc. Further, the training set and test set data of the corresponding prediction model can be adaptively adjusted according to the above scheme.
[0106] Optionally, the second target warning strategy is generated according to the second expected event occurrence probability, including:
[0107] In a case where the second expected event occurrence probability is less than or equal to a first preset occurrence probability, no prewarning is performed.
[0108] In a case where the second expected event occurrence probability is greater than the first preset occurrence probability and less than or equal to a second preset occurrence probability, the video monitoring data and the real-time monitoring data are combined into suspected event monitoring data, and the suspected event monitoring data is sent to a third-party artificial review platform to realize artificial review of the suspected event monitoring data.
[0109] In a case where the second expected event occurrence probability is greater than the second preset occurrence probability, a target event personnel portrait is captured by using the video monitoring data and the real-time monitoring data, and the target event personnel portrait is sent to a third-party prewarning platform.
[0110] Optionally, the embodiment describes a specific process of generating a second target prewarning strategy according to a second expected event occurrence probability. Specifically, different prewarning measures are taken according to different ranges of the second expected event occurrence probability, including no prewarning, sending suspected event monitoring data to a third-party artificial review platform, and sending a target event personnel portrait to a third-party prewarning platform. When the second expected event occurrence probability is less than or equal to a first preset occurrence probability, the system considers that the event is less likely to occur, and therefore no prewarning measure is taken. When the second expected event occurrence probability is greater than the first preset occurrence probability and less than or equal to a second preset occurrence probability, the system combines the video monitoring data and the real-time monitoring data into suspected event monitoring data, and sends the suspected event monitoring data to a third-party artificial review platform for review by an artificial person to confirm whether the event actually occurs. When the second expected event occurrence probability is greater than the second preset occurrence probability, the system considers that the event is more likely to occur, and the system captures a target event personnel portrait by using the video monitoring data and the real-time monitoring data, and sends the target event personnel portrait to a third-party prewarning platform to take prewarning measures in a timely manner.
[0111] The application takes different prewarning measures according to different ranges of the second expected event occurrence probability, avoids unnecessary prewarning and false positives, improves the accuracy of prewarning, sends a target event personnel portrait to a third-party prewarning platform in a timely manner when the event is more likely to occur, helps to quickly take emergency response measures, reduces the loss and influence caused by the event, and further confirms whether the event actually occurs by reviewing suspected event monitoring data by an artificial person, thereby improving the reliability and effectiveness of security protection.
[0112] The application more comprehensively acquires various types of information related to the event by fusing multi-modal data (including video monitoring data, audio data, infrared data) and integrating user's active alarm information, thereby more accurately predicting potential safety events, effectively avoiding the inaccurate early warning problem that may be caused by a single data source, and significantly improving the accuracy of early warning; real-time monitoring data can be acquired and analyzed, corresponding early warning strategies are quickly generated according to different expected event types, early warning can be timely issued in the early stage of event occurrence, and the timeliness of early warning is significantly improved;
[0113] Different early warning strategies are adopted for different types of expected events in the application, the differential early warning strategy design makes the early warning measures more close to the actual event demand, improves the pertinence and effectiveness of early warning, the automatic processing of active alarm information and the intelligent analysis of real-time monitoring data are realized by introducing voice recognition, machine learning and other technical means, not only the burden of manual monitoring is reduced, but also the efficiency and accuracy of the system processing massive data are improved; by accurately predicting the probability and location of event occurrence, monitoring resources and emergency forces are reasonably allocated, waste of resources and repeated investment are avoided, which helps to improve the resource utilization efficiency and reduce the monitoring cost.
[0114] Figure 2 It is a structural schematic diagram of a video monitoring event early warning system based on multi-modal data fusion provided by the application, the video monitoring event early warning system based on multi-modal data fusion comprises a determination unit 1, the determination unit 1 is used for determining an expected event occurrence location and an expected event type according to event state information corresponding to a current period, the event state information comprises active alarm information from any user and video monitoring data determined by any passive monitoring device associated with the active alarm information, the expected event type comprises a fall event, a fight event, a pickpocketing event, a fire event and a theft event, the working principle of the determination unit 1 can refer to the foregoing step 101, and details are not described herein.
[0115] The video monitoring event early warning system based on multi-modal data fusion further comprises an acquisition unit 2, the acquisition unit 2 is used for determining an actual monitoring device according to the expected event occurrence location and the expected event type, and acquiring real-time monitoring data of a next period from the actual monitoring device, the working principle of the acquisition unit 2 can refer to the foregoing step 102, and details are not described herein.
[0116] The video monitoring event early warning system based on multi-modal data fusion further comprises a first generation unit 3, which is configured to determine a first expected event occurrence probability according to the video monitoring data and the real-time monitoring data, and generate a first target early warning strategy according to the first expected event occurrence probability, when the expected event type is a fire event or a theft event.
[0117] The video monitoring event early warning system based on multi-modal data fusion further comprises a second generation unit 4, which is configured to input the real-time monitoring data into a preset event prediction model, determine a second expected event occurrence probability output by the preset event prediction model, and generate a second target early warning strategy according to the second expected event occurrence probability, when the expected event type is a fall event, a fight event or a pickpocketing event.
[0118] The present application can more accurately predict potential security events by fusing multi-modal data (including video monitoring data, audio data, infrared data) and integrating user-initiated alarm information, effectively avoiding the inaccuracy of early warning caused by a single data source, and significantly improving the accuracy of early warning.
[0119] The present application adopts different early warning strategies for different types of expected events, and uses differentiated early warning strategy design to make the early warning measures more closely meet the actual event needs, thereby improving the pertinence and effectiveness of early warning.
[0120] Figure 3 It is a structural schematic diagram of the electronic device provided by the present application. Figure 3As shown, the electronic device can include a processor 110, a communications interface 120, a memory 130, and a communications bus 140, wherein the processor 110, the communications interface 120, and the memory 130 complete mutual communication through the communications bus 140. The processor 110 can invoke a logical instruction in the memory 130 to execute a video monitoring event early warning method based on multi-modal data fusion, the method comprising: determining an expected event occurrence location and an expected event type according to event state information corresponding to a current period, the event state information including active alarm information from any user and video monitoring data determined by any passive monitoring device associated with the active alarm information, the expected event type including a fall event, a fight event, a pickpocketing event, a fire event, and a theft event; determining an actual monitoring device according to the expected event occurrence location and the expected event type, and acquiring real-time monitoring data of a next period from the actual monitoring device; in the case that the expected event type is a fire event or a theft event, determining a first expected event occurrence probability according to the video monitoring data and the real-time monitoring data, and generating a first target early warning strategy according to the first expected event occurrence probability; in the case that the expected event type is a fall event, a fight event, or a pickpocketing event, inputting the real-time monitoring data to a preset event prediction model, determining a second expected event occurrence probability output by the preset event prediction model, and generating a second target early warning strategy according to the second expected event occurrence probability.
[0121] In addition, the logical instruction in the memory 130 described above can be implemented in the form of a software functional unit and sold or used as an independent product, and can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the method described in various embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various program code storage media.
[0122] In another aspect, the present application also provides a computer program product comprising a computer program, which can be stored on a non-transitory computer-readable storage medium, and the computer program can be executed by a processor to enable a computer to perform a video monitoring event early warning method based on multi-modal data fusion, which comprises: determining an expected event occurrence location and an expected event type according to event state information corresponding to a current time period, wherein the event state information comprises active alarm information from any user and video monitoring data determined by any passive monitoring device associated with the active alarm information, and the expected event type comprises a fall event, a fight event, a pickpocketing event, a fire event and a theft event; determining an actual monitoring device according to the expected event occurrence location and the expected event type, and obtaining real-time monitoring data of a next time period from the actual monitoring device; in the case that the expected event type is a fire event or a theft event, determining a first expected event occurrence probability according to the video monitoring data and the real-time monitoring data, and generating a first target early warning strategy according to the first expected event occurrence probability; in the case that the expected event type is a fall event, a fight event or a pickpocketing event, inputting the real-time monitoring data into a preset event prediction model, determining a second expected event occurrence probability output by the preset event prediction model, and generating a second target early warning strategy according to the second expected event occurrence probability.
[0123] In another aspect, the present application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, and the computer program can be executed by a processor to implement a video monitoring event early warning method based on multi-modal data fusion, which comprises: determining an expected event occurrence location and an expected event type according to event state information corresponding to a current time period, wherein the event state information comprises active alarm information from any user and video monitoring data determined by any passive monitoring device associated with the active alarm information, and the expected event type comprises a fall event, a fight event, a pickpocketing event, a fire event and a theft event; determining an actual monitoring device according to the expected event occurrence location and the expected event type, and obtaining real-time monitoring data of a next time period from the actual monitoring device; in the case that the expected event type is a fire event or a theft event, determining a first expected event occurrence probability according to the video monitoring data and the real-time monitoring data, and generating a first target early warning strategy according to the first expected event occurrence probability; in the case that the expected event type is a fall event, a fight event or a pickpocketing event, inputting the real-time monitoring data into a preset event prediction model, determining a second expected event occurrence probability output by the preset event prediction model, and generating a second target early warning strategy according to the second expected event occurrence probability.
[0124] The device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed to multiple network units. Part or all of the modules can be selected to achieve the purposes of the embodiments according to actual needs. Those skilled in the art can understand and implement without creative labor.
[0125] Through the description of the above embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software and the necessary universal hardware platform, and of course can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of software products, and the computer software products can be stored in a computer readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and include a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.
[0126] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement to part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1.A video monitoring event pre-warning method based on multi-modal data fusion, characterized in that, The method comprises the following steps: determining an expected event occurrence location and an expected event type according to event state information corresponding to a current period, wherein the event state information comprises active alarm information from any user and video monitoring data determined by any passive monitoring device associated with the active alarm information, and the expected event type comprises a fall event, a fight event, a pickpocketing event, a fire event and a theft event; determining an actual monitoring device according to the expected event occurrence location and the expected event type, and obtaining real-time monitoring data of a next period from the actual monitoring device; in a case where the expected event type is the fire event or the theft event, determining a first expected event occurrence probability according to the video monitoring data and the real-time monitoring data, and generating a first target early warning strategy according to the first expected event occurrence probability; in a case where the expected event type is the fall event, the fight event or the pickpocketing event, inputting the real-time monitoring data into a preset event prediction model, determining a second expected event occurrence probability output by the preset event prediction model, and generating a second target early warning strategy according to the second expected event occurrence probability; the method of determining the expected event occurrence location and the expected event type according to the event state information corresponding to the current period comprises the following steps: recognizing the active alarm information by voice to obtain an alarm event location and an alarm event type, calling all passive monitoring devices in a region where the alarm event location is located, and obtaining video monitoring data corresponding to a current period of each passive monitoring device, wherein different passive monitoring devices correspond to different event occurrence locations, and the passive monitoring device is a video monitoring device; for any video monitoring data, obtaining a current period target analysis result corresponding to the video monitoring data, wherein the current period target analysis result comprises initial occurrence probabilities of different events, and in a case where an initial occurrence probability of an event corresponding to any alarm event type is greater than a preset occurrence probability, the alarm event type is determined as the expected event type; all passive monitoring devices whose initial occurrence probabilities of events corresponding to the alarm event type are greater than the preset occurrence probability are determined as candidate monitoring devices, the initial occurrence probabilities of all candidate monitoring devices are sorted in descending order, and an event occurrence location corresponding to a candidate monitoring device with the largest initial occurrence probability is determined as the expected event occurrence location. 2.The video monitoring event pre-warning method based on multi-modal data fusion according to claim 1, characterized in that, before obtaining the current period target analysis result corresponding to the video monitoring data, the method further comprises the following steps: for any passive monitoring device, obtaining historical monitoring data corresponding to the passive monitoring device in different historical periods, for historical monitoring data of any historical period, inputting the historical monitoring data into a preset event classification model to obtain a historical period event analysis result corresponding to the historical period output by the preset event classification model; iterating through all historical periods to obtain each historical period event analysis result corresponding to all historical periods, and for any event type, determining the current period target analysis result according to each historical period event analysis result and a current period event analysis result; The current period event analysis result is determined by inputting current monitoring data of the current period into the preset event classification model; the preset event classification model is determined by training all sample monitoring data and sample classification labels corresponding to each sample monitoring data, and the sample classification labels include occurrence probability of each event type corresponding to the sample monitoring data. 3.The video monitoring event pre-warning method based on multi-modal data fusion according to claim 2, characterized in that, The current period target analysis result is determined according to the historical period event analysis result and the current period event analysis result, and the method comprises the following steps: The three periods before the current period are sequentially determined as a first target period, a second target period and a third target period, and the first analysis result corresponding to the first target period, the second analysis result corresponding to the second target period and the third analysis result corresponding to the third target period are determined from the historical period event analysis result; For any event type, the target event result corresponding to the event type is determined according to the first event result, the second event result, the third event result and the current event result, all event types are traversed, and the current period target analysis result is determined; The first event result is the event occurrence probability corresponding to the event type in the first analysis result, the second event result is the event occurrence probability corresponding to the event type in the second analysis result, the third event result is the event occurrence probability corresponding to the event type in the third analysis result, and the current event result is the event occurrence probability corresponding to the event type in the current period event analysis result. 4.The video monitoring event pre-warning method based on multi-modal data fusion according to claim 3, characterized in that, The target event result corresponding to the event type is determined according to the first event result, the second event result, the third event result and the current event result, and the method comprises the following steps: ; wherein T is a target event outcome, is a first event outcome, is a second event outcome, is a third event outcome, is a current event outcome. 5.The video monitoring event pre-warning method based on multi-modal data fusion according to claim 1, characterized in that, The actual monitoring device is determined according to the expected event occurrence location and the expected event type, real-time monitoring data of the next period is obtained from the actual monitoring device, and the method comprises the following steps: In the case that the expected event type is a fire event, a video monitoring device and an infrared sensing device associated with the expected event occurrence location are determined as the actual monitoring device, and video data and infrared data of the next period are obtained from the actual monitoring device; In the case that the expected event type is a fall event, a fight event or a pickpocketing event, a video monitoring device and an audio monitoring device associated with the expected event occurrence location are determined as the actual monitoring device, and video data and audio data of the next period are obtained from the actual monitoring device; In the case that the expected event type is a theft event, a video monitoring device associated with the expected event occurrence location is determined as the actual monitoring device, and video data of the next period is obtained from the actual monitoring device. 6.The video monitoring event pre-warning method based on multi-modal data fusion according to claim 5, characterized in that, The first expected event occurrence probability is determined according to the video monitoring data and the real-time monitoring data, and the first target warning strategy is generated according to the first expected event occurrence probability, and the method comprises the following steps: In a case where the expected event type is a fire event, a luminance average value and a red saturation average value of a video frame are determined according to the video data, a highest temperature value is determined according to the infrared data, a first fire event occurrence probability is determined according to the luminance average value of the video frame, a preset luminance weight, the red saturation average value, a preset saturation weight, the highest temperature value, and a preset temperature weight, a fire event occurrence probability is determined according to the first fire event occurrence probability, a first fire weight, a second fire event occurrence probability corresponding to the video monitoring data, and a second fire weight, in a case where the fire event occurrence probability is greater than a preset fire probability, a fire indication instruction is generated, and the fire indication instruction is used to instruct to go to the expected event occurrence location to extinguish fire; In a case where the expected event type is a theft event, all first articles in a current period are determined according to the video monitoring data, all second articles in a next period are determined according to the real-time monitoring data, a lost article is determined by comparing the first articles and the second articles, in a case where the lost article exists in a preset monitoring list, it is determined that a theft event occurs, a target theft personnel portrait is captured by using the video monitoring data and the real-time monitoring data, and the target theft personnel portrait is sent to a third-party early warning platform. 7.The video monitoring event pre-warning method based on multi-modal data fusion according to claim 1, characterized in that, The preset event prediction model includes a fall event prediction model, a fight event prediction model, and a pickpocketing event prediction model; In a case where the expected event type is a fall event, the real-time monitoring data is input into the fall event prediction model, and a fall event occurrence probability output by the fall event prediction model is determined; In a case where the expected event type is a fight event, the real-time monitoring data is input into the fight event prediction model, and a fight event occurrence probability output by the fight event prediction model is determined; In a case where the expected event type is a pickpocketing event, the real-time monitoring data is input into the pickpocketing event prediction model, and a pickpocketing event occurrence probability output by the pickpocketing event prediction model is determined. 8.The video monitoring event pre-warning method based on multi-modal data fusion according to claim 1, characterized in that, The second target early warning strategy is generated according to the second expected event occurrence probability, including: In a case where the second expected event occurrence probability is less than or equal to a first preset occurrence probability, no early warning occurs; In a case where the second expected event occurrence probability is greater than the first preset occurrence probability and less than or equal to a second preset occurrence probability, the video monitoring data and the real-time monitoring data are merged into suspected event monitoring data, the suspected event monitoring data is sent to a third-party artificial review platform, so as to realize artificial review of the suspected event monitoring data; In a case where the second expected event occurrence probability is greater than the second preset occurrence probability, a target event personnel portrait is captured by using the video monitoring data and the real-time monitoring data, and the target event personnel portrait is sent to a third-party early warning platform. 9.A video monitoring event pre-warning system based on multi-modal data fusion, characterized in that, including: The determination unit is configured to determine an expected event occurrence location and an expected event type according to event state information corresponding to a current time period, the event state information including active alarm information from any user and video monitoring data determined by any passive monitoring device associated with the active alarm information, and the expected event type including a fall event, a fight event, a pickpocketing event, a fire event, and a theft event; The acquisition unit is configured to determine an actual monitoring device according to the expected event occurrence location and the expected event type, and acquire real-time monitoring data of a next time period from the actual monitoring device; The first generation unit is configured to, in a case where the expected event type is the fire event or the theft event, determine a first expected event occurrence probability according to the video monitoring data and the real-time monitoring data, and generate a first target early warning strategy according to the first expected event occurrence probability; The second generation unit is configured to, in a case where the expected event type is the fall event, the fight event, or the pickpocketing event, input the real-time monitoring data into a preset event prediction model, determine a second expected event occurrence probability output by the preset event prediction model, and generate a second target early warning strategy according to the second expected event occurrence probability; The determination of the expected event occurrence location and the expected event type according to the event state information corresponding to the current time period includes: Voice recognition of the active alarm information, acquisition of an alarm event location and an alarm event type, calling of all passive monitoring devices in a region where the alarm event location is located, and acquisition of video monitoring data corresponding to a current time period of each passive monitoring device, wherein different passive monitoring devices correspond to different event occurrence locations, and the passive monitoring device is a video monitoring device; For any video monitoring data, acquisition of a current time period target analysis result corresponding to the video monitoring data, the current time period target analysis result including initial occurrence probabilities of different events, and determination of the alarm event type as the expected event type in a case where an initial occurrence probability of an event corresponding to the alarm event type is greater than a preset occurrence probability; Determination of all passive monitoring devices, in which an initial occurrence probability of an event corresponding to the alarm event type is greater than a preset occurrence probability, as candidate monitoring devices, sorting of initial occurrence probabilities corresponding to all candidate monitoring devices in descending order, and determination of an event occurrence location corresponding to a candidate monitoring device with the largest initial occurrence probability as the expected event occurrence location.
Citation Information
Patent Citations
Event early warning method and device, electronic equipment and storage medium
CN112257546A
Urban rail transit security and protection integrated monitoring method and security and protection integrated platform
CN115802011A
Security monitoring method and system for GSM-R communication system
CN119012160A