Abnormal Event Prediction Method and Device

By using a variety of identification models and risk calculation methods in public places to generate warning information, the problems of slow response speed and limited coverage in the prior art are solved, and more efficient abnormal event response is achieved.

CN119557822BActive Publication Date: 2025-05-30富盛科技股份有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510131853.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-05-30
Estimated Expiration
2045-02-06

AI Technical Summary

Technical Problem

When handling abnormal events in public places, the prior art has slow response speed, limited coverage, and high labor costs. It is impossible to respond to abnormal events in a timely and effective manner, resulting in an increase in losses.

Method used

A method and device for predicting abnormal events is provided. By receiving learning materials from the business platform, model training is carried out, face recognition model, item recognition model, behavior recognition model and sound recognition model are determined, environmental images and sounds of multiple cameras are received, abnormal risk weight values ​​are calculated, and warning information is generated based on the comprehensive risk value.

Benefits of technology

It improves the response efficiency and flexibility of abnormal events, and can respond in a timely manner when events occur, reducing losses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119557822B_ABST
    Figure CN119557822B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides an abnormal event prediction method and device. The method includes: performing model training by receiving learning materials from a service platform to respectively obtain a face recognition model, an object recognition model, a behavior recognition model, and a voice recognition model; receiving environmental images captured by multiple cameras, inputting the environmental images into the face recognition model, the object recognition model, and the behavior recognition model to respectively determine abnormal image risk weight values corresponding to each camera, receiving environmental sounds captured by multiple cameras, inputting the environmental sounds into the voice recognition model to respectively determine abnormal sound risk weight values corresponding to each camera; receiving environmental noise captured by a noise source camera, if the environmental noise is greater than a preset noise threshold, calculating a comprehensive risk value of the noise source camera, and generating a corresponding warning message according to the judgment result of the comprehensive risk value. The present application can improve the response efficiency and flexibility of abnormal events.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, and particularly to a method and device for predicting abnormal events. Background Art

[0002] In social public enclosed places, such as buses, trains, railway stations, etc., abnormal problems occur from time to time, especially public security problems and some behaviors that endanger public safety. These problems not only affect public order but also may pose a threat to public safety. Common abnormal problems include quarrels, fights, brawls caused by disputes, as well as potential safety hazards such as theft, carrying contraband, violence, and terrorist incidents.

[0003] To prevent and respond to these abnormal problems, existing technical means rely on manual monitoring and post-event processing. Cameras are installed in public places such as subways, railway stations, and bus carriages. Once the above abnormal events occur, post-event processing can be carried out based on the video pictures recorded by the cameras. However, this method faces problems such as slow response speed, limited coverage, and high labor costs, and cannot respond in a timely manner during the occurrence of abnormal events. The post-event processing method alone also expands the losses caused by abnormal events. For example, for theft, the best time to stop it is during the theft. If it cannot be monitored and stopped during the theft, a large amount of manpower and resources will be wasted during subsequent recovery of stolen goods. Another example is that for secret photography, if it cannot be discovered and processed in a timely manner during secret photography, and the subsequent secretly photographed images are uploaded to the network, it will also waste a large amount of public resources.

[0004] Therefore, the existing methods cannot meet the requirements of abnormal events for response efficiency, and there is an urgent need for a method to increase the response flexibility of abnormal events and improve the response efficiency of abnormal events. Summary of the Invention

[0005] In view of the problems in the prior art, this application provides a method and device for predicting abnormal events, which can improve the response efficiency and flexibility of abnormal events.

[0006] To solve at least one of the above problems, this application provides the following technical solutions:

[0007] In a first aspect, this application provides a method for predicting abnormal events, including:

[0008] Receiving learning materials from the service platform, performing model training operations according to the learning materials, and respectively determining corresponding face recognition models, item recognition models, behavior recognition models, and voice recognition models, where the learning materials include face recognition pictures, item recognition pictures, behavior recognition pictures, and voice corpus information;

[0009] Receive the environmental images captured by the multiple cameras, input the environmental images into the face recognition model, the item recognition model, and the behavior recognition model, respectively determine the face attribute information, item type information, and behavior pattern information corresponding to each camera, perform a weighted calculation operation on the face attribute information, the item type information, and the behavior pattern information according to the preset abnormal event weights, and respectively determine the abnormal image risk weight values corresponding to each camera. Receive the environmental sounds captured by the multiple cameras, input the environmental sounds into the sound recognition model, perform a matching operation on the recognition result output by the sound recognition model and the preset abnormal negative vocabulary corpus, and perform a weighted calculation operation according to the negative semantics and the preset semantic weights obtained after the matching operation to respectively determine the abnormal sound risk weight values corresponding to each camera;

[0010] Receive the environmental noise captured by the first camera, determine whether the environmental noise is greater than the preset noise threshold. If it is greater, perform a weight accumulation operation according to the abnormal image risk weight value corresponding to the first camera and the abnormal sound risk weight value corresponding to the first camera to determine the corresponding first comprehensive risk value. Determine whether the first comprehensive risk value is greater than the preset first risk threshold, and generate the corresponding warning information according to the judgment result after the judgment operation, and send the warning information to the service platform so that the service platform performs an abnormal response according to the warning information, where the warning information includes alarm information, early warning information, and prompt information, and the first camera is the camera closest to the noise source among the multiple cameras.

[0011] Further, the method further includes:

[0012] According to the preset time interval, poll and determine whether the comprehensive risk value corresponding to each camera is greater than the preset second risk threshold, where the comprehensive risk value corresponding to each camera is determined by performing a weight accumulation operation according to the abnormal image risk weight value corresponding to each camera and the abnormal sound risk weight value corresponding to each camera;

[0013] If it is greater, perform an alarm information generation operation and send the alarm information to the service platform so that the service platform performs an abnormal response according to the alarm information;

[0014] If it is not greater, perform an accumulation operation on the comprehensive risk values corresponding to each camera, determine whether the comprehensive risk value corresponding to all cameras obtained after the accumulation operation is greater than the preset third risk threshold. If it is greater, perform an alarm information generation operation and send the alarm information to the service platform so that the service platform performs an abnormal response according to the alarm information

[0015] Further, the model training operation is performed according to the learning materials to respectively determine the corresponding face recognition model, object recognition model, behavior recognition model, and voice recognition model, including:

[0016] Perform a model training operation on the preset Fast RCNN network according to the face recognition pictures in the learning materials to determine the corresponding face recognition model. Among them, the face recognition model includes a face detection module, a face normalization module, a face attribute recognition module, a face quality assessment module, and a face tracking module. The face recognition model is used to recognize faces and analyze the negative emotions of people to obtain face attribute information;

[0017] Perform a model training operation on the preset convolutional neural network according to the object recognition pictures in the learning materials to determine the corresponding object recognition model. Among them, the object recognition model is used to recognize objects that endanger environmental safety;

[0018] Perform a model training operation on the preset Crowd Counting algorithm according to the behavior recognition pictures in the learning materials to determine the corresponding behavior recognition model. Among them, the behavior recognition model is used to detect crowd gatherings and behaviors that endanger environmental safety;

[0019] Perform a model training operation on the preset semantic recognition initial model according to the voice corpus information in the learning materials to determine the corresponding voice recognition model. Among them, the voice recognition model is used to convert voices into semantic information.

[0020] Further, the weighted calculation operation is performed on the face attribute information, the object type information, and the behavior pattern information according to the preset abnormal event weights to respectively determine the abnormal image risk weight values corresponding to each camera, including:

[0021] Perform a weight assignment operation on the face attribute information, the object type information, and the behavior pattern information according to the preset abnormal event weights to determine the face attribute information weight, the object type information weight, and the behavior pattern information weight corresponding to each camera;

[0022] Perform a weight accumulation operation according to the face attribute information weight, the object type information weight, and the behavior pattern information weight to respectively determine the abnormal image risk weight values corresponding to each camera.

[0023] Further, the matching operation is performed on the recognition results output by the voice recognition model and the preset abnormal negative vocabulary corpus, and the weighted calculation operation is performed according to the negative semantics obtained after the matching operation and the preset semantic weights to respectively determine the abnormal voice risk weight values corresponding to each camera, including:

[0024] Perform word segmentation on the recognition result output by the voice recognition model according to the preset word segmentation algorithm to determine the corresponding segmented words;

[0025] Match the segmented words with a preset abnormal negative word corpus, and perform weighted calculation according to the negative semantics and preset semantic weights obtained after the matching operation to respectively determine the abnormal sound risk weight values corresponding to each camera.

[0026] Further, determine whether the environmental noise is greater than a preset noise threshold. If it is greater, perform a weight accumulation operation according to the abnormal image risk weight value corresponding to the first camera and the abnormal sound risk weight value corresponding to the first camera to determine the corresponding first comprehensive risk value, including:

[0027] Determine whether the environmental noise is greater than a preset noise threshold. If it is greater, determine whether the abnormal image risk weight value corresponding to the first camera and the abnormal sound risk weight value corresponding to the first camera exist simultaneously;

[0028] If they exist simultaneously, perform a weight accumulation operation on the abnormal image risk weight value corresponding to the first camera and the abnormal sound risk weight value corresponding to the first camera, and perform a multiplication operation on the accumulation result obtained after the weight accumulation operation to determine the corresponding first comprehensive risk value;

[0029] If they do not exist simultaneously, perform a weight accumulation operation on the abnormal image risk weight value corresponding to the first camera and the abnormal sound risk weight value corresponding to the first camera to determine the corresponding first comprehensive risk value.

[0030] Further, determine whether the first comprehensive risk value is greater than a preset first risk threshold, and generate a corresponding warning message according to the judgment result after the judgment operation, including:

[0031] Determine whether the first comprehensive risk value is greater than a preset first risk threshold. If it is greater, perform a warning message generation operation to determine the corresponding alarm message;

[0032] If it is less, determine whether it is greater than half of the preset first risk threshold. If it is greater, perform a warning message generation operation to determine the corresponding early warning message. If it is less, perform a warning message generation operation to determine the corresponding prompt message.

[0033] In a second aspect, the present application provides an abnormal event prediction device, including:

[0034] A multi-model training module, configured to receive learning materials from a business platform, perform model training operations according to the learning materials, and respectively determine corresponding face recognition models, item recognition models, behavior recognition models, and voice recognition models, wherein the learning materials include face recognition pictures, item recognition pictures, behavior recognition pictures, and voice corpus information;

[0035] A risk weight calculation module, configured to receive environmental images captured by multiple cameras, input the environmental images into the face recognition model, the item recognition model, and the behavior recognition model, respectively determine face attribute information, item type information, and behavior pattern information corresponding to each camera, perform weighted calculation operations on the face attribute information, the item type information, and the behavior pattern information according to preset abnormal event weights, respectively determine abnormal image risk weight values corresponding to each camera, receive environmental sounds captured by multiple cameras, input the environmental sounds into the voice recognition model, perform matching operations according to the recognition results output by the voice recognition model and a preset abnormal negative vocabulary corpus, and perform weighted calculation operations according to the negative semantics obtained after the matching operations and preset semantic weights, respectively determine abnormal sound risk weight values corresponding to each camera;

[0036] A warning information generation module, configured to receive the environmental noise captured by a first camera, determine whether the environmental noise is greater than a preset noise threshold, if it is greater, perform a weight accumulation operation according to the abnormal image risk weight value corresponding to the first camera and the abnormal sound risk weight value corresponding to the first camera, determine a corresponding first comprehensive risk value, determine whether the first comprehensive risk value is greater than a preset first risk threshold, and generate a corresponding warning information according to the judgment result after the judgment operation, and send the warning information to the business platform, so that the business platform performs an abnormal response according to the warning information, wherein the warning information includes alarm information, early warning information, and prompt information, and the first camera is the camera closest to the noise source among the multiple preset cameras.

[0037] In a third aspect, the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, the steps of the abnormal event prediction method are implemented.

[0038] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the abnormal event prediction method are implemented.

[0039] Fifth aspect, the present application provides a computer program product, including computer programs / instructions, which when executed by a processor implement the steps of the abnormal event prediction method described above.

[0040] As can be seen from the above technical solutions, the present application provides an abnormal event prediction method and device. By receiving learning materials from the service platform for model training, a face recognition model, an item recognition model, a behavior recognition model, and a voice recognition model are respectively determined; environmental images captured by multiple cameras are received, and the environmental images are input into the face recognition model, the item recognition model, and the behavior recognition model to respectively determine the abnormal image risk weight values corresponding to each camera. The environmental sounds captured by multiple cameras are received, and the environmental sounds are input into the voice recognition model to respectively determine the abnormal sound risk weight values corresponding to each camera; the environmental noise captured by the noise source camera is received. If the environmental noise is greater than the preset noise threshold, the comprehensive risk value of the noise source camera is calculated, and a corresponding warning message is generated according to the judgment result of the comprehensive risk value, thereby improving the response efficiency and flexibility of abnormal events. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0042] Figure 1 is one of the flow diagrams of the abnormal event prediction method in the embodiments of the present application;

[0043] Figure 2 is another flow diagram of the abnormal event prediction method in the embodiments of the present application;

[0044] Figure 3 is yet another flow diagram of the abnormal event prediction method in the embodiments of the present application;

[0045] Figure 4 is still another flow diagram of the abnormal event prediction method in the embodiments of the present application;

[0046] Figure 5 is one of the flow diagrams of the abnormal event prediction method in the embodiments of the present application;

[0047] Figure 6 is another flow diagram of the abnormal event prediction method in the embodiments of the present application;

[0048] Figure 7It is the seventh flowchart schematic diagram of the abnormal event prediction method in the embodiments of the present application;

[0049] Figure 8 It is the structural diagram of the abnormal event prediction device in the embodiments of the present application;

[0050] Figure 9 It is the structural schematic diagram of the electronic device in the embodiments of the present application.

[0051] Reference numerals:

[0052] Electronic device 9600, central processing unit 9100, memory 9140, communication module 9110, input unit 9120, audio processor 9130, display 9160, power supply 9170, buffer memory 9141, application / function storage unit 9142, data storage unit 9143, driver program storage unit 9144, antenna 9111, speaker 9131, microphone 9132. Detailed implementation manners

[0053] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0054] The acquisition, storage, use, processing, etc. of data in the technical solutions of the present application all comply with the relevant regulations of national laws and regulations.

[0055] Considering the problem that the existing abnormal event handling methods cannot meet the requirements for efficiency and flexibility. The present application provides an abnormal event prediction method and device. By receiving learning materials from the service platform for model training, a face recognition model, an item recognition model, a behavior recognition model, and a voice recognition model are respectively determined; environmental images captured by multiple cameras are received, and the environmental images are input into the face recognition model, the item recognition model, and the behavior recognition model to respectively determine the abnormal image risk weight values corresponding to each camera. The environmental sounds captured by multiple cameras are received, and the environmental sounds are input into the voice recognition model to respectively determine the abnormal sound risk weight values corresponding to each camera; the environmental noise captured by the noise source camera is received. If the environmental noise is greater than the preset noise threshold, the comprehensive risk value of the noise source camera is calculated, and a corresponding warning message is generated according to the judgment result of the comprehensive risk value, thereby being able to improve the response efficiency and flexibility of abnormal events.

[0056] To improve the response efficiency and flexibility of abnormal events, this application provides an embodiment of a method for predicting abnormal events. Refer to Figure 1 The method for predicting abnormal events specifically includes the following content:

[0057] Step S101: Receive the learning materials of the service platform, perform model training operations according to the learning materials, and respectively determine the corresponding face recognition model, item recognition model, behavior recognition model, and voice recognition model. Among them, the learning materials include face recognition pictures, item recognition pictures, behavior recognition pictures, and voice corpus information;

[0058] Optionally, in this embodiment, the purpose of this step is to configure a pre-trained algorithm model on the edge computing platform. The edge computing platform receives the algorithm adjustment pictures sent by the service platform for self-learning to identify abnormal behaviors.

[0059] Optionally, in this embodiment, in this step, the learning materials of the service platform include two categories: image materials and voice materials. It can be understood that the learning materials of the service platform are also the training sets used by the edge computing platform to train models.

[0060] Specifically, the image materials are divided into face images, item images, and behavior images.

[0061] More specifically, for face images, mainly based on face recognition, on the basis of the images for training face detection, image labels of face attributes are added, including analyzing face expressions (such as fear, anger, anxiety, etc.), gender, age, glasses (sunglasses), masks, hats, etc.

[0062] More specifically, for item images, they mainly refer to items that endanger environmental safety, especially public security, such as controlled knives, guns, etc. In the item images, the category and position of each prohibited item need to be clearly pointed out (for example, the position box of the gun or the bounding box of the controlled knife).

[0063] More specifically, for behavior images, a behavior feature library is used. The behavior feature library contains a variety of labeled bad behavior data, which is used to compare the actions and scenes in the images, such as theft behavior, voyeurism behavior, etc.

[0064] More specifically, for voice materials, an interface is used to dock with the speech recognition training set of iFlytek, which is used to convert speech into text and lay a foundation for subsequent negative word analysis.

[0065] Optionally, in this embodiment, the above learning materials are used to train the corresponding models respectively.

[0066] Specifically, a face recognition algorithm Fast RCNN is trained with face images. This algorithm consists of a face detection module, a face normalization module, a face attribute module, a face quality module, and a face tracking module.

[0067] More specifically, the face detection module first detects whether there is a face in the image. Through the features in the image, Fast RCNN can accurately extract the face region from the background, mark the position of the face in the image, and crop it out for subsequent processing. Secondly, the face normalization module normalizes the detected face image, adjusts the size of the image, ensures that the face is in a standardized coordinate system, reduces the influence of external factors such as different lighting and angles, and unifies it to a fixed size and ratio. This process helps to improve the accuracy of subsequent emotion analysis and attribute recognition. Next, the face attribute module performs attribute recognition on the normalized face image, uses a convolutional neural network (CNN) for multi-task learning, and simultaneously performs expression classification (such as fear, anger, anxiety), accessory recognition (such as sunglasses, masks, hats), etc., to analyze the attributes of the face. Then, based on the specific results output by the face attribute module, the face quality module determines whether further image repair or discarding is required. Based on the image quality assessment model, it judges whether the image is clear, whether there is too much noise or blur. Finally, the face tracking module maintains the tracking of the target face based on the correlation between video frames, using a face tracking algorithm (such as the KLT tracking algorithm).

[0068] In terms of the effect, by analyzing the captured images and the recognized face attributes, the system can identify potential negative emotions, such as fear, anger, anxiety, etc., and the wearing conditions of sunglasses, masks, hats. Subsequently, by adding weights to these situations, the probability of abnormal events occurring can be evaluated.

[0069] Specifically, a convolutional neural network is trained with item images. Through image recognition technology, contraband (such as guns, controlled knives, etc.) is detected. A target detection model is trained using a dataset containing various contraband images (such as guns, knives, etc.), and the trained model annotates the contraband images to learn how to identify contraband items from a complex environment.

[0070] Specifically, in the behavior recognition model, the crowd gathering behavior is realized through the crowd counting algorithm. For behaviors such as stealing, secretly taking pictures, spitting, littering, and physical conflicts, on the basis of the traditional behavior feature library, to capture the dynamics of the behavior, a time dimension is added. By analyzing multiple images within 1 - 2 consecutive seconds (for example, stealing, secretly taking pictures, etc.), a "time series set" of behavior features is formed, thereby increasing the accurate recognition of behaviors.

[0071] For example, using camera data and combining with the crowd counting algorithm, the system can analyze the number of people in a region in real time. When the number of people exceeds the set threshold, the system will trigger a "crowd gathering" event. When a crowd gathering event is detected, the behavior analysis system will be activated, and the currently captured image set (a time series set containing multiple images) will be compared with the behavior images in the feature library. Using image similarity metrics (such as cosine similarity, Euclidean distance, etc.), it is determined whether the current behavior matches the known behaviors in the feature library. Based on the comparison results, the corresponding behavior types (such as theft, secret photography, etc.) are classified. If the behavior conforms to the pattern in the feature library, it is marked as a relevant behavior.

[0072] In terms of the effect, in the crowd gathering area, the behavior analysis will be more accurate. Utilizing the characteristics of the time dimension, the system can accurately identify specific behavior patterns, especially in the case of a large number of gathered people, to avoid missed detections.

[0073] Specifically, the speech recognition model directly interfaces with iFlytek. The speech file is transmitted to the call interface, and the recognition result is returned.

[0074] Step S102: Receive the environmental images captured by the multiple cameras, input the environmental images into the face recognition model, the item recognition model, and the behavior recognition model respectively, and determine the face attribute information, item type information, and behavior pattern information corresponding to each camera. Perform a weighted calculation operation on the face attribute information, the item type information, and the behavior pattern information according to the preset abnormal event weights, and respectively determine the abnormal image risk weight values corresponding to each camera. Receive the environmental sounds captured by the multiple cameras, input the environmental sounds into the sound recognition model, perform a matching operation according to the recognition result output by the sound recognition model and the preset abnormal negative vocabulary corpus, and perform a weighted calculation operation according to the negative semantics obtained after the matching operation and the preset semantic weights, and respectively determine the abnormal sound risk weight values corresponding to each camera.

[0075] Optionally, in this embodiment, the purpose of this step is that the edge computing platform receives the image and voice information captured by the cameras in the environment, analyzes them through the models trained in the above step S101, obtains the analysis results, and assigns weights to the analysis results according to the probability of easy-occurring public security events, and finally obtains the image risk weight values and the sound risk weight values.

[0076] Optionally, in this embodiment, all the cameras in the environment transmit the environmental images to the edge computing platform, and the edge computing platform performs calculations simultaneously. The face recognition model, the item recognition model, and the behavior recognition model respectively calculate the image abnormal event prediction results of each camera output by them. Subsequently, one camera is taken as an example for illustration.

[0077] Specifically, the face recognition model outputs face attribute information, including emotional information (fear, anger, anxiety, etc.) and accessory information (hats, sunglasses, masks, etc.).

[0078] Specifically, the object recognition model outputs object type information, including (guns, controlled knives, etc.).

[0079] Specifically, the behavior recognition model outputs behavior pattern information, including (secret photography, theft, robbery, etc.).

[0080] Optionally, after obtaining the prediction results output by the above model, weights are assigned to the above abnormal results according to the probability of abnormal events. The sub-items of each major item (such as facial expressions, hats, sunglasses, and masks in facial attribute analysis) have weight ratios. The weight of the major item is calculated based on the weights of the sub-items, and then the major items also have weight ratios, and the comprehensive weight is calculated based on the weight ratios between the major items.

[0081] For example:

[0082] (1) Facial attribute analysis, identifying specific emotions (such as fear, anger, and anxiety) and assigning corresponding weight values ​​to each emotion.

[0083] Negative emotions: Identify fearful expressions (weight 1), angry expressions (weight 2), and anxious expressions (weight 1).

[0084] Wearing condition: Identify wearing sunglasses (weight 0.5), wearing a mask (weight 0.5), and wearing a hat (weight 0.5).

[0085] Assume that in the captured image, there is a person showing an angry expression and wearing a mask at the same time. Then, the comprehensive weight of the facial attributes of this person is calculated as follows:

[0086] Angry emoticon weight: 2

[0087] Mask wearing weight: 0.5

[0088] The combined weight of facial attributes = 2 (anger) + 0.5 (mask) = 2.5

[0089] (2) Item type analysis

[0090] Identify guns (weight 10) and controlled knives (weight 5).

[0091] Assuming that no contraband is found in the captured image, then the comprehensive weight of contraband is:

[0092] Contraband Weight: 0

[0093] (3) Behavior recognition analysis

[0094] Through the crowd counting algorithm, it is identified that the number of people gathering exceeds a certain threshold (assuming the threshold is 5 people and the weight is 2).

[0095] Suppose there are 6 people gathering in the front door area in the captured picture. Then, the comprehensive weight of people gathering is as follows:

[0096] Weight of people gathering: 2

[0097] Stealing behavior (weight 3), voyeurism behavior (weight 3), and robbery behavior (weight 5) are identified.

[0098] Suppose there is a person stealing in the captured picture. Then, the comprehensive weight of the behavior action is: Weight of stealing behavior: 3

[0099] In this way, the weight of the large item of face attributes is 2.5, the weight of the large item of item type is 0, the weight of the large item of behavior recognition is 5, and the risk weight value of the abnormal image is 7.5.

[0100] Optionally, in this embodiment, all cameras in the environment transmit the environmental sound to the edge computing platform, and the edge computing platform performs calculations simultaneously. The sound recognition model calculates the prediction results of the sound abnormal event of each camera output respectively. Subsequently, one camera is used as an example for illustration.

[0101] Specifically, a third-party speech recognition platform (such as iFlytek) is used to convert the speech into text, and the converted text is matched with the negative emotion word library to identify whether it contains emotionally intense words (such as "kill", "curse people", etc.). The emotionally intense words contained in the text are classified, such as anger, threat, intimidation, etc.

[0102] After identifying the negative emotion words, one way is to give different risk weights according to the emotion type, such as anger: weight 2, threat / violence: weight 5. Another way is to assign a weight of 1 to each negative word. Finally, the total risk weight value of the abnormal sound is the accumulation of the sound risks of each item.

[0103] For example,

[0104] According to the first method, in the environmental scenario, the word "violence" appears. No matter how many times it appears, it is counted as a weight value of 5, and no other negative words appear, so the weight value is 0. The risk weight value of the abnormal sound is 5 + 0 = 5.

[0105] According to the second method, in the environmental scenario, assume that within the recent 10 seconds, the sound recognition system recognizes 5 negative words. Then, regardless of the types of negative words, the comprehensive weight of the sound is 5.

[0106] After the above-mentioned step S102, we calculated the abnormal image risk weight value and abnormal sound risk weight value of each camera through the edge computing platform, laying a foundation for timely response to risks in the follow-up.

[0107] Step S103: Receive the environmental noise recorded by the first camera, determine whether the environmental noise is greater than a preset noise threshold. If it is greater, perform a weight accumulation operation according to the abnormal image risk weight value corresponding to the first camera and the abnormal sound risk weight value corresponding to the first camera, determine the corresponding first comprehensive risk value, determine whether the first comprehensive risk value is greater than a preset first risk threshold, and generate a corresponding warning message according to the judgment result after the judgment operation, and send the warning message to the service platform, so that the service platform performs an abnormal response according to the warning message. Among them, the warning message includes an alarm message, a warning message, and a prompt message, and the first camera is the camera closest to the noise source among the multiple cameras.

[0108] Optionally, in this embodiment, the purpose of this step is to construct a risk response algorithm for timely responding to environmental abnormal risks.

[0109] Specifically, generally, abnormal safety incidents usually start with a quarrel and then escalate to a fight. Thus, a method is defined. Based on the noise sensor configured on the camera, the environmental noise is detected. The edge computing platform receives the environmental noise and determines whether the noise is greater than 60 decibels (the sound of a quarrel is greater than 60 decibels). If it is greater, first locate the camera closest to the noise and obtain the abnormal image risk weight value and abnormal sound risk weight value of this camera within the last 10 seconds (at this time, the edge computing platform has already calculated the abnormal image risk weight value and abnormal sound risk weight value of each camera). At the same time, perform a judgment. If there is both a voice risk weight value X and a sound risk weight value Y, then calculate the weight as (X + Y) 2. Because there are both sound risks and image recognition risks, the risk must be high, so a multiple processing is done. If there is only an image risk or only a sound risk, then the comprehensive risk weight is (X + Y).

[0110] If the accumulated value of the comprehensive risk weight is greater than 10, generate an alarm and send it to the service platform, including alarm information (alarm pictures, alarm information, alarm points).

[0111] If the accumulated value of the comprehensive risk weight is less than 10 and greater than 5, generate a warning message and send it to the service platform, including warning information (warning pictures, warning information, warning points).

[0112] If the comprehensive risk weight is less than 5, send a prompt voice to this camera (eg: Keep quiet in public places, etc.).

[0113] For example, suppose there is an edge computing box on a bus. Two cameras are installed on the bus, one facing the interior of the front door and the other facing the interior of the rear door. The edge box retrieves the video stream from the cameras via RTSP (Real-Time Streaming Protocol) and captures one frame every 1 second.

[0114] First, the edge box analyzes the face attributes, item types, and behavior patterns of each second of video images in real time, and accumulates the identified sub-categories with weights (such as the sub-categories illustrated in step S102) to obtain the corresponding major-category weight values. The above three major-category weight values are accumulated to obtain the abnormal image risk weight value. The edge box performs semantic recognition on the sound in real time, and compares the recognized sentences with the negative vocabulary library. If negative words or semantics appear, the recognized risk types are accumulated according to a predetermined weight value accumulation method to obtain the abnormal sound risk weight value.

[0115] Then, when an abnormal environmental event occurs, there is usually noise. The camera noise sensor identifies the noise, and the edge box obtains the camera noise data, locates the nearest camera (assumed to be the front door), and accumulates the abnormal image risk weight value and the abnormal sound risk weight value of the front door camera in the edge box. If both have numerical values, the accumulated value is processed by a multiple to obtain the comprehensive risk weight value. It is judged whether the comprehensive risk weight value is greater than 10. If it is greater than 10, an alarm message is generated and sent to the service platform. If the weight accumulation value is less than 10 and greater than 5, a warning is generated and sent to the platform. If the weight is less than 5 and does not reach the warning threshold, the box sends a prompt voice file to this camera (eg: Keep quiet in public places. Harmony is precious among people. Classmate, being angry not only harms your body but also makes you age easily, etc.). If a person wants to report a law-breaking or public security event on the bus and shouts directly with a voice exceeding 60 decibels, the warning strategy can be triggered directly, which is convenient and efficient.

[0116] This example demonstrates how this embodiment constructs an edge computing platform by deploying a multi-functional model and fusing the weights of the predictions output by the model, through which the environmental anomalies can be quickly judged and the corresponding warning information can be sent, improving the efficiency of abnormal response.

[0117] As can be seen from the above description, the abnormal event prediction method provided by the embodiments of the present application can perform model training by receiving learning materials from the service platform, and respectively determine a face recognition model, an object recognition model, a behavior recognition model, and a voice recognition model; receive environmental images captured by multiple cameras, input the environmental images into the face recognition model, the object recognition model, and the behavior recognition model, and respectively determine the abnormal image risk weight values corresponding to each camera, receive the environmental sounds captured by multiple cameras, input the environmental sounds into the voice recognition model, and respectively determine the abnormal sound risk weight values corresponding to each camera; receive the environmental noise captured by the noise source camera, if the environmental noise is greater than the preset noise threshold, calculate the comprehensive risk value of the noise source camera, and generate a corresponding warning message according to the judgment result of the comprehensive risk value, thereby improving the response efficiency and flexibility of abnormal events.

[0118] In an embodiment of the abnormal event prediction method of the present application, refer to Figure 2 , it may specifically include the following content:

[0119] Step S201: Poll and judge whether the comprehensive risk value corresponding to each camera is greater than a preset second risk threshold at a preset time interval, where the comprehensive risk value corresponding to each camera is determined by performing a weight accumulation operation according to the abnormal image risk weight value corresponding to each camera and the abnormal sound risk weight value corresponding to each camera;

[0120] Step S202: If it is greater, perform an alarm information generation operation, and send the alarm information to the service platform, so that the service platform performs an abnormal response according to the alarm information;

[0121] Step S203: If it is not greater, perform an accumulation operation on the comprehensive risk value corresponding to each camera, and judge whether the comprehensive risk value corresponding to all cameras obtained after the accumulation operation is greater than a preset third risk threshold. If it is greater, perform an alarm information generation operation, and send the alarm information to the service platform, so that the service platform performs an abnormal response according to the alarm information.

[0122] Optionally, in this embodiment, this step is a supplementary step to the above step S103. Step S103 provides one way of environmental abnormal response, and this step is another way. When no noise is generated in the environmental abnormal event, we use a timed polling task to timely detect the environmental abnormal event.

[0123] Specifically, the edge box sets a timed polling task to scan the comprehensive risk weights of all cameras connected within 10 seconds (i.e., the cumulative value of the abnormal image risk weight value and the abnormal sound risk weight value). If the cumulative weight value of a single camera is greater than 10, a single-camera alarm is generated and sent to the platform. If it is less than 10, the cumulative weights of all cameras connected to the edge box are judged. If the weight value is greater than 5 N (where N represents the number of working cameras connected to this edge box), a combined alarm is generated and sent to the platform. If the weight value is less than 10 or 5 N, the edge box does not report warning information to the platform, but can directly send a prompt voice to the camera, such as "Please take good care of your belongings and be careful not to lose them" and so on.

[0124] For example, assume that a robbery occurs on a bus without any noise. At this time, the camera face detects a hat (0.5), a mask (0.5), the behavior detects robbery (5), the item detects a controlled knife (5), and the sound detects negative words (2). The comprehensive risk weight is 13. At this time, an alarm message for this camera will be generated, and the alarm picture, alarm information, and alarm location will be sent to the business platform. The platform can obtain the real-time location coordinates of this alarm camera (the camera installed on the bus or train will move in real time), which serves as a basis for the next response and handling.

[0125] For example, assume that a theft occurs on a bus without any noise. At this time, the front-door camera face detects a hat (0.5), a mask (0.5), the behavior detects theft (3), and the comprehensive risk weight is 4. At this time, the edge computing platform calculates the cumulative weight value of the two cameras at the front door and the back door. Assume that the back-door camera detects an angry expression (2) and the comprehensive risk weight is 2. Then the cumulative weight value is 6, and 6 < 10 (5 2). At this time, the edge computing platform does not interact with the business platform, but through the built-in voice file sent to the camera (the camera closest to this behavior for capture) intercom function, plays a pre-prepared sound file ("Please take good care of your belongings and be careful not to lose them", "Pay attention to safety when going out" and so on). If a person wants to report an illegal or public security incident on a bus, just shout a voice exceeding 60 decibels to trigger the warning strategy, which is convenient and efficient. At this time, the real-time location of the vehicle will be uploaded to the platform, laying a foundation for subsequent timely handling.

[0126] Through step S203, this embodiment supplements another environmental anomaly response method, increasing the flexibility of the response to environmental anomalies.

[0127] In an embodiment of the abnormal event prediction method of the present application, see Figure 3 , it may specifically include the following content:

[0128] Step S301: Perform model training operations on a preset Fast RCNN network according to the face recognition pictures in the learning materials to determine the corresponding face recognition model. Among them, the face recognition model includes a face detection module, a face normalization module, a face attribute recognition module, a face quality assessment module, and a face tracking module. The face recognition model is used to recognize faces and analyze the negative emotions of people to obtain face attribute information;

[0129] Step S302: Perform model training operations on a preset convolutional neural network according to the object recognition pictures in the learning materials to determine the corresponding object recognition model. Among them, the object recognition model is used to recognize objects that endanger environmental safety;

[0130] Step S303: Perform model training operations on a preset Crowd Counting algorithm according to the behavior recognition pictures in the learning materials to determine the corresponding behavior recognition model. Among them, the behavior recognition model is used to detect crowd gathering and behaviors that endanger environmental safety;

[0131] Step S304: Perform model training operations on a preset initial semantic recognition model according to the voice corpus information in the learning materials to determine the corresponding voice recognition model. Among them, the voice recognition model is used to convert voices into semantic information.

[0132] Optionally, the face recognition model uses the Fast RCNN algorithm and consists of a face detection module, a face normalization module, a face attribute module, a face quality module, and a face tracking module.

[0133] More specifically, the face detection module first detects whether there is a face in the image. Through the features in the image, Fast RCNN can accurately extract the face region from the background, calibrate the face position in the image, and crop it for subsequent processing. Secondly, the face normalization module normalizes the detected face image, adjusts the image size, ensures that the face is in a standardized coordinate system, reduces the influence of external factors such as different lighting and angles, and unifies it to a fixed size and ratio. This process helps to improve the accuracy of subsequent emotion analysis and attribute recognition. Next, the face attribute module performs attribute recognition on the normalized face image, uses a convolutional neural network (CNN) for multi-task learning, and simultaneously performs expression classification (such as fear, anger, anxiety), accessory recognition (such as sunglasses, masks, hats), etc., to analyze the attributes of the face. Then, based on the specific results output by the face attribute module, the face quality module determines whether further image repair or discarding is required. Based on the image quality assessment model, it judges whether the image is clear, whether there is too much noise or blur. Finally, the face tracking module, based on the correlation between video frames, uses a face tracking algorithm (such as the KLT tracking algorithm) to keep tracking the target face.

[0134] In terms of effects, by analyzing the captured images and the recognized face attributes, the system can identify potential negative emotions such as fear, anger, anxiety, etc., and the wearing conditions of sunglasses, masks, and hats. Subsequently, by adding weights to these situations, the probability of abnormal events occurring can be evaluated.

[0135] Optionally, for the item recognition model, the training set contains a large number of images labeled with contraband (such as guns, controlled knives, etc.). The platform or a third-party platform will send these high-quality images through an interface for learning.

[0136] Use a convolutional neural network (CNN) model to perform object detection on the image and train the model to recognize different types of contraband. During the training process, the model optimizes the parameters by minimizing the loss function, enabling the model to more accurately identify the objects in the image. The loss function usually includes classification loss and localization loss. Use the validation set to test the model and evaluate its accuracy, recall rate, and F1 value to ensure that the model can effectively identify contraband in the actual environment. After the model is trained and deployed, the edge box will analyze the captured image to identify whether there is contraband.

[0137] Optionally, for the behavior recognition model, the crowd gathering behavior is implemented through the crowd counting algorithm. For behaviors such as theft, secretly taking pictures, spitting, littering, and physical conflicts, based on the traditional behavior feature library, in order to capture the dynamics of the behavior, the time dimension is added. By analyzing multiple images within 1-2 consecutive seconds (for example, theft, spitting, etc.), a "time series set" of behavior features is formed, thereby increasing the accurate recognition of behaviors. In terms of effect, in the area where people gather, the behavior analysis will be more accurate. Utilizing the characteristics of the time dimension, the system can accurately identify specific behavior patterns, especially in the case of a large number of gathered people, to avoid missed detections.

[0138] Optionally, for the voice recognition model, it directly interfaces with iFlytek, transmits the voice file to the call interface, and returns the recognition result.

[0139] Through step S304, this embodiment respectively constructs four intelligent models for face detection, behavior detection, item detection, and voice detection, laying a foundation for subsequent identification of environmental anomalies.

[0140] In an embodiment of the abnormal event prediction method of this application, refer to Figure 4 and it may specifically include the following content:

[0141] Step S401: Perform a weight assignment operation on the face attribute information, the item type information, and the behavior pattern information according to the preset abnormal event weights, and determine the face attribute information weight, the item type information weight, and the behavior pattern information weight corresponding to each camera;

[0142] Step S402: Perform a weight accumulation operation according to the face attribute information weight, the item type information weight, and the behavior pattern information weight, and respectively determine the abnormal image risk weight value corresponding to each camera.

[0143] Optionally, in this embodiment, the abnormal results are weighted according to the probability of easily occurring abnormal events. Each sub-item of each major item (for example, in the face attribute analysis, the expression, hat, sunglasses, and mask of the face) has a weight ratio. The weight of the major item is calculated according to the sub-item weights, and then there is also a weight ratio between the major items. The comprehensive weight is calculated according to the weight ratio between the major items.

[0144] Illustrate with examples:

[0145] (1) For face attribute analysis, identify specific emotions (such as fear, anger, anxiety), and assign corresponding weight values to each emotion.

[0146] Negative emotions: Identify a fearful expression (weight 1), an angry expression (weight 2), and an anxious expression (weight 1).

[0147] Wearing situation: It is recognized that sunglasses are worn (weight 0.5), a mask is worn (weight 0.5), and a hat is worn (weight 0.5).

[0148] Suppose in the captured picture, a person shows an angry expression and wears a mask at the same time. Then, the comprehensive weight of this facial attribute is calculated as follows:

[0149] Weight of angry expression: 2

[0150] Weight of mask wearing: 0.5

[0151] Comprehensive weight of facial attribute = 2 (angry) + 0.5 (mask) = 2.5

[0152] (2) Analysis of item types

[0153] It is recognized that a firearm (weight 10) and a controlled knife (weight 5) are present.

[0154] Suppose no contraband is found in the captured picture. Then, the comprehensive weight of contraband is:

[0155] Weight of contraband: 0

[0156] (3) Behavior recognition analysis

[0157] Through the crowd counting algorithm, it is recognized that the number of people gathering exceeds a certain threshold (assuming the threshold is 5 people, weight 2).

[0158] Suppose in the captured picture, 6 people are gathering in the front door area. Then, the comprehensive weight of people gathering is:

[0159] Weight of people gathering: 2

[0160] It is recognized that a theft behavior (weight 3), a voyeurism behavior (weight 3), and a robbery behavior (weight 5) are present.

[0161] Suppose in the captured picture, a person is performing a theft behavior. Then, the comprehensive weight of the behavior action is: Weight of theft behavior: 3

[0162] In this way, the weight of the major item of facial attributes is 2.5, the weight of the major item of item types is 0, the weight of the major item of behavior recognition is 5, and the risk weight value of the abnormal image is 7.5.

[0163] Through step S402, this embodiment successfully calculates the risk weight value of the abnormal image for each camera, laying a solid foundation for calculating the comprehensive risk weight and identifying environmental abnormalities in a timely manner.

[0164] In an embodiment of the abnormal event prediction method of this application, refer to Figure 5, it may specifically include the following content:

[0165] Step S501: Perform a word segmentation operation on the recognition result output by the voice recognition model according to a preset word segmentation algorithm to determine the corresponding segmented words;

[0166] Step S502: Match the segmented words with a preset abnormal negative word corpus, and perform a weighted calculation operation according to the negative semantics and preset semantic weights obtained after the matching operation to respectively determine the abnormal sound risk weight values corresponding to each camera.

[0167] Optionally, in this embodiment, a third-party speech recognition platform (such as iFlytek) is used to convert speech into text, and the converted text is matched with a negative emotion word library to identify whether it contains emotionally intense words (such as "kill", "scold people", etc.). Classify the emotionally intense words included in the text, such as anger, threat, intimidation, etc.

[0168] After identifying negative emotion words, one way is to assign different risk weights according to the emotion type, such as anger: weight 2, threat / violence: weight 5. Another way is to assign a weight of 1 to each negative word. Finally, the total abnormal sound risk weight value is the accumulation of the sound risk of each item.

[0169] For example,

[0170] According to the first method, in the environmental scenario, words related to "violence" appear. No matter how many times they appear, the weight value is counted as 5. No other negative words appear, and the weight value is 0. The abnormal sound risk weight value is 5 + 0 = 5.

[0171] According to the second method, in the environmental scenario, assume that within the last 10 seconds, the voice recognition system recognizes 5 negative words. Then, regardless of the type of negative words, the comprehensive weight of the sound is 5.

[0172] Through step S502, this embodiment successfully calculates the abnormal sound risk weight values of each camera, laying a solid foundation for subsequent calculation of the comprehensive risk weight and timely response to identify environmental abnormalities.

[0173] In an embodiment of the abnormal event prediction method of the present application, refer to Figure 6 , it may specifically include the following content:

[0174] Step S601: Determine whether the environmental noise is greater than a preset noise threshold. If it is greater, determine whether the abnormal image risk weight value corresponding to the first camera and the abnormal sound risk weight value corresponding to the first camera exist simultaneously;

[0175] Step S602: If both exist, perform a weight accumulation operation on the abnormal image risk weight value corresponding to the first camera and the abnormal sound risk weight value corresponding to the first camera, and perform a multiplication operation on the accumulation result obtained after the weight accumulation operation to determine the corresponding first comprehensive risk value;

[0176] Step S603: If they do not exist simultaneously, perform a weight accumulation operation on the abnormal image risk weight value corresponding to the first camera and the abnormal sound risk weight value corresponding to the first camera to determine the corresponding first comprehensive risk value.

[0177] Optionally, in this embodiment, a noise sensor configured based on the camera is used to detect the ambient noise. The edge computing platform receives the ambient noise and determines whether the noise is greater than 60 decibels (the sound of arguing is greater than 60 decibels). If it is greater, first locate the camera closest to the noise, and obtain the abnormal image risk weight value and abnormal sound risk weight value of this camera within the most recent 10 seconds (at this time, the edge computing platform has already calculated the abnormal image risk weight value and abnormal sound risk weight value of each camera), and at the same time make a judgment. If there is both an image risk weight value X and a sound risk weight value Y, then calculate the weight as (X + Y) 2. Because there is both a sound risk and an image recognition risk, the risk must be high, so a multiplication process is performed. If there is only an image risk X or only a sound risk Y, then the comprehensive risk weight is (X + Y).

[0178] Through step S603, this embodiment successfully obtains the comprehensive risk weight, laying a foundation for subsequent risk response based on this comprehensive risk weight value.

[0179] In an embodiment of the abnormal event prediction method of the present application, refer to Figure 7 and it may specifically include the following content:

[0180] Step S701: Determine whether the first comprehensive risk value is greater than a preset first risk threshold. If it is greater, perform an operation to generate a warning message to determine the corresponding alarm message;

[0181] Step S702: If it is less, determine whether it is greater than half of the preset first risk threshold. If it is greater, perform an operation to generate a warning message to determine the corresponding early warning message. If it is less, perform an operation to generate a warning message to determine the corresponding prompt message.

[0182] Optionally, in this embodiment, the comprehensive risk value in step S603 above is judged and corresponding measures are defined.

[0183] Specifically, if the comprehensive risk value is greater than 10, an alarm is generated and sent to the service platform, including alarm information (alarm pictures, alarm messages, alarm points).

[0184] Specifically, if the comprehensive risk value is less than 10 and greater than 5, warning information (warning pictures, warning information, warning points) will be generated and sent to the service platform.

[0185] Specifically, if the comprehensive risk value is less than 5, a prompt voice will be sent to the camera (eg: keep quiet in public places, etc.).

[0186] Through step S702, this embodiment successfully makes a real-time response according to the comprehensive environmental risk, improving the efficiency of abnormal event response.

[0187] To improve the response efficiency and flexibility of abnormal events, this application provides an embodiment of an abnormal event prediction device for implementing all or part of the content of the abnormal event prediction method. See Figure 8 The abnormal event prediction device specifically includes the following content:

[0188] The multi-model training module 10 is used to receive the learning materials of the service platform, perform model training operations according to the learning materials, and respectively determine the corresponding face recognition model, object recognition model, behavior recognition model, and voice recognition model. Among them, the learning materials include face recognition pictures, object recognition pictures, behavior recognition pictures, and voice corpus information;

[0189] The risk weight calculation module 20 is used to receive the environmental images captured by the multiple cameras, input the environmental images into the face recognition model, the object recognition model, and the behavior recognition model, respectively determine the face attribute information, object type information, and behavior pattern information corresponding to each camera, perform weighted calculation operations on the face attribute information, the object type information, and the behavior pattern information according to the preset abnormal event weights, and respectively determine the abnormal image risk weight values corresponding to each camera. Receive the environmental sounds captured by the multiple cameras, input the environmental sounds into the voice recognition model, match the recognition results output by the voice recognition model with the preset abnormal negative vocabulary corpus, and perform weighted calculation operations according to the negative semantics obtained after the matching operation and the preset semantic weights to respectively determine the abnormal sound risk weight values corresponding to each camera;

[0190] The warning information generation module 30 is configured to receive the environmental noise recorded by the first camera, determine whether the environmental noise is greater than a preset noise threshold. If it is greater, a weight accumulation operation is performed according to the abnormal image risk weight value corresponding to the first camera and the abnormal sound risk weight value corresponding to the first camera to determine the corresponding first comprehensive risk value. It is determined whether the first comprehensive risk value is greater than a preset first risk threshold, and a corresponding warning information is generated according to the determination result after the determination operation, and the warning information is sent to the service platform so that the service platform performs an abnormal response according to the warning information. Wherein, the warning information includes an alarm message, a warning message, and a prompt message, and the first camera is the camera closest to the noise source among the multiple cameras.

[0191] As can be seen from the above description, the abnormal event prediction device provided by the embodiment of the present application can perform model training by receiving learning materials from the service platform, and respectively determine a face recognition model, an item recognition model, a behavior recognition model, and a sound recognition model; receive environmental images recorded by multiple cameras, and input the environmental images into the face recognition model, the item recognition model, and the behavior recognition model to respectively determine the abnormal image risk weight value corresponding to each camera. Receive the environmental sounds recorded by multiple cameras, input the environmental sounds into the sound recognition model, and respectively determine the abnormal sound risk weight value corresponding to each camera; receive the environmental noise recorded by the noise source camera. If the environmental noise is greater than the preset noise threshold, calculate the comprehensive risk value of the noise source camera, and generate a corresponding warning information according to the judgment result of the comprehensive risk value, thereby improving the response efficiency and flexibility of abnormal events.

[0192] From a hardware perspective, in order to improve the response efficiency and flexibility of abnormal events, the present application provides an embodiment of an electronic device for implementing all or part of the content in the abnormal event prediction method. The electronic device specifically includes the following content:

[0193] A processor, a memory, a communication interface, and a bus; wherein, the processor, the memory, and the communication interface complete mutual communication through the bus; the communication interface is used to implement information transmission between the abnormal event prediction method and related devices such as the core business system, the user terminal, and the related database. The logic controller can be a desktop computer, a tablet computer, a mobile terminal, etc., and this embodiment is not limited thereto. In this embodiment, the logic controller can be implemented with reference to the embodiments of the abnormal event prediction method and the embodiments of the abnormal event prediction method, and the content is incorporated herein, and the repeated parts are not described again.

[0194] It can be understood that the user terminal may include a smart phone, a tablet electronic device, a network set-top box, a portable computer, a desktop computer, a personal digital assistant (PDA), a vehicle-mounted device, a smart wearable device, etc. Among them, the smart wearable device may include smart glasses, a smart watch, a smart bracelet, etc.

[0195] In practical applications, part of the abnormal event prediction method may be executed on the electronic device side as described above, or all operations may be completed in the client device. Specifically, it can be selected according to the processing capacity of the client device and the limitations of the user usage scenario, etc. This application does not make a limitation in this regard. If all operations are completed in the client device, the client device may further include a processor.

[0196] The above-mentioned client device may have a communication module (i.e., a communication unit), and may be communicatively connected to a remote server to implement data transmission with the server. The server may include a server on the task scheduling center side, and in other implementation scenarios, it may also include a server of an intermediate platform, such as a server of a third-party server platform communicatively linked to the task scheduling center server. The server may include a single computer device, or may include a server cluster composed of multiple servers, or a server structure of a distributed device.

[0197] Figure 9 It is a schematic block diagram of the system composition of the electronic device 9600 according to an embodiment of the present application. As Figure 9 shown, the electronic device 9600 may include a central processing unit 9100 and a memory 9140; the memory 9140 is coupled to the central processing unit 9100. It should be noted that this Figure 9 is exemplary; other types of structures may also be used to supplement or replace this structure to implement telecommunication functions or other functions.

[0198] In one embodiment, the abnormal event prediction method function may be integrated into the central processing unit 9100. Among them, the central processing unit 9100 may be configured to perform the following controls:

[0199] Step S101: Receive the learning materials of the service platform, perform model training operations according to the learning materials, and respectively determine corresponding face recognition models, item recognition models, behavior recognition models, and voice recognition models, where the learning materials include face recognition pictures, item recognition pictures, behavior recognition pictures, and voice corpus information;

[0200] Step S102: Receive the environmental images captured by the multiple cameras, input the environmental images into the face recognition model, the item recognition model, and the behavior recognition model, respectively determine the face attribute information, item type information, and behavior pattern information corresponding to each camera, perform a weighted calculation operation on the face attribute information, the item type information, and the behavior pattern information according to the preset abnormal event weights, respectively determine the abnormal image risk weight values corresponding to each camera, receive the environmental sounds captured by the multiple cameras, input the environmental sounds into the sound recognition model, perform a matching operation according to the recognition results output by the sound recognition model and the preset abnormal negative vocabulary corpus, and perform a weighted calculation operation according to the negative semantics and preset semantic weights obtained after the matching operation, respectively determine the abnormal sound risk weight values corresponding to each camera;

[0201] Step S103: Receive the environmental noise captured by the first camera, determine whether the environmental noise is greater than the preset noise threshold. If it is greater, perform a weight accumulation operation according to the abnormal image risk weight value corresponding to the first camera and the abnormal sound risk weight value corresponding to the first camera, determine the corresponding first comprehensive risk value, determine whether the first comprehensive risk value is greater than the preset first risk threshold, and generate a corresponding warning message according to the judgment result after the judgment operation, and send the warning message to the service platform so that the service platform performs an abnormal response according to the warning message, where the warning message includes an alarm message, a warning message, and a prompt message, and the first camera is the camera closest to the noise source among the multiple cameras.

[0202] As can be seen from the above description, the electronic device provided in the embodiment of the present application performs model training by receiving learning materials from the service platform, and respectively determines a face recognition model, an item recognition model, a behavior recognition model, and a sound recognition model; receives the environmental images captured by multiple cameras, inputs the environmental images into the face recognition model, the item recognition model, and the behavior recognition model, respectively determines the abnormal image risk weight values corresponding to each camera, receives the environmental sounds captured by multiple cameras, inputs the environmental sounds into the sound recognition model, and respectively determines the abnormal sound risk weight values corresponding to each camera; receives the environmental noise captured by the noise source camera. If the environmental noise is greater than the preset noise threshold, calculate the comprehensive risk value of the noise source camera, and generate a corresponding warning message according to the judgment result of the comprehensive risk value, thereby improving the response efficiency and flexibility of abnormal events.

[0203] In another embodiment, the abnormal event prediction method can be separately configured from the central processing unit 9100. For example, the abnormal event prediction method can be configured as a chip connected to the central processing unit 9100, and the function of the abnormal event prediction method is realized through the control of the central processing unit.

[0204] As Figure 9 shown, the electronic device 9600 may further include: a communication module 9110, an input unit 9120, an audio processor 9130, a display 9160, and a power supply 9170. It should be noted that the electronic device 9600 does not necessarily have to include Figure 9 all the components shown in; in addition, the electronic device 9600 may further include Figure 9 components not shown in, and reference may be made to the prior art.

[0205] As Figure 9 shown, the central processing unit 9100, sometimes also referred to as a controller or operation control, may include a microprocessor or other processor device and / or logic device. The central processing unit 9100 receives inputs and controls the operations of the various components of the electronic device 9600.

[0206] Among them, the memory 9140, for example, may be one or more of a buffer, a flash memory, a hard drive, a removable medium, a volatile memory, a non-volatile memory, or other suitable devices. The above information related to failures can be stored, and in addition, programs for executing relevant information can also be stored. And the central processing unit 9100 can execute the program stored in the memory 9140 to implement information storage or processing, etc.

[0207] The input unit 9120 provides inputs to the central processing unit 9100. The input unit 9120 is, for example, a key or a touch input device. The power supply 9170 is used to supply power to the electronic device 9600. The display 9160 is used to display display objects such as images and texts. The display may be, for example, an LCD display, but is not limited thereto.

[0208] The memory 9140 may be a solid-state memory, for example, a read-only memory (ROM), a random access memory (RAM), a SIM card, etc. It may also be such a memory that stores information even when powered off, can be selectively erased and has more data. Examples of such a memory are sometimes referred to as EPROMs, etc. The memory 9140 may also be some other type of device. The memory 9140 includes a buffer memory 9141 (sometimes referred to as a buffer). The memory 9140 may include an application / function storage unit 9142, and the application / function storage unit 9142 is used to store application programs and function programs or the processes for operating the electronic device 9600 through the central processing unit 9100.

[0209] The memory 9140 may further include a data storage unit 9143 for storing data such as contacts, digital data, pictures, sounds, and / or any other data used by the electronic device. The driver storage unit 9144 of the memory 9140 may include various drivers of the electronic device for communication functions and / or for performing other functions of the electronic device (such as messaging applications, address book applications, etc.).

[0210] The communication module 9110 is a transmitter / receiver that transmits and receives signals via the antenna 9111. The communication module 9110 is coupled to the central processor 9100 to provide input signals and receive output signals, which may be the same as in the case of a conventional mobile communication terminal.

[0211] Based on different communication technologies, multiple communication modules 9110 may be provided in the same electronic device, such as a cellular network module, a Bluetooth module, and / or a wireless local area network module, etc. The communication module 9110 is also coupled to the speaker 9131 and the microphone 9132 via the audio processor 9130 to provide an audio output via the speaker 9131 and receive an audio input from the microphone 9132, thereby implementing normal telecommunication functions. The audio processor 9130 may include any suitable buffers, decoders, amplifiers, etc. In addition, the audio processor 9130 is also coupled to the central processor 9100, so that recording can be performed on the local device through the microphone 9132, and the sound stored on the local device can be played through the speaker 9131.

[0212] Embodiments of the present application also provide a computer-readable storage medium capable of implementing all steps of the abnormal event prediction method with the execution subject being a server or a client in the above embodiments. A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, all steps of the abnormal event prediction method with the execution subject being a server or a client in the above embodiments are implemented. For example, when the processor executes the computer program, the following steps are implemented:

[0213] Step S101: Receive the learning materials of the service platform, perform model training operations according to the learning materials, and respectively determine the corresponding face recognition model, object recognition model, behavior recognition model, and voice recognition model, where the learning materials include face recognition pictures, object recognition pictures, behavior recognition pictures, and voice corpus information;

[0214] Step S102: Receive the environmental images captured by the multiple cameras, input the environmental images into the face recognition model, the item recognition model, and the behavior recognition model, respectively determine the face attribute information, item type information, and behavior pattern information corresponding to each camera, perform a weighted calculation operation on the face attribute information, the item type information, and the behavior pattern information according to a preset abnormal event weight, respectively determine the abnormal image risk weight values corresponding to each camera, receive the environmental sounds captured by the multiple cameras, input the environmental sounds into the sound recognition model, perform a matching operation based on the recognition result output by the sound recognition model and a preset abnormal negative vocabulary corpus, and perform a weighted calculation operation based on the negative semantics and preset semantic weights obtained after the matching operation, respectively determine the abnormal sound risk weight values corresponding to each camera;

[0215] Step S103: Receive the environmental noise captured by the first camera, determine whether the environmental noise is greater than a preset noise threshold. If it is greater, perform a weight accumulation operation according to the abnormal image risk weight value corresponding to the first camera and the abnormal sound risk weight value corresponding to the first camera, determine the corresponding first comprehensive risk value, determine whether the first comprehensive risk value is greater than a preset first risk threshold, and generate a corresponding warning message according to the judgment result after the judgment operation, and send the warning message to the service platform so that the service platform performs an abnormal response according to the warning message, where the warning message includes an alarm message, a warning message, and a prompt message, and the first camera is the camera closest to the noise source among the multiple cameras.

[0216] As can be seen from the above description, the computer-readable storage medium provided by the embodiment of the present application trains a model by receiving learning materials from a service platform, and respectively determines a face recognition model, an item recognition model, a behavior recognition model, and a sound recognition model; receives environmental images captured by multiple cameras, inputs the environmental images into the face recognition model, the item recognition model, and the behavior recognition model, and respectively determines the abnormal image risk weight values corresponding to each camera, receives environmental sounds captured by multiple cameras, inputs the environmental sounds into the sound recognition model, and respectively determines the abnormal sound risk weight values corresponding to each camera; receives the environmental noise captured by the noise source camera, if the environmental noise is greater than a preset noise threshold, calculates the comprehensive risk value of the noise source camera, and generates a corresponding warning message according to the judgment result of the comprehensive risk value, thereby being able to improve the response efficiency and flexibility of abnormal events.

[0217] Embodiments of the present application further provide a computer program product capable of implementing all steps in the abnormal event prediction method where the execution subject in the above embodiments is a server or a client. When the computer program / instructions are executed by a processor, the steps of the abnormal event prediction method are implemented. For example, the computer program / instructions implement the following steps:

[0218] Step S101: Receive the learning materials of the service platform, perform model training operations according to the learning materials, and respectively determine the corresponding face recognition model, item recognition model, behavior recognition model, and voice recognition model. The learning materials include face recognition pictures, item recognition pictures, behavior recognition pictures, and voice corpus information;

[0219] Step S102: Receive the environmental images captured by the multiple cameras, input the environmental images into the face recognition model, the item recognition model, and the behavior recognition model, respectively determine the face attribute information, item type information, and behavior pattern information corresponding to each camera, perform weighted calculation operations on the face attribute information, the item type information, and the behavior pattern information according to the preset abnormal event weights, respectively determine the abnormal image risk weight values corresponding to each camera, receive the environmental sounds captured by the multiple cameras, input the environmental sounds into the voice recognition model, perform matching operations according to the recognition results output by the voice recognition model and the preset abnormal negative vocabulary corpus, and perform weighted calculation operations according to the negative semantics obtained after the matching operations and the preset semantic weights, respectively determine the abnormal sound risk weight values corresponding to each camera;

[0220] Step S103: Receive the environmental noise captured by the first camera, determine whether the environmental noise is greater than the preset noise threshold. If it is greater, perform a weight accumulation operation according to the abnormal image risk weight value corresponding to the first camera and the abnormal sound risk weight value corresponding to the first camera to determine the corresponding first comprehensive risk value, determine whether the first comprehensive risk value is greater than the preset first risk threshold, and generate the corresponding warning information according to the judgment result after the judgment operation, and send the warning information to the service platform so that the service platform performs an abnormal response according to the warning information. The warning information includes alarm information, early warning information, and prompt information. The first camera is the camera closest to the noise source among the multiple cameras.

[0221] As can be seen from the above description, the computer program product provided by the embodiments of the present application performs model training by receiving learning materials from the service platform, and respectively determines a face recognition model, an object recognition model, a behavior recognition model, and a voice recognition model; receives environmental images captured by multiple cameras, inputs the environmental images into the face recognition model, the object recognition model, and the behavior recognition model, and respectively determines the abnormal image risk weight values corresponding to each camera, receives the environmental sounds captured by multiple cameras, inputs the environmental sounds into the voice recognition model, and respectively determines the abnormal sound risk weight values corresponding to each camera; receives the environmental noise captured by the noise source camera, if the environmental noise is greater than the preset noise threshold, calculates the comprehensive risk value of the noise source camera, and generates a corresponding warning message according to the judgment result of the comprehensive risk value, thereby being able to improve the response efficiency and flexibility of abnormal events.

[0222] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a device, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0223] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (apparatus), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in Figure 1 one or more of the flows Figure 1 or multiple flows and / or blocks

[0224] These computer program instructions can also be stored in a computer-readable memory capable of guiding a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device realizes the functions specified in Figure 1 one or more of the flows Figure 1 or multiple flows and / or blocks

[0225] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions for implementing the steps of the function specified in one process or a plurality of processes and / or blocks Figure 1 in one block or a plurality of blocks Figure 1 in the process.

[0226] Specific embodiments of the present invention are used to illustrate the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A method for predicting abnormal events, characterized in that: Applied to an edge computing platform, the edge computing platform is communicatively connected with a preset business platform and a plurality of preset cameras, wherein the plurality of cameras are configured with a voice device and a noise sensor for interacting with the environment, the method comprising: Receive learning materials from the business platform, perform model training operations according to the learning materials, and respectively determine corresponding face recognition models, object recognition models, behavior recognition models, and voice recognition models, wherein the learning materials include face recognition pictures, object recognition pictures, behavior recognition pictures, and voice corpus information; Receive the environmental images captured by the multiple cameras, input the environmental images into the face recognition model, the object recognition model and the behavior recognition model, respectively determine the facial attribute information, object type information and behavior pattern information corresponding to each camera, perform weighted calculation operations on the facial attribute information, the object type information and the behavior pattern information according to preset abnormal event weights, respectively determine the abnormal image risk weight value corresponding to each camera, receive the environmental sounds captured by the multiple cameras, input the environmental sounds into the sound recognition model, perform a matching operation based on the recognition results output by the sound recognition model and a preset abnormal negative vocabulary corpus, and perform a weighted calculation operation based on the negative semantics obtained after the matching operation and the preset semantic weights, respectively determine the abnormal sound risk weight value corresponding to each camera; Receive the environmental noise recorded by the first camera, and determine whether the environmental noise is greater than a preset noise threshold; if so, perform a weight accumulation operation according to the abnormal image risk weight value corresponding to the first camera and the abnormal sound risk weight value corresponding to the first camera to determine the corresponding first comprehensive risk value; determine whether the first comprehensive risk value is greater than the preset first risk threshold, and generate corresponding warning information according to the judgment result after the judgment operation; send the warning information to the business platform, so that the business platform performs an abnormal response according to the warning information, wherein the warning information includes alarm information, early warning information and prompt information, and the first camera is the camera closest to the noise source among the multiple cameras.

2. The abnormal event prediction method according to claim 1, characterized in that: The method further comprises: According to a preset time interval, polling is performed to determine whether the comprehensive risk value corresponding to each camera is greater than a preset second risk threshold, wherein the comprehensive risk value corresponding to each camera is determined by performing a weighted accumulation operation based on the abnormal image risk weight value corresponding to each camera and the abnormal sound risk weight value corresponding to each camera; If it is greater, an alarm information generation operation is performed, and the alarm information is sent to the business platform, so that the business platform performs an abnormal response according to the alarm information; If not, the comprehensive risk values ​​corresponding to each camera are accumulated to determine whether the comprehensive risk values ​​corresponding to all cameras obtained after the accumulation operation are greater than the preset third risk threshold; if so, an alarm information generation operation is performed and the alarm information is sent to the business platform so that the business platform responds to the exception according to the alarm information.

3. The abnormal event prediction method according to claim 1, characterized in that: The model training operation is performed according to the learning data to respectively determine the corresponding face recognition model, object recognition model, behavior recognition model and sound recognition model, including: Performing a model training operation on a preset Fast RCNN network according to the face recognition images in the learning material to determine a corresponding face recognition model, wherein the face recognition model includes a face detection module, a face normalization module, a face attribute recognition module, a face quality assessment module, and a face tracking module, and the face recognition model is used to recognize faces and analyze people's negative emotions to obtain face attribute information; Performing a model training operation on a preset convolutional neural network according to the object recognition images in the learning material to determine a corresponding object recognition model, wherein the object recognition model is used to identify objects that endanger environmental safety; Performing a model training operation on a preset Crowd Counting algorithm according to the behavior recognition pictures in the learning material to determine a corresponding behavior recognition model, wherein the behavior recognition model is used to detect crowd gathering and behaviors that endanger environmental safety; A model training operation is performed on a preset semantic recognition initial model according to the sound corpus information in the learning material to determine a corresponding sound recognition model, wherein the sound recognition model is used to convert sound into semantic information.

4. The abnormal event prediction method according to claim 1, characterized in that: The step of performing a weighted calculation operation on the face attribute information, the object type information, and the behavior pattern information according to the preset abnormal event weights to respectively determine the abnormal image risk weight value corresponding to each camera includes: Performing a weighting operation on the facial attribute information, the item type information, and the behavior pattern information according to the preset abnormal event weights, and determining the facial attribute information weight, the item type information weight, and the behavior pattern information weight corresponding to each camera; A weight accumulation operation is performed according to the facial attribute information weight, the item type information weight, and the behavior pattern information weight, to respectively determine the abnormal image risk weight value corresponding to each camera.

5. The abnormal event prediction method according to claim 1, characterized in that: The matching operation is performed based on the recognition result output by the sound recognition model and the preset abnormal negative vocabulary corpus, and the weighted calculation operation is performed based on the negative semantics obtained after the matching operation and the preset semantic weight, to respectively determine the abnormal sound risk weight value corresponding to each camera, including: Performing a word segmentation operation on the recognition result output by the sound recognition model according to a preset word segmentation algorithm to determine the corresponding word segmentation vocabulary; The segmented vocabulary and the preset abnormal negative vocabulary corpus are matched, and a weighted calculation operation is performed according to the negative semantics obtained after the matching operation and the preset semantic weight, so as to determine the abnormal sound risk weight value corresponding to each camera respectively.

6. The abnormal event prediction method according to claim 1, characterized in that: The determining whether the environmental noise is greater than a preset noise threshold, and if so, performing a weight accumulation operation according to the abnormal image risk weight value corresponding to the first camera and the abnormal sound risk weight value corresponding to the first camera to determine the corresponding first comprehensive risk value, includes: Determine whether the environmental noise is greater than a preset noise threshold, and if so, determine whether the abnormal image risk weight value corresponding to the first camera and the abnormal sound risk weight value corresponding to the first camera exist at the same time; If both exist, a weighted accumulation operation is performed on the abnormal image risk weight value corresponding to the first camera and the abnormal sound risk weight value corresponding to the first camera, and a multiplication operation is performed on the accumulated result after the weighted accumulation operation to determine the corresponding first comprehensive risk value; If they do not exist at the same time, the abnormal image risk weight value corresponding to the first camera and the abnormal sound risk weight value corresponding to the first camera are weighted and accumulated to determine the corresponding first comprehensive risk value.

7. The abnormal event prediction method according to claim 1, characterized in that: The determining whether the first comprehensive risk value is greater than a preset first risk threshold, and generating corresponding warning information according to the determination result after the determination operation, includes: Determine whether the first comprehensive risk value is greater than a preset first risk threshold, and if so, generate a warning message to determine the corresponding warning message; If it is less than, determine whether it is greater than half of the preset first risk threshold. If it is greater, perform a warning information generation operation to determine the corresponding early warning information. If it is less than, perform a warning information generation operation to determine the corresponding prompt information.

8. An abnormal event prediction device, characterized in that: The device comprises: A multi-model training module, used to receive learning materials from the business platform, perform model training operations according to the learning materials, and respectively determine corresponding face recognition models, object recognition models, behavior recognition models, and voice recognition models, wherein the learning materials include face recognition pictures, object recognition pictures, behavior recognition pictures, and voice corpus information; A risk weight calculation module, for receiving environmental images captured by multiple cameras, inputting the environmental images into the face recognition model, the object recognition model and the behavior recognition model, respectively determining the facial attribute information, object type information and behavior pattern information corresponding to each camera, performing weighted calculation operations on the facial attribute information, the object type information and the behavior pattern information according to preset abnormal event weights, respectively determining the abnormal image risk weight value corresponding to each camera, receiving environmental sounds captured by multiple cameras, inputting the environmental sounds into the sound recognition model, performing a matching operation based on the recognition results output by the sound recognition model and a preset abnormal negative vocabulary corpus, and performing a weighted calculation operation based on the negative semantics obtained after the matching operation and the preset semantic weights, respectively determining the abnormal sound risk weight value corresponding to each camera; A warning information generation module is used to receive the environmental noise recorded by the first camera, determine whether the environmental noise is greater than a preset noise threshold, and if so, perform a weight accumulation operation according to the abnormal image risk weight value corresponding to the first camera and the abnormal sound risk weight value corresponding to the first camera to determine the corresponding first comprehensive risk value, determine whether the first comprehensive risk value is greater than the preset first risk threshold, and generate corresponding warning information according to the judgment result after the judgment operation, and send the warning information to the business platform so that the business platform performs an abnormal response according to the warning information, wherein the warning information includes alarm information, early warning information and prompt information, and the first camera is the camera closest to the noise source among the preset multiple cameras.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the abnormal event prediction method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the abnormal event prediction method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Noise source positioning method, device and system based on sound and image

    CN113239913A

  • Risk prediction method based on multi-modal data fusion

    CN117708746A