Emergency detection method and device and computer readable medium
By performing preliminary category recognition and object detection on the input images, combining CLIP and YOLOv3 models to calculate the similarity of feature information and using visual language model verification, the problem of inaccurate recognition in the prior art is solved, and high-accurate emergency detection is achieved.
Patent Information
- Application Number
- CN202510531468.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-08-19
AI Technical Summary
The existing technology cannot accurately identify emergencies, resulting in high economic losses.
By performing preliminary category recognition and object detection on the input image, combining the CLIP model and the YOLOv3 model, the feature information similarity is calculated and the categories are updated, and the visual language model verification results are used to achieve detailed judgment of emergencies.
It improves the accuracy of emergency detection and provides more accurate emergency detection results.
Smart Images

Figure CN120510418A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a method, device, and computer-readable medium for detecting an emergency. Background Art
[0002] As society continues to develop, various emergencies may occur in people's daily lives, such as severe weather, natural disasters, and traffic accidents. According to statistics, the global economic losses caused by various emergencies are over US$1 billion each year. How to quickly identify various emergencies is a hot topic in current research.
[0003] Image processing technology has become one of the most advanced and effective information processing techniques. It can parse and extract diverse information from images, then identify corresponding features and patterns in events through processing, analysis, classification, and reasoning, ultimately enabling the accurate location and resolution of emergencies. However, there is currently no image processing solution that can accurately identify emergencies. Summary of the Invention
[0004] One purpose of the present application is to provide an emergency event detection method, device and computer-readable medium to solve the problems existing in existing solutions.
[0005] To achieve the above objectives, an embodiment of the present application provides a method for detecting an emergency, the method comprising:
[0006] Performing preliminary category recognition on the overall image content of the input image to obtain a preliminary event category corresponding to the input image;
[0007] Performing target detection on the input image to obtain the object target contained in the input image;
[0008] The emergency event category corresponding to the input image is determined according to the preliminary event category corresponding to the input image and the object target contained in the input image.
[0009] Furthermore, the preliminary event categories include specific event categories and other categories;
[0010] If the acquired preliminary event category corresponding to the input image is a specific event category, the method further includes:
[0011] Calculating first feature information of the input image using an image encoder of a CLIP model;
[0012] Calculating second feature information of the category name text of the preliminary event category using a text encoder of the CLIP model;
[0013] A first similarity value between the first feature information and the second feature information is calculated, and a preliminary event category corresponding to the input image is updated according to the first similarity value.
[0014] Furthermore, calculating a similarity value between the first feature information and the second feature information, and updating the preliminary event category according to the similarity value, includes:
[0015] Calculating a first similarity value between the first feature information and the second feature information, and comparing the first similarity value with a first threshold;
[0016] If the first similarity value is greater than the first threshold, the preliminary event category corresponding to the input image is updated from a specific event category to another category.
[0017] Furthermore, after performing target detection on the input image and obtaining the object target contained in the input image, the method further includes:
[0018] intercepting a target image containing the object from the input image;
[0019] Calculating the third feature information of the target image using an image encoder of a CLIP model;
[0020] Calculating fourth feature information of the target name text of the object target using a text encoder of the CLIP model;
[0021] A second similarity value between the third feature information and the fourth feature information is calculated, and the object target included in the input image is updated according to the second similarity value.
[0022] Further, calculating a second similarity value between the third feature information and the fourth feature information, and updating the object target contained in the input image according to the second similarity value, includes:
[0023] Calculating a second similarity value between the third feature information and the fourth feature information, and comparing the second similarity value with a second threshold;
[0024] If the second similarity value is greater than the second threshold, determining the object target as an invalid detection result, and deleting the invalid detection result;
[0025] If the second similarity value is greater than the second threshold, the object target is determined as a valid detection result, and the valid detection result is retained.
[0026] Furthermore, performing preliminary category recognition based on the overall image content of the input image to obtain a preliminary event category corresponding to the input image includes:
[0027] A ResNet model is used to perform preliminary category recognition on the overall image content of the input image to obtain a preliminary event category corresponding to the input image.
[0028] Furthermore, performing target detection on the input image to obtain the object target contained in the input image includes:
[0029] The YOLOv3 model is used to perform target detection on the input image to obtain the object target contained in the input image.
[0030] Furthermore, after determining the emergency event category corresponding to the input image according to the preliminary event category corresponding to the input image and the object target contained in the input image, the method further includes:
[0031] According to the emergency event category, construct prompt words for verifying the emergency event category;
[0032] Inputting the input image and the prompt word into a visual language model, and obtaining a verification result determined by the visual language model based on the input image and the prompt word;
[0033] If the verification result is correct, the emergency event category corresponding to the input image is output.
[0034] Some embodiments of the present application also provide an emergency event detection device, wherein the device includes a memory for storing computer program instructions and a processor for executing the computer program instructions, wherein when the computer program instructions are executed by the processor, the device is triggered to execute the aforementioned emergency event detection method.
[0035] Other embodiments of the present application further provide a computer-readable medium having computer program instructions stored thereon, wherein the computer program instructions can be executed by a processor to implement the emergency event detection method.
[0036] Compared to the prior art, the embodiment of the present application provides a sudden event detection solution that can perform preliminary category recognition on the overall image content of the input image to obtain the preliminary event category corresponding to the input image. At the same time, it can perform target detection on the input image to obtain the object targets contained in the input image. Then, based on the preliminary event category corresponding to the input image and the object targets contained in the input image, the sudden event category corresponding to the input image is determined. This solution performs preliminary detection based on the overall image content of the input image as a whole, and further combines the object targets contained in the input image to make a more detailed judgment on the sudden event category, thereby improving the accuracy of sudden event detection and obtaining more accurate sudden event detection results. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Other features, objects and advantages of the present application will become more apparent upon reading the detailed description of non-limiting embodiments made with reference to the following drawings:
[0038] Figure 1 A processing flow chart of an emergency event detection method provided in an embodiment of the present application;
[0039] Figure 2 This is a flowchart of the process of detecting an emergency using the solution of an embodiment of the present application;
[0040] The same or similar reference numerals in the drawings represent the same or similar components. DETAILED DESCRIPTION
[0041] The present application is described in further detail below with reference to the accompanying drawings.
[0042] In a typical configuration of the present application, the terminal and the equipment of the service network each include one or more processors (CPUs), input / output interfaces, network interfaces and memories.
[0043] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0044] Computer-readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology for information storage. The information can be computer program instructions, data structures, program devices, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc-read only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information that can be accessed by a computing device.
[0045] The embodiment of the present application provides a method for detecting sudden events. The method can perform preliminary category recognition on the overall image content of the input image to obtain the preliminary event category corresponding to the input image. At the same time, the method can perform target detection on the input image to obtain the object targets contained in the input image. Then, based on the preliminary event category corresponding to the input image and the object targets contained in the input image, the sudden event category corresponding to the input image is determined. This solution performs preliminary detection based on the overall image content of the input image as a whole, and further combines the object targets contained in the input image to make a more detailed judgment on the sudden event category, thereby improving the accuracy of sudden event detection and obtaining more accurate sudden event detection results.
[0046] In practical scenarios, the execution subject of this method can be a user device, a network device, or a device formed by integrating a user device and a network device via a network, or it can also be an application running on the above devices. The user device includes but is not limited to various terminal devices such as computers, mobile phones, and tablets; the network device includes but is not limited to network hosts, single network servers, multiple network server sets, or cloud computing-based computer collections. Here, the cloud is composed of a large number of hosts or network servers based on cloud computing (Cloud Computing). Cloud computing is a type of distributed computing, a virtual computer composed of a group of loosely coupled computers.
[0047] Figure 1 The following is a processing flow of an emergency event detection method provided by an embodiment of the present application, which includes at least the following processing steps:
[0048] Step S101 : performing preliminary category recognition on the overall image content of the input image to obtain a preliminary event category corresponding to the input image.
[0049] The preliminary event category recognition is used to make a preliminary judgment on the category of the emergency. For example, in this embodiment, the preliminary event category recognition results can be divided into two categories, including specific event categories and other categories. The specific event categories can be any emergency event categories that will produce a large-scale coverage effect in the image, such as blizzards and heavy fog, while the other categories are other emergency event categories that will not produce a large-scale coverage effect in the image, such as traffic accidents and fires.
[0050] Since the main difference between the different recognition results of preliminary event categories is whether there is a specific image content with large coverage in the image, a classification detection method can achieve better results. Therefore, when performing preliminary category recognition, the solution of the embodiment of the present application can use the ResNet (Residual Network) model to perform preliminary category recognition on the overall image content of the input image to obtain the preliminary event category corresponding to the input image. Among them, the ResNet model, due to the introduction of residual connections, can effectively alleviate the problem of gradient disappearance or gradient explosion that occurs as the network depth increases, so that the model can build a very deep network, so that it can learn richer and more discriminative image features, thereby improving the accuracy of classification, and the training is relatively stable and efficient.
[0051] In some embodiments of the present application, in order to further improve the accuracy of preliminary event category recognition, if the preliminary event category corresponding to the acquired input image is a specific event category, the solution of the embodiments of the present application may further confirm the result of the preliminary category recognition, specifically including the following processing:
[0052] First, the image encoder of the CLIP (Contrastive Language-Image Pre-Training) model is used to calculate first feature information of the input image. Simultaneously, the text encoder of the CLIP model is used to calculate second feature information of the category name text of the preliminary event category. A first similarity value between the first feature information and the second feature information can then be calculated, and the preliminary event category corresponding to the input image can be updated based on the first similarity value.
[0053] Among them, the CLIP model can map the feature information of the image and text to the same vector space through the image encoder and the text encoder, respectively, so that the feature information under different modalities can be calculated and compared in the same vector space. Thus, the first feature information can represent the global image features of the input image, while the second feature information can represent the text semantic features of the category name text of the preliminary event category. Since both are encoded by the image encoder and the text encoder of the CLIP model and mapped to the same vector space, the similarity of the semantic relationship between the input image and the category name text can be calculated based on the first feature information and the second feature information. This similarity can be quantitatively represented based on a first similarity value. For example, in this embodiment, the Euclidean distance can be used as the first similarity value. By calculating the Euclidean distance between the first feature information and the second feature information, the similarity of the semantic relationship between the input image and the category name text can be obtained, thereby updating the preliminary event category corresponding to the input image.
[0054] When updating the preliminary event category based on the similarity value, the solution in this embodiment may first calculate a first similarity value between the first feature information and the second feature information, and compare the first similarity value with a first threshold. If the first similarity value is greater than the first threshold, the preliminary event category corresponding to the input image is updated from the specific event category to another category.
[0055] Among them, the first threshold value can be set according to the needs of the actual scene. For example, when the first similarity value is the Euclidean distance, the larger its value is, the lower the similarity between the two is. On the contrary, the closer the Euclidean distance is to zero, the higher the similarity between the two is. In the actual scene, the first threshold value can be set to 0.5. If the calculated Euclidean distance is greater than 0.5, it means that the semantic relationship between the specific event category obtained by the preliminary category recognition and its corresponding category name text is not close enough, and the result of the preliminary category recognition may be wrong. Therefore, the preliminary event category can be corrected and updated from the specific event category to other categories. On the contrary, if the calculated Euclidean distance is less than or equal to 0.5, it means that the degree of closeness of the semantic relationship between the specific event category obtained by the preliminary category recognition and its corresponding category name text meets the requirements of this solution, which is sufficient to consider that the result of the preliminary category recognition is correct. Therefore, the preliminary event category can be confirmed as the identified specific event category.
[0056] Step S102: performing target detection on the input image to obtain the object target contained in the input image.
[0057] Among them, the object targets can be pre-set targets related to various emergencies, such as flames, white smoke, black smoke, firefighters, fire trucks, fallen people, fallen non-motor vehicles, overturned cars, traffic police, damaged cars, rescue workers, sandstorms, people wearing life jackets, warning cones, warning signs, maintenance vehicles, cleaners, rollers, lifeboats, umbrellas, etc.
[0058] To accurately detect objects in the input image, the solution of the embodiment of the present application can use the YOLOv3 (You Only Look Once version 3) model to perform object detection on the input image and obtain the objects contained in the input image. The YOLOv3 model uses a multi-scale prediction method when performing object detection. It can detect on feature maps of different scales to better cope with objects of different sizes, improve the detection effect of small objects, and achieve better detection accuracy.
[0059] In addition, in order to further improve the accuracy of target detection, the solution of the embodiment of the present application can also further confirm the identified object target, specifically including the following processing:
[0060] After performing target detection on the input image and obtaining the object target contained in the input image, a target image containing the object target can be captured from the input image, and the image encoder of the CLIP model is used to calculate the third feature information of the target image. At the same time, the text encoder of the CLIP model is used to calculate the fourth feature information of the target name text of the object target, and then the second similarity value between the third feature information and the fourth feature information is calculated, and the object target contained in the input image is updated according to the second similarity value.
[0061] The CLIP model's image encoder and text encoder can respectively obtain third feature information regarding the image features of the target image and fourth feature information regarding the textual semantic features of the target name text of the object target. Therefore, the similarity of the semantic relationship between the target image and the target name text can be calculated based on the third feature information and the fourth feature information. This similarity can be quantified based on a second similarity value. For example, in this embodiment, the Euclidean distance can also be used as the second similarity value. By calculating the Euclidean distance between the third feature information and the fourth feature information, the degree of similarity of the semantic relationship between the target image and the target name text can be determined, thereby updating the object target contained in the input image.
[0062] When updating the preliminary event category based on the similarity value, the solution in this embodiment may first calculate a first similarity value between the first feature information and the second feature information, and compare the first similarity value with a first threshold. If the first similarity value is greater than the first threshold, the preliminary event category corresponding to the input image is updated from the specific event category to another category.
[0063] Among them, the first threshold value can be set according to the needs of the actual scene. For example, when the first similarity value is the Euclidean distance, the larger its value is, the lower the similarity between the two is. On the contrary, the closer the Euclidean distance is to zero, the higher the similarity between the two is. In the actual scene, the first threshold value can be set to 0.5. If the calculated Euclidean distance is greater than 0.5, it means that the semantic relationship between the specific event category obtained by the preliminary category recognition and its corresponding category name text is not close enough, and the result of the preliminary category recognition may be wrong. Therefore, the preliminary event category can be corrected and updated from the specific event category to other categories. On the contrary, if the calculated Euclidean distance is less than or equal to 0.5, it means that the degree of closeness of the semantic relationship between the specific event category obtained by the preliminary category recognition and its corresponding category name text meets the requirements of this solution, which is sufficient to consider that the result of the preliminary category recognition is correct. Therefore, the preliminary event category can be confirmed as the identified specific event category.
[0064] When calculating a second similarity value between the third feature information and the fourth feature information, and updating the object target included in the input image based on the second similarity value, the second similarity value between the third feature information and the fourth feature information can be first calculated, and the second similarity value can be compared with a second threshold. If the second similarity value is greater than the second threshold, the object target is determined to be an invalid detection result and the invalid detection result is deleted. If the second similarity value is greater than the second threshold, the object target is determined to be a valid detection result and the valid detection result is retained.
[0065] Among them, the second threshold value can also be set according to the needs of the actual scene. For example, when the second similarity value is the Euclidean distance, the larger its value is, the lower the similarity between the two is. On the contrary, the closer the Euclidean distance is to zero, the higher the similarity between the two is. In the actual scene, the second threshold value can be set to 0.5. If the calculated Euclidean distance is greater than 0.5, it means that the target name text corresponding to the object target obtained by the target detection is not close enough to the target image of the object target in the input image in terms of semantic relationship, and the target detection result may be wrong. Therefore, the object target previously detected and determined can be determined as an invalid detection result and deleted from the detection result set. On the contrary, if the calculated Euclidean distance is less than or equal to 0.5, it means that the target name text corresponding to the object target obtained by the target detection and the target image of the object target in the input image in terms of semantic relationship meet the requirements of this solution, which is sufficient to consider that the target detection result is correct. Therefore, the object target previously detected and determined can be determined as a valid detection result and retained as a result in the detection result set.
[0066] Step S103 : determining the emergency event category corresponding to the input image according to the preliminary event category corresponding to the input image and the object target contained in the input image.
[0067] In actual scenarios, for some specific emergencies, the corresponding images generally contain some corresponding object targets. For example, for the emergency category of "fire", the preliminary event category corresponding to the input image should be other categories, and the object targets contained in the input image should include at least one of flames, white smoke, black smoke, firefighters, and fire trucks. For example, for the emergency category of "snowstorm", the preliminary event category corresponding to the input image can be "snowstorm" in the specific event category, and the object targets contained in the input image should be empty, that is, no object target can be detected in the input image. Therefore, the emergency corresponding to the input image can be further identified based on the preliminary event category corresponding to the input image and the object targets contained in the input image, so as to finally determine the emergency category corresponding to the input image, thereby improving the accuracy of emergency detection.
[0068] This solution performs preliminary detection based on the overall image content of the input image, and further combines the objects contained in the input image to make a more detailed judgment on the category of the emergency. This can improve the accuracy of emergency detection and obtain more accurate emergency detection results.
[0069] In other embodiments of the present application, in order to further improve the accuracy of the emergency detection results, after determining the emergency category corresponding to the input image, a visual language model (VLM) can be used to verify the detection results. Specifically, the solution of this embodiment can first construct a prompt word (prompt) for verifying the emergency category based on the emergency category after determining the emergency category corresponding to the input image based on the preliminary event category corresponding to the input image and the object target contained in the input image.
[0070] For example, the prompt in this embodiment can be structured as follows: "Has [emergency event category] occurred in the input image? Please answer yes or no." [Emergency event category] is the detection result determined based on steps S101 to S103, such as a snowstorm or fire. Therefore, the prompt can be "Has a fire occurred in the input image? Please answer yes or no," or "Has a snowstorm occurred in the input image? Please answer yes or no," etc.
[0071] After the prompt word is generated, the input image and the prompt word can be input into the visual language model to obtain a verification result determined by the visual language model based on the input image and the prompt word. If the verification result is correct, the emergency category corresponding to the input image is output. For example, taking the aforementioned scenario as an example, when the answer given by the VLM is yes, it means that the verification result is correct, the emergency category previously detected is credible, and the emergency category corresponding to the input image can be finally output, such as blizzard, fire, etc. Otherwise, if the answer given by the VLM is no, it means that the verification result is wrong. At this time, the credibility of the emergency category corresponding to the input image determined in the aforementioned step S103 is insufficient. At this time, other methods can be used to correct or reconfirm the emergency category to avoid outputting wrong detection results to the user.
[0072] Figure 2 The following is a flowchart of an emergency event detection process using the solution provided in the embodiment of the present application, including the following steps:
[0073] In step S201 , the ResNet model gives a preliminary event category of the input image, which may specifically include snowstorm, heavy fog, and other categories.
[0074] In step S202, the input image whose preliminary event category is blizzard or heavy fog is processed through the image encoder of the CLIP model to obtain first feature information. Simultaneously, the category name text of the preliminary event category output by the ResNet model is processed through the text encoder of the CLIP model to obtain second feature information. The Euclidean distance between the first and second feature information is then calculated. If the Euclidean distance is greater than a first threshold of 0.5, the preliminary event category is updated from blizzard or heavy fog to another category. If the Euclidean distance is less than or equal to the first threshold of 0.5, the current preliminary event category is considered correct.
[0075] In step S203, the YOLOv3 model performs object detection on the input image to obtain objects contained in the input image. Detectable objects include flames, white smoke, black smoke, firefighters, fire trucks, fallen people, fallen non-motor vehicles, overturned cars, traffic police, damaged cars, rescue workers, sandstorms, people wearing life jackets, warning cones, warning signs, maintenance vehicles, cleaners, road rollers, lifeboats, umbrellas, etc.
[0076] In step S204, the YOLOv3 model detects the target position of the object in the input image and extracts the target image from the input image based on the target position. The target image is processed by the image encoder of the CLIP model to obtain third feature information. At the same time, the target name text output by the YOLOv3 model is processed by the text encoder of the CLIP model to obtain fourth feature information. Then, the Euclidean distance between the third feature information and the fourth feature information is calculated. If the Euclidean distance is greater than the second threshold of 0.5, the detected object is considered invalid. If the Euclidean distance is less than or equal to the first threshold of 0.5, the detected object is considered valid and retained as the correct detection result.
[0077] Step S205: Determine the emergency event category corresponding to the input image based on the preliminary event category and the object target. The following is the correspondence between the emergency event category and the detected preliminary event category and the object target in this embodiment:
[0078] (1) Heavy fog: The preliminary event category output by the ResNet model is heavy fog, and the YOLOv3 model does not detect any targets.
[0079] (2) Snowstorm: The preliminary event category output by the ResNet model is snowstorm, and the YOLOv3 model does not detect any targets.
[0080] (3) Fire: The preliminary event category output by the ResNet model is other categories. The YOLOv3 model detects any one of flames, white smoke, black smoke, firefighters, and fire trucks, but does not detect traffic police.
[0081] (4) Traffic accidents: The preliminary event category output by the ResNet model is other categories. The YOLOv3 model detects any one of the following: a fallen person, a fallen non-motor vehicle, an overturned car, a traffic policeman, or a severely damaged car, but does not detect a sandstorm.
[0082] (5) Earthquake: The preliminary event category output by the ResNet model is other categories. The YOLOv3 model detects any of the following: firefighters, fire trucks, fallen people, overturned cars, fallen non-motor vehicles, severely damaged cars, and rescue workers. It does not detect white smoke, black smoke, sandstorms, lifeboats, or umbrellas.
[0083] (6) Flood: The preliminary event category output by the ResNet model is other categories. The YOLOv3 model detects any of the following: lifeboat, rescuer, umbrella, person wearing life jacket, firefighter, and no sandstorm, white smoke, black smoke, or flame.
[0084] (7) Sandstorm: The preliminary event category output by the ResNet model is other categories. The YOLOv3 model detects sandstorms, but does not detect white smoke, black smoke, or flames.
[0085] (8) Collapse: The preliminary event category output by the ResNet model is other categories. The YOLOv3 model detects any one of the following: people wearing life jackets, firefighters, rescue workers, and fire trucks, but does not detect sandstorms, white smoke, black smoke, flames, umbrellas, or lifeboats.
[0086] (9) Highway maintenance: The preliminary event category output by the ResNet model is other categories. The YOLOv3 model detects any one of the following: cleaner, roller, maintenance vehicle, warning sign, and warning cone, but does not detect sandstorm, white smoke, black smoke, flame, umbrella, lifeboat, or overturned car.
[0087] In step S206, the VLM verifies the emergency event category obtained in the previous step. Specifically, a prompt can be set, with the structure of "Has [emergency event category] occurred in this input image? Please answer yes or no," for example, "Has a collapse occurred in this input image? Please answer yes or no." The prompt and the input image are then provided to the VLM, which outputs a yes or no verification result.
[0088] If the VLM answers yes, the verification result is correct and the previously detected emergency category is credible. The emergency category corresponding to the input image can be output. If the VLM answers no, the verification result is incorrect and the previously determined emergency category corresponding to the input image is not credible enough. Other methods can be used to correct or reconfirm the emergency category to avoid outputting an erroneous detection result to the user. This effectively improves the accuracy of emergency detection.
[0089] Based on another aspect of the present application, an embodiment of the present application also provides an emergency event detection device, which includes a memory for storing computer program instructions and a processor for executing computer program instructions, wherein when the computer program instructions are executed by the processor, the device is triggered to execute the aforementioned emergency event detection method.
[0090] In particular, the methods and / or embodiments of the present application can be implemented as computer software programs. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for executing the method shown in the flowchart. When the computer program is executed by a processing unit, the above-mentioned functions defined in the method of the present application are performed.
[0091] It should be noted that the computer-readable medium described in this application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. Computer-readable media may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, a computer-readable medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device.
[0092] In the present application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries a computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0093] Computer program code for performing the operations of the present application can be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0094] The flow chart or block diagram in the accompanying drawings illustrate the possible architecture, functions and operations of the equipment, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code include one or more executable instructions for realizing the logical function of the specification. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented with a dedicated system for hardware that performs the function or operation of the specification, or can be implemented with a combination of dedicated hardware and computer instructions.
[0095] As another aspect, the present application further provides a computer-readable medium, which may be included in the device described in the above embodiments, or may exist independently without being incorporated into the device. The computer-readable medium carries one or more computer program instructions, which can be executed by a processor to implement the methods and / or technical solutions of the above embodiments of the present application.
[0096] It should be noted that the present application can be implemented in a combination of software and / or software and hardware, for example, can be implemented using an application specific integrated circuit (ASIC), a general purpose computer or any other similar hardware device. In certain embodiments, the software program of the present application can be executed by a processor to implement the above steps or functions. Similarly, the software program of the present application (including related data structures) can be stored in a computer-readable recording medium, for example, a RAM memory, a magnetic or optical drive or a floppy disk and similar devices. In addition, some steps or functions of the present application can be implemented using hardware, for example, as a circuit that cooperates with a processor to perform each step or function.
[0097] It is obvious to those skilled in the art that the present application is not limited to the details of the above-mentioned exemplary embodiments, and that the present application can be implemented in other specific forms without departing from the spirit or basic characteristics of the present application. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-restrictive, and the scope of the present application is limited by the appended claims rather than the above description, and it is intended that all changes that fall within the meaning and scope of the equivalent elements of the claims are included in the present application. Any figure mark in the claims should not be regarded as limiting the claims involved. In addition, it is obvious that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices stated in the device claim can also be implemented by one unit or device through software or hardware. Words such as first and second are used to indicate names and do not indicate any specific order. The numbers corresponding to the steps are used to mark and distinguish different steps, and the size of the numbers does not limit any specific execution order.
Claims
1. A method for detecting an emergency, characterized in that: The method comprises: Performing preliminary category recognition on the overall image content of the input image to obtain a preliminary event category corresponding to the input image; Performing target detection on the input image to obtain the object target contained in the input image; The emergency event category corresponding to the input image is determined according to the preliminary event category corresponding to the input image and the object target contained in the input image.
2. The method according to claim 1, characterized in that The preliminary event categories include specific event categories and other categories; If the acquired preliminary event category corresponding to the input image is a specific event category, the method further includes: Calculating first feature information of the input image using an image encoder of a CLIP model; Calculating second feature information of the category name text of the preliminary event category using a text encoder of the CLIP model; A first similarity value between the first feature information and the second feature information is calculated, and a preliminary event category corresponding to the input image is updated according to the first similarity value.
3. The method according to claim 2, characterized in that Calculating a similarity value between the first feature information and the second feature information, and updating the preliminary event category according to the similarity value, includes: Calculating a first similarity value between the first feature information and the second feature information, and comparing the first similarity value with a first threshold; If the first similarity value is greater than the first threshold, the preliminary event category corresponding to the input image is updated from a specific event category to another category.
4. The method according to claim 1, wherein After performing target detection on the input image and obtaining the object target contained in the input image, the method further includes: intercepting a target image containing the object from the input image; Calculating the third feature information of the target image using an image encoder of a CLIP model; Calculating fourth feature information of the target name text of the object target using a text encoder of the CLIP model; A second similarity value between the third feature information and the fourth feature information is calculated, and the object target included in the input image is updated according to the second similarity value.
5. The method according to claim 4, characterized in that Calculating a second similarity value between the third feature information and the fourth feature information, and updating the object target included in the input image according to the second similarity value, including: Calculating a second similarity value between the third feature information and the fourth feature information, and comparing the second similarity value with a second threshold; If the second similarity value is greater than the second threshold, determining the object target as an invalid detection result, and deleting the invalid detection result; If the second similarity value is greater than the second threshold, the object target is determined as a valid detection result, and the valid detection result is retained.
6. The method according to claim 1, wherein Performing preliminary category recognition based on the overall image content of the input image to obtain a preliminary event category corresponding to the input image includes: A ResNet model is used to perform preliminary category recognition on the overall image content of the input image to obtain a preliminary event category corresponding to the input image.
7. The method according to claim 1, characterized in that Performing target detection on the input image to obtain the object target contained in the input image includes: The YOLOv3 model is used to perform target detection on the input image to obtain the object target contained in the input image.
8. The method according to claim 1, characterized in that After determining the emergency event category corresponding to the input image according to the preliminary event category corresponding to the input image and the object target contained in the input image, the method further includes: Constructing prompt words for verifying the emergency event category according to the emergency event category; Inputting the input image and the prompt word into a visual language model, and obtaining a verification result determined by the visual language model based on the input image and the prompt word; If the verification result is correct, the emergency event category corresponding to the input image is output.
9. An emergency detection device, wherein: The device comprises a memory for storing computer program instructions and a processor for executing the computer program instructions, wherein when the computer program instructions are executed by the processor, the device is triggered to execute the method according to any one of claims 1 to 8.
10. A computer-readable medium having computer program instructions stored thereon, wherein the computer program instructions can be executed by a processor to implement the method according to any one of claims 1 to 8.