Event early warning method, electronic equipment, storage medium and program product
By acquiring images in the factory for target object detection and semantic understanding, using open domain object detection model and multimodal large language model, the problem of inaccurate event warning in the existing technology is solved, and efficient safety hazard identification and early warning is achieved.
Patent Information
- Application Number
- CN202510703531.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-07-01
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
It is difficult for the prior art to capture potential safety hazards in the factory in real time and provide accurate incident warnings, resulting in possible shutdowns, property losses or casualties.
By obtaining the image to be identified, the object detection is detected, the image semantics is understood using the open domain object detection model and the multimodal large language model, determining whether there are abnormal events, and issuing an early warning when an abnormality is detected.
It improves the accuracy of incident warning, reduces false alarms, ensures timely safety intervention, and avoids the occurrence of safety accidents.
Smart Images

Figure CN120236234A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to an event warning method, an electronic device, a storage medium, and a program product. Background Art
[0002] In the daily production of factories, safety production management is always of utmost importance. On the one hand, although some high-risk events (such as operation safety, fire, etc.) are low-probability events, once these events occur, they often have a serious negative impact on personnel safety, production equipment, and corporate reputation. On the other hand, when these dangerous events are detected, they usually have already caused a certain degree of impact on the production line or the factory area, and may even lead to shutdowns, property losses, or casualties.
[0003] Therefore, factories hope to be able to capture potential safety hazard events in real time and issue warnings to assist management personnel in making quick intervention decisions, thereby avoiding the further deterioration of the situation. Summary of the Invention
[0004] Embodiments of this application provide an event warning method, an electronic device, a storage medium, and a program product to achieve timely discovery of abnormal events and accurate event warnings.
[0005] In a first aspect, embodiments of this application provide an event warning method, including: obtaining an image to be recognized, and performing target object detection in the image to be recognized; if a target object is detected in the image to be recognized, then performing image semantic understanding on the image to be recognized to obtain a semantic understanding result; the semantic understanding result indicates whether the target object appears abnormally; if it is determined according to the semantic understanding result that the image to be recognized contains an abnormal event, then an event warning is issued.
[0006] In a second aspect, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored on the memory, and the processor implements the method of any one of the above when executing the computer program.
[0007] In a third aspect, embodiments of this application provide a computer-readable storage medium, in which a computer program is stored, and the computer program implements the method of any one of the above when executed by a processor.
[0008] In a fourth aspect, embodiments of this application provide a computer program product, the computer program product includes a computer program, and the computer program implements the method of any one of the above when executed by a processor.
[0009] Compared with the prior art, this application has the following advantages: The present application provides an event warning method, an electronic device, a storage medium, and a program product. An image to be recognized is acquired, and target objects are detected in the image to be recognized. If a target object is detected in the image to be recognized, then the image to be recognized is subjected to image semantic understanding to obtain a semantic understanding result. The semantic understanding result indicates whether the target object is abnormal. If it is determined according to the semantic understanding result that the image to be recognized contains an abnormal event, then an event warning is issued. In this embodiment, target object detection is first performed in the image to be recognized. If a target object is detected, then image semantic understanding is performed. If it is determined according to the semantic understanding result that the image to be recognized contains an abnormal event, then an event warning is issued. By using semantic recognition to further determine whether the image to be recognized contains an abnormal event, false alarms can be reduced and the accuracy of event warning can be improved.
[0010] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of the present application more obvious and understandable, the following specifically describes the specific embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In the drawings, unless otherwise specified, the same reference numerals throughout the several views denote the same or similar components or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings only depict some embodiments in accordance with the present application and should not be regarded as limiting the scope of the present application.
[0012] Figure 1 It is a schematic diagram of an application scenario of the event warning method according to an embodiment of the present application.
[0013] Figure 2 It is a flowchart of the event warning method according to an embodiment of the present application.
[0014] Figure 3 It is a schematic structural diagram of an initial open-domain target detection model according to an embodiment of the present application.
[0015] Figure 4 It is a flowchart of the training method of the open-domain target detection model according to an embodiment of the present application.
[0016] Figure 5 It is a flowchart of the event warning method according to an embodiment of the present application.
[0017] Figure 6 It is a schematic block diagram of the event warning device according to an embodiment of the present application.
[0018] Figure 7 It is a block diagram of an electronic device for implementing the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] In the following text, only some exemplary embodiments are simply described. As those skilled in the art can recognize, the described embodiments can be modified in various different ways without departing from the concept or scope of the present application. Therefore, the drawings and the description are considered to be exemplary in nature and not restrictive.
[0020] To facilitate the understanding of the technical solutions of the embodiments of the present application, the related technologies of the embodiments of the present application are described below. The following related technologies can be arbitrarily combined with the technical solutions of the embodiments of the present application as optional solutions, and all of them fall within the protection scope of the embodiments of the present application.
[0021] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data that have been authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for the user to choose to authorize or reject.
[0022] Figure 1Schematic diagram of an application scenario of the event warning method provided for this application. First, obtain the image to be recognized. The image to be recognized can be an image input by the user, an image obtained from a preset storage space, or an image collected by an image acquisition device in a preset environment. Input the image to be recognized into the open-domain object detection model to detect whether there is a target object. If not, no processing is performed, thereby filtering out a large number of images irrelevant to abnormal events and improving the performance of image processing. If so, input the target object detected by the open-domain object detection model, the category information of the target object, and the image to be recognized into the multimodal large language model, and use the multimodal large language model to perform image semantic understanding to determine whether the image to be recognized contains an abnormal event. If so, send the image to be recognized and the semantic understanding result to the control center to issue an event warning prompt; otherwise, no processing is performed. Among them, the open-domain object detection model refers to a model that can identify and locate objects in images that are not limited to the object categories known during training, but also include objects that have not been seen or are unknown. This ability enables the model to work in a wider range of scenarios, rather than being limited to a predefined set of closed categories. Traditional object detection models usually need to be trained on specific data sets, and these data sets limit the types of objects that can be detected, while open-domain object detection aims to break through this limitation and enable the model to adapt to unseen object categories. The open-domain object detection model can include, but is not limited to, the open-domain You Only Look Once-world (YOLO-world) model. The multimodal large language model is based on the large language model and combines multimodal inputs such as images and videos to complete the fusion of visual elements, and then realizes cross-modal understanding and generation. Among them, the large language model can be a language model based on deep learning, and this model contains more than one billion parameters. The underlying transformer of the large language model contains a series of neural networks. In one or more embodiments, it can include an encoder and / or a decoder and has a self-attention function. These two types of modules extract meanings from the input text and clarify the structural relationships of the words and phrases in the context. The model has been pre-trained on a public data set containing more than 1TB of text data to learn general language representations.
[0023] In this embodiment, the open-domain object detection model is a small model with relatively few parameters, and the multimodal large language model is a large model with relatively many parameters. Combining the small model and the large model gives full play to the high-efficiency reasoning of the small model and the accurate judgment of the large model, meeting the online usage requirements.
[0024] The embodiment of this application provides an event warning method. The method in this embodiment can be applied to servers, terminal devices, platforms, devices, etc. with computing and processing capabilities. Among them, the server can be a server cluster or a single server, and can be a server deployed in the cloud or a local server.
[0025] As shown Figure 2 in the flowchart of the event warning method according to an embodiment of the present application, which includes: Step S201: Obtain an image to be recognized, and perform target object detection on the image to be recognized.
[0026] Among them, the specific ways to obtain the image to be recognized include at least one of the following: receiving the image to be recognized input by the user, obtaining the image to be recognized from a preset storage space, or collecting the image to be recognized by using an image acquisition device in a preset environment.
[0027] The target object can be any preset type of target, including but not limited to people, objects, etc. In practical applications, a neural network model can be used for target object detection.
[0028] Step S202: If a target object is detected in the image to be recognized, perform image semantic understanding on the image to be recognized to obtain a semantic understanding result; the semantic understanding result represents whether the target object appears abnormally.
[0029] Through target object detection, images irrelevant to abnormal events can be filtered out, improving the performance of image processing. If a target object is detected, further perform image semantic understanding, and determine whether the target object appears abnormally by identifying the relationship between the target object and other objects in the image. In practical applications, a neural network model can be used for image semantic understanding, and the semantic understanding result can represent whether the target object appears abnormally.
[0030] Optionally, if the semantic understanding result represents that the target object appears abnormally, the semantic understanding result may further include the specific category of the abnormality, such as, fire, liquid leakage, etc.
[0031] Step S203: If it is determined according to the semantic understanding result that the image to be recognized contains an abnormal event, perform event warning.
[0032] If it is determined according to the semantic understanding result that the image to be recognized contains an abnormal event, send the image to be recognized and the semantic understanding result to the control center, and issue an event warning prompt, which is convenient for taking timely measures to avoid safety accidents.
[0033] Among them, the abnormal event can be determined according to specific needs, including but not limited to dangerous events in the daily production of factories, such as, fire, operation safety, etc.
[0034] In the embodiment of the present application, the event warning method provided obtains an image to be recognized, and performs target object detection on the image to be recognized; if a target object is detected in the image to be recognized, the image to be recognized is subjected to image semantic understanding to obtain a semantic understanding result; the semantic understanding result represents whether the target object appears abnormally; if it is determined according to the semantic understanding result that the image to be recognized contains an abnormal event, an event warning is performed. In this embodiment, target object detection is first performed on the image to be recognized, which can filter out images irrelevant to abnormal events and improve the performance of image processing. If a target object is detected, image semantic understanding is performed. If it is determined according to the semantic understanding result that the image to be recognized contains an abnormal event, an event warning is performed. By semantic recognition to further determine whether the image to be recognized contains an abnormal event, false alarms can be reduced and the accuracy of event warning can be improved.
[0035] The following introduces the specific implementation processes of the above steps through various implementation methods: In one implementation method, performing target object detection on the image to be recognized includes: using an open-domain target detection model to perform target object detection on the image to be recognized. The open-domain target detection model is obtained by fine-tuning training with training samples. The training samples include positive samples, and the positive samples are obtained through at least one of the following methods: obtaining images containing target objects on the Internet; using an image acquisition device to obtain images containing target objects in a preset environment; using a neural network model to generate images containing target objects based on text information.
[0036] In the related art, when the open-domain target detection model faces rare scenarios such as industrial production safety, due to the lack of samples, it may perform poorly and have poor generalization ability. Therefore, in this embodiment, the training samples of the open-domain target detection model are expanded, positive samples are obtained through various methods, and by enriching the sample data, the generalization ability of the model is enhanced and the accuracy of the recognition result is improved.
[0037] Specifically, an image containing a target can be obtained from the Internet using a data acquisition tool that meets safety requirements. Images containing target objects can be collected in environments such as factory areas or workshops using devices such as cameras. Virtual data can also be constructed using a neural network model through text-to-image generation. For example, a generative artificial intelligence model based on the Diffusion Models is used to generate high-quality images. The generative artificial intelligence model based on the Diffusion Models maps the image to a low-dimensional space through an encoder-decoder, and then performs diffusion and inverse diffusion processes in this low-dimensional space. This method significantly reduces the computational cost and makes it possible to process higher-resolution images.
[0038] Use an open-domain object detection model to detect target objects in the image to be recognized. The detection result is accurate. Moreover, only positive samples are needed to fine-tune the open-domain object detection model, without retraining the model, and the cost is relatively low.
[0039] In one implementation, before using the open-domain object detection model to detect target objects in the image to be recognized, the method further includes: obtaining misdetected images of the open-domain object detection model, determining negative samples based on the misdetected images; and fine-tuning the open-domain object detection model using the positive samples and the negative samples.
[0040] In practical applications, to avoid the problem that the open-domain object detection model has "data bias" due to the lack of negative samples, resulting in inaccurate detection results, negative sample augmentation is performed. By adding data samples that are likely to cause false alarms or missed detections to fine-tune the open-domain object detection model, the open-domain object detection model can better learn to distinguish positive samples (including target objects) and negative samples (excluding target objects), improving the accuracy of object detection.
[0041] For example, in the liquid leakage detection scenario, it may include various types of shadows, plates, and other object images in the factory that may cause misjudgment. Collect these misdetected images as negative samples and fine-tune the open-domain object detection model to improve the accuracy of the detection results.
[0042] In one implementation, determining negative samples based on the misdetected images includes: performing data augmentation operations on the misdetected images to obtain negative samples; where the data augmentation operations include at least one of the following: rotation, flipping, scaling, cropping, or color adjustment.
[0043] In practical applications, in addition to directly using the misdetected images as negative samples, to increase the richness and quantity of negative samples, data augmentation operations can also be performed on the misdetected images to obtain more and richer negative samples, training the open-domain object detection model, and improving the generalization ability of the open-domain object detection model.
[0044] In one implementation, before using the open-domain object detection model to detect target objects in the image to be recognized, the method further includes: freezing the parameters of the image encoder of the initial open-domain object detection model, and using the text encoder of the Contrastive Language–Image Pre-training (CLIP) model as the text encoder of the initial open-domain object detection model to fine-tune the initial open-domain object detection model to obtain the open-domain object detection model.
[0045] In one example, the YOLO-world model is used as the initial open-domain object detection model. The text encoder part of the YOLO-world model is replaced with the text encoder of the CLIP model, and the parameters of the image encoder of the YOLO-world model are frozen. Since the image encoder of the YOLO-world model has been fully trained on a large-scale dataset and has strong generality, there is no need to readjust it. On the other hand, freezing the parameters of the image encoder not only greatly reduces the computational cost but also avoids the overfitting problem caused by insufficient data volume. Since the CLIP model is pre-trained on a large-scale text-image dataset, its text encoder can map natural language to a semantic space aligned with visual features. In this way, it can provide high-quality initial representations for the object categories in scenarios with fewer sample data, improving the model's adaptability to rare scenarios.
[0046] Figure 3 Schematic diagram of the structure of the initial open-domain object detection model. As Figure 3 shown, during the training process, the training samples are text and images. The input text is encoded by the text encoder of the CLIP model to obtain text embedding features, and the input image is encoded by the image encoder of the YOLO-world model to obtain multi-scale image features. Among them, the input image includes one or more objects, for example, men, women, and dogs. The text embedding features and multi-scale image features are input into the visual-language information fusion module, for example, the Vision-Language Pyramid Attention Network (Vision-Language PAN). After the Vision-Language PAN fuses the visual features (multi-scale image features) and text features (text embedding features), it enhances the model's ability to understand cross-modal information, involving multi-level information interaction and integration, and maps the visual features and text features to a common semantic space. The Text Contrastive Head module is used for text contrastive learning to help the model better understand the relationship between text and image and obtain object embedding features. The BoxHead module is used to predict the bounding boxes of the objects in the input image to help the model locate the positions of the objects in the image. The region-text matching module is used to perform region-text matching between the text embedding features and the object embedding features to determine whether the regions (objects) in the image match the given text description, and outputs the objects and the corresponding category information, for example, men, women, dogs, and the corresponding bounding boxes respectively.
[0047] In one implementation, an open-domain object detection model is used to detect target objects in an image to be recognized, including: merging multiple neural network layers of the open-domain object detection model to obtain a merged open-domain object detection model; using the merged open-domain object detection model to detect target objects in the image to be recognized; wherein, the multiple neural network layers include any one of the following: a normalization layer and an activation layer, a convolutional layer and a normalization layer, a convolutional layer and an activation layer, a linear transformation layer and an activation layer.
[0048] When performing inference using the open-domain object detection model, by merging multiple neural network layers into one operator, the number of times of saving and loading intermediate results can be reduced, the number of memory accesses can be reduced, and the data processing performance can be improved.
[0049] Specifically, an activation layer (such as the activation function ReLU) usually follows the normalization layer immediately. These two layers can also be merged into a single operation, directly applying the activation function to the normalized data, reducing unnecessary data transmission and temporary storage. Normalization processing usually follows the convolution operation. These two operations can be fused into one step, so that the parameters of normalization can be directly applied to the result of convolution during actual calculation, reducing the storage requirements and read-write operations of intermediate results. If an activation function is directly connected after the convolutional layer, these two operations can also be attempted to be merged, and this approach may depend on specific hardware support or the capabilities of the deep learning framework. After the fully connected layer or other forms of linear transformation (such as matrix multiplication), if an activation function follows, these layers can also be considered for merging to reduce the number of calculation steps and accelerate the inference process.
[0050] In addition, the data processing process can also be optimized in other ways to improve the data processing speed of the open-domain object detection model: Optionally, inference requests are processed in a multi-way parallel manner. Specifically, multiple processes are started, each process independently loads the model, and concurrently processes multiple object detection requests to improve the throughput of data processing.
[0051] Optionally, multi-batch processing is performed. Specifically, the model processes multiple images in one forward pass, thereby improving the data processing speed.
[0052] Optionally, before the image to be recognized is input into the open-domain object detection model, the size of the image to be recognized is adjusted to the size required by the open-domain object detection model, thereby accelerating the speed of the open-domain object detection model for object detection.
[0053] In one implementation, an open-domain object detection model is used to detect target objects in an image to be recognized, including: quantizing multiple model parameters of the open-domain object detection model to obtain a quantized open-domain object detection model; using the quantized open-domain object detection model to detect target objects in the image to be recognized.
[0054] In practical applications, the inference process of the open-domain object detection model can also be accelerated by quantizing model parameters. Among them, quantization refers to the process of converting high-precision numerical values (such as 32-bit floating-point numbers FP32) into low-precision numerical values (such as 8-bit integers INT8). For example, int8 quantization is performed using the quantization tool TFLite to quantize the model parameters into 8-bit integers.
[0055] Model parameter quantization can reduce the model size, and the model size is usually reduced to 1 / 4 of the original because each parameter is reduced from 4 bytes (FP32) to 1 byte (INT8). It can improve the inference speed because integer operations are faster than floating-point operations.
[0056] In one implementation, image semantic understanding is performed on the image to be recognized to obtain a semantic understanding result, including: inputting the target objects detected in the image to be recognized, the category information of the target objects, and the image to be recognized into a multi-modal large language model to obtain the semantic understanding result output by the multi-modal large language model.
[0057] In practical applications, prompt words can be constructed according to specific needs, and then the prompt words, the target objects detected in the image to be recognized, the category information, and the image to be recognized are input into the multi-modal large language model together for semantic understanding.
[0058] Among them, the multi-modal large language model is based on the large language model, combines multi-modal inputs such as images and texts to complete the fusion of visual elements, and then realizes cross-modal understanding and generation, which can more comprehensively understand and explain complex scenes. Using the multi-modal large language model for image semantic understanding can improve the accuracy of recognition results.
[0059] The embodiments of the present application provide a training method for an open-domain object detection model. The method in this embodiment can be applied to servers, terminal devices, platforms, devices, etc. with computing and processing capabilities. Among them, the server can be a server cluster or a single server, and can be a server deployed in the cloud or a local server.
[0060] As Figure 4 shown in the flowchart of the training method for the open-domain object detection model according to an embodiment of the present application, it includes: Step S401, obtain images containing target objects on the Internet.
[0061] Step S402: Use an image acquisition device to obtain an image containing the target object in a preset environment.
[0062] Step S403: Use a neural network model to generate an image containing the target object based on text information.
[0063] Step S404: Obtain misdetected images of the open-domain object detection model.
[0064] Step S405: Use the images containing the target object on the Internet, the images containing the target object obtained in the preset environment, and the images containing the target object generated based on text information as positive samples, and use the misdetected images as negative samples to fine-tune the open-domain object detection model.
[0065] In this embodiment, positive samples are obtained in multiple ways, misdetected images are used as negative samples, and the open-domain object detection model is trained with positive and negative samples, which can improve the inference accuracy of the open-domain object detection model.
[0066] The embodiment of the present application provides an event warning method. The method in this embodiment can be applied to servers, terminal devices, platforms, devices, etc. with computing and processing capabilities. Among them, the server can be a server cluster or a single server, and can be a server deployed in the cloud or a local server.
[0067] As Figure 5 shown in the flowchart of the event warning method according to an embodiment of the present application, it includes: Step S501: Freeze the parameters of the image encoder of the initial open-domain object detection model, and use the text encoder of the contrastive language-image pre-training CLIP model as the text encoder of the initial open-domain object detection model to fine-tune the initial open-domain object detection model to obtain an open-domain object detection model.
[0068] Step S502: Obtain the image to be recognized, and perform quantization processing on multiple model parameters of the open-domain object detection model to obtain a quantized open-domain object detection model.
[0069] Step S503: Use the quantized open-domain object detection model to detect the target object in the image to be recognized.
[0070] Step S504: If a target object is detected in the image to be recognized, input the target object, the category information of the target object, and the image to be recognized into the multi-modal large language model to obtain the semantic understanding result output by the multi-modal large language model.
[0071] Step S505: If it is determined that the image to be recognized contains an abnormal event according to the semantic understanding result, an event warning is issued.
[0072] Corresponding to the application scenario and method of the method provided in the embodiments of the present application, the embodiments of the present application further provide an event warning device. As Figure 6 shown in the structural block diagram of the event warning device according to an embodiment of the present application, the device includes: A target detection module 601, configured to obtain an image to be recognized and perform target object detection in the image to be recognized.
[0073] A semantic recognition module 602, configured to, if a target object is detected in the image to be recognized, perform image semantic understanding on the image to be recognized to obtain a semantic understanding result; the semantic understanding result represents whether the target object appears abnormally.
[0074] An event warning module 603, configured to perform event warning if it is determined according to the semantic understanding result that the image to be recognized contains an abnormal event.
[0075] For the event warning method provided in the embodiments of the present application, an image to be recognized is obtained, and target object detection is performed in the image to be recognized; if a target object is detected in the image to be recognized, image semantic understanding is performed on the image to be recognized to obtain a semantic understanding result; the semantic understanding result represents whether the target object appears abnormally; if it is determined according to the semantic understanding result that the image to be recognized contains an abnormal event, event warning is performed. In this embodiment, target object detection is first performed in the image to be recognized, which can filter out images irrelevant to abnormal events and improve the performance of image processing. If a target object is detected, image semantic understanding is performed, and if it is determined according to the semantic understanding result that the image to be recognized contains an abnormal event, event warning is performed. By further determining whether the image to be recognized contains an abnormal event through semantic recognition, false alarms can be reduced and the accuracy of event warning can be improved.
[0076] In one implementation, when the target detection module 601 performs target object detection in the image to be recognized, it is configured to: Use an open-domain target detection model to perform target object detection in the image to be recognized. The open-domain target detection model is obtained by fine-tuning training with training samples, and the training samples include positive samples, and the positive samples are obtained through at least one of the following methods: Obtain images containing target objects on the Internet; Use an image acquisition device to obtain images containing target objects in a preset environment; Use a neural network model to generate images containing target objects based on text information.
[0077] In one implementation, the device is further configured to: Before using the open-domain target detection model to perform target object detection in the image to be recognized, obtain misdetected images of the open-domain target detection model, and determine negative samples based on the misdetected images; Fine-tune the open-domain object detection model using positive and negative samples for training.
[0078] In one implementation, when the device determines negative samples based on misdetected images, it is used for: Perform data augmentation operations on the misdetected images to obtain negative samples; Among them, the data augmentation operations include at least one of the following: rotation, flipping, scaling, cropping, or color adjustment.
[0079] In one implementation, the device is also used for: Before using the open-domain object detection model to detect target objects in the image to be recognized, freeze the parameters of the image encoder of the initial open-domain object detection model, and use the text encoder of the contrastive language-image pre-training CLIP model as the text encoder of the initial open-domain object detection model to fine-tune the initial open-domain object detection model to obtain the open-domain object detection model.
[0080] In one implementation, when the target detection module 601 uses the open-domain object detection model to detect target objects in the image to be recognized, it is used for: Merge multiple neural network layers of the open-domain object detection model to obtain the merged open-domain object detection model; Use the merged open-domain object detection model to detect target objects in the image to be recognized; Among them, the multiple neural network layers include any one of the following: normalization layer and activation layer, convolutional layer and normalization layer, convolutional layer and activation layer, linear transformation layer and activation layer.
[0081] In one implementation, when the target detection module 601 uses the open-domain object detection model to detect target objects in the image to be recognized, it is used for: Quantize multiple model parameters of the open-domain object detection model to obtain the quantized open-domain object detection model; Use the quantized open-domain object detection model to detect target objects in the image to be recognized.
[0082] In one implementation, the semantic recognition module 602 is used for: Input the target objects detected in the image to be recognized, the category information of the target objects, and the image to be recognized into the multi-modal large language model to obtain the semantic understanding result output by the multi-modal large language model.
[0083] The functions of the modules in the embodiments of this application can refer to the corresponding descriptions in the above methods and have the corresponding beneficial effects, which will not be elaborated here.
[0084] Figure 7 It is a block diagram of an electronic device for implementing the embodiments of the present application. As Figure 7 shown, the electronic device includes: a memory 710 and a processor 720. The memory 710 stores a computer program that can run on the processor 720. When the processor 720 executes the computer program, the method in the above embodiments is implemented. The number of the memory 710 and the processor 720 can be one or more.
[0085] The electronic device further includes: a communication interface 730, configured to communicate with external devices and perform data interaction and transmission.
[0086] If the memory 710, the processor 720, and the communication interface 730 are implemented independently, the memory 710, the processor 720, and the communication interface 730 can be interconnected through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity in representation, Figure 7 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0087] Optionally, in specific implementation, if the memory 710, the processor 720, and the communication interface 730 are integrated on a chip, the memory 710, the processor 720, and the communication interface 730 can communicate with each other through an internal interface.
[0088] The embodiments of the present application provide a computer-readable storage medium that stores a computer program, and when the program is executed by a processor, the method provided in the embodiments of the present application is implemented.
[0089] The embodiments of the present application provide a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the method provided in the embodiments of the present application is implemented.
[0090] The embodiments of the present application further provide a chip. The chip includes a processor, configured to call and run instructions stored in a memory from the memory, so that a communication device installed with the chip executes the method provided in the embodiments of the present application.
[0091] An embodiment of the present application further provides a chip, including: an input interface, an output interface, a processor, and a memory. The input interface, the output interface, the processor, and the memory are connected through an internal connection path. The processor is configured to execute the code in the memory. When the code is executed, the processor is configured to execute the method provided by the embodiment of the application.
[0092] It should be understood that the above-mentioned processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. It is worth noting that the processor may be a processor supporting the advanced reduced instruction set machine (ARM) architecture.
[0093] Further, optionally, the above-mentioned memory may include a read-only memory and a random access memory. The memory may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may include a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may include a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available. For example, static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM).
[0094] In the above embodiments, it may be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it may be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium.
[0095] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0096] In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of this application, "a plurality of" means two or more unless otherwise specifically defined.
[0097] Any process or method described in the flowchart or otherwise described herein can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a specific logical function or process. And the scope of the preferred embodiments of this application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in a reverse order according to the involved functions, rather than in the order shown or discussed.
[0098] The logic and / or steps described in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing a logical function, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in connection with these instruction execution systems, apparatus, or devices.
[0099] It should be understood that each part of this application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. All or part of the steps of the method in the above embodiments can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0100] In addition, each functional unit in various embodiments of the present application may be integrated into one processing module, or each unit may exist physically alone, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the above-mentioned integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium. The storage medium may be a read-only memory, a magnetic disk, an optical disc, or the like.
[0101] As described above, only the exemplary embodiments of the present application are provided, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope recorded in the present application can easily think of various changes or substitutions, and these should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. An event early warning method, characterized in that, Including: Obtain the image to be recognized and perform target object detection in the image to be recognized; If the target object is detected in the image to be recognized, perform image semantic understanding on the image to be recognized to obtain a semantic understanding result; The semantic understanding result indicates whether the target object is abnormal; If it is determined according to the semantic understanding result that the image to be recognized contains an abnormal event, an event warning is issued.
2. The method according to claim 1, wherein The performing target object detection in the image to be recognized includes: Use an open-domain target detection model to perform target object detection in the image to be recognized. The open-domain target detection model is obtained by fine-tuning training with training samples. The training samples include positive samples, and the positive samples are obtained by at least one of the following methods: Obtain images containing the target object on the Internet; Use an image acquisition device to obtain images containing the target object in a preset environment; Use a neural network model to generate images containing the target object based on text information.
3. The method according to claim 2, wherein Before using the open-domain target detection model to perform target object detection in the image to be recognized, the method further includes: Obtain misdetection images of the open-domain target detection model, and determine negative samples based on the misdetection images; Use the positive samples and the negative samples to perform fine-tuning training on the open-domain target detection model.
4. The method according to claim 3, wherein The determining negative samples based on the misdetection images includes: Perform data augmentation operations on the misdetection images to obtain the negative samples; Wherein, the data augmentation operations include at least one of the following: rotation, flipping, scaling, cropping, or color adjustment.
5. The method according to claim 2, wherein Before using the open-domain target detection model to perform target object detection in the image to be recognized, the method further includes: Freeze the parameters of the image encoder of the initial open-domain target detection model, and use the text encoder of the contrastive language-image pre-training CLIP model as the text encoder of the initial open-domain target detection model to perform fine-tuning training on the initial open-domain target detection model to obtain the open-domain target detection model.
6. The method according to claim 2, wherein, The using the open-domain target detection model to perform target object detection in the image to be recognized includes: Merge multiple neural network layers of the open-domain target detection model to obtain a merged open-domain target detection model; Use the merged open-domain target detection model to perform target object detection in the image to be recognized; Wherein, the multiple neural network layers include any one of the following: a normalization layer and an activation layer, a convolutional layer and a normalization layer, a convolutional layer and an activation layer, a linear transformation layer and an activation layer.
7. The method according to claim 2, wherein The using the open-domain target detection model to perform target object detection in the image to be recognized includes: Quantize multiple model parameters of the open-domain target detection model to obtain a quantized open-domain target detection model; Use the quantized open-domain target detection model to perform target object detection in the image to be recognized.
8. The method according to any one of claims 1 to 7, characterized in that The performing image semantic understanding on the image to be recognized to obtain a semantic understanding result includes: Input the target object detected in the image to be recognized, the category information of the target object, and the image to be recognized into the multimodal large language model, and obtain the semantic understanding result output by the multimodal large language model.
9. An electronic device, characterized in that, The electronic device includes a memory, a processor, and a computer program stored on the memory. When the processor executes the computer program, it implements the method according to any one of claims 1-8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the method according to any one of claims 1-8.
11. A computer program product, characterized in that, The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the method according to any one of claims 1-8.
Citation Information
Patent Citations
Anchor recommendation method and device and storage medium
CN108683927A
A method and system for detecting legitimacy of early warning information based on intelligent semantic perception
CN109543764A
Event detection model construction method and device, electronic equipment and storage medium
CN111813931A
Track line foreign matter invasion monitoring and risk early warning method and system
CN115205796A
Small sample sensitive information identification method based on fine tuning prototype network
CN115409124A