Event identification method and device, storage medium and program product

By combining the target detection model and the target multimodal model, identifying event types in the work scenarios is solved, and traditional methods cannot accurately identify security events and risk events in complex operation environments are achieved, achieving more efficient event identification and risk warning.

CN120220034AActive Publication Date: 2025-06-27ALIBABA CLOUD FEITIAN (HANGZHOU) CLOUD COMPUTING TECH CO LTD

Patent Information

Application Number
CN202510694135.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-06-27
Estimated Expiration
2045-05-27

AI Technical Summary

Technical Problem

Traditional event recognition methods cannot accurately identify different types of security events and risk events in work operation scenarios, especially in complex operation environments.

Method used

Combining the object detection model and the target multimodal model of computer vision, we use the multimodal model to identify event types by obtaining scene image data, detecting target objects, determining work scenarios, and using the multimodal model to identify event types.

Benefits of technology

It realizes more accurate identification of event types in work operation scenarios, and improves the efficiency of operation safety and risk warning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220034A_ABST
    Figure CN120220034A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an event recognition method and device, a storage medium and a program product. In the event identification method, after scene image data in a public work process is obtained, a target detection model can be utilized to perform target detection on the scene image data to obtain attribute information of a target object contained in the scene image data. According to the attribute information of the target object, a target operation scene corresponding to the scene image data can be determined, and a target event type corresponding to the scene image data in the target operation scene can be identified by using the target multi-modal model according to the attribute information of the target object and knowledge information corresponding to different event types in the target operation scene. In the implementation mode, the target detection model based on computer vision is combined with the target multi-modal model, the advantages of the target detection model in the aspect of fine-grained target detection and the powerful ability of the multi-modal model in the aspects of advanced semantic comprehension and reasoning are fully played, and more accurate event recognition is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular, to an event recognition method, device, storage medium, and program product. Background Art

[0002] During the operation of work in fields such as power, transportation, and industrial production, the operation environment where operators and equipment are located is relatively complex. Therefore, there are relatively high requirements for ensuring operation safety and avoiding risks. Identifying safety events and risk events in the operation scene and timely warning or intervening in risk events is a very important link. For example, in the industrial production field, it can be identified whether a welding operator wears safety goggles when performing welding operations. In the power field, it can be identified whether maintenance workers use compliant operation tools when performing power maintenance operations.

[0003] In different fields, the definitions of safety events and risk events in the operation scene are relatively complex, and traditional event recognition means cannot accurately identify different types of events in the operation scene. Therefore, there is a need to propose a new solution. Summary of the Invention

[0004] Embodiments of this application provide an event recognition method, device, storage medium, and program product for accurately identifying the event types in the work operation scene.

[0005] Embodiments of this application provide an event recognition method, including: obtaining scene image data during the work operation; using a target detection model to perform target detection on the scene image data to obtain the attribute information of the target objects included in the scene image data; determining the target operation scene corresponding to the scene image data according to the attribute information of the target objects; using a target multi-modal model to identify the target event type corresponding to the scene image data in the target operation scene according to the attribute information of the target objects and the knowledge information corresponding to different event types in the target operation scene.

[0006] Optionally, determining the target operation scene corresponding to the scene image data according to the attribute information of the target objects includes: obtaining the category information of the target objects according to the attribute information of the target objects; comparing the category information of the target objects with the category information of the necessary objects included in different operation scenes respectively; if the category information of the target objects matches the category information of the necessary objects included in any one of the different operation scenes, determining the any one of the operation scenes as the target operation scene corresponding to the scene image data.

[0007] Optionally, using the target multi-modal model, according to the attribute information of the target object and the knowledge information corresponding to different event types in the target operation scenario, identify the target event type corresponding to the scene image data in the target operation scenario, including: according to the event recognition rule of the target operation scenario, determine whether the target event type corresponding to the scene image data in the target operation scenario can be determined according to the attribute information of the target object; if not, use the target multi-modal model to identify the target event type corresponding to the scene image data in the target operation scenario according to the attribute information of the target object and the knowledge information corresponding to different event types in the target operation scenario.

[0008] Optionally, it further includes: if the target event type corresponding to the scene image data in the target operation scenario can be determined according to the attribute information of the target object, then determine the category of the target object according to the attribute information of the target object; match the category of the target object with the necessary object categories relied on by different event types in the target operation scenario respectively; if the category of the target object is consistent with the necessary object category relied on by any event type in the target operation scenario, then determine the any event type as the target event type and output the target event type.

[0009] Optionally, according to the event recognition rule of the target operation scenario, determine whether the target event type corresponding to the scene image data in the target operation scenario can be determined according to the attribute information of the target object, including: if the event recognition rule of the target operation scenario relies on the category of the object, then determine that the target event type corresponding to the scene image data in the target operation scenario can be determined according to the attribute information of the target object; if the event recognition rule of the target operation scenario relies on the category of the object and the relationship between objects, then determine that the target event type corresponding to the scene image data in the target operation scenario cannot be determined according to the attribute information of the target object.

[0010] Optionally, the attribute information of the target object includes: the coordinate information and category information of the target object; using the target multi-modal model, according to the attribute information of the target object and the knowledge information corresponding to different event types in the target operation scenario, identifying the target event type corresponding to the scene image data in the target operation scenario includes: cropping the scene image data according to the coordinate information of the target object to obtain a scene sub-graph corresponding to the target object; constructing a target prompt according to the scene sub-graph, the coordinate information and category information of the target object, and the knowledge information corresponding to different events in the target operation scenario; calling the target multi-modal model according to the target prompt to use the target multi-modal model to identify the target event type corresponding to the scene image data in the target operation scenario.

[0011] Optionally, calling the target multi-modal model according to the target prompt to use the target multi-modal model to identify the target event type corresponding to the scene image data in the target operation scenario includes: inputting the target prompt into the target multi-modal model; in the target multi-modal model, obtaining the coordinate information and category information of each of the multiple objects in the target object from the target prompt, and extracting the semantic features of the scene sub-graphs of the multiple objects; identifying the logical relationship between the multiple objects according to the coordinate information, category information, and semantic features of the multiple objects; respectively matching the knowledge information corresponding to different event types according to the logical relationship between the multiple objects to obtain target knowledge information adapted to the logical relationship between the multiple objects; determining the event type corresponding to the target knowledge information as the event type corresponding to the scene image data in the target operation scenario.

[0012] Optionally, the target multi-modal model is obtained by lightweight compression of a pre-trained multi-modal large model, and the target multi-modal model is deployed on a terminal device or an edge device.

[0013] The embodiment of the present application further provides an electronic device, including: a memory and a processor; the memory is used to store one or more computer instructions; the processor is used to execute the one or more computer instructions for: executing the steps in the method provided by the embodiment of the present application.

[0014] The embodiment of the present application further provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, it can implement the steps in the method provided by the embodiment of the present application.

[0015] An embodiment of the present application also provides a computer program product, including: computer programs / instructions, which can implement the steps in the method provided by the embodiment of the present application when executed by a processor.

[0016] In the embodiment of the present application, after obtaining the scene image data during the track maintenance operation, the target detection model can be used to perform target detection on the scene image data to obtain the attribute information of the target objects included in the scene image data. According to the attribute information of the target objects, the target operation scene corresponding to the scene image data can be determined, and the target multi-modal model can be used to identify the target event type corresponding to the scene image data in the target operation scene according to the attribute information of the target objects and the knowledge information corresponding to different event types in the target operation scene. In this implementation manner, the combination of the target detection model based on computer vision and the target multi-modal model gives full play to the advantages of the target detection model in fine-grained target detection and the powerful capabilities of the multi-modal model in advanced semantic understanding and reasoning, realizing more accurate event recognition.

[0017] Secondly, taking the target detection model as the pre-model of the target multi-modal model can enable the target multi-modal model to focus on advanced semantic understanding and reasoning, which is beneficial to reducing the number of parameters of the target multi-modal model, can significantly reduce the engineering implementation complexity while ensuring the semantic understanding ability, and thus facilitate the engineering deployment of the target multi-modal model. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The drawings described herein are used to provide a further understanding of the present application, form a part of the present application, and the schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings: Figure 1 is a schematic flow chart of an event recognition method provided by an exemplary embodiment of the present application; Figure 2 is a schematic flow chart of an event recognition method provided by another exemplary embodiment of the present application; Figure 3 is a schematic structural diagram of an electronic device provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] To make the purpose, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0020] The terms used in the embodiments of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The singular forms "a", "the", and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. "Plural" generally includes at least two, but does not exclude the case of including at least one.

[0021] It should be understood that the term "and / or" used herein is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after.

[0022] It should also be noted that the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, such that a product or system including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such product or system. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the product or system including the said element.

[0023] During the operation process of public works in fields such as electricity, transportation, and industrial production, the operation environment where the operator and the equipment are located is relatively complex. Therefore, the requirements for ensuring operation safety and avoiding risks are relatively high. Identifying safety events and risk events in the operation scenario and timely warning or intervening in risk events are very important links. For example, in the industrial production field, it can be identified whether a welding operator wears safety goggles when performing welding operations. In the electricity field, it can be identified whether maintenance workers use compliant operation tools when performing electricity maintenance operations. In different fields, the definitions of safety events and risk events in the operation scenario are relatively complex, and traditional event identification means cannot accurately identify different types of events in the operation scenario.

[0024] In one solution, the event type in the operation scenario can be identified based on a computer vision (CV) model and event discrimination rules. However, in this solution, the scale of the CV model is small, the generalization ability is poor, and the determination accuracy is not high. When facing an operation scenario with relatively complex semantic requirements, it cannot accurately identify the event type in the operation scenario.

[0025] In other solutions, a multimodal large model with a relatively large number of parameters can be directly used to understand the pictures taken during the operation process and determine events. However, when multimodal large models perform fine-grained object recognition tasks for specific scenarios, their performance may be inferior to that of small models specifically designed for such tasks. In addition, multimodal large models have a large number of parameters, require more data for model fine-tuning, and need more machine resources and higher costs for deploying such multimodal large models, which is not conducive to engineering deployment in actual production environments with limited resources.

[0026] In view of the above technical problems, in some embodiments of the present application, a solution is provided. The following will, in conjunction with the accompanying drawings, detail the technical solutions provided by each embodiment of the present application.

[0027] Figure 1 is a schematic flowchart of an event recognition method provided by an exemplary embodiment of the present application. The method may include steps as Figure 1 shown below: Step 101: Obtain scene image data during the maintenance operation process.

[0028] Step 102: Use an object detection model to perform object detection on the scene image data to obtain the attribute information of the target objects included in the scene image data.

[0029] Step 103: Determine the target operation scene corresponding to the scene image data according to the attribute information of the target objects.

[0030] Step 104: Use a target multimodal model to identify the target event type corresponding to the scene image data in the target operation scene according to the attribute information of the target objects and the knowledge information corresponding to different event types in the target operation scene.

[0031] This embodiment is applicable to the maintenance operation scene and is used to analyze the scene image data during the maintenance operation process to identify the event types during the maintenance operation process. Among them, the maintenance operation refers to the operations involved in industrial production and engineering construction, and is mainly used for tasks related to infrastructure construction, equipment maintenance, process operation, and technical support, etc. The maintenance operation may include, but is not limited to: installation and commissioning of production equipment, construction of engineering structures (such as welding, cutting), infrastructure operation and maintenance (such as railway traction, power inspection), safety hazard investigation, and implementation of process standards, etc. The operation environment of the maintenance operation is usually relatively complex and relies on professional technologies and standardized processes to ensure safe and efficient implementation.

[0032] For example, in the scenario of installation and commissioning of production equipment, operators need to wear appropriate personal protective equipment, such as safety helmets, safety shoes, etc., especially when dealing with heavy machinery. In the welding operation scenario, operators need to wear goggles. In the cutting operation scenario, operators need to wear goggles, and an appropriate safety distance should be maintained between the operator and the cutting point to avoid injury caused by spark splashing. For example, in the power inspection operation scenario, operators should wear gloves and insulating boots to protect the whole body from the risk of electric shock.

[0033] This embodiment can be executed by an event recognition device, which can run on a single computer device or in a distributed manner on multiple computer devices, and this embodiment is not limited. In some engineering deployment scenarios, the event recognition device can run on a local terminal device or an edge device close to the railway maintenance operation environment for quick response.

[0034] In step 101, the railway maintenance operation process can be any kind of railway maintenance operation process. During the railway maintenance operation, scene image data can be captured, and the scene image data is used to describe the operation scene from a visual dimension. In some embodiments, the event recognition device can obtain the video or image during the railway maintenance operation through the image acquisition device deployed in the railway maintenance operation environment. If the video is obtained, some static frames can be extracted from the video to obtain the scene image data.

[0035] In step 102, the event recognition device can use the object detection model to perform object detection on the scene image data to obtain the attribute information of the target object included in the scene image data. Among them, the attribute information of the target object can include the position information and / or category information of the target object in the scene image data.

[0036] Among them, the object detection model is a machine learning model used to identify specific targets (such as people, vehicles, animals, actions, etc.) in images or videos. Its core task is to locate the position of the target (output the bounding box), or classify the category of the target, or classify the category of the target while locating the position of the target. In some engineering deployment scenarios, the object detection model can run on a terminal device or an edge device close to the railway maintenance operation environment, and thus can quickly respond to the call of the event recognition device to meet the real-time warning requirements in the railway maintenance operation scenario. Of course, in some other alternative embodiments, the object detection model can also be deployed on a cloud server for remote call by the event recognition device.

[0037] The event recognition device can send a detection request to the object detection model through the interface opened by the object detection model, and the detection request carries scene image data. The object detection model can take the scene image data as input and perform feature extraction operations on the scene image. Among them, the feature extraction operations can include: using a convolutional neural network (CNN) to extract features in the image and capture information such as the shape and texture of the object. Based on the extracted features, candidate regions (i.e., potential bounding boxes) that may contain the target object are generated, or the bounding box is directly predicted based on the extracted features. If candidate regions that may contain the target object are generated, classification and regression algorithms can be used to classify each candidate region, that is, to determine the category of the target object corresponding to the candidate region. At the same time, the position and size of the bounding box can be adjusted to more accurately enclose the target object. Then, the non-maximum suppression (NMS) algorithm can be used to remove redundant overlapping bounding boxes and retain more accurate detection results. Based on the above embodiments, the object detection model can output the category label and the corresponding bounding box coordinates of each target object in the scene image data to the event recognition device. The event recognition device can use the category label and / or the bounding box coordinates of any target object as the attribute information of the target object and perform step 103 based on this attribute information.

[0038] In step 103, the event recognition device can determine the target operation scene corresponding to the scene image data according to the attribute information of the target object. In some alternative embodiments, the event recognition device can call a neural network model capable of understanding context semantics and determine the target operation scene corresponding to the scene image data according to the attribute information of the target object.

[0039] In other alternative embodiments, the event recognition device can use a rule engine to perform rule matching on the attribute information of the target object based on the knowledge in the field of railway maintenance operations to determine the target operation scene corresponding to the scene image data. Optionally, different operation scenes have different necessary conditions to ensure the safety, efficiency, and quality of operations. For example, the necessary conditions for the production equipment installation and commissioning scene are: the scene contains target objects such as operators, safety helmets, and safety shoes. The necessary conditions for the welding operation scene are that the scene contains target objects such as operators, welding tools, goggles, masks, and sparks. The necessary conditions for the power inspection operation scene are: operators, gloves, insulating boots, and other target objects.

[0040] The rule engine can traverse the necessary conditions corresponding to different operation scenarios based on the attribute information of the target object. When traversing any operation scenario, the rule engine can determine whether the attribute information of the target object can meet the necessary conditions corresponding to that operation scenario. If it can be met, it is determined that the operation scenario is the target operation scenario corresponding to the scenario image data. If it cannot be met, the next operation scenario can be traversed, and the necessary condition judgment operation can be performed based on the attribute information of the target object, which will not be elaborated further.

[0041] In some optional embodiments, the necessary condition corresponding to any operation scenario refers to the category of objects that should be included in that operation scenario. Based on this, a way for the event recognition device to determine the target operation scenario corresponding to the scenario image data may include: obtaining the category information of the target object according to the attribute information of the target object; comparing the category information of the target object with the category information of the necessary objects included in different operation scenarios respectively; if the category information of the target object matches the category information of the necessary objects included in any one of the different operation scenarios, it is determined that the operation scenario corresponding to the scenario image data is the target operation scenario. For example, the necessary conditions for the cutting operation scenario are that the scenario needs to include objects such as a cutting machine and an operator. If the target objects identified in the scenario image data include a cutting machine and an operator, etc., it is determined that the target operation scenario corresponding to the scenario image data is the cutting operation scenario. Based on this condition matching method, the scenario image data can be accurately associated with the target operation scenario.

[0042] After determining the target operation scenario, in step 104, the event recognition device can use the target multi-modal model to identify the target event type corresponding to the scenario image data in the target operation scenario according to the attribute information of the target object and the knowledge information corresponding to different event types in the target operation scenario.

[0043] In this embodiment, the target multi-modal model is a machine learning model involving multiple input sources or data types (such as vision, audition, text), and can comprehensively use various information for analysis and decision-making. The multi-modal model is a neural network model that can process and understand various types of data (such as text, image, audio, video, etc.), and can capture the associations and semantics between different modalities by fusing information from different modalities, thereby enhancing the ability to understand complex scenarios and the ability to interact. The target multi-modal model can be a neural network model based on the Transformer architecture. In this embodiment, the information in the text modality may at least include: the text information in the detection result obtained by target detection and the knowledge information corresponding to different event types in the target operation scenario; the information in the image modality may include: the detection frame of the target object extracted by target detection and / or the original scenario image data, which is not limited in this embodiment.

[0044] In this embodiment, the target multimodal model can be a small or medium-sized multimodal model suitable for deployment on the terminal side or the edge side. Compared with large-scale multimodal models, small and medium-sized multimodal models have fewer parameters. For example, the number of parameters usually ranges from several million to hundreds of millions, and they are suitable for deployment on devices with limited resources. The training process of the target multimodal model will be introduced in subsequent embodiments and will not be elaborated here.

[0045] Among them, an event refers to one or a group of behaviors that occur in the operation scenario. The event types corresponding to different operation scenarios may be different. In some embodiments, the event types in the operation scenario may include, but are not limited to: equipment failure events, safe operation events, and illegal operation events, etc.

[0046] Among them, the knowledge information corresponding to different event types is used to describe the tools, equipment, materials, safety equipment and skill requirements of operators, and specific operation environment conditions required for performing operations of different event types in the railway maintenance operation scenario. For example, in the excavation operation scenario, the knowledge information corresponding to the safe operation event is: excavation tools such as excavators or shovels and excavation location markers should be available in the excavation operation environment; the operator should wear a safety helmet, gloves and protective boots, and wear a reflective vest as needed. In the excavation operation scenario, the knowledge information corresponding to the illegal operation event is: the excavation tools or excavation location markers are missing in the excavation operation environment; the operator does not wear at least one of a safety helmet, gloves, protective boots and a reflective vest.

[0047] For example, in the high-altitude operation scenario, the knowledge information corresponding to the safe operation event is: scaffolding, ladders or aerial work platforms should be equipped in the high-altitude operation environment; the operator must wear a safety belt and a helmet, and ensure that the safety belt is connected to a reliable fixed point; weather conditions also need to be considered for high-altitude operations, such as strong winds or rain and snow weather may not be suitable for operations. In the high-altitude operation scenario, the knowledge information corresponding to the illegal operation event is: at least one of scaffolding, ladders and aerial work platforms is not equipped in the high-altitude operation environment; the operator does not wear a safety belt or a helmet, or the safety belt is not connected to a reliable fixed point; or, the operation environment is exposed to strong winds or rain and snow weather.

[0048] Another example is that in the electrical installation operation scenario, the knowledge information corresponding to the safe operation event is: electrical components such as wires, sockets, switches, and distribution boxes should be prepared at the electrical installation operation site; the operator should carry an electrician's tool kit (screwdrivers, wire strippers, etc.), and wear insulating shoes and use insulating gloves. In the electrical installation operation scenario, the knowledge information corresponding to the illegal operation event is: at least one of wires, sockets, switches and distribution boxes is not prepared at the electrical installation operation site; the operator does not carry an electrician's tool kit (screwdrivers, wire strippers, etc.), or the operator does not wear insulating shoes or does not use insulating gloves.

[0049] In this embodiment, the knowledge information corresponding to different event types can be retrieved by the event recognition device from a specified knowledge base using Retrieval Augmented Generation (RAG). The specified knowledge base stores standardized knowledge information corresponding to different event types in different railway maintenance operation scenarios, and this knowledge information can be obtained by sedimenting relevant documents such as industry standards and specifications. In some embodiments, the event recognition device can obtain the input data of the target multi-modal model at least based on the attribute information of the target object and the knowledge information corresponding to different event types in the target operation scenario. The target multi-modal model can combine the learned semantic understanding ability to understand the attribute information of the target object and the knowledge information corresponding to different event types in the target operation scenario, so as to identify the target event type corresponding to the scene image data in the target operation scenario.

[0050] In this embodiment, by combining the object detection model based on computer vision with the target multi-modal model, the advantages of the object detection model in fine-grained object detection and the powerful capabilities of the multi-modal model in advanced semantic understanding and reasoning are fully utilized to achieve more accurate event recognition. Secondly, taking the object detection model as the pre-model of the target multi-modal model can enable the target multi-modal model to focus on advanced semantic understanding and reasoning, which is beneficial to reducing the number of parameters of the target multi-modal model. While ensuring the semantic understanding ability, it can significantly reduce the complexity of engineering implementation, thus facilitating the engineering deployment of the target multi-modal model. In addition, the object detection model and the target multi-modal model can be optimized independently, which is convenient for flexible version iteration.

[0051] In some alternative embodiments, the multi-modal model is trained on a dataset in the field of railway maintenance operations. The following will give an exemplary description of the alternative training methods.

[0052] Optionally, in the data collection and preprocessing stage, a dataset containing multiple modalities can be collected from railway maintenance operation scenarios, such as video data, picture data captured in railway maintenance operation scenarios, and text data corresponding to different operation scenarios. The text data in any operation scenario can be used to describe concepts, definitions, recognition rules, knowledge information, etc. of different event types in that operation scenario. For each modality of data, preprocessing can be performed separately, and this preprocessing can include operations such as text tokenization, image normalization, and audio feature extraction.

[0053] During the encoding stage, a dedicated encoder can be used to convert data of different modalities into a unified representation. For example, a Transformer or other encoder can be used to encode text data, a Convolutional Neural Network (CNN) or a Vision Transformer (ViT) can be used to encode image data, and a temporal model (such as an RNN or WaveNet) can be used to extract the spectral features of audio data.

[0054] During the cross-modal alignment and fusion stage, specific mechanisms (such as the attention mechanism and contrastive learning) can be used to align the features of different modalities, making the features of different modalities comparable in a shared semantic space. After aligning the features of different modalities, the multi-modal features can be fused to form a joint representation, and the fused multi-modal features can be input into a multi-modal model.

[0055] During the model training stage, the task objective can be designed as an event type recognition task, and a loss function can be defined. After obtaining the prediction results of the multi-modal model for the input data, the prediction loss of the multi-modal model can be calculated according to the loss function. Then, the model parameters can be optimized with the goal of minimizing the prediction loss until the prediction loss converges to a specified range.

[0056] In some other alternative embodiments, the target multi-modal model can be obtained by lightweight compression of a pre-trained large-scale multi-modal model. A large multi-modal model refers to a large-scale model that can process and understand data from multiple input types. These inputs can include text, images, audio, video, etc. The large multi-modal model has strong language understanding capabilities and can integrate information from other modalities. Optionally, the lightweight target multi-modal model can be deployed on a terminal device or an edge device.

[0057] Among them, the optional ways to compress the large-scale multi-modal model can include: pruning, quantization, or knowledge distillation of the large-scale multi-modal model. Among them, pruning refers to removing unimportant weights or neurons in the large-scale multi-modal model to simplify the model structure. Quantization refers to converting the parameters in the model from high precision (such as 32-bit floating-point numbers) to low precision (such as 8-bit integers or lower) to reduce storage requirements and accelerate the inference process. Knowledge distillation refers to using the knowledge of a large teacher model to guide the learning of a small student model, so that the student model can have a smaller scale while maintaining a high accuracy. The following will take knowledge distillation as an example for specific illustrative description.

[0058] Optionally, knowledge distillation technology can be used to perform knowledge transfer on the large multi-modal model to obtain the target multi-modal model. The embodiments of the present application do not limit the number of model parameters supported by the large multi-modal model.

[0059] Specifically, a large-scale multi-modal model in a general domain can be selected as the teacher model. This teacher model is usually trained on a large-scale dataset and performs well on multiple tasks. In addition, a small and medium-sized model can be designed as the student model, which has fewer parameters and lower computational complexity, making it easier to be deployed in resource-constrained environments. To perform effective distillation, relevant data in the railway maintenance operation scenario can be collected to obtain a domain training dataset, ensuring that the student model can learn knowledge in the railway maintenance operation domain. During the knowledge transfer process, the teacher model can be used to predict the training samples to obtain the "soft" labels of each sample. These soft labels not only contain the correct class information but also the confidence scores of the teacher model for other classes. During the training of the student model, the cross-entropy loss between the predicted labels of the student model and the true labels of the training samples can be minimized, and at the same time, the difference between the predicted labels output by the student model and the soft labels provided by the teacher model can be minimized. After training, the training dataset in the railway maintenance operation scenario can be used to further fine-tune the student model obtained by knowledge transfer to obtain a target multi-modal model that is more adapted to the railway maintenance operation domain.

[0060] Based on this implementation, the knowledge of the large model can be effectively transferred to the small model, thereby obtaining a domain-specific model that not only has a certain accuracy but also can run in an environment with limited resources. On the one hand, the deployability of the target multi-modal model can be improved to meet the requirements of engineering deployment; on the other hand, this small model focuses on a specific task domain and can provide more accurate and efficient recognition results.

[0061] In some alternative embodiments, before using the target multi-modal model for event recognition, the event recognition device can further determine whether the recognition of the event type corresponding to the scene image data strongly depends on the target multi-modal model. If it strongly depends on the target multi-modal model, the target multi-modal model is used to recognize the target event type corresponding to the scene image data in the target operation scenario. Otherwise, other algorithms can be used to recognize the target event type corresponding to the scene image data in the target operation scenario. This other algorithm refers to an algorithm with fewer parameters, less computational complexity, and less required resources than the target multi-modal model, such as a logical judgment algorithm or a decision tree algorithm. An exemplary illustration will be provided below.

[0062] Optionally, after determining the target operation scenario corresponding to the scene image data, the event recognition device can judge whether the target event type corresponding to the scene image data in the target operation scenario can be determined according to the attribute information of the target object according to the event recognition rules of the target operation scenario. The event recognition rules refer to a set of predefined logical conditions used to describe the conditions on which the recognition of different event types in the target operation scenario depends.

[0063] In some alternative embodiments, the event recognition rules are constructed based on the categories of the objects dependent on the target operation scenario and / or the relationships between the objects. For example, the safety operation event recognition rules in the tractor operation scenario include: the operation scenario includes an "operator" and a "tractor". The safety operation event recognition rules in the welding operation scenario are: the operation scenario includes objects such as an "operator", a "cutting machine", "goggles", and a "protective mask", the distance between the "operator" and the "cutting machine" is greater than 50 cm, and the "operator" needs to wear "goggles" and a "protective mask".

[0064] Optionally, if the event recognition rules of the target operation scenario depend on the categories of the objects, the event recognition device can determine the target event type corresponding to the scene image data in the target operation scenario according to the attribute information of the target object. For example, the safety operation event recognition rules in the tractor operation scenario only depend on the categories of the objects and do not depend on the relationships between the objects. The event recognition device can directly determine the target event type corresponding to the scene image data in the tractor operation scenario according to the attribute information of the target object.

[0065] Optionally, if the event recognition rules of the target operation scenario depend on the categories of the objects and the relationships between the objects, it is determined that the target event type corresponding to the scene image data in the target operation scenario cannot be determined according to the attribute information of the target object. For example, the safety operation event recognition rules in the welding operation scenario not only depend on the categories of the objects but also depend on the relationships between the objects. The event recognition device can determine that the target event type corresponding to the scene image data in the tractor operation scenario cannot be directly determined according to the attribute information of the target object. Based on the above embodiments, according to the event recognition rules of the target operation scenario, a logical judgment method can be used to judge whether the target event type can be directly determined according to the attribute information of the target object. Compared with the neural network model, the logical judgment method is relatively lightweight and reduces the consumption of computing resources.

[0066] Optionally, if the target event type corresponding to the image data in the target operation scenario cannot be determined based on the attribute information of the target object, then based on the attribute information of the target object and the knowledge information corresponding to different event types in the target operation scenario, identify the target event type corresponding to the scenario image data in the target operation scenario. Optionally, if the target event type corresponding to the scenario image data in the target operation scenario can be determined based on the attribute information of the target object, the event recognition device may output the target event type. Specifically, the category of the target object can be determined according to the attribute information of the target object, and the category of the target object is respectively matched with the necessary object categories relied on by different event types in the target operation scenario. If the category of the target object is consistent with the necessary object category relied on by any event type in the target operation scenario, then determine that the event type is the target event type. For example, in the tractor operation scenario, the necessary object categories relied on by the safe operation event include: the operator and the tractor, and the necessary object categories relied on by the illegal operation event include: the operator or the tractor. If the category of the target object includes: the operator and the tractor, then determine that the target event type corresponding to the scenario image data in the tractor operation scenario is the safe operation type; conversely, if the category of the target object does not include the operator or the tractor, then determine that the target event type corresponding to the scenario image data in the tractor operation scenario is the illegal operation type.

[0067] That is to say, in the case where the target event type can be directly determined based on the attribute information of the target object, there is no need to initiate a call to the target multi-modal model. On the one hand, the overall event recognition efficiency can be improved, and on the other hand, the unnecessary calls to the target multi-modal model can be effectively reduced, the number of runs of the target multi-modal model can be reduced, thereby reducing resource consumption.

[0068] In some alternative embodiments, the target multi-modal model is a multi-modal model, and the attribute information of the target object includes: the coordinate information and category information of the target object. Correspondingly, one way for the event recognition device to use the target multi-modal model to identify the target event type corresponding to the scene image data in the target operation scenario according to the attribute information of the target object and the knowledge information corresponding to different event types in the target operation scenario may include: cropping the scene image data according to the coordinate information of the target object to obtain a scene sub-graph corresponding to the target object, and constructing a target prompt word according to the scene sub-graph, the coordinate information and category information of the target object, and the knowledge information corresponding to different events in the target operation scenario. In this embodiment, a prompt word template may be preset, and the prompt word template may include fields such as a task type field, a knowledge information field, a scene sub-graph field of the target object, a coordinate field, and a category field, etc. The event recognition device may set the value of the task type field to the event type recognition task, and fill in the obtained above information into the fields of the prompt word template correspondingly to obtain the target prompt word. The event recognition device may call the target multi-modal model according to the target prompt word to use the target multi-modal model to identify the target event type corresponding to the scene image data in the target operation scenario.

[0069] In this implementation manner, the input data of the target multi-modal model includes not only text data but also image data such as scene sub-graphs, enabling the target multi-modal model to not only understand text descriptions but also more accurately identify event types by analyzing the logical relationships and context information in the images. In addition, inputting the scene sub-graph into the target multi-modal model can prevent errors from being directly passed to the inference stage of the target multi-modal model when errors occur in the visual feature extraction stage of the target detection model, which helps reduce error accumulation.

[0070] Optionally, when the event recognition device calls the target multi-modal model according to the target prompt word, it may input the target prompt word into the target multi-modal model.

[0071] In this target multi-modal model, coordinate information and category information of each of multiple objects in the target object are obtained from the target prompt, and semantic features of scene sub-graphs of the multiple objects are extracted. In some embodiments, if the target multi-modal model is a model based on the Transformer architecture or its variant, the target multi-modal model can capture semantic information in the scene sub-graph through the attention mechanism. In other embodiments, the target multi-modal model may include a convolutional neural network or a vision transformer, and may extract semantic features in the scene sub-graph based on the convolutional neural network or the vision transformer. According to the coordinate information, category information, and semantic features of each of the multiple objects, the logical relationship between the multiple objects is identified. Among them, the logical relationship may include, but is not limited to: a holding relationship, a traction relationship, a wearing relationship, or a using relationship, etc. Specifically, the target multi-modal model can extract high-dimensional semantic features in the scene sub-graph of each object, and determine the logical relationship between different objects according to the high-dimensional semantic features and the coordinate information of each object. For example, when it is determined that the head and the safety helmet overlap in position according to the coordinate information of the operator's head and the safety helmet, and the feature matching degree between the scene sub-graph of the head and the scene sub-graph of the safety helmet is relatively high, the relationship between the head and the safety helmet can be determined to be a wearing relationship.

[0072] The target multi-modal model can match the knowledge information corresponding to different event types according to the logical relationship between the multiple objects, obtain target knowledge information adapted to the logical relationship between the multiple objects, and use the event type corresponding to the target knowledge information as the event type corresponding to the scene image data in the target operation scene.

[0073] For example, if the target operation scene is an excavation operation scene, the target multi-modal model can match the knowledge information corresponding to the safety operation events and violation operation events in the excavation operation scene according to the logical relationship between the multiple objects. If the logical relationship between the multiple objects is: the operator (object a1) wears (relationship r1) a safety helmet (object a2), gloves (object a3), and safety boots (object a4), and wears (relationship r2) a reflective vest (object a5), then it is determined that the event type corresponding to the scene image data in the excavation operation scene is a safety operation event.

[0074] For another example, if the target operation scene is a high-altitude operation scene, the target multi-modal model can match the knowledge information corresponding to the safety operation events and violation operation events in the high-altitude operation scene according to the logical relationship between the multiple objects. If the logical relationship between the multiple objects is: the operator (object b1) wears (relationship r3) a safety belt (object b2) and a helmet (object b3), but the safety belt (object b2) is not connected (relationship r4) to a reliable fixing point (object b4), then it is determined that the event type corresponding to the scene image data in the excavation operation scene is a violation operation event.

[0075] Based on the above embodiments, the multi-modal target multi-modal model can understand the logical relationships between multiple objects according to the input data, and accurately map the input data to the corresponding event types according to the logical relationships and the knowledge information of different events, realizing the automatic understanding of complex operation scenarios.

[0076] The following will combine Figure 2 , taking the identification of abnormal events in the power safety monitoring operation scenario as an example, to further exemplarily illustrate the event identification method provided by the embodiments of the present application.

[0077] As Figure 2 shown, in the step of inputting an image, image data taken during the power inspection process can be obtained. In the step of target detection, a target detection model can be used to identify the target objects in the image, such as key devices (such as welding equipment, safety helmets), staff and their actions. In the step of associated scenario judgment, according to the information of the target objects detected in the previous step, it can be judged whether the current operation scenario conforms to a certain common event pattern, such as a device failure event, a violation operation event. In the step of event clear judgment, if the event type cannot be directly determined based on the results of target detection, the scene sub-graph generation and prompt word construction steps are executed. In the scene sub-graph generation and prompt word construction steps, a scene sub-graph of the target object is further generated, and based on the information such as the scene sub-graph, category, coordinates of the target object, and the operation scenario to which the target object belongs, a prompt word is generated. In the judgment step of the multi-modal model, the generated prompt word can be input into the multi-modal model for event type judgment, and the judgment result output by the multi-modal model can be obtained. If it is determined that the event type is a violation event, relevant personnel can be prompted to handle it in time.

[0078] In this embodiment, the target detection model based on computer vision is combined with the multi-modal model, giving full play to the advantages of the target detection model in fine-grained target detection and the powerful capabilities of the multi-modal large model in cross-modal understanding and reasoning, to achieve more accurate event identification. Secondly, the multi-modal model can be a multi-modal large model with a relatively small number of parameters, which can reduce the consumption of computing resources while ensuring the understanding and reasoning effects, and improve the real-time performance and deployability of the system. In addition, in this embodiment, by executing the steps of associated scenario judgment and event clear judgment, the unnecessary calls to the multi-modal model can be effectively reduced, improving the overall identification efficiency.

[0079] It should be noted that the execution entity of each step of the method provided in the above embodiments can be the same device, or the method can also be executed by different devices as the execution entity. For example, the execution entity of steps 101 to 104 can be device A; for another example, the execution entity of steps 101 and 102 can be device A, and the execution entity of step 103 can be device B; and so on.

[0080] In addition, in some of the processes described in the above embodiments and the accompanying drawings, there are multiple operations that appear in a specific order. However, it should be clearly understood that these operations can be executed not in the order in which they appear in this document or in parallel. The operation numbers such as 101, 102, etc. are only used to distinguish different operations, and the numbers themselves do not represent any execution order. In addition, these processes can include more or fewer operations, and these operations can be executed sequentially or in parallel. It should be noted that the descriptions such as "first" and "second" in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequence, nor do they limit that "first" and "second" are of different types.

[0081] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.

[0082] Figure 3 The figure schematically shows a structural diagram of an electronic device provided by an exemplary embodiment of the present application, and the event recognition device provided in the foregoing embodiment can run on this electronic device. As Figure 3 shown, the electronic device includes: a memory 301 and a processor 302.

[0083] The memory 301 is used to store computer programs and can be configured to store various other data to support operations on the electronic device. Examples of these data include instructions for any application program or method for operating on the electronic device, data structures, contact data, phone book data, messages, pictures, videos, etc.

[0084] A processor 302, coupled to a memory 301, is configured to execute a computer program in the memory 301 for: obtaining scene image data during a maintenance operation; using an object detection model to perform object detection on the scene image data to obtain attribute information of target objects included in the scene image data; determining a target operation scene corresponding to the scene image data according to the attribute information of the target objects; and using a target multi-modal model to identify a target event type corresponding to the scene image data in the target operation scene according to the attribute information of the target objects and knowledge information corresponding to different event types in the target operation scene.

[0085] Optionally, when determining the target operation scene corresponding to the scene image data according to the attribute information of the target objects, the processor 302 is specifically configured to: obtain category information of the target objects according to the attribute information of the target objects; compare the category information of the target objects with the category information of necessary objects included in different operation scenes respectively; and if the category information of the target objects matches the category information of necessary objects included in any one of the different operation scenes, determine the any one of the operation scenes as the target operation scene corresponding to the scene image data.

[0086] Optionally, when using the target multi-modal model to identify the target event type corresponding to the scene image data in the target operation scene according to the attribute information of the target objects and the knowledge information corresponding to different event types in the target operation scene, the processor 302 is specifically configured to: judge whether it is possible to determine the target event type corresponding to the scene image data in the target operation scene according to the attribute information of the target objects according to the event recognition rule of the target operation scene; if not, use the target multi-modal model to identify the target event type corresponding to the scene image data in the target operation scene according to the attribute information of the target objects and the knowledge information corresponding to different event types in the target operation scene.

[0087] Optionally, the processor 302 is further configured to: if it is possible to determine the target event type corresponding to the scene image data in the target operation scene according to the attribute information of the target objects, determine the category of the target objects according to the attribute information of the target objects; match the category of the target objects with the necessary object categories relied on by different event types in the target operation scene respectively; and if the category of the target objects is consistent with the necessary object category relied on by any one of the event types in the target operation scene, determine the any one of the event types as the target event type and output the target event type.

[0088] Optionally, when the processor 302 determines whether it is possible to determine the target event type corresponding to the scene image data in the target operation scenario according to the event recognition rule of the target operation scenario, it specifically is used for: if the event recognition rule of the target operation scenario depends on the category of the object, determining that it is possible to determine the target event type corresponding to the scene image data in the target operation scenario according to the attribute information of the target object; if the event recognition rule of the target operation scenario depends on the category of the object and the relationship between objects, determining that it is not possible to determine the target event type corresponding to the scene image data in the target operation scenario according to the attribute information of the target object.

[0089] Optionally, the attribute information of the target object includes: the coordinate information and category information of the target object; when the processor 302 uses the target multimodal model to identify the target event type corresponding to the scene image data in the target operation scenario according to the attribute information of the target object and the knowledge information corresponding to different event types in the target operation scenario, it specifically is used for: cropping the scene image data according to the coordinate information of the target object to obtain a scene subgraph corresponding to the target object; constructing a target prompt according to the scene subgraph, the coordinate information and category information of the target object, and the knowledge information corresponding to different events in the target operation scenario; calling the target multimodal model according to the target prompt to use the target multimodal model to identify the target event type corresponding to the scene image data in the target operation scenario.

[0090] Optionally, when the processor 302 calls the target multimodal model according to the target prompt to use the target multimodal model to identify the target event type corresponding to the scene image data in the target operation scenario, it specifically is used for: inputting the target prompt into the target multimodal model; in the target multimodal model, obtaining the coordinate information and category information of each of the multiple objects in the target object from the target prompt, and extracting the semantic features of the scene subgraphs of each of the multiple objects; identifying the logical relationship between the multiple objects according to the coordinate information, category information, and semantic features of each of the multiple objects; respectively matching the knowledge information corresponding to the different event types according to the logical relationship between the multiple objects to obtain target knowledge information adapted to the logical relationship between the multiple objects; determining the event type corresponding to the target knowledge information as the event type corresponding to the scene image data in the target operation scenario.

[0091] Optionally, the target multimodal model is obtained by performing lightweight compression on a pre-trained multimodal large model, and the target multimodal model is deployed on a terminal device or an edge device.

[0092] Furthermore, as Figure 3 shown, the electronic device further includes: other components such as a communication component 303, a power supply component 304, a display component 305, and an audio component 306. Figure 3 Only some components are schematically shown in Figure 3 and it does not mean that the electronic device only includes Figure 3 the components shown. Figure 3 In Figure 3 , the components within the dashed box are optional components rather than essential components, and can be determined according to the product form of the electronic device. The electronic device in this embodiment can be implemented as a terminal device such as a desktop computer, a laptop computer, a smart phone, or an IOT device, or can also be a server device such as a conventional server, a cloud server, or a server array. If the electronic device in this embodiment is implemented as a terminal device such as a desktop computer, a laptop computer, or a smart phone, it may include

[0093] the components within the dashed box; if the electronic device in this embodiment is implemented as a server device such as a conventional server, a cloud server, or a server array, it may not include

[0094] Among them, the communication component 303 is configured to facilitate communication between the device where the communication component is located and other devices in a wired or wireless manner. The device where the communication component is located can access a wireless network based on communication standards, such as 2G (such as Global System for Mobile Communications (GSM), etc.), 3G (such as Wideband Code Division Multiple Access (WCDMA)), 4G (such as Long Term Evolution (LTE), etc.), 4G+ (such as LTE-Advanced (LTE-A), etc.) or 5G (5th Generation Mobile Communication Technology), or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel.

[0095] Among them, the power supply component 304 is used to provide power for various components of the device where the power supply component is located. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device where the power supply component is located.

[0096] The display component includes a screen, and the screen may include a Liquid Crystal Display (LCD) and a Touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from users. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operations.

[0097] The audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC), and when the device where the audio component is located is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive external audio signals. The received audio signals can be further stored in the memory or sent via the communication component. In some embodiments, the audio component further includes a speaker for outputting audio signals.

[0098] In this embodiment, the object detection model based on computer vision is combined with the object multi-modal model, giving full play to the advantages of the object detection model in fine-grained object detection and the powerful capabilities of the multi-modal model in advanced semantic understanding and reasoning, and achieving more accurate event recognition. Secondly, taking the object detection model as the pre-model of the object multi-modal model enables the object multi-modal model to focus on advanced semantic understanding and reasoning, which is conducive to reducing the number of parameters of the object multi-modal model. While ensuring the semantic understanding ability, it can significantly reduce the complexity of engineering implementation, thus facilitating the engineering deployment of the object multi-modal model. In addition, the object detection model and the object multi-modal model can be optimized independently, facilitating flexible version iteration.

[0099] Correspondingly, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to be able to implement the steps in the above method embodiments. Among them, the computer-readable storage medium can be implemented by volatile or non-volatile or a combination thereof, and can be removable or non-removable. Examples of computer-readable storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices or any other non-transmission medium Correspondingly, an embodiment of the present application further provides a computer program product, which includes a computer program or instruction, which, when executed by a processor, causes the processor to be able to implement the steps in the above method embodiments. It should be understood that each process or a combination of multiple processes in the above method flow can be implemented by the computer program or instruction. In addition, these computer programs or instructions can be applied to the processors of general-purpose computers, special-purpose computers, embedded processors or other programmable data processing devices, so that the processors of general-purpose computers, special-purpose computers, embedded processors or other programmable data processing devices can be used as devices to implement the corresponding functions in the above method embodiments.

[0100] It should also be noted that the term "comprise", "include" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, product or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, product or apparatus. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, method, product or apparatus comprising said element.

[0101] The above are only examples of the present application and are not intended to limit the present application. For those skilled in the art, various modifications and changes can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. An event recognition method, characterized in that, Including: Obtain scene image data during the track maintenance operation; Use an object detection model to perform object detection on the scene image data to obtain the attribute information of the target objects included in the scene image data; Determine the target operation scene corresponding to the scene image data according to the attribute information of the target objects; Use a target multimodal model to identify the target event type corresponding to the scene image data in the target operation scene according to the attribute information of the target objects and the knowledge information corresponding to different event types in the target operation scene.

2. The method according to claim 1, wherein Determine the target operation scene corresponding to the scene image data according to the attribute information of the target objects, including: Obtain the category information of the target objects according to the attribute information of the target objects; Compare the category information of the target objects with the category information of the necessary objects included in different operation scenes respectively; If the category information of the target objects matches the category information of the necessary objects included in any one of the different operation scenes, determine the any one of the operation scenes as the target operation scene corresponding to the scene image data.

3. The method according to claim 1, characterized in that Use a target multimodal model to identify the target event type corresponding to the scene image data in the target operation scene according to the attribute information of the target objects and the knowledge information corresponding to different event types in the target operation scene, including: Judge whether it is possible to determine the target event type corresponding to the scene image data in the target operation scene according to the attribute information of the target objects according to the event recognition rule of the target operation scene; If not, use the target multimodal model to identify the target event type corresponding to the scene image data in the target operation scene according to the attribute information of the target objects and the knowledge information corresponding to different event types in the target operation scene.

4. The method according to claim 3, characterized in that, It also includes: If it is possible to determine the target event type corresponding to the scene image data in the target operation scene according to the attribute information of the target objects, determine the category of the target objects according to the attribute information of the target objects; Match the category of the target objects with the necessary object categories relied on by different event types in the target operation scene respectively; If the category of the target objects is consistent with the necessary object category relied on by any one of the event types in the target operation scene, determine the any one of the event types as the target event type and output the target event type.

5. The method according to claim 3, wherein Judge whether it is possible to determine the target event type corresponding to the scene image data in the target operation scene according to the attribute information of the target objects according to the event recognition rule of the target operation scene, including: If the event recognition rule of the target operation scene relies on the category of objects, determine that it is possible to determine the target event type corresponding to the scene image data in the target operation scene according to the attribute information of the target objects; If the event recognition rule of the target operation scenario depends on the categories of the objects and the relationships between the objects, it is determined that the target event type corresponding to the scenario image data in the target operation scenario cannot be determined based on the attribute information of the target object.

6. The method according to any one of claims 1-5, characterized in that, The attribute information of the target object includes: the coordinate information and category information of the target object; Using a target multi-modal model, according to the attribute information of the target object and the knowledge information corresponding to different event types in the target operation scenario, identifying the target event type corresponding to the scenario image data in the target operation scenario includes: According to the coordinate information of the target object, cropping the scenario image data to obtain a scenario sub-graph corresponding to the target object; According to the scenario sub-graph, the coordinate information and category information of the target object, and the knowledge information corresponding to different events in the target operation scenario, constructing a target prompt; Invoking the target multi-modal model according to the target prompt, so as to use the target multi-modal model to identify the target event type corresponding to the scenario image data in the target operation scenario.

7. The method according to claim 6, characterized in that, Invoking the target multi-modal model according to the target prompt, so as to use the target multi-modal model to identify the target event type corresponding to the scenario image data in the target operation scenario includes: Inputting the target prompt into the target multi-modal model; In the target multi-modal model, obtaining the coordinate information and category information of each of the multiple objects in the target object from the target prompt, and extracting the semantic features of the scenario sub-graphs of the multiple objects; According to the coordinate information, category information, and semantic features of each of the multiple objects, identifying the logical relationships between the multiple objects; According to the logical relationships between the multiple objects, respectively matching the knowledge information corresponding to different event types to obtain target knowledge information adapted to the logical relationships between the multiple objects; Determining the event type corresponding to the target knowledge information as the event type corresponding to the scenario image data in the target operation scenario.

8. The method according to any one of claims 1-5, characterized in that The target multi-modal model is obtained by performing lightweight compression on a pre-trained multi-modal large model, and the target multi-modal model is deployed on a terminal device or an edge device.

9. An electronic device, characterized in that, Including: A memory and a processor; The memory is used to store one or more computer instructions; The processor is used to execute the one or more computer instructions for: executing the steps in the method according to any one of claims 1-8.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it can implement the event recognition method according to any one of claims 1-8.

11. A computer program product, characterized in that, Including: A computer program / instructions, when the computer program / instructions are executed by the processor, it can implement the event recognition method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Risk identification method and system based on operation scene, and storage medium

    CN118230048A

  • Multi-modal scene adaptive prompt open vocabulary target detection method and device

    CN118865379A

  • Scene understanding information generation method and device, equipment and medium

    CN119251657A

  • Operation site safety management and control method and system based on multi-modal large model

    CN119625477A

  • Abnormal event identification method and device based on multi-modal large model, and storage medium

    CN119810627A

Cited By

  • Event reasoning method and system based on multi-modal characteristics

    CN121366385A

  • Intelligent review method for compliance of electric power security supervision document

    CN121414303A

  • An intelligent review method for power safety supervision document compliance

    CN121414303B