Event detection method, device and storage medium

By extracting, fusion and extracting the image feature of the target video, combining the matching degree of video features with text features, the problem of insufficient adaptability and generalization capabilities of existing event detection methods in complex scenes is solved, and the accuracy and efficiency of event detection are improved.

CN119693861BActive Publication Date: 2025-06-06ZHEJIANG DAHUA TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510204571.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-06
Estimated Expiration
2045-02-24

AI Technical Summary

Technical Problem

The existing event detection methods have weak adaptability and generalization capabilities in complex scenarios, and have low real-time processing efficiency, resulting in low detection efficiency and accuracy.

Method used

By extracting the image features of several image frames in the target video, feature fusion and extraction are performed, and event detection is achieved by combining the matching degree between the video features and the text features of the preset event.

Benefits of technology

It improves the accuracy and efficiency of event detection, can better capture the dynamic changes of events, and quickly determine the detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119693861B_ABST
    Figure CN119693861B_ABST
Patent Text Reader

Abstract

The present application discloses an event detection method, device and storage medium, the method comprising: extracting image features of several first image frames in a target video; performing feature fusion on the image features of each of the first image frames to obtain fusion features of the target video; performing feature extraction on the fusion features to obtain video features; obtaining event detection results based on the matching degree between the video features and several text features, wherein the several text features are respectively extracted from text descriptions of several preset events, and the event detection results characterize whether any of the preset events occurs in the target video. In the above manner, the present application can effectively improve the efficiency and accuracy of event detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to an event detection method, device and storage medium. Background Art

[0002] With the rapid development of video surveillance technology, event detection has become increasingly important in the fields of security detection, traffic management, public safety, etc. Traditional video surveillance systems usually rely on manual monitoring, but they are difficult to identify potential abnormal events in real time and accurately, and are easily affected by human factors.

[0003] However, existing event detection methods still face some challenges, including low adaptability to complex scenarios, weak generalization ability to different types of abnormal events, and slow real-time processing, which leads to low efficiency and accuracy of event detection. Summary of the invention

[0004] In order to solve the above technical problems, the technical solution adopted by the present application is: to provide an event detection method, device and storage medium to at least solve the problem of low efficiency and accuracy of event detection in related technologies.

[0005] According to an embodiment of the present invention, there is provided an event detection method, comprising:

[0006] Extracting image features of a plurality of first image frames in a target video;

[0007] Performing feature fusion on the image features of each of the first image frames to obtain fusion features of the target video;

[0008] Performing feature extraction on the fused features to obtain video features;

[0009] Based on the matching degree between the video features and a number of text features, an event detection result is obtained, wherein the several text features are extracted from text descriptions of a number of preset events, and the event detection result represents whether any of the preset events occurs in the target video.

[0010] To solve the above technical problems, a technical solution adopted in the present application is: to provide an electronic device, including a memory and a processor, wherein the memory is used to store a computer program, and when the computer program is executed by the processor, it is used to implement the event detection method in the above technical solution.

[0011] In order to solve the above technical problems, a technical solution adopted in the present application is: to provide a computer-readable storage medium, which is used to store a computer program. When the computer program is executed by a processor, it is used to implement the event detection method in the above technical solution.

[0012] Through the above scheme, the beneficial effect of the present application is: the event detection method provided by the present application extracts image features of several first image frames in the target video, performs feature fusion on the image features of each first image frame to obtain fusion features of the target video, performs feature extraction on the fusion features to obtain video features, and obtains event detection results based on the matching degree between the video features and several text features; in this way, the present application can capture the dynamic changes of events by performing feature fusion on the image frame sequence of the target video, thereby improving the accuracy of event detection, and further combines the matching results between the video features extracted by the fusion features and several text features to quickly determine the events detected in the current target video, thereby improving the efficiency of event detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. Among them:

[0014] Figure 1 It is a flowchart of an embodiment of an event detection method provided by the present application;

[0015] Figure 2 It is a structural diagram of a video feature fusion module provided by this application;

[0016] Figure 3 It is a structural diagram of a visual text fusion model provided by this application;

[0017] Figure 4 is a flow chart of another embodiment of the event detection method provided by the present application;

[0018] Figure 5 It is a structural diagram of a feature similarity module provided by this application;

[0019] Figure 6 It is a structural schematic diagram of a road abnormal event detection model provided by the present application;

[0020] Figure 7 It is a structural schematic diagram of an embodiment of an electronic device provided by the present application;

[0021] Figure 8 It is a structural schematic diagram of an embodiment of a computer-readable storage medium provided by the present application. DETAILED DESCRIPTION

[0022] The present application is further described in detail below in conjunction with the accompanying drawings and examples. It is particularly noted that the following examples are only used to illustrate the present application, but are not intended to limit the scope of the present application. Similarly, the following examples are only some embodiments of the present application rather than all embodiments, and all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application.

[0023] Reference to "embodiments" in this application means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0024] It should be noted that the terms "first", "second", etc. in this application are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features. Thus, the features defined as "first", "second", etc. can explicitly or implicitly include at least one of the features. In the description of this application, the meaning of "multiple" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally also includes steps or units that are not listed, or optionally also includes other steps or units inherent to these processes, methods, products or devices.

[0025] See also Figure 1 , Figure 1 is a flow chart of an embodiment of the event detection method provided by the present application. It should be noted that if there are substantially the same results, this embodiment does not Figure 1 The process sequence shown is limited. Figure 1 As shown, this embodiment includes:

[0026] S110: Extracting image features of a plurality of first image frames in a target video.

[0027] The target video is a sequence of video frames to be detected for an event, and the first image frame is selected from the sequence of video frames. In one example, the target video can be analyzed frame by frame, and each frame can be extracted from the target video as the first image frame. In another example, a preset time period or a preset number of interval frames can also be set, and image frames can be extracted from the target video according to the preset period or the preset number of interval frames, for example, one frame is extracted every 5 seconds or one frame is extracted every 5 frames. The method of extracting the first image frame can be set according to actual needs and is not limited here.

[0028] In one embodiment, before extracting image features of several first image frames in the target video, the several first image frames are also preprocessed, such as denoising, contrast enhancement, size normalization, etc., to improve the accuracy and robustness of feature extraction.

[0029] The image features of the first image frame can be extracted by extracting color features, texture features and / or shape features of the image frame based on traditional computer vision methods, or by deep learning based methods such as convolutional neural networks and autoencoders, which are not limited here.

[0030] In one embodiment, a multi-scale feature fusion technology combining image segmentation, feature extraction and position coding may be used. In one example, for each first image frame, the first image frame is divided into a number of image blocks, the block features of each image block are obtained, and the block features of each image block in the first image frame are spliced ​​to obtain the image features of the first image frame. In this way, by dividing the image into small blocks and extracting features, the detailed information and local features of the image frame can be captured.

[0031] In one example, before the block features of each image block in the first image frame are spliced ​​to obtain the image features of the first image frame, the method further includes the step of determining the spatial position coding of each image block based on the position of each image block in the first image frame, splicing the block features of each image block in the first image frame according to the spatial position coding of each image block in the first image frame, and obtaining the image features of the first image frame. The addition of the spatial position coding enables the feature extraction model to understand the relative or absolute position of the image block in the image, thereby enhancing the spatial semantics of the feature.

[0032] In another example, after determining the spatial position code of each image block, the block features of each image block in the first image frame may be spliced ​​with the corresponding spatial position code to obtain the splicing features of each image block, and then the splicing features of each image block in the first image frame may be spliced ​​to obtain the image features of the first image frame. In the above example, the extracted block features and position codes are combined to form splicing features, which integrate local features and spatial information and can provide rich data representation for subsequent analysis.

[0033] S120: Perform feature fusion on the image features of each first image frame to obtain fusion features of the target video.

[0034] Feature fusion refers to combining features from one or more sensors or data sources to obtain a more comprehensive and stable representation of the data. In video processing, feature fusion usually involves combining information from different frames or different types of features to obtain an accurate understanding of the video content.

[0035] In complex and changing environments, combining information from multiple image frames is more accurate than using information from a single image frame for event detection. For example, when detecting road collapse, the shadow of a vehicle on a single frame may cause a false alarm. Multi-frame images can make more accurate judgments because they have temporal information about the vehicle and the shadow.

[0036] Feature fusion methods include but are not limited to feature concatenation, weighted averaging, deep learning methods, and attention mechanisms. By fusing the image features of each first image frame extracted from the target video to obtain the fused features of the target video, a more comprehensive description of the video content can be provided, thereby improving the accuracy of event detection and video understanding. At the same time, the fused features can help the event detection method to better generalize to new and unseen data, thereby improving the generalization ability of the event detection method.

[0037] In one embodiment, the process of performing feature fusion on the image features of each first image frame to obtain the fusion feature of the target video includes: selecting the image feature of a first image frame from a number of first image frames as the first feature to be fused, and selecting the image feature of another first image frame as the second feature to be fused. After adding the first feature to be fused and the second feature to be fused, input them into the self-attention module for feature fusion to obtain an intermediate fusion feature. The intermediate fusion feature is used as the new first feature to be fused, and the image feature of a first image frame is selected from the remaining unselected first image frames as the new second feature to be fused, and the new first feature to be fused and the new second feature to be fused are used to repeatedly perform the process of adding the first feature to be fused and the second feature to be fused, and then inputting them into the self-attention module for feature fusion to obtain the intermediate fusion feature and subsequent steps, until all first image frames are selected. Finally, the intermediate fusion feature obtained at the end is used as the fusion feature of the target video.

[0038] In one embodiment, see Figure 2 , Figure 2This is a structural diagram of a video feature fusion module provided by the present application. The steps of extracting the image features of several first image frames in the target video and fusing the image features of each first image frame to obtain the fusion features of the target video are all performed by the video feature fusion module. Among them, the video feature fusion module is trained using the first contrast learning. Through the first contrast learning, the video feature fusion module can distinguish different events and scenes in the video, enhance the ability to distinguish features, and help the model learn a more generalized feature representation, thereby reducing the dependence on a large amount of labeled data and making the training process more efficient.

[0039] S130: Extract the fused features to obtain video features.

[0040] After obtaining the fusion features of the target video, the fusion features are further extracted to obtain video features. By extracting higher-level and more abstract features from the fusion features, the semantic information of the target video can be captured, which helps to use the video features to understand the content of the target video more deeply.

[0041] Methods for extracting features from fused features include but are not limited to using deep learning models, autoencoders, attention mechanisms, and multimodal fusion, etc. In one embodiment, extracting features from fused features to obtain video features is implemented using an image encoder of a visual text fusion model.

[0042] In one example, the visual text fusion model is fine-tuned based on the Clip (Contrastive Language-Image Pre-training) model. The Clip model maps images and texts to the same embedding space through contrastive learning, making the associated images and texts close to each other in the vector space. The core architecture of the Clip model includes two submodules, an image encoder and a text encoder. The image encoder is responsible for converting images into corresponding vector representations, and the text encoder is used to process natural language text.

[0043] The visual text fusion model is obtained by first contrastive learning training, and the visual text fusion model has been pre-trained before contrastive learning training. Specifically, based on the contrastive learning method, pre-training is performed through large-scale image and text data. During training, the model embeds images and texts into the same vector space respectively, and then maximizes the similarity between correct image-text pairs and minimizes the similarity between incorrect image-text pairs through contrastive learning.

[0044] In one example, the visual text fusion model is obtained by further fine-tuning the pre-trained clip model through the first contrastive learning training. Specifically, during the first contrastive learning training, several videos are input, image frames are extracted, and after temporal feature fusion, the obtained fusion features are input into the image encoder to obtain video features; several text descriptions are randomly extracted and input into the text encoder to obtain text features, and then the distance between the text features and the video features is calculated. If the text description and the video are corresponding, the model will optimize this distance to the minimum. For each pair of positive samples V i , T j , the loss function of contrastive learning can be shown as the following formula (1), where sim represents similarity calculation, which can be cosine similarity or Euclidean distance, τ is the temperature parameter used to control the degree of difference. l [k≠i] is a confidence function, when k Not equal to i , it is equal to 1, otherwise it is 0, and the final loss is equal to the loss of all positive sample pairs.

[0045] (1)

[0046] By optimizing this objective function, the contrastive learning network can automatically learn the similarities between video features and text features, so that it can determine whether the video description is correct.

[0047] S140: Obtaining event detection results based on the matching degrees between the video features and the plurality of text features.

[0048] By obtaining the matching degree between the video features of the target video and a number of text features, the event results detected in the target video are obtained based on the matching results. Among them, the several text features are extracted from the text descriptions of several preset events respectively, and the event detection results represent whether any preset event occurs in the target video. The method of obtaining the matching degree between the video features and the several text features includes but is not limited to calculating cosine similarity, Euclidean distance or dot product, etc. By calculating the matching degree between the video features and the text features of several preset events, it is possible to automatically detect whether the corresponding event occurs in the target video, thereby providing timely warning and decision support.

[0049] In one embodiment, the visual text fusion model is used to obtain the matching degree between the video features of the target video and several text features, see Figure 3 , Figure 3 is a structural diagram of a visual text fusion model provided by this application. Figure 3As shown in the figure, the video features of the target video are obtained by the image encoder of the visual text fusion model, and the text encoder of the visual text fusion model is used to extract the text descriptions of several preset events to obtain several text features. The video features are matrix multiplied with each text feature to obtain the multiplication result corresponding to each text feature, and the multiplication result corresponding to each text feature is processed by the activation function to obtain the matching probability of each text feature and the video feature.

[0050] The event detection result can be expressed in a variety of forms, which are set according to the needs of the actual application scenario and are not limited here. In one example, the matching degree between the video features and several text features is output as the event detection result; in another example, the preset event corresponding to the text feature whose matching degree meets the matching requirements is determined as the occurrence event of the target video, and the information of the occurrence event is output.

[0051] See also Figure 4 , Figure 4 FIG. 1 is a flow chart of another embodiment of the event detection method provided by the present application. It should be noted that if there are substantially the same results, this embodiment does not necessarily refer to the event detection method of the present application. Figure 4 The process sequence shown is limited.

[0052] When performing event detection in real scenes, since events may occur suddenly and unpredictably, it is impossible to perform resource-intensive event detection on every frame of video considering the limitations of computing resources and costs. Therefore, in this embodiment, a feature similarity module is introduced as a pre-step to determine whether a significant change has occurred in the scene. This significantly reduces the demand for computing resources and avoids invalid event detection in static or unchanged scenes, thereby saving computing resources and improving efficiency. Figure 4 As shown, this embodiment includes:

[0053] S210: Determine whether a target scene in a target video changes.

[0054] A target scene is preset as the main detection object, such as a specific area in the target video or the foreground or background in the target video, etc. The specific settings can be selected according to the actual application and are not limited here.

[0055] Whether the target scene in the target video has changed can be determined based on whether changes occur between image frames in the target video. For example, each image frame in the target video or image frames are sequentially extracted from the target video at intervals of a preset number of frames or a preset time as the object of determining whether the target scene has changed.

[0056] In one example, a plurality of second image frames are sequentially extracted from a target video, and image features of each second image frame are extracted. Similarities between image features of adjacent second image frames are sequentially obtained in accordance with the extraction order. In response to the similarity between image features of adjacent second image frames satisfying a similarity requirement, it is determined that a scene change occurs in the target video. The similarity requirement may be preset as a similarity greater than a similarity threshold.

[0057] In one embodiment, detecting whether a scene change occurs in a target video can be achieved by a feature similarity module. The feature similarity module is determined by a trained deep learning module. The trained deep learning module is obtained by performing a second comparative learning training using features extracted by an image encoder. During the second comparative learning training process, the parameters of the image encoder are in a frozen state.

[0058] In one example, the subsequent step of extracting features from the fused features to obtain video features is based on the implementation of the image encoder of the visual-text fusion model. In order to save computing resources, the feature similarity module is also implemented based on the image encoder of the visual-text fusion model. For example, the image encoder of the visual-text fusion model is used to extract image features of each second image frame, and the feature similarity module is used to perform similarity detection on the extracted image features of the second image frames.

[0059] See also Figure 5 , Figure 5 It is a structural diagram of a feature similarity module provided by the present application. The backbone model of the feature similarity module uses the image encoder of the visual-text fusion model. When training the feature similarity module, all the parameters of the image encoder and the text encoder in the visual-text fusion model after the second contrast learning training and fine-tuning of the corresponding scene data are frozen. On this basis, a second contrast learning module is connected after the image encoder, and the second contrast learning module learns to perform contrast learning training using the image features extracted by the image encoder of the visual-text fusion model. In this way, not only can the feature extraction capability of the image encoder of the visual-text fusion model be reused, but also the parameter amount of the entire event detection model can be reduced.

[0060] The training method of the feature similarity module is as follows: randomly select several videos, randomly extract some frames from each video, and in each video, the image frames with unchanged background are used as positive sample pairs, and the image frames with changed background are used as negative sample pairs. The image frames are input into the image encoder of the visual text fusion model to extract the corresponding features, and then the features are input into a multi-head attention module to obtain processed features, and the processed features are input into the contrast learning loss function to calculate the loss. The loss function calculation formula is the same as formula (1) in the above step S130.

[0061] S220: In response to a change in a target scene in a target video, determining a number of first image frames in the target video.

[0062] In response to a change in a target scene in a target video, several first image frames in the target video are determined for subsequent event detection. In one example, the next second image frame among adjacent second image frames that meet the similarity requirement can be used as a reference image frame, and the first image frame is located after the reference image frame or is the reference image frame. Exemplarily, the reference image frame and at least one second image frame extracted thereafter are used as the first image frame; or, several second image frames extracted after the reference image frame are used as the first image frame.

[0063] After obtaining several first image frames in the target video, in one embodiment, the trained feature similarity module can be combined with the visual text fusion model to obtain an event detection model for event detection of the target video, and the following steps S230 to S260 are performed in sequence to obtain the event detection result corresponding to the target video. It should be noted that the specific methods of steps S230 to S260 can refer to the above steps S110 to S140, which will not be repeated here.

[0064] S230: Extracting image features of a plurality of first image frames in the target video.

[0065] S240: Perform feature fusion on the image features of each first image frame to obtain fusion features of the target video.

[0066] S250: extracting the fused features to obtain video features.

[0067] S260: Obtaining event detection results based on the matching degrees between the video features and the plurality of text features.

[0068] In order to better illustrate the event detection method of the present application, a specific implementation method applied to abnormal road event detection scenarios is provided below.

[0069] Road abnormal events refer to changes in the road surface that seriously affect traffic safety, mainly road collapse, mudslides, rolling stones, fog, ice, etc. The current road abnormal event detection method based on pure vision can only detect pre-set road abnormal event types. If new road abnormal event types need to be supported, materials must be collected and annotated again, and then the model must be trained and deployed. The process is complicated and the time and financial costs are high.

[0070] The technical solution provided in this embodiment is a road abnormal event detection method based on multi-modal fusion. Figure 6 , Figure 6It is a structural diagram of a road abnormal event detection model provided by the present application. The road abnormal event detection model of this embodiment combines a feature similarity module, a video feature fusion module and a visual text fusion module, wherein the feature similarity module is used to determine whether there is a scene change in the target video. When the target scene in the target video changes, an image frame sequence that may have changed is obtained, and the image frame sequence is input into the video feature fusion module to obtain the video features of the target video. Finally, based on the visual text fusion module, the obtained video features are matched with the text features of several events to obtain the final event detection results. In this way, when detecting abnormal road events, more accurate judgments can be made by combining the information of multiple frames of video images. The specific process is as follows:

[0071] Before performing event detection based on the application scenario of road abnormal events, it is necessary to build and train a road abnormal event detection model. Before training, it is necessary to first collect as many types of road abnormal event data as possible in the application scenario, and provide a text description of the road abnormal event that occurred in each video.

[0072] This embodiment uses the clip model as the visual text fusion model because clip is pre-trained on a large number of image-text pair datasets and has strong generalization for image tasks. However, the clip model lacks the ability to extract features between video frames. Therefore, the visual feature fusion module proposed in this proposal is a model fine-tuned based on clip that can utilize video features.

[0073] For an input video V and text description T, we first extract N frames of images from the video V, and then cut each frame into multiple P×P patches. Assuming the resolution of the image is H×W, the image will be cut into P size = (H*W) / (P×P) patches of shape P×P. Then each patch is input into a feature mapping layer (which can be a convolutional layer or a fully connected layer) to obtain a feature of length L, and then the spatial position encoding of each patch is added. After this operation, each frame can be transformed into a P size ×L feature matrix or P size ×(L+m) feature matrix, where m is the size of the positional encoding.

[0074] The feature matrix F generated by the first picture 1 The feature matrix F generated by the feature map of the second image 2 Add them together and input them into the self-attention module to perform temporal feature fusion. Then add the output features to the feature matrix F in the third figure. 3, input into self-attention for feature fusion until the last frame. The calculation process of the video feature fusion module can be as follows Figure 2 shown.

[0075] The fused visual features F N Input to the clip image encoder to get the video feature F v . Calculate the token embedding (word vector) corresponding to the text, and input the token embedding into the text encoder of the clip to obtain the text feature F t , F v and F t Perform matrix multiplication, and then pass the sigmiod function to output the probability of text description and video feature matching. The calculation process of the visual text fusion module can be as follows: Figure 3 shown.

[0076] The road abnormal event detection model that combines the video feature fusion module and the visual text fusion module is trained by contrastive learning to optimize the loss function. Further, on this basis, all parameters of the currently trained road abnormal event detection model are frozen, and a contrastive learning module is connected after the image encoder of the road abnormal event detection model. After contrastive learning training, the loss function is optimized.

[0077] When detecting abnormal road events in the video to be detected, Figure 6 As shown in the figure, first set the text description of the disaster type that may need to be detected, input these text descriptions into the text encoder of the clip to extract text features, and store them in the text feature cache list, which can save the model's reasoning time. If you need to add or reduce disaster types later, you only need to add the corresponding video description features to the list or delete the corresponding video description features.

[0078] For the input video, extract a frame every t frames and cache the extracted image features into an image frame queue. Use the image features of the current frame and the first feature of the queue to input into the feature similarity detection module to detect whether the background of the two frames has changed. If there is no change, the current frame is queued and the first element is dequeued.

[0079] If the similarity between the two frames is lower than a certain threshold, it means that the background image has changed. At this time, the image frames collected starting from the current frame image or the next frame of the current frame are input into the video feature fusion module to extract the video features of the video frame sequence, and then compare them with several stored text features to determine what type of road abnormality event it is and issue an alarm.

[0080] In the road abnormal event detection model provided in this embodiment, when deployed, the weight of the image encoder of the visual text fusion model is combined with a contrast learning module to obtain a feature similarity module to determine whether the target scene has changed. In addition, there is no need to frequently use the text encoding module during use, which reduces the hardware requirements for model deployment. At the same time, the road abnormal event detection model is combined with a video feature fusion module that can fuse multi-frame information, but there is no need to call the video feature fusion module in real time. The module is only called when it is determined that the background of the previous and next frames has changed, thereby reducing the computational cost of the entire road abnormal event detection model.

[0081] See also Figure 7 , Figure 7 It is a structural diagram of an embodiment of an electronic device provided in the present application. The electronic device 60 includes a memory 61 and a processor 62 connected to each other. The memory 61 is used to store a computer program. When the computer program is executed by the processor 62, it is used to implement the event detection method in the above embodiment.

[0082] The above-mentioned method is an embodiment, which can exist in the form of a computer program. Therefore, the present application proposes a computer-readable storage medium. Figure 8 , Figure 8 It is a schematic diagram of the structure of an embodiment of a computer-readable storage medium provided in the present application. The computer-readable storage medium 80 is used to store a computer program 81, which can be executed to implement the event detection method in the above embodiment.

[0083] The computer-readable storage medium 80 may be a server, a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc., which may store program codes.

[0084] If the technical solution of this application involves personal information, the product using the technical solution of this application has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using the technical solution of this application has obtained the individual's separate consent before processing the sensitive personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, clear and prominent signs are set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that he or she agrees to the collection of his or her personal information; or on the device that processes personal information, the personal information processing rules are notified by obvious signs / information, and the individual's authorization is obtained through pop-up information or by asking the individual to upload his or her personal information; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.

[0085] The above description is only an embodiment of the present application and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. An event detection method, characterized in that: The method comprises: Determine whether the target scene in the target video has changed; In response to a change in the target scene in the target video, determining a plurality of first image frames in the target video; Extracting image features of a plurality of first image frames in a target video; Performing feature fusion on the image features of each of the first image frames to obtain fusion features of the target video; Performing feature extraction on the fused features to obtain video features; Based on the matching degree between the video features and a plurality of text features, an event detection result is obtained, wherein the plurality of text features are respectively extracted from text descriptions of a plurality of preset events, and the event detection result indicates whether any of the preset events occurs in the target video; Among them, detecting whether the scene change occurs in the target video is implemented by a feature similarity module; the step of extracting image features of several first image frames in the target video and the step of fusing the image features of each of the first image frames to obtain the fused features of the target video are performed by a video feature fusion module; the step of extracting features from the fused features to obtain video features is implemented by using an image encoder of a visual text fusion model; the feature similarity module is obtained by using the weights of the image encoder of the visual text fusion model in combination with a contrast learning module; in response to the feature similarity module determining that the scene change occurs in the target video, the video feature fusion module is called to fuse the multi-frame image frame information and the image encoder of the visual text fusion model to extract the video features of the multi-frame image frame information, so as to obtain an event detection result based on the matching degree between the video features and several text features respectively.

2. The event detection method according to claim 1, characterized in that: The step of extracting image features of a plurality of first image frames in the target video includes: For each of the first image frames, dividing the first image frame into a plurality of image blocks; Acquire block features of each of the image blocks; The block features of the image blocks in the first image frame are spliced ​​to obtain the image features of the first image frame.

3. The event detection method according to claim 2, characterized in that: Before the step of splicing the block features of the image blocks in the first image frame to obtain the image features of the first image frame, the step further includes: Based on the position of each image block in the first image frame, respectively determine the spatial position code of each image block; The step of splicing the block features of the image blocks in the first image frame to obtain the image features of the first image frame includes: According to the spatial position coding of each image block in the first image frame, the block features of each image block in the first image frame are spliced ​​to obtain the image features of the first image frame; or, The block features of each image block in the first image frame are spliced ​​with the corresponding spatial position code to obtain the splicing features of each image block, and the splicing features of each image block in the first image frame are spliced ​​to obtain the image features of the first image frame.

4. The event detection method according to claim 1, characterized in that: The step of fusing the image features of each of the first image frames to obtain the fused features of the target video includes: From the plurality of first image frames, select an image feature of a first image frame as a first feature to be fused, and select an image feature of another first image frame as a second feature to be fused; After adding the first feature to be fused and the second feature to be fused, the features are input into the self-attention module for feature fusion to obtain an intermediate fusion feature; The intermediate fusion feature is used as a new first feature to be fused, and an image feature of a first image frame is selected from the remaining unselected first image frames as a new second feature to be fused, and the first feature to be fused and the second feature to be fused are used to repeatedly perform the step of adding the first feature to be fused and the second feature to be fused, and then inputting the first feature to be fused and the second feature to be fused into the self-attention module for feature fusion to obtain the intermediate fusion feature and subsequent steps, until all first image frames are selected; The final intermediate fusion feature is used as the fusion feature of the target video.

5. The event detection method according to claim 1, characterized in that: The plurality of text features are extracted by using the text encoder of the visual-text fusion model to extract text descriptions of a plurality of preset events respectively.

6. The event detection method according to claim 5, characterized in that: The video feature fusion module and the visual text fusion model are obtained by using the first contrastive learning training, and the visual text fusion model has been pre-trained before using the contrastive learning training.

7. The event detection method according to claim 1, characterized in that: Before extracting the image features of the first plurality of image frames in the target video, the method includes: Sequentially extracting a plurality of second image frames from the target video, and extracting image features of each of the second image frames; Sequentially obtaining similarities between image features of adjacent second image frames according to an extraction order; In response to similarities between image features of adjacent second image frames satisfying a similarity requirement, determining that a scene change occurs in the target video; The latter second image frame among the adjacent second image frames that meet the similarity requirement is a reference image frame, and the first image frame is located after the reference image frame or is the reference image frame.

8. The event detection method according to claim 7, characterized in that: The similarity requirement is that the similarity is greater than a similarity threshold; And / or, the extracting of image features of each of the second image frames is achieved by using an image encoder of a visual-text fusion model; And / or, extracting a plurality of second image frames in sequence from the target video comprises: Sequentially extracting a second image frame from the target video at intervals of a preset number of frames; And / or, before extracting the image features of the first image frames in the target video, the method further includes: Using the reference image frame and at least one of the second image frames extracted later as the first image frame; or, A plurality of the second image frames extracted after the reference image frame are used as the first image frames.

9. The event detection method according to claim 7, characterized in that: The feature similarity module is determined by a trained deep learning module, and the trained deep learning module is obtained by performing a second contrastive learning training using the features extracted by the image encoder. During the second contrastive learning training process, the parameters of the image encoder are in a frozen state.

10. The event detection method according to claim 1, characterized in that: Before obtaining the event detection result based on the matching degree between the video feature and the plurality of text features, the method further includes: Performing matrix multiplication on the video feature and each of the text features to obtain a multiplication result corresponding to each of the text features; The multiplication results corresponding to each of the text features are processed using an activation function to obtain a matching probability between each of the text features and the video features.

11. The event detection method according to claim 1, characterized in that: The event detection result is obtained based on the matching degree between the video feature and the plurality of text features, including: Outputting the matching degree between the video feature and a plurality of text features as the event detection result; or, The preset event corresponding to the text feature whose matching degree meets the matching requirement is determined as the occurrence event of the target video, and the information of the occurrence event is output.

12. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the processor is coupled to the memory, and the processor is configured to execute one or more steps of the event detection method according to any one of claims 1 to 11 based on instructions stored in the memory.

13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps of the event detection method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Method and equipment for generating video content description information

    CN112749660A

  • Real-time detection method and system for abnormal event in monitoring scene

    CN118133198A

  • Capsule gastroscope stomach part identification method and system based on deep learning

    CN118430024A