Event Detection Method, Apparatus, Device, and Storage Medium
Through cross-modal semantic feature extraction and multimodal prediction tasks of visual language large models, the problems of low event detection efficiency and insufficient accuracy in security monitoring equipment are solved, and efficient and accurate event automation detection is achieved, which is suitable for automated monitoring of quality inspection, security, traffic and home safety.
Patent Information
- Application Number
- CN202310812928.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-04
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2043-07-04
AI Technical Summary
In video data processing, existing security monitoring equipment has low event detection efficiency and inaccurate detection results, making it difficult to achieve automated and efficient abnormal situation recognition.
Through cross-modal semantic feature extraction, visual semantic features are used to match the event semantic features of candidate events, and multimodal prediction task learning is performed in combination with visual language big models to improve the accuracy and efficiency of event detection.
It realizes the automation and efficiency of event detection, improves the accuracy of detection results, and is suitable for automated monitoring of various application scenarios such as quality inspection, security, traffic and home safety.
Smart Images

Figure CN116824455B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, particularly to the fields of deep learning and computer vision, and especially to large model technology, and can be used in the field of Internet of Things. Background Art
[0002] With the increasing emphasis on security by people, the use of security monitoring devices has gradually become popular. Usually, monitoring devices are set in the monitored area to collect video data, and based on the collected video data, the abnormal situations in the monitored area are grasped to facilitate taking effective measures in a timely manner. Summary of the Invention
[0003] The present disclosure provides an event detection method, device, equipment and storage medium.
[0004] According to one aspect of the present disclosure, there is provided an event detection method, including:
[0005] Obtaining a video frame to be detected of a video to be detected;
[0006] Performing cross-modal semantic feature extraction on the video frame to be detected to obtain visual semantic features;
[0007] Matching the visual semantic features with the event semantic features of different candidate events; wherein, the event semantic features of the candidate events are the extraction results of performing cross-modal semantic feature extraction on the event description data corresponding to the respective candidate events;
[0008] Determining the target event included in the video to be detected according to the matching result.
[0009] According to another aspect of the present disclosure, there is also provided an electronic device, including:
[0010] At least one processor; and
[0011] A memory communicatively connected to the at least one processor; wherein,
[0012] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute any one of the event detection methods provided by the embodiments of the present disclosure.
[0013] According to another aspect of the present disclosure, there is also provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute any one of the event detection methods provided by the embodiments of the present disclosure.
[0014] According to the technology of the present disclosure, the event detection efficiency and the accuracy of the detection result are improved.
[0015] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily apparent from the following description. Description of the Drawings
[0016] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0017] Figure 1 is a flowchart of an event detection method provided by an embodiment of the present disclosure;
[0018] Figure 2 is a flowchart of another event detection method provided by an embodiment of the present disclosure;
[0019] Figure 3A is an architecture diagram of an event detection system provided by an embodiment of the present disclosure;
[0020] Figure 3B is a flowchart of another event detection method provided by an embodiment of the present disclosure;
[0021] Figure 4 is a structural diagram of an event detection device provided by an embodiment of the present disclosure;
[0022] Figure 5 is a block diagram of an electronic device for implementing the event detection method of the embodiment of the present disclosure. Detailed Embodiments
[0023] The following describes exemplary embodiments of the present disclosure in conjunction with the drawings. Various details of the embodiments of the present disclosure are included to assist understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted below.
[0024] The event detection method and event detection device provided by the embodiments of the present disclosure are applicable to scenarios for automatically detecting interesting events in videos. Each event detection method provided by the embodiments of the present disclosure can be executed by an event detection device, which can be implemented by software and / or hardware and is specifically configured in an electronic device. The electronic device can be a server or a terminal device, and the present disclosure makes no limitation thereto.
[0025] For ease of understanding, the event detection method is first described in detail.
[0026] See Figure 1 shown in an event detection method, including:
[0027] S101. Obtain a video frame to be detected from the video to be detected.
[0028] Among them, the video to be detected can be understood as a video clip with event detection requirements. The video frame to be detected can be understood as a picture frame that constitutes the video to be detected. It should be noted that the video frame to be detected can be at least part of the picture frames that constitute the video to be detected, and the present disclosure does not impose any limitation on the number of video frames to be detected.
[0029] Optionally, the video to be detected can be pre-stored locally in the execution device that executes the event detection method or in other storage devices communicatively connected to the execution device; and when event detection is required, the video to be detected is obtained from the corresponding storage location; at least one picture frame in the video to be detected is extracted as the video frame to be detected for subsequent processing. In a specific implementation, the video to be detected can be directly obtained from a video acquisition device for use.
[0030] To reduce the data computation amount of the execution device that executes the event detection method, thereby further improving the event detection efficiency, optionally, the video frames to be detected of the video to be detected can also be pre-stored locally in the execution device that executes the event detection method or in other storage devices communicatively connected to the execution device, and when event detection is required, the video frames to be detected are obtained from the corresponding storage location.
[0031] In an optional embodiment, the video frames to be detected of the video to be detected can be determined in the following manner: determine each original video frame in the video to be detected; select at least one video frame to be detected from each original video frame according to the difference situation between adjacent original video frames.
[0032] Among them, the original video frame is each picture frame that constitutes the video to be detected. Correspondingly, for any two adjacent original video frames, the change situation of the picture content between the adjacent original video frames is reflected by the difference situation between the adjacent original video frames; if the difference is greater, it indicates that the picture change situation between the adjacent original video frames is larger, that is, the picture content difference is significant. At this time, the subsequent original video frame can be added to the video frames to be detected.
[0033] By determining the video frames to be detected through the above-mentioned differential selection method, it is possible to screen out the video frames to be detected with obvious picture changes, avoid missing key information, thereby avoiding missing the target event, and further improving the comprehensiveness of the target event determination result.
[0034] In another alternative embodiment, frame extraction processing can be directly performed on the video to be detected according to a preset frame extraction frequency, to obtain at least one video to be detected. The preset frame extraction frequency can be set by a technician according to requirements or empirical values, or determined by adjusting through a large number of experiments.
[0035] By determining the video frames to be detected through the above frame extraction and selection method, the operation process is convenient and fast, improving the efficiency of selecting the video frames to be detected.
[0036] In yet another alternative embodiment, each original video frame in the video to be detected can be determined; according to the difference situation between adjacent original video frames, at least one first candidate video frame is selected from each original video frame; according to the preset frame extraction frequency, at least one second candidate video frame is extracted from each original video frame; according to the union of the first candidate video frame and the second candidate video, the video frames to be detected are determined. The preset frame extraction frequency can be set by a technician according to requirements or empirical values, or determined by adjusting through a large number of experiments.
[0037] It can be understood that by adopting the method of differential selection and comprehensive frame extraction selection to determine different candidate video frames, the comprehensiveness of the candidate video frames is ensured, the omission of key information is avoided, and thus the omission of subsequent event detection results is avoided. At the same time, by taking the intersection of the first candidate video frame and the second candidate video frame to determine the video frames to be detected, the repeated selection of candidate video frames is avoided, which increases the computational amount for subsequent event detection and saves the occupied amount of computing resources.
[0038] S102. Extract cross-modal semantic features from the video frames to be detected, to obtain visual semantic features.
[0039] Among them, the cross-modal semantic features can be understood as the semantic features across different domains and modalities, aiming to utilize the complementarity between the semantic features in different modalities, eliminate the redundancy of the semantic features in different modalities, and through the mutual cooperation and complementarity of the semantic features in different modalities, realize the extraction of more rich, comprehensive and concise semantic features carrying information.
[0040] Exemplarily, the same or different feature extraction methods can be adopted to extract semantic features of different modalities from the video frames to be detected; by performing duplicate removal processing on the semantic features of different modalities, the redundancy of the semantic features in different modalities is eliminated; by performing feature fusion on the de-duplicated semantic features in each modality, visual semantic features are obtained.
[0041] S103. Match the visual semantic features with the event semantic features of different candidate events; among them, the event semantic features of the candidate events are the extraction results of performing cross-modal semantic feature extraction on the event description data of the corresponding candidate events.
[0042] Among them, candidate events can be understood as events of interest, such as events related to quality inspection, safety, or violations, etc., which can be set or adjusted according to requirements or experience.
[0043] Among them, the event description data of candidate events is used to describe the content of the corresponding candidate events in at least one dimension. Among them, the event description data can be presented in at least one of the forms of text, pictures, audio, and video. The present disclosure does not make any limitation on the specific presentation form of the event description data.
[0044] Exemplarily, for the event description data of any candidate event, the same or different feature extraction methods can be used to extract the semantic features of different modalities of the event description data; by removing duplicates from the semantic features of different modalities, the redundancy of the semantic features in different modalities is eliminated; by fusing the semantic features of each modality after removing duplicates, the event semantic features of the candidate event are obtained.
[0045] It should be noted that, for any candidate event, the execution device that generates the event semantic features of the candidate event and the execution device that executes the event detection method may be the same or different. The present disclosure does not make any limitation on this, and only needs to ensure that the event semantic features of different candidate events can be obtained before event detection.
[0046] For the convenience of data search and matching, the event semantic features of different candidate events can be pre-stored in the event retrieval library, and the visual semantic features can be searched and matched in the event retrieval library, thereby avoiding the situation of low search efficiency caused by scattered data or unstable matching results caused by inconsistent search ranges.
[0047] It should be noted that the event retrieval library can be stored in the device or cluster that executes the event detection method, or in other storage devices or clusters that are communicatively connected to the execution device of the event detection method. The present disclosure does not make any limitation on the specific storage location of the event retrieval library.
[0048] Optionally, the similarity between the visual semantic features and the event semantic features of different candidate events can be determined; according to the similarity, the matching result can be determined.
[0049] S104. Determine the target event included in the video to be detected according to the matching result.
[0050] Exemplarily, select the target event of the video frame to be detected from at least one candidate event that matches; use the target events corresponding to different video frames to be detected in the video to be detected as the target events included in the video to be detected.
[0051] Optionally, if the similarity matching method is adopted, at least one candidate event with a relatively high similarity (such as the highest) is selected as the target event corresponding to the video frame to be detected; the target events corresponding to different video frames to be detected in the video to be detected are used as the target events included in the video to be detected.
[0052] Exemplarily, an alarm reminder can be given when a target event is detected; alternatively, the target event can be marked at the position of the corresponding video frame to be detected in the video to be detected; or, the relevant information of the target event can be added to a preset queue for the data requester to consume as needed (for example, it can be consumed in a subscription manner). Among them, the preset queue can be implemented by at least one queue in the prior art, for example, it can be a Kafka queue.
[0053] It should be noted that the present disclosure does not make any limitations on the specific reminder method of the alarm reminder, the specific marking method of the marked target event, and the data consumption method of the data requester.
[0054] In the embodiment of the present disclosure, by introducing the visual semantic features of the video to be detected and searching and matching with the event semantic features corresponding to different candidate events, the target events included in the video to be detected are determined according to the matching results. The above searching and matching process is automatically implemented without manual intervention, improving the matching efficiency and thus the event detection efficiency. Since both the visual semantic features and the event semantic features are cross-modal semantic features, the semantic information carried is richer, more comprehensive and concise, and there are fewer redundant features. Therefore, the accuracy of the target events determined based on the visual semantic features and the event semantic features is higher.
[0055] In an optional embodiment, the video to be detected may include a product monitoring video of a quality inspection product, and the corresponding candidate events may include quality inspection compliance events, such as different quality inspection problem events, so as to adapt to the application scenario of automatically inspecting the quality inspection product.
[0056] In another optional embodiment, the video to be detected may include a security monitoring video, and the corresponding candidate events may include security abnormal events, such as fighting events, or controlled item carrying events, etc., so as to adapt to the application scenario of automatically identifying security problems within the monitoring range.
[0057] In still another optional embodiment, the video to be detected may include a traffic monitoring video, and the corresponding candidate events may include traffic violation events, such as lane crossing driving events, reverse driving events or red light running events, etc., so as to adapt to the application scenario of automatically identifying traffic safety problems within the monitoring range.
[0058] In yet another alternative embodiment, the video to be detected may include a home monitoring video, and the corresponding candidate events may include home security events, such as a guardian falling event, or a home burglary event, etc., so as to adapt to the application scenario of automatically monitoring home security problems.
[0059] It can be understood that by refining the video to be detected and the corresponding candidate events, different application scenarios can be adapted, automatic monitoring of security problems or abnormal problems in the corresponding scenarios can be realized, the scope of use of the event detection method is broadened, and the universality is good.
[0060] Based on the above technical solutions, the present disclosure also provides an alternative embodiment. In this alternative embodiment, the determination mechanism of visual semantic features and event semantic features is optimized and improved. It should be noted that for the parts not detailed in the embodiments of the present disclosure, reference may be made to the relevant descriptions in other embodiments, which will not be elaborated here.
[0061] See Figure 2 An event detection method shown in
[0062] S201. Obtain the video frames to be detected of the video to be detected.
[0063] S202. Based on the vision-language large model, extract cross-modal semantic features from the video frames to be detected to obtain visual semantic features.
[0064] S203. Match the visual semantic features with the event semantic features of different candidate events; wherein, the event semantic features of the candidate events are the extraction results of cross-modal semantic feature extraction based on the event description data of the corresponding candidate events by the vision-language large model.
[0065] S204. Determine the target event included in the video to be detected according to the matching result.
[0066] Among them, the vision-language large model is obtained by performing multi-modal prediction task learning based on the scene graph constructed from different dimensional information of sample objects in sample pictures.
[0067] Among them, the scene graph is used to represent the association relationship between different sample objects and between the attribute information of the same sample object in different sample pictures. Exemplarily, the scene graph of the sample picture can be constructed based on the sample objects, the attribute information of the sample objects, and the association relationship between different sample objects in the sample picture.
[0068] Among them, the multi-modal prediction task may include prediction tasks in at least two modalities, such as at least one of an object prediction task, an attribute information prediction task, and a relationship prediction task. Among them, the large model is used to represent a neural network model with a large number of parameters (such as hundreds of millions of scales).
[0069] It can be understood that since the visual semantic large model is learned for multi-modal prediction tasks based on the scene graph constructed from multi-dimensional information, the semantic understanding ability of the visual semantic large model will be greatly improved on the basis of the conventional neural network model. At the same time, the introduction of the scene graph enables the model to more accurately grasp the fine-grained semantic alignment between visual semantics and cross-modalities.
[0070] In a specific embodiment, the vision-language large model can be ERNIE-VIL (Knowledge Enhanced Vision-Language Representations Through Scene Graph).
[0071] In the embodiments of the present disclosure, by introducing the same vision-language large model to extract cross-modal semantic features from the video frame to be detected and the event description data of the candidate event, the corresponding visual semantic features and event semantic features are obtained, ensuring the modal consistency and multi-modal fusion of the features extracted from the video frame to be detected and the event description data, and avoiding the situation where the semantic features cannot be matched due to the use of semantic features in a single modality or different modalities. Therefore, when event matching is performed through visual semantic features and event semantic features, false matching and missed matching are avoided, and the accuracy of the event matching result is improved.
[0072] In an alternative embodiment, the determination of whether there is a target event can be implemented according to the matching result. For example, if a certain candidate event is included in the matching result, the candidate event is used as the target event included in the video to be detected; if a certain candidate event is not included in the matching result, the candidate event is prohibited from being used as the target event included in the video to be detected.
[0073] In another alternative embodiment, the determination of the fine-grained information of the target event can also be implemented according to the matching result.
[0074] Exemplarily, the event attributes of the matched candidate event can be directly used as the event attributes of the target event. Among them, the event attributes of the candidate event can include at least one of event category, event content label, event severity, etc. Among them, the event category is used to classify events with the same category attributes. It should be noted that the number of candidate events under the same event category can be at least one. Among them, the event content label is used to reflect the content summary or theme of the corresponding candidate event, etc. Among them, the event severity is used to characterize the severity of the corresponding candidate event, which can be presented by a severity level or a degree score.
[0075] Among them, the event attributes of the candidate events can be set manually by technicians or obtained by extracting attribute features from the event description data. The present disclosure does not impose any limitation on the acquisition method of the event attributes.
[0076] Optionally, the event severity of the target event can also be determined according to the duration of the target event in the video to be detected. Among them, the event severity can be positively correlated with the duration, that is, the longer the duration, the more serious the event.
[0077] Alternatively, optionally, the event severity of the target event can also be determined according to at least one of the degree of crowd density of the people appearing in the video frames to be detected of the target event and the item categories of the items associated with the event. For example, the greater the degree of crowd density, the more serious the event; the higher the control level of the item category, the more serious the event, etc.
[0078] It can be understood that by introducing the event attributes of the candidate events to assist in determining the event attributes such as the event category and / or event severity of the target event, it is possible to further determine the event attributes with finer granularity during the process of determining the target event included in the video to be detected, realizing event classification and / or event severity division, and improving the richness of the information carried by the event detection results.
[0079] Exemplarily, the present disclosure can use the same or different alarm methods to give alarm reminders for target events of different event categories, or use the same or different marking methods to mark events. Among them, different alarm methods can be distinguished by alarm categories such as sound, light, text, or vibration, or configuration attributes under the same alarm method. For example, the pitch, timbre, or frequency of sound; the color or frequency of light; the font, thickness, foreground color, or background color of text; the intensity or frequency of vibration. Among them, different marking methods can be distinguished by the category of the marking symbol or configuration attributes such as the size, color, thickness, foreground color, or background color of the same marking symbol.
[0080] Based on the technical solutions of the above embodiments, historical event retrieval and positioning can also be assisted through event detection. Exemplarily, the video to be detected can be the video to which the event to be queried belongs; correspondingly, obtaining the video frames to be detected of the video to be detected can be obtaining the video frames to be detected of the video to which the event to be queried belongs; correspondingly, after determining the target event in the video to which the event to be queried belongs, the corresponding target event can be quickly located directly in the historical event detection results, improving the event retrieval efficiency.
[0081] Based on the technical solutions of the above embodiments, the relevant data of the candidate events can also be assisted in being enriched through event detection.
[0082] In an alternative embodiment, the event semantic features of different candidate events can be stored in an event retrieval library; correspondingly, the event attributes of the corresponding target event in the event retrieval library can be supplemented according to the attribute data of the video frame to be detected that matches the target event.
[0083] Among them, the attribute data of the video frame to be detected can be semantic labels determined based on the description data of the video frame to be detected (such as at least one of text and picture, etc.).
[0084] It can be understood that through the above method, the richness and comprehensiveness of the event attributes of the target event in the event retrieval library can be gradually improved, laying a foundation for improving the comprehensiveness of the event attributes in the event detection result obtained in subsequent event detection.
[0085] In another alternative embodiment, the event semantic features of different candidate events can be stored in an event retrieval library; correspondingly, the visual semantic features of the video frame to be detected that matches the target event are added to the event retrieval library as the event semantic features of other candidate events under the same event category as the target event.
[0086] It can be understood that through the above method, the richness and comprehensiveness of different candidate events under the same event category in the event retrieval library can be gradually improved, providing data support for the matching of more fine-grained candidate events under the same event category in the future.
[0087] Based on the above technical solutions, the present disclosure also provides a preferred embodiment, which will be described in detail below in combination with Figure 3A the architecture diagram of the event detection system shown in Figure 3B the event detection method shown in
[0088] Exemplarily, Figure 3A the PAAS (Platform as a Service) layer of the event detection system architecture shown in can be implemented based on a Kubernetes (abbreviation: K8s, which is an open-source system for managing containerized applications on multiple hosts in a cloud platform) cluster; the IAAS (Infrastructure as a Service) layer can be implemented based on cloud infrastructure.
[0089] Referring to Figure 3B the event detection method shown in, including:
[0090] S301. Configure alarm events: In response to a monitoring event configuration operation, pre-configure the event description data of at least one alarm event under different categories.
[0091] Among them, the event description data may include at least one of descriptive text, pictures, audio clips, video clips, etc.
[0092] Among them, the alarm event configuration operation can be implemented by calling the event registration interface. Exemplarily, it can be through Figure 3A the event management module in the service layer in Figure 3A to perform alarm event configuration; furthermore, the configured alarm events can also be queried through the event management module. Optionally, it can be through
[0093] the interface / gateway in the application layer in Figure 3A to provide interface call and data transmission services.
[0094] S302. Event semantic feature extraction: Based on the visual semantic large model, extract the event semantic features in the event description data.
[0095] Among them, the visual semantic large model is obtained by performing multi-modal prediction task learning based on the scene graph constructed from different dimensional information of each sample object in the sample images. For example, the visual semantic large model can be the ERNIE_VIL model.
[0096] S303. Event retrieval library construction / updating: Correlate and store the event semantic features of different alarm events, their alarm categories, and alarm severity levels into the event retrieval library.
[0097] Among them, the alarm categories and alarm severity levels corresponding to different alarm events can be set by technicians according to needs or experience.
[0098] Exemplarily, it can be through Figure 3A the semantic library in
[0099] to realize the construction and updating of the event retrieval library.
[0100] S304. Obtaining the video to be detected: Obtain the video to be detected. Figure 3A Exemplarily, it can be through Figure 3A the media asset management module in the service layer in
[0101] to perform storage management on the video to be detected; it can be through
[0102] Exemplarily, the similarity between adjacent picture frames constituting the video to be detected can be determined, and at least one of the adjacent picture frames can be used as the video frame to be detected when the similarity difference is large (e.g., less than a preset difference threshold); the video frames to be detected in the video to be detected are extracted according to a preset frame extraction frequency; and duplicate removal processing is performed on the video frames to be detected to update the video frames to be detected. Among them, the preset difference threshold and frame extraction frequency can be set by technicians according to empirical values or determined through a large number of experiments.
[0103] S306. Visual semantic feature extraction: Based on a visual semantic large model, extract the visual semantic features in the video frame to be detected.
[0104] Exemplarily, it can be deployed in Figure 3A the model layer of to extract the event semantic features corresponding to the alarm event and the visual semantic features corresponding to the video frame to be detected.
[0105] Exemplarily, it can be through Figure 3A the task management module in the service layer of to perform concurrent control when batch extracting visual semantic features. For example, task queues can be set up in advance, and feature extraction is performed sequentially according to the arrangement order of each visual semantic feature extraction task in the task queue.
[0106] Exemplarily, it can be through Figure 3A the semantic library in to store the visual semantic features, which only need to be distinguished from the aforementioned event retrieval library. For example, the semantic library can be implemented based on FAISS (Facebook AI Similarity Search, a similarity search library).
[0107] S307. Semantic feature retrieval: Match the alarm event from the event retrieval library according to the similarity between the visual semantic features and each event semantic feature in the event retrieval library.
[0108] Exemplarily, it can be through Figure 3A the semantic query module in the service layer of to query and match the alarm event in the event retrieval library.
[0109] S308A. Trigger alarm: If there is an alarm event with a similarity greater than the preset similarity threshold, trigger the alarm.
[0110] Among them, the preset similarity threshold can be set by technicians according to empirical values or determined through a large number of experiments.
[0111] Exemplarily, event alarms can be performed based on the alarm event, its event category, and event severity. Among them, the alarm methods for alarm events with different event categories or different event severities are the same or different.
[0112] Exemplarily, it can be triggered by Figure 3A the alarm management module in the service layer in
[0113] Optionally, it is also possible to extract the text tags of the video segment to which the video frame to be detected that matches the alarm event belongs, and determine the event category of the corresponding alarm event according to the extraction result.
[0114] Exemplarily, it can be through Figure 3A the event library in
[0115] S308B, Queue update: Generate queue messages for the alarm events with a similarity greater than the preset similarity threshold, their event categories, and event severity levels, and add them to the preset queue for data requesters to subscribe and consume as needed.
[0116] Among them, the preset similarity threshold can be set by technicians according to empirical values or determined through a large number of experiments.
[0117] Among them, the preset queue can be a Kafka queue.
[0118] It should be noted that the device for executing S301 to S303 and the device for executing S304 to S308A / S308B can be the same or different, and the present disclosure does not make any limitation on this.
[0119] As an implementation of the above event detection methods, the present disclosure also provides an optional embodiment of an execution device for implementing the above event detection methods.
[0120] See Figure 4 An event detection device 400 shown in
[0121] The module 401 for obtaining video frames to be detected is used to obtain the video frames to be detected of the video to be detected;
[0122] The module 402 for obtaining visual semantic features is used to extract cross-modal semantic features of the video frames to be detected to obtain visual semantic features;
[0123] The visual semantic feature matching module 403 is configured to match the visual semantic features with the event semantic features of different candidate events; wherein, the event semantic features of the candidate events are the extraction results of cross-modal semantic feature extraction performed on the event description data of the corresponding candidate events.
[0124] The target event determination module 404 is configured to determine the target event included in the video to be detected according to the matching result.
[0125] In the embodiment of the present disclosure, by introducing the visual semantic features of the video to be detected and searching and matching them with the event semantic features corresponding to different candidate events, the target event included in the video to be detected is determined according to the matching result. The above-mentioned searching and matching process is automatically implemented without human intervention, which improves the matching efficiency and further improves the event detection efficiency. Since both the visual semantic features and the event semantic features are cross-modal semantic features, the semantic information carried is richer, more comprehensive and concise, and there are fewer redundant features. Therefore, the target event determined based on the visual semantic features and the event semantic features has higher accuracy.
[0126] In an alternative embodiment, the visual semantic feature obtaining module 402 is specifically configured to:
[0127] Extract cross-modal semantic features from the video frames to be detected based on a visual language large model to obtain visual semantic features.
[0128] The event semantic features of the candidate events are the extraction results of cross-modal semantic feature extraction performed on the event description data of the corresponding candidate events based on the visual language large model.
[0129] Wherein, the visual language large model is obtained by performing multi-modal prediction task learning based on a scene graph constructed from different dimensional information of sample objects in sample pictures.
[0130] In an alternative embodiment, the target event determination module 404 is specifically configured to:
[0131] Use the event attributes of the matched candidate events as the event attributes of the target event.
[0132] Wherein, the event attributes include event categories and / or event severity levels.
[0133] In an alternative embodiment, the video to be detected is a video belonging to the event to be queried; the apparatus further includes:
[0134] The target event location module is configured to locate the target event in the historical event detection results.
[0135] In an alternative embodiment, the event semantic features of different candidate events are stored in an event retrieval library; the apparatus further includes:
[0136] An event attribute supplement module, configured to supplement the event attributes of a corresponding target event in the event retrieval library according to the attribute data of a to-be-detected video frame that matches the target event; and / or,
[0137] A candidate event addition module, configured to add the visual semantic features of a to-be-detected video frame that matches the target event as the event semantic features of other candidate events under the same event category as the target event to the event retrieval library.
[0138] In an alternative embodiment, the apparatus further includes a to-be-detected video frame determination module, which specifically includes:
[0139] An original video frame determination unit, configured to determine each original video frame in the to-be-detected video;
[0140] A first candidate video frame determination unit, configured to select at least one first candidate video frame from each of the original video frames according to the difference situation between adjacent original video frames;
[0141] A second candidate video frame determination unit, configured to extract at least one second candidate video frame from each of the original video frames according to a preset frame extraction frequency;
[0142] A to-be-detected video frame determination unit, configured to determine the to-be-detected video frame according to the union of the first candidate video frame and the second candidate video frame.
[0143] In an alternative embodiment, the to-be-detected video and the corresponding candidate events under the to-be-detected video include at least one of the following:
[0144] The to-be-detected video includes a product monitoring video of a quality inspection product, and the corresponding candidate event includes a quality inspection compliance event;
[0145] The to-be-detected video includes a security monitoring video, and the corresponding candidate event includes a security anomaly event;
[0146] The to-be-detected video includes a traffic monitoring video, and the corresponding candidate event includes a traffic violation event;
[0147] The to-be-detected video includes a home monitoring video, and the corresponding candidate event includes a home security event.
[0148] The above event detection apparatus can execute the event detection method provided in any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects for executing each event detection method.
[0149] In the technical solution of the present disclosure, the processing of collection, storage, use, processing, transmission, provision, and disclosure of the video to be detected, the video frame to be detected, the event semantic features, etc. all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0150] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0151] Figure 5 FIG. shows a schematic block diagram of an exemplary electronic device 500 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0152] As Figure 5 shown, the device 500 includes a computing unit 501 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0153] A plurality of components in the device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a disk, an optical disc, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0154] The computing unit 501 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 executes the various methods and processes described above, such as the event detection method. For example, in some embodiments, the event detection method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the event detection method described above can be executed. Alternatively, in other embodiments, the computing unit 501 can be configured to execute the event detection method by any other suitable means (e.g., by means of firmware).
[0155] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0156] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0157] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0158] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0159] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0160] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The client-server relationship is generated by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services. The server can also be a server of a distributed system or a server combined with blockchain.
[0161] Artificial intelligence is a discipline that studies how to make a computer simulate certain thinking processes and intelligent behaviors of humans (such as learning, reasoning, thinking, planning, etc.), and it has technologies at both the hardware level and the software level. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, machine learning / deep learning technology, big data processing technology, and knowledge graph technology.
[0162] Cloud computing refers to a technical system that accesses an elastic and scalable shared physical or virtual resource pool through a network. The resources can include servers, operating systems, networks, software, applications, and storage devices, etc., and the resources can be deployed and managed in a demand-based and self-service manner. Through cloud computing technology, it can provide efficient and powerful data processing capabilities for the application and model training of technologies such as artificial intelligence and blockchain.
[0163] It should be understood that various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recorded in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions provided by this disclosure can be achieved, and no limitations are imposed herein.
[0164] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. An event detection method, comprising: Obtaining a video frame to be detected of a video to be detected; Extracting cross-modal semantic features from the video frame to be detected to obtain visual semantic features; Matching the visual semantic features with event semantic features of different candidate events; wherein, the event semantic features of the candidate events are extraction results of performing cross-modal semantic feature extraction on event description data of corresponding candidate events; Determining a target event included in the video to be detected according to the matching result; Wherein, the extracting cross-modal semantic features from the video frame to be detected to obtain visual semantic features includes: based on a vision-language large model, extracting cross-modal semantic features from the video frame to be detected to obtain visual semantic features; the event semantic features of the candidate events are extraction results of performing cross-modal semantic feature extraction on event description data of corresponding candidate events based on the vision-language large model; wherein, the vision-language large model is obtained by performing multi-modal prediction task learning based on a scene graph constructed from different-dimensional information of sample objects in sample pictures; Wherein, the video frame to be detected of the video to be detected is determined by the following method: determining each original video frame in the video to be detected; selecting at least one first candidate video frame from each of the original video frames according to the difference situation between adjacent original video frames; extracting at least one second candidate video frame from each of the original video frames according to a preset frame extraction frequency; determining the video frame to be detected according to the union of the first candidate video frame and the second candidate video frame.
2. The method according to claim 1, wherein The determining the target event included in the video to be detected according to the matching result includes: Taking the event attribute of the matched candidate event as the event attribute of the target event; Wherein, the event attribute includes an event category and / or an event severity level.
3. The method according to claim 1, wherein The video to be detected is a video to which a query event belongs; the method further includes: Locating the target event in historical event detection results.
4. The method according to claim 1, wherein, The event semantic features of different candidate events are stored in an event retrieval library; the method further includes: Supplementing the event attribute of the corresponding target event in the event retrieval library according to the attribute data of the video frame to be detected that matches the target event; and / or, Taking the visual semantic features of the video frame to be detected that matches the target event as the event semantic features of other candidate events under the same event category as the target event, and adding them to the event retrieval library.
5. The method according to claim 1, wherein The video to be detected and the corresponding candidate events thereof include at least one of the following: The video to be detected includes a product monitoring video of a quality-inspected product, and the corresponding candidate event includes a quality inspection compliance event; The video to be detected includes a security monitoring video, and the corresponding candidate event includes a security anomaly event; The video to be detected includes a traffic monitoring video, and the corresponding candidate event includes a traffic violation event; The video to be detected includes a home monitoring video, and the corresponding candidate event includes a home safety event.
6. An event detection device, comprising: A video frame to be detected obtaining module, configured to obtain a video frame to be detected of a video to be detected; A visual semantic feature obtaining module, configured to extract cross-modal semantic features from the video frame to be detected, so as to obtain visual semantic features; A visual semantic feature matching module, configured to match the visual semantic features with the event semantic features of different candidate events; wherein, the event semantic features of the candidate events are the extraction results of performing cross-modal semantic feature extraction on the event description data of the corresponding candidate events; A target event determining module, configured to determine the target event included in the video to be detected according to the matching result; Wherein, the visual semantic feature obtaining module is specifically configured to: based on a vision-language large model, extract cross-modal semantic features from the video frame to be detected, so as to obtain visual semantic features; the event semantic features of the candidate events are the extraction results of performing cross-modal semantic feature extraction on the event description data of the corresponding candidate events based on the vision-language large model; wherein, the vision-language large model is obtained by performing multi-modal prediction task learning based on a scene graph constructed from different-dimensional information of sample objects in sample pictures; Wherein, the device further includes a video frame to be detected determining module, which specifically includes: an original video frame determining unit, configured to determine each original video frame in the video to be detected; a first candidate video frame determining unit, configured to select at least one first candidate video frame from each of the original video frames according to the difference between adjacent original video frames; a second candidate video frame determining unit, configured to extract at least one second candidate video frame from each of the original video frames according to a preset frame extraction frequency; and a video frame to be detected determining unit, configured to determine the video frame to be detected according to the union of the first candidate video frames and the second candidate video frames.
7. The apparatus according to claim 6, wherein, The target event determining module is specifically configured to: Use the event attributes of the matched candidate event as the event attributes of the target event; Wherein, the event attributes include event category and / or event severity.
8. The device according to claim 6, wherein The video to be detected is a video to which the event to be queried belongs; the device further includes: A target event locating module, configured to locate the target event in the historical event detection results.
9. The device according to claim 6, wherein The event semantic features of different candidate events are stored in an event retrieval library; the device further includes: An event attribute supplementing module, configured to supplement the event attributes of the corresponding target event in the event retrieval library according to the attribute data of the video frame to be detected that matches the target event; and / or, A candidate event adding module, configured to add the visual semantic features of the video frame to be detected that matches the target event as the event semantic features of other candidate events under the same event category as the target event to the event retrieval library.
10. The apparatus according to claim 6, wherein, The video to be detected and the corresponding candidate events thereof include at least one of the following: The video to be detected includes a product monitoring video of a quality inspection product, and the corresponding candidate event includes a quality inspection compliance event; The video to be detected includes a security monitoring video, and the corresponding candidate event includes a security anomaly event; The video to be detected includes a traffic monitoring video, and the corresponding candidate event includes a traffic violation event; The video to be detected includes a home surveillance video, and the corresponding candidate events include home security events.
11. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the event detection method according to any one of claims 1-5.
12. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause a computer to execute the event detection method according to any one of claims 1-5.
13. A computer program product, comprising a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the event detection method according to any one of claims 1-5 are implemented.
Citation Information
Patent Citations
Video event identification method and device, equipment and storage medium
CN113361344A
Video searching method and device, electronic device and storage medium
CN116127127A