Event retrieval method and system in mixed scene, electronic equipment and storage medium
By applying multimodal large models and preset event databases in massive video streaming data, the problem of traditional CV models identifying events in multiple scenarios is solved, and efficient and automated event retrieval is achieved, which is suitable for IoT video monitoring in smart cities.
Patent Information
- Application Number
- CN202411923436.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-05-13
AI Technical Summary
It is difficult for the prior art to efficiently retrieve and identify event information in massive video stream data, especially in multiple scenarios, traditional CV models cannot quickly identify undefined event types, and the search process lacks automation.
An event retrieval method in a mixed scenario is proposed. By receiving video stream information, dividing it into image sets, and using a multimodal large model combined with a preset event library, it generates prompt words for recognition and storage, and finally searches the target event through tags and description information.
It alleviates the limitations of CV model recognition in specific scenarios, realizes efficient retrieval of multi-scene events, improves retrieval efficiency, and reduces manual participation through automated processes.
Smart Images

Figure CN119992405A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of Internet of Things technology, and in particular to an event retrieval method, system, electronic device and storage medium in a hybrid scenario. Background Art
[0002] As the construction of smart cities continues to advance, the number of IoT devices such as cameras is growing exponentially. Faced with massive amounts of video stream data, how to efficiently retrieve and accurately identify the information collected by cameras places extremely high demands on urban IoT platforms. The efficiency of video stream retrieval relies on large-scale storage and computing power, which cannot meet the needs of fast and efficient business development.
[0003] In the related art, manual video review is used to identify the current events, archive them, and provide feedback based on the events that need to be retrieved. This method is time-consuming and costly, and it is unable to process massive amounts of video stream data. Another embodiment uses AI event recognition to identify targets in the fields of security and fire protection by installing traditional CV models on camera devices. This method reduces manual participation to a certain extent, and business personnel can view them according to the event label classification of different cameras, thereby improving retrieval efficiency. However, event types that are not defined in advance by the system cannot be quickly retrieved. At the same time, the retrieval process is not automated. Summary of the invention
[0004] The main purpose of the embodiments of the present application is to provide an efficient event retrieval method, system, electronic device and storage medium in a hybrid scenario.
[0005] To achieve the above purpose, one aspect of the embodiment of the present application proposes an event retrieval method in a mixed scene, the method comprising: receiving video stream information in several scenes, segmenting the video stream information, and obtaining several image sets; generating the first prompt word of the first image set according to the preset event library; the first image set is any one of the image sets in the image set; the preset event library includes the description information and label information of the image set; inputting the first prompt word and the first image set into the multimodal large model, obtaining the first cognitive result corresponding to the first image set, and storing the first cognitive result in the preset event library; the first cognitive result includes the description information and label information of the first image set; obtaining the question information of the target event to be retrieved; searching the preset event library according to the question information to obtain the target event. The embodiment of the present application recognizes the image set corresponding to the video stream through the multimodal large model, which alleviates the limitation of the CV model corresponding to the recognition of specific scenes; at the same time, the present application can retrieve the target event through the label information and description information, which improves the retrieval efficiency. The present application can retrieve events in multiple scenes, which is conducive to improving the retrieval efficiency.
[0006] In some embodiments, the method provided by the embodiments of the present application, generating the first prompt word of the first image set according to the preset event library includes:
[0007] If the first image set is not matched in the preset event library, generating a first prompt word for the first image set according to a first prompt word template;
[0008] Alternatively, if the first image set is matched in the preset event library, a first prompt word of the first image set is generated according to a second prompt word template; and the label parts of the first prompt word template and the second prompt word template are different.
[0009] In some embodiments, in the method provided by the embodiments of the present application, the prompt word template is determined by the following steps:
[0010] If the first image set is not matched in the preset event library, determining attribute information corresponding to the label in the first prompt word template; the attribute information includes character length;
[0011] Alternatively, if the first image set is matched in the preset event library, a range of optional tags corresponding to tags in the second prompt word template is determined.
[0012] In some embodiments, the method provided in the embodiments of the present application, storing the first cognitive result in the preset event library includes:
[0013] storing the label information, the description information and the numbering information corresponding to the first image set in a vector database;
[0014] The label information and the serial number information corresponding to the first image set are stored in a relational database.
[0015] In some embodiments, the method provided by the embodiments of the present application, searching the preset event library according to the question information to obtain the target event, includes:
[0016] If the question information includes a tag, a search is performed in the preset event library based on the tag to obtain target number information corresponding to the target tag; the target number information is the number corresponding to the target image set;
[0017] According to the target number information, a video clip is obtained to determine the target event.
[0018] In some embodiments, the method provided by the embodiments of the present application, searching the preset event library according to the question information to obtain the target event, includes:
[0019] According to the question information, based on the search enhancement generation technology, a search is performed in the preset event library to obtain a target event.
[0020] In some embodiments, the method provided by the embodiments of the present application, wherein the video stream information is segmented to obtain a plurality of image sets, includes:
[0021] The video stream is segmented according to a preset time length, and video frames are extracted according to a preset frame rate to obtain an image set;
[0022] According to the start time, end time, camera code, and unique value code of the current image set, the numbering information of the current image set is added.
[0023] To achieve the above purpose, another aspect of the embodiment of the present application provides an event retrieval system in a hybrid scenario, the system comprising:
[0024] The first module is used to receive video stream information in several scenes, and segment the video stream information to obtain several image sets;
[0025] The second module is used to generate a first prompt word of the first image set according to a preset event library; the first image set is any image set in the image set; the preset event library includes description information and label information of the image set;
[0026] A third module is used to input the first prompt word and the first image set into the multimodal large model to obtain a first cognitive result corresponding to the first image set, and store the first cognitive result in the preset event library; the first cognitive result includes description information and label information of the first image set;
[0027] The fourth module is used to obtain question information about the target event to be retrieved;
[0028] The fifth module is used to search the preset event library according to the question information to obtain the target event.
[0029] To achieve the above objective, another aspect of an embodiment of the present application provides an electronic device, the electronic device comprising a memory and a processor, the memory storing a computer program, and the processor implementing the above method when executing the computer program.
[0030] To achieve the above objective, another aspect of an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program implements the above method when executed by a processor.
[0031] The embodiment of the present application includes at least the following beneficial effects: The method provided by the embodiment of the present application includes: receiving video stream information under several scenes, segmenting the video stream information, and obtaining several image sets; generating the first prompt word of the first image set according to the preset event library; the first image set is any one of the image sets in the image set; the preset event library includes the description information and label information of the image set; inputting the first prompt word and the first image set into the multimodal large model, obtaining the first cognitive result corresponding to the first image set, and storing the first cognitive result in the preset event library; the first cognitive result includes the description information and label information of the first image set; obtaining the question information of the target event to be retrieved; searching in the preset event library according to the question information to obtain the target event. The embodiment of the present application recognizes the image set corresponding to the video stream through the multimodal large model, which alleviates the limitation of the CV model corresponding to the recognition of specific scenes; at the same time, the present application can retrieve the target event through the label information and description information, which improves the retrieval efficiency. The present application can retrieve events in multiple scenes, which is conducive to improving the retrieval efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 is a flowchart of an embodiment of an event retrieval method in a hybrid scenario provided by the present application;
[0033] Figure 2 is a flowchart of another embodiment of the event retrieval method in a hybrid scenario provided by the present application;
[0034] Figure 3 is a flowchart of an embodiment of a process for generating a first prompt word provided by the present application;
[0035] Figure 4 is a flowchart of an embodiment of a prompt word template generation process provided by the present application;
[0036] Figure 5 is a flowchart of an embodiment of the data storage process provided by the present application;
[0037] Figure 6 is a flowchart of an embodiment of the retrieval process provided by the present application;
[0038] Figure 7 This is a schematic diagram of a scenario of an embodiment corresponding to a scenario provided by the present application;
[0039] Figure 8 is a schematic diagram of a scenario of an embodiment corresponding to another scenario provided by the present application;
[0040] Fig. 9 It is a structural diagram of an event retrieval system in a hybrid scenario provided by an embodiment of the present application;
[0041] Fig.10 It is a schematic diagram of the hardware structure of the electronic device provided in the embodiment of the present application. DETAILED DESCRIPTION
[0042] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described in detail below in conjunction with the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present application. They are only examples of devices and methods consistent with some aspects of the embodiments of the present application as detailed in the attached claims.
[0043] It is understood that the terms "first", "second", etc. used in this application can be used to describe various concepts in this article, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another concept. For example, without departing from the scope of the embodiment of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "if" as used herein can be interpreted as "at the time of" or "when" or "in response to determination".
[0044] The terms "at least one", "multiple", "each", "any", etc. used in this application, at least one includes one, two or more, multiple includes two or more, each refers to each of the corresponding multiple, and any refers to any one of the multiple.
[0045] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0046] Before describing the embodiments of the present application in detail, some nouns and terms involved in the embodiments of the present application are first described. The nouns and terms involved in the embodiments of the present application are subject to the following explanations.
[0047] Multimodal Large Language Models (MM-LLMs);
[0048] Retrieval-Augmented Generation (RAG);
[0049] Prompt Engineering (PE).
[0050] As the construction of smart cities continues to advance, the number of IoT devices such as cameras is growing exponentially. Faced with massive amounts of video stream data, how to efficiently retrieve and accurately identify the information collected by cameras places extremely high demands on urban IoT platforms. The efficiency of video stream retrieval relies on large-scale storage and computing power, which cannot meet the needs of fast and efficient business development.
[0051] The currently widely used event retrieval methods in the field of urban IoT videos mainly include the following two:
[0052] 1. Manual video review;
[0053] Calling past video data to identify current events manually for archiving and sorting, and providing feedback based on the events to be retrieved. This method is time-consuming and costly, and cannot process massive amounts of video stream data.
[0054] 2. AI event recognition;
[0055] Mainly by installing traditional CV models on camera devices, it can identify targets in the fields of security, fire protection, etc. Among them, the computer vision (CV) model mainly performs time retrieval through CV retrieval. Specifically, CV retrieval usually refers to computer vision retrieval, which refers to the process of using computer vision technology to find images or videos similar to a given query image or video from an image or video data set. Common types are as follows: Image retrieval: Text search: The user enters a text description, and the system retrieves related images from the image database based on the description. For example, if you enter "red flowers", the system will find pictures containing red flowers. Image search: The user uploads or enters a picture, and the system searches for similar pictures in the database. It can be used to find photos of the same object from different angles and different scenes, or pictures of similar styles and content. Video retrieval: Content-based video retrieval: Retrieval of videos is achieved by analyzing the visual content of the video, such as color, texture, shape, motion trajectory, etc. For example, finding video clips containing specific people, specific scenes, or specific actions.
[0056] In actual scene monitoring, AI cameras are used to monitor violations and abnormal phenomena. When an abnormality occurs, the event is verified and confirmed by retrieving the relevant camera video to verify the reported event, and the video clips are saved for future viewing. This method reduces manual participation to a certain extent, and business personnel can view it according to the event labels of different cameras, which improves retrieval efficiency. However, event types that are not defined in advance by the system cannot be quickly retrieved. At the same time, the retrieval process is not automated.
[0057] It can be seen that the relevant technology has the following shortcomings:
[0058] 1. The CV model installed in the traditional AI algorithm camera can only achieve "seeing" and "seeing clearly" but cannot solve the problem of "understanding". It is difficult for a single CV model to meet the requirements for complex multi-scene discrimination.
[0059] 2. Manual retrieval is inefficient and consumes a lot of resources. The alarms generated by monitoring need to be manually pulled to check the video stream, and fuzzy retrieval is not possible for a certain alarm event. Therefore, in actual applications, due to the lack of sufficient manpower for verification, a large amount of video alarm data is not fully utilized.
[0060] In view of this, an event retrieval method in a hybrid scenario is provided in an embodiment of the present application, aiming to improve the application breadth of the model and improve the retrieval efficiency.
[0061] The event retrieval method in a hybrid scenario provided by the embodiment of the present application relates to the field of smart city and Internet of Things technology. The event retrieval method in a hybrid scenario provided by the embodiment of the present application can be applied to a terminal, can be applied to a server, or can be software running in a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, and a car terminal, etc., but is not limited to this; the server side can be configured as an independent physical server, or a server cluster or distributed system composed of multiple physical servers, and can also be configured to provide cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and cloud servers for basic cloud computing services such as big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application that implements the event retrieval method in a hybrid scenario, etc., but is not limited to the above forms.
[0062] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0063] It should be noted that in each specific implementation of the present application, when it comes to the need to perform relevant processing based on data related to user identity or characteristics such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first, and the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiment of the present application needs to obtain the user's sensitive personal information, the user's separate permission or consent will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the user's separate permission or consent, the necessary user-related data for the normal operation of the embodiment of the present application will be obtained.
[0064] Figure 1 This is an optional flowchart of the event retrieval method in a hybrid scenario provided by an embodiment of the present application; Figure 1 The method may include but is not limited to steps S100 to S400.
[0065] Step S100, receiving video stream information in several scenes, segmenting the video stream information to obtain several image sets;
[0066] Step S200, generating a first prompt word of a first image set according to a preset event library; the first image set is any image set in the image set; the preset event library includes description information and label information of the image set;
[0067] Step S300, inputting the first prompt word and the first image set into the multimodal large model, obtaining a first cognitive result corresponding to the first image set, and storing the first cognitive result in a preset event library; the first cognitive result includes description information and label information of the first image set;
[0068] Step S400, obtaining question information about the target event to be retrieved;
[0069] Step S500: searching in a preset event library according to the question information to obtain a target event.
[0070] The video stream information in the embodiment of the present application can be video information in any scene or field, and the present application does not limit the specific attributes of the video stream. Figure 2 As shown, the preset event library is used to store the video content and corresponding labels of the identified video. The present application prompts the multimodal large model through the first prompt word. The first prompt word corresponds to the first image set. It can be understood that for any image set obtained by the video stream, a corresponding prompt word is generated. The first prompt word can be generated by a prompt word template. Specifically, the specific content of the first prompt word is determined according to whether there is information related to the first image set in the preset event library. The first prompt word contains preset event label information, which is convenient for defining event labels for the image set using the multimodal large model. The embodiment of the present application identifies the first image set through the multimodal large model to obtain cognitive information of the first image set. Through the processing of the multimodal large model, the information corresponding to the video is converted into cognitive information and stored in the preset event library, which is convenient for subsequent retrieval. The question information in the present application can be information containing labels or other information.
[0071] It is understandable that in other embodiments, if the event label cannot be matched, the event label is generated according to the image set using the multimodal large model through the second prompt word, and the preset event library and the first prompt word are updated. That is, the present application is divided into two cases: if the event label is matched, the first prompt word template generates the first image set recognition result; if the event library event label cannot be matched, the event label is generated using the second prompt word and the database is updated.
[0072] This application can cope with complex scenarios. This method combines a multimodal model for monitoring, which makes up for the fact that traditional CV can only understand single event detection scenarios with trained models, and cannot recognize untrained scenarios. However, there are often many sudden unknown events in actual scenarios, and CV models have limitations in dealing with such scenarios. This method is based on the image recognition capabilities of a large multimodal model, combined with NLP data analysis algorithms to solve the identification of undefined sudden events.
[0073] This method supports both precise retrieval and question retrieval. This method can quickly retrieve data by mapping tags and video streams stored in a relational database. At the same time, this method can also understand the automatically extracted video key frames through a multimodal large model, map the video content description with event tags, and store them in a vector database. It uses RAG, prompt and other technologies to realize efficient retrieval and query of video events. Since the storage form is converted from video stream data to vector data, it can save a lot of storage resources and improve retrieval efficiency.
[0074] In some embodiments, reference Figure 3 As shown, the method provided in the embodiment of the present application generates a first prompt word of a first image set according to a preset event library, including:
[0075] Step S210, if the first image set is not matched in the preset event library, generating a first prompt word of the first image set according to a first prompt word template;
[0076] Alternatively, in step S220, if the first image set is matched in the preset event library, a first prompt word for the first image set is generated according to the second prompt word template; the label part of the first prompt word template and the second prompt word template are different.
[0077] In some possible implementations, if there is information related to the first image set in the preset event library, it is considered that the preset event library matches the first image set; the related information may be a label or other information. At this time, the second prompt word template is determined according to the label, and then the first prompt word is determined according to the second prompt word template. In other embodiments, if the preset event library does not match the first image set, the first prompt word of the first image set is generated according to the first prompt word template.
[0078] In some embodiments, reference Figure 4 As shown, in the method provided in the embodiment of the present application, the prompt word template is determined by the following steps:
[0079] Step S230, if the first image set is not matched in the preset event library, determining the attribute information corresponding to the label in the first prompt word template; the attribute information includes the character length;
[0080] Alternatively, in step S240, if the first image set is matched in the preset event library, a range of optional tags corresponding to the tags in the second prompt word template is determined.
[0081] In some possible implementations, if the first image set is not matched in the preset event library, the attribute information of the label information corresponding to the label part in the first prompt word template is determined. The attribute information can be a character length or a preset character. In another embodiment, if the first image set is matched in the preset event library, the optional label range corresponding to the label in the second prompt word template is determined according to the scene information to which the first image set belongs or according to the matched label.
[0082] In some embodiments, reference Figure 5 As shown, the method provided in the embodiment of the present application stores the first cognitive result in a preset event library, including:
[0083] Step S310, storing the label information, description information and number information corresponding to the first image set into a vector database;
[0084] Step S320: store the label information and the serial number information corresponding to the first image set into a relational database.
[0085] In some possible implementations, the vector database can be used to retrieve images based on description information by query information; the relational database can be used to retrieve images by tags. It can be understood that the preset event library in the present application includes a vector database and a relational database. In other embodiments, if the first image set is matched in the preset event library, the matched tag information and number information are stored in the relational database.
[0086] In some embodiments, reference Figure 6 As shown, the method provided in the embodiment of the present application searches the preset event library according to the question information to obtain the target event, including:
[0087] Step S510: If the question information contains a tag, a search is performed in the preset event library based on the tag to obtain target number information corresponding to the target tag; the target number information is the number corresponding to the target image set;
[0088] Step S520, obtaining a video clip according to the target number information, and determining a target event.
[0089] In some embodiments, the method provided in the embodiments of the present application searches a preset event library according to the question information to obtain a target event, including:
[0090] According to the question information, based on the retrieval enhancement generation technology, a search is performed in the preset event library to obtain the target event.
[0091] In some embodiments, the method provided in the embodiments of the present application divides the video stream information into several image sets, including:
[0092] The video stream is segmented according to a preset time length, and video frames are extracted according to a preset frame rate to obtain an image set;
[0093] Add the numbering information of the current image set according to the start time, end time, camera code, and unique value code of the current image set.
[0094] Below, refer to Figure 2 , combined with specific application examples, the scheme of the embodiment of the present invention is described in detail and explained:
[0095] On the one hand, this application deeply embeds image recognition capabilities into the traditional video monitoring business through a multimodal large model. Make full use of cross-modal multidimensional data understanding capabilities. Improve video scene monitoring capabilities and be able to cope with the identification and detection of undefined sudden disaster events in complex and sudden scenes. On the other hand, combined with RAG technology, question-and-answer event scene retrieval can be achieved.
[0096] Reference Figure 2 As shown, this application includes: cloud model training, data set processing lora fine-tuning, training, Lora and model merging and saving end-side deployment. Specifically, see the following steps:
[0097] Step S21, video stream segmentation and numbering.
[0098] The video stream is divided into time length T, T (that is, the preset time length in this application) is less than or equal to 24h, and video frames are extracted at 15 frames / s and packaged into a group of image sets images, and the image sets are numbered according to the rules: start time + end time + camera code + unique value number. For example, images-202409111420-111520-vabc-123. Of course, the specific values of the above time lengths and preset frame rates are illustrative examples, and this application does not make specific limitations.
[0099] Step S22: initial recognition of the video image.
[0100] The event library SJ is preset, the images set is identified, and the images set is marked with event tags. If no corresponding tag is found in the event library SJ, step S23 is directly performed to recognize the video image.
[0101] For example, if the fire recognition finds that there is a fire in the images, the image set is labeled "fire". The recognition result (for example, images-202409111420-111520-vabc-123, "fire") is stored in a relational database.
[0102] Step S23: video image recognition.
[0103] Preset event library SJ for video image recognition. It is divided into the following two situations:
[0104] Case 1:
[0105] Feed the images set (i.e., the first image set) to the multimodal large model, edit the preset prompt word (i.e., the first prompt word), guide the model to output its understanding of the video, and generate a detailed description of the video. The images number, label, and cognitive results are stored in the vector database. For example: the system inputs a Figure 7 The image set extracted from the fire video shown (it can be understood that Figure 7 and Figure 8 is an illustrative example of a relevant scene. In practical applications, technicians in this field input a set of images of relevant real situations according to actual conditions. This application does not specifically limit the scenes shown in the image set), generate and input the following prompt words: ""msgs":[{"role":"user","content":"Classify the events in the video as neutral, negative or positive\nAnswer strictly in the following format. For example: [Video content]: A fight occurred; [Event type]: Fire; [Event classification]: Negative. \n[Video content]: XXXXXXXX (describe the video content in detail); [Label]: XXXX (only one can be selected from: a fire, b traffic congestion, c flooding, d personnel injury, e garbage accumulation, f cannot be defined); [Event classification]: XX (can only answer positive, negative or neutral). "." After the multimodal large model is processed, the system returns the following information: ""result":["[Video content]: A fire occurred, grass and trees beside the road burned, and no casualties occurred; [Label]: Fire; [Event classification]: Negative. "," takes "images-202409111420-111520-vabc-123" as the text number, and "result" returns the result and stores it in the vector database. It is understandable that only one can be selected from the following: a fire, b traffic congestion, c flooding, d injuries, e garbage accumulation, f cannot be defined, which is the range of optional tags in this application. In the above example, "[Video content]: XXXXXXXX (detailed description of video content); [Labels]: XXXX (only one can be selected from the following: a fire, b traffic congestion, c flooding, d injuries, e garbage accumulation, f cannot be defined); [Event classification]: XX (can only answer positive, negative or neutral)" is the content related to the second prompt word template in this application. This application does not limit the specific content and form of the prompt word template.
[0106] Case 2:
[0107] Feed the unlabeled images set to the multimodal large model, edit the preset prompt words, guide the model to output the understanding of the video, and generate a detailed description of the video. And store the image number, label and cognitive results together in the vector database. For example: suppose the image set has a road collapse such as Figure 8 As shown, the system does not have a corresponding label, and generates and enters the prompt word: ""msgs":[{"role":"user","content":"Classify the events in the video as neutral, negative or positive\nAnswer strictly in the following format. For example: [Video content]: A fight occurred; [Event type]: Fire; [Event classification]: Negative. \n[Video content]: XXXXXXXX (describe the video content in detail); [Label]: XXXX (please generate a four-character label based on the event content); [Event classification]: XX (can only answer positive, negative or neutral). "." The system returns information: ""result":["[Video content]: Road collapsed and traffic signs were damaged; [Label]: Road collapse (due to the nature of the event, it is not suitable to give a specific classification that is usually associated with other categories); [Event classification]: Negative. ","". Update the event library SJ and add the "road collapse" label. Similarly, "images-202409111420-111520-vabc-123" is used as the text number, and the result returned by "result" is stored in the vector database.
[0108] Step S24, search. Specifically, there are two ways:
[0109] Method 1: Precise search, perform precise search according to the event library SJ tag. For example: query "fire" returns the image set number images-202409111420-111520-vabc-123, and retrieve the corresponding video clip according to the number.
[0110] Method 2, question and answer retrieval, search the knowledge base according to the event library SJ tag, return the video content question and answer results, and the corresponding imges number to retrieve the corresponding video clip according to the number. For example: [{"role":"user","content":"Please help me check all the video clips of road collapse." The system returns ""result":["The following content may appear in the video that matches the road collapse, [Video content-1]: Road collapse occurred, and the traffic sign was damaged;..., [reference number]; images-202409111420-111520-vabc-123; [Video content-2]: Road collapse occurred, and the traffic sign was damaged;..., [reference number]; images-202409***; [Video content-3]: Road collapse occurred, and the traffic sign was damaged;..., [reference number]; images-202409***"
[0111] The pseudo code involved in this application is as follows:
[0112]
[0113]
[0114]
[0115]
[0116] This application introduces a multimodal model coupled with an NLP processing method to solve the problem that traditional AI image detection cannot recognize untrained scenes, and realizes the detection of sudden untrained scene targets. This application creatively proposes to use a multimodal large model to translate event-marked video streams into text information such as event labels and descriptions, convert them into matrices and store them in a vector library, and use RAG technology to solve the rapid retrieval of event fuzzy questions and answers in massive video streams and the automatic labeling of scenes that have not been trained by the model.
[0117] This application uses multimodal large-model image recognition capabilities to deeply embed into traditional video monitoring services. Make full use of cross-modal multidimensional data understanding capabilities. Improve video scene monitoring capabilities and be able to cope with the identification and detection of undefined sudden disaster events in complex and sudden scenes. The method provided in this application can be applied to application scenarios such as video monitoring and analysis in the field of the Internet of Things.
[0118] See also Fig. 9 The embodiment of the present application further provides an event retrieval system in a hybrid scenario, which can implement the event retrieval method in the hybrid scenario, and the system includes:
[0119] The first module 810 is used to receive video stream information in several scenes, and segment the video stream information to obtain several image sets;
[0120] The second module 820 is used to generate a first prompt word for a first image set according to a preset event library; the first image set is any image set in the image set; the preset event library includes description information and label information of the image set;
[0121] The third module 830 is used to input the first prompt word and the first image set into the multimodal large model, obtain a first cognitive result corresponding to the first image set, and store the first cognitive result in a preset event library; the first cognitive result includes description information and label information of the first image set;
[0122] The fourth module 840 is used to obtain question information about the target event to be retrieved;
[0123] The fifth module 850 is used to search in a preset event library according to the question information to obtain a target event.
[0124] In some embodiments, the system provided by the embodiments of the present application further includes a sixth module for:
[0125] If the first image set is not matched in the preset event library, determining attribute information corresponding to the label in the first prompt word template; the attribute information includes character length;
[0126] Alternatively, if the first image set is matched in the preset event library, a range of optional tags corresponding to tags in the second prompt word template is determined.
[0127] It can be understood that the contents of the above method embodiments are all applicable to the present system embodiments, the functions specifically implemented by the present system embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0128] The embodiment of the present application also provides an electronic device, the electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the event retrieval method in the above hybrid scenario when executing the computer program. The electronic device can be any smart terminal including a tablet computer, a car computer, etc.
[0129] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0130] See also Fig.10 , Fig.10The hardware structure of an electronic device of another embodiment is illustrated, and the electronic device includes:
[0131] The processor 901 may be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;
[0132] The memory 902 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store an operating system and other applications. When the technical solution provided in the embodiment of this specification is implemented by software or firmware, the relevant program code is stored in the memory 902, and the processor 901 calls and executes the event retrieval method in the hybrid scenario of the embodiment of this application;
[0133] Input / output interface 903, used to implement information input and output;
[0134] Communication interface 904, used to realize communication interaction between the device and other devices, which can be realized by wired mode (such as USB, network cable, etc.) or wireless mode (such as mobile network, WIFI, Bluetooth, etc.);
[0135] A bus 905 that transmits information between the various components of the device (e.g., the processor 901, the memory 902, the input / output interface 903, and the communication interface 904);
[0136] The processor 901 , the memory 902 , the input / output interface 903 and the communication interface 904 are connected to each other in communication within the device via a bus 905 .
[0137] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the event retrieval method in the above-mentioned hybrid scenario.
[0138] It can be understood that the contents of the above method embodiments are all applicable to the present storage medium embodiments, the functions specifically implemented by the present storage medium embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0139] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0140] The embodiments described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0141] Those skilled in the art will appreciate that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.
[0142] The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0143] Those skilled in the art will appreciate that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices may be implemented as software, firmware, hardware, or a suitable combination thereof.
[0144] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0145] It should be understood that in the present application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0146] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the above units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0147] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0148] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0149] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including multiple instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, referred to as ROM), random access memory (Random Access Memory, referred to as RAM), disk or optical disk and other media that can store programs.
[0150] The preferred embodiments of the present application are described above with reference to the accompanying drawings, but the scope of the rights of the present application is not limited thereto. Any modification, equivalent substitution and improvement made by a person skilled in the art without departing from the scope and essence of the present application should be within the scope of the rights of the present application.
Claims
1. An event retrieval method in a mixed scenario, characterized in that: The method comprises: Receiving video stream information in a plurality of scenes, and segmenting the video stream information to obtain a plurality of image sets; According to the preset event library, a first prompt word of a first image set is generated; the first image set is any image set in the image set; the preset event library includes description information and label information of the image set; Inputting the first prompt word and the first image set into the multimodal large model to obtain a first cognitive result corresponding to the first image set, and storing the first cognitive result in the preset event library; the first cognitive result includes description information and label information of the first image set; Obtaining query information for the target event to be retrieved; According to the question information, a search is performed in the preset event library to obtain a target event.
2. The method according to claim 1, characterized in that The step of generating a first prompt word for the first image set according to a preset event library includes: If the first image set is not matched in the preset event library, generating a first prompt word for the first image set according to a first prompt word template; Alternatively, if the first image set is matched in the preset event library, a first prompt word of the first image set is generated according to a second prompt word template; and the label parts of the first prompt word template and the second prompt word template are different.
3. The method according to claim 2, characterized in that The prompt word template is determined by the following steps: If the first image set is not matched in the preset event library, determining attribute information corresponding to the label in the first prompt word template; the attribute information includes character length; Alternatively, if the first image set is matched in the preset event library, a range of optional tags corresponding to tags in the second prompt word template is determined.
4. The method according to claim 1, characterized in that: The storing the first recognition result into the preset event library includes: storing the label information, the description information and the numbering information corresponding to the first image set in a vector database; The label information and the serial number information corresponding to the first image set are stored in a relational database.
5. The method according to claim 1, characterized in that The step of searching the preset event library according to the question information to obtain the target event includes: If the question information includes a tag, a search is performed in the preset event library based on the tag to obtain target number information corresponding to the target tag; the target number information is the number corresponding to the target image set; According to the target number information, a video clip is obtained to determine the target event.
6. The method according to claim 1, characterized in that The step of searching the preset event library according to the question information to obtain the target event includes: According to the question information, based on the search enhancement generation technology, a search is performed in the preset event library to obtain a target event.
7. The method according to claim 1, characterized in that The video stream information is segmented to obtain a plurality of image sets, including: The video stream is segmented according to a preset time length, and video frames are extracted according to a preset frame rate to obtain an image set; According to the start time, end time, camera code, and unique value code of the current image set, the numbering information of the current image set is added.
8. An event retrieval system in a mixed scenario, characterized in that: The system comprises: The first module is used to receive video stream information in several scenes, and segment the video stream information to obtain several image sets; The second module is used to generate a first prompt word for a first image set according to a preset event library; the first image set is any image set in the image set; the preset event library includes description information and label information of the image set; A third module is used to input the first prompt word and the first image set into the multimodal large model to obtain a first cognitive result corresponding to the first image set, and store the first cognitive result in the preset event library; the first cognitive result includes description information and label information of the first image set; The fourth module is used to obtain question information about the target event to be retrieved; The fifth module is used to search the preset event library according to the question information to obtain the target event.
9. An electronic device, characterized in that: The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.