Multi-modal retrieval method and system for monitoring video, electronic equipment and storage medium

Through multimodal retrieval methods, industry graphs and knowledge vector databases are used to generate target retrieval reports for surveillance videos, which solves the problem that existing systems are difficult to meet users' precise search needs, and realizes the accurate fusion of images and text in surveillance videos and the intelligent presentation of information.

CN120632155APending Publication Date: 2025-09-12CHINA UNICOM (GUANGDONG) IND INTERNET CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510702176.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing surveillance video retrieval systems are unable to meet users' precise search needs. Traditional methods are inefficient and prone to missing key information.

Method used

Through multimodal retrieval methods, industry graphs and knowledge vector databases are used to generate target retrieval reports, combining search text and target images to achieve accurate fusion of image and text descriptions.

Benefits of technology

The accuracy and efficiency of surveillance video retrieval are improved. The generated retrieval report contains detailed structured information and supports intelligent fusion and interactive presentation of multimodal data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120632155A_ABST
    Figure CN120632155A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a multi-mode retrieval method and system for a monitoring video, electronic equipment and a storage medium. The method comprises the steps that a search text which is input by a user and aims at the monitoring video is obtained; obtaining a target image from the monitoring video according to the search text; performing industry classification on the search text, determining a derivative relationship corresponding to the search text according to an industry map corresponding to a target industry type to which the search text belongs, and generating a target text description of a target retrieval report based on the derivative relationship; the target retrieval report is generated on the basis of the target text description and the target image, and in the technical scheme, the target text description is generated through the derivative relation corresponding to the search text, the target image is obtained from the monitoring video according to the search text, and the more accurate retrieval report combining the image and the text description is generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of video retrieval, and in particular to a multimodal retrieval method, system, electronic device, and storage medium for surveillance videos. Background Art

[0002] With the widespread adoption of video surveillance systems, the storage and analysis of massive amounts of surveillance video data have become a critical requirement. Traditional methods often rely on manual review of video clips to locate specific objects or events, which is inefficient and prone to missing key information. In recent years, technologies based on computer vision and natural language processing have been applied to video retrieval, such as extracting key frames or events through object detection, behavior recognition, or semantic analysis. However, most systems only produce simple retrieval results, which are difficult to meet users' needs for precise searches. Therefore, improving the accuracy of retrieval results is an urgent problem that needs to be solved. Summary of the Invention

[0003] The embodiments of the present application provide a multimodal retrieval method, system, electronic device, and storage medium for surveillance videos, which can improve the accuracy of retrieving corresponding target images and searching for text-related content in surveillance videos.

[0004] The first embodiment of the present application provides a multimodal retrieval method for surveillance videos, including:

[0005] Obtaining the search text for surveillance video input by the user;

[0006] Obtaining a target image from the surveillance video according to the search text;

[0007] Classifying the search text by industry, determining a derivative relationship corresponding to the search text according to an industry map corresponding to the target industry type to which the search text belongs, and generating a target text description of a target retrieval report based on the derivative relationship;

[0008] The target retrieval report is generated based on the target text description and the target image.

[0009] In some possible embodiments, determining a derivative relationship corresponding to the search text according to an industry graph corresponding to a target industry type to which the search text belongs includes:

[0010] Obtaining a search vector corresponding to the search text;

[0011] Obtaining a plurality of keywords corresponding to a plurality of related vectors related to the search vector in a pre-stored knowledge vector database, wherein the plurality of keywords are used to characterize the target industry type;

[0012] Mapping the multiple keywords to corresponding multiple nodes of the industry graph;

[0013] Derivative relationships corresponding to the search text are determined based on the multiple nodes.

[0014] In some possible embodiments, determining, based on the multiple nodes, a derivative relationship corresponding to the search text includes:

[0015] Obtaining semantic weights of the multiple nodes;

[0016] According to the semantic weights of the plurality of nodes and a first semantic weight threshold, screening out at least one first target node having a semantic weight greater than or equal to the first semantic weight threshold;

[0017] For each first target node, the maximum number of hops corresponding to the first target node is determined according to the semantic weight of the first target node. Starting from the first target node, the industry graph is traversed according to the maximum number of hops to determine the derivative relationship corresponding to the search text, wherein the maximum number of hops is positively correlated with the semantic weight, and the maximum number of hops is used to indicate the maximum number of association steps when performing multi-hop reasoning in the industry graph starting from the first target node.

[0018] In some possible embodiments, the target retrieval report is a retrieval report that meets preset conditions, and the generating of a target text description of the target retrieval report based on the derivative relationship includes:

[0019] generating a first text description of a target retrieval report based on the derived relationship;

[0020] If the first text description satisfies the preset condition, outputting the first text description as the target text description;

[0021] If the first text description does not meet the preset condition, an Internet search is performed to generate a second text description, and if the second text description meets the preset condition, the second text description is output as the target text description.

[0022] In some possible embodiments, the method further includes:

[0023] If the second text description does not meet the preset condition, the sub-target node corresponding to the maximum number of hops in the first text description is used as the second target node;

[0024] For each second target node, the maximum number of hops corresponding to the second target node is determined according to the semantic weight of the second target node. Starting from the second target node, the industry graph is traversed according to the maximum number of hops to determine the derivative relationship corresponding to the search text, wherein the maximum number of hops is positively correlated with the semantic weight, and the maximum number of hops is used to indicate the maximum number of association steps when performing multi-hop reasoning in the industry graph starting from the second target node.

[0025] In some possible embodiments, obtaining a target image from the surveillance video according to the search text includes:

[0026] Obtaining key frame images in the surveillance video;

[0027] Performing image preprocessing on the key frame image to obtain an image vector corresponding to the key frame image, and storing the image vector in an image vector database;

[0028] Performing a preprocessing operation on the search text to obtain a search vector corresponding to the search text;

[0029] The target image is obtained from the surveillance video according to the similarity between the search vector and the image vectors in the image vector database.

[0030] In some possible embodiments, the image vector includes a global vector and a local vector, and obtaining the target image from the surveillance video based on the similarity between the search vector and the image vectors in the image vector database includes:

[0031] Obtaining local similarity according to the search vector and the local vector;

[0032] Obtaining a global similarity based on the search vector and the global vector;

[0033] The similarity between the search vector and the image vector is obtained by summing the product of the local similarity, the local feature weight corresponding to the local similarity, and the global feature weight corresponding to the global similarity, wherein the sum of the local feature weight and the global feature weight is a preset value, and the target image is an image whose similarity is greater than a preset similarity threshold.

[0034] The second embodiment of the present application provides a multimodal retrieval system for surveillance videos, including:

[0035] An acquisition module is configured to obtain a search text input by a user for a surveillance video; and obtain a target image from the surveillance video according to the search text;

[0036] a processing module configured to classify the search text by industry, and determine a derivative relationship corresponding to the search text based on an industry map corresponding to a target industry type to which the search text belongs, wherein the derivative relationship is used to generate a text description of a target retrieval report;

[0037] A generating module is used to generate the target retrieval report based on the text description and the target image.

[0038] The third aspect embodiment of the present application proposes an electronic device, including a processor and a memory, wherein the memory stores a computer program, and when the processor executes the computer program, it implements the steps of any one of the methods proposed in the first aspect embodiment of the present application.

[0039] The fourth aspect embodiment of the present application proposes a computer-readable storage medium storing computer instructions, which, when executed by a processor, implement the steps of any one of the methods proposed in the first aspect embodiment of the present application.

[0040] The technical solutions provided by the embodiments of the present application include at least the following beneficial effects:

[0041] The embodiments of the present application propose a multimodal retrieval method, system, electronic device and storage medium for surveillance videos, the method comprising: obtaining a search text for a surveillance video input by a user; obtaining a target image from the surveillance video based on the search text; classifying the search text by industry, determining a derivative relationship corresponding to the search text based on an industry map corresponding to a target industry type to which the search text belongs, and generating a target text description of a target retrieval report based on the derivative relationship; generating the target retrieval report based on the target text description and the target image. In the above technical solution, a more accurate retrieval report combining image and text description is generated by generating a target text description through the derivative relationship corresponding to the search text and obtaining a target image from the surveillance video based on the search text. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 A schematic diagram of an application scenario proposed in an embodiment of the present application;

[0043] Figure 2 A flowchart of a multimodal retrieval method for surveillance videos proposed in an embodiment of the present application;

[0044] Figure 3 A flowchart of another multimodal retrieval method for surveillance videos proposed in an embodiment of the present application;

[0045] Figure 4 A flowchart of another multimodal retrieval method for surveillance videos proposed in an embodiment of the present application;

[0046] Figure 5 A flowchart of another multimodal retrieval method for surveillance videos proposed in an embodiment of the present application;

[0047] Figure 6 A flowchart of another multimodal retrieval method for surveillance videos proposed in an embodiment of the present application;

[0048] Figure 7 A flowchart of another multimodal retrieval method for surveillance videos proposed in an embodiment of the present application;

[0049] Figure 8 A flowchart of another multimodal retrieval method for surveillance videos proposed in an embodiment of the present application;

[0050] Figure 9 A flowchart of another multimodal retrieval method for surveillance videos proposed in an embodiment of the present application;

[0051] Figure 10 A schematic diagram of the structure of a multimodal retrieval system for surveillance videos proposed in an embodiment of the present application;

[0052] Figure 11 This is a schematic diagram of the structure of the electronic device proposed in an embodiment of the present application. DETAILED DESCRIPTION

[0053] The technical solutions in the embodiments of the present application will be described below with reference to the accompanying drawings.

[0054] In order to facilitate the clear description of the technical solutions of the embodiments of the present application, in the embodiments of the present application, words such as "first" and "second" are used to distinguish between identical or similar items with substantially the same functions and effects. For example, the first instruction and the second instruction are intended to distinguish different user instructions and do not limit their order. Those skilled in the art will understand that words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not necessarily limit them to be different.

[0055] It should be noted that in the embodiments of this application, words such as "exemplarily" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplarily" or "for example" in the embodiments of this application should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplarily" or "for example" is intended to present the relevant concepts in a concrete manner.

[0056] In addition, "at least one" means one or more, and "more" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b and c can mean: a, or b, or c, or a and b, or a and c, or b and c, or a, b and c, where a, b, c can be single or multiple.

[0057] It should be noted that, in the embodiments of the present application, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or also includes elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0058] With the widespread adoption of video surveillance systems, the storage and analysis of massive amounts of surveillance video data has become a critical requirement. Traditional methods often rely on manual review of video clips to locate specific objects or events, which is inefficient and prone to missing critical information. In recent years, technologies based on computer vision and natural language processing have been applied to video retrieval, such as extracting key frames or events through object detection, behavior recognition, or semantic analysis. However, most systems only produce simple retrieval results, which are difficult to meet users' needs for precise searches. Therefore, improving the accuracy of retrieval results has become a key technical challenge that needs to be addressed.

[0059] In view of this, an embodiment of the present application proposes a multimodal retrieval method, system, electronic device and storage medium for surveillance videos, the method comprising: obtaining a search text for a surveillance video input by a user; obtaining a target image from the surveillance video based on the search text; classifying the search text by industry, determining a derivative relationship corresponding to the search text based on an industry map corresponding to the target industry type to which the search text belongs, and generating a target text description of a target retrieval report based on the derivative relationship; generating the target retrieval report based on the target text description and the target image. In the above technical solution, a more accurate retrieval report combining image and text description is generated by generating a target text description through the derivative relationship corresponding to the search text and the target image obtained from the surveillance video based on the search text.

[0060] The following is a detailed explanation of an application scenario example of the embodiment of the present application. For example, Figure 1 As shown, it includes: a terminal device 101 and a server 102.

[0061] In the application scenario of the embodiment of the present application, the terminal device and the server work together through an encrypted network, where the terminal device 101 (such as a smart camera equipped with an AI chip, a 5G law enforcement recorder or an edge computing box) is responsible for real-time video acquisition and receives user retrieval requests through voice / touch interaction.

[0062] For example, server 102 usually adopts a high-performance cloud computing platform. These servers process the search requests uploaded by the terminal through multimodal models such as CLIP, use distributed vector databases (such as Milvus) to quickly match target images, and combine industry maps to generate structured reports containing information such as spatiotemporal trajectories and related events, and finally push the search results to the terminal device for display in real time.

[0063] For example, in another possible scenario, the processing of the above-mentioned terminal device and the server can be processed on the terminal device, which is not limited in this embodiment of the present application.

[0064] For example, the method proposed in the embodiments of the present application is described below by taking execution on a terminal device as an example.

[0065] After understanding the application scenario of the embodiment of the present application, we will now introduce in detail the execution steps of a multimodal retrieval method for surveillance videos proposed in the embodiment of the present application in the system, such as Figure 2 As shown:

[0066] Step 201: Obtain the search text for surveillance video input by the user.

[0067] Exemplarily, obtaining the search text for surveillance video input by the user means that the system obtains the query request described by the user in natural language or in a structured manner through a human-computer interaction interface, aiming to locate a specific target or event from the surveillance video.

[0068] For example, in terms of input form, users can enter text queries through the keyboard or voice (such as "suspicious person wearing a red jacket and carrying a black bag"), or select structured query options provided by the system (such as a time + location + feature combination), or even use a mixed input method (such as text combined with sketch annotations to specify "person wearing a hat in the circled area"). The input form is not limited in the embodiments of the present application.

[0069] Step 202: Obtain a target image from the surveillance video according to the search text.

[0070] Exemplarily, the surveillance video may be acquired through a network camera and stored in an image database.

[0071] For example, based on the similarity between the search text and each video frame image in the surveillance video, an image with a similarity greater than a preset similarity threshold may be obtained as a target image.

[0072] In order to effectively reduce redundant information and resource consumption, for example, dynamic sampling can be used to obtain key frame images with smaller picture changes, calculate the similarity between the search text and the continuous frame images, and obtain an image with a similarity greater than a preset similarity threshold as the target image.

[0073] It should be understood that the similarity threshold can be set and adjusted according to needs.

[0074] Step 203: Classify the search text by industry, and determine the derivative relationship corresponding to the search text based on the industry map corresponding to the target industry type to which the search text belongs.

[0075] For example, a multi-label text classification model may be used to identify the industry type of the search text.

[0076] For example, an input query of finding workers in a factory without helmets can be categorized as industrial safety. For example, an input query of counting customers who stay at a shopping mall entrance for more than 10 minutes can be categorized as retail analysis. For example, a wider range of industry types can be obtained from a pre-stored knowledge vector database.

[0077] For example, in an industry graph, derived relationships refer to implicit associations automatically inferred from the user's original search text based on industry-specific logic and business rules. These relationships not only include direct literal matches but also tap into industry knowledge to uncover relevant content that users may not have explicitly expressed but actually need.

[0078] Suppose a user inputs "Find workers on a construction site who are not wearing hard hats." For example, a derived relationship might be "Associated with industry standards" to obtain "Safety protection equipment wearing standards." Another example derived relationship might be "Associated with penalty procedures" to obtain "Automatically mark violations and trigger rectification notices."

[0079] Step 204: Generate a target text description of the target retrieval report based on the derivative relationship.

[0080] In the embodiment of the present application, the derivative relationship generates a target text description of the search text input by the user, and the target text should meet the preset conditions of the target retrieval report.

[0081] In this embodiment of the present application, step 204 generates a target text description for the target search report by analyzing derivative relationships in the industry knowledge graph. These derivative relationships include multi-dimensional information chains such as factual associations, causal associations, and dispositional associations. Structured information is extracted from the derivative relationships, including key elements such as the violation facts, the basis clauses, and the disposition recommendations. Fields such as time, location, and responsible party are annotated through named entity recognition. A template with pre-set conditions that is suitable for the target industry type is selected for content organization. For example, a large language model can also be used to convert the structured data into a natural language description that conforms to industry standards.

[0082] Exemplarily, the generated text description needs to meet preset conditions such as completeness, accuracy, operability and readability, and be verified through technical means such as rule engines, industry graph version comparison and semantic evaluation. Exemplarily, if the text description generated for the first time fails the verification, an exception handling mechanism will be triggered. Exemplarily, the exception handling mechanism includes supplementing missing information by expanding the number of derivative relationship hops of the knowledge graph, expanding the target type (the semantic weight of the node in the industry graph), using the Internet to search and update expired terms, or re-optimizing the semantic expression in combination with visual features.

[0083] Step 205: Generate the target retrieval report based on the target text description and the target image.

[0084] For example, a template for a target retrieval report with preset conditions is selected according to the industry type (such as security / retail / medical), and then natural language polishing is performed through a fine-tuned large language model (such as Llama3-8B) to ensure a balance between professionalism and readability.

[0085] Exemplarily, the output target retrieval report includes interactive elements.

[0086] For example, clicking on an image reveals the original video clip, while a floating timeline displays a heat map of target movement. Key conclusions are accompanied by traceability links pointing to the analytical basis. The report utilizes a structured Markdown format, supports automatic export to various office documents such as PDF and PPT, and can be connected to the enterprise work order system via an API to directly generate processing tasks. By combining the professional standardization of structured templates with the semantic understanding capabilities of LLM, the report ensures that the content complies with industry standards while achieving intelligent fusion and interactive presentation of multimodal data, significantly improving the accuracy of decision-making information.

[0087] The following describes in detail how to obtain the derivative relationship corresponding to the search text, thereby obtaining the target text description of the target retrieval report.

[0088] Exemplarily, the terminal device determines the derivative relationship corresponding to the search text according to the industry map corresponding to the target industry type to which the search text belongs, such as Figure 3 As shown, the following steps are included:

[0089] Step 301: Obtain a search vector corresponding to the search text.

[0090] Exemplarily, the search vector of the search text may be obtained through a CLIP text encoder or through Sentence-BERT.

[0091] For example, the CLIP text encoder, as a multimodal model, has the advantage of mapping text and images into a unified vector space, which is particularly suitable for cross-modal retrieval scenarios of images and texts, and can effectively capture visual-related semantic features. Sentence-BERT, as a pure text encoding model, optimizes sentence-level representations through a twin network structure, performs well in semantic similarity calculations, and is more suitable for text retrieval tasks that require fine semantic matching. It should be understood that the embodiments of the present application do not limit the method of obtaining the search vector for the search text, and can be selected or used in combination according to the needs of the specific application scenario.

[0092] Step 302: Obtain multiple keywords corresponding to multiple related vectors related to the search vector in a pre-stored knowledge vector database.

[0093] In the embodiment of the present application, the multiple keywords are used to characterize the target industry type.

[0094] Exemplarily, the knowledge vector database is a structured data storage system based on artificial intelligence technology, which stores and manages industry knowledge, entity relationships and semantic concepts through vectorized encoding. Exemplarily, in this knowledge base, each industry category (such as industrial safety) and its subcategories (such as construction safety, factory inspection) are represented as feature vectors in a high-dimensional space. These vectors not only contain classification labels, but also embed rich semantic association information. When the user enters the search text, the system can convert it into a query vector, and then perform a nearest neighbor search in the knowledge vector database to discover potentially related industry types and their derived concepts, even if these concepts are not explicitly mentioned in the original query. For example, a search for "failure to wear safety equipment on construction sites" may automatically be associated with the subcategory of personal protective equipment compliance inspection under construction safety, and further expanded to relevant local safety regulations.

[0095] Based on the search vector generated in step 301, an approximate search (such as using FAISS or Milvus) is performed in the knowledge vector database to quickly find multiple industry knowledge vectors that are most relevant thereto.

[0096] For example, the search vector input of "using mobile phones at gas stations" may be associated with knowledge vectors such as "ban on electronic equipment in flammable areas", "electrostatic protection regulations" and "safety inspection procedures", each of which corresponds to a keyword (such as "open flame risk" or "explosion-proof regulations").

[0097] For example, keywords are essentially standardized terms extracted from a knowledge vector database to accurately characterize the target industry type.

[0098] For example, the retrieval process may use cosine similarity to retain only highly relevant results whose similarity exceeds a similarity threshold (eg, 0.75).

[0099] Compared with traditional keyword matching, this vector-based knowledge representation and retrieval method can more flexibly capture the deep semantic connections between industry knowledge and significantly improve the recall and accuracy of domain retrieval.

[0100] Step 303: Map the multiple keywords to the corresponding multiple nodes of the industry graph.

[0101] For example, an industry graph is a domain knowledge base organized in a graph structure, where nodes represent entities (such as "safety helmet", "gas station") or concepts (such as "violation", "protection standard"), and edges represent derived relationships (such as "belongs to", "need to comply").

[0102] The keywords obtained in step 302 are used as query anchors to locate the corresponding nodes in the industry graph. For example, the keyword "open flame risk" might be mapped to the [High-risk Operations] node in the graph, while "explosion-proof regulations" might be mapped to the [Industry Standards] node. This ensures that even if a keyword is ambiguous, it can be accurately linked to the correct node in the industry context.

[0103] Step 304: Determine a derivative relationship corresponding to the search text based on the multiple nodes.

[0104] Exemplarily, the derivative relationship corresponding to the search text is determined based on the multiple nodes, such as Figure 4 As shown, the following steps are included:

[0105] Step 401: Obtain the semantic weights of the multiple nodes.

[0106] The core of this step is to quantify the importance of multiple nodes mapped in the industry map.

[0107] For example, the semantic weight is calculated by comprehensively considering the following factors: industry relevance and query relevance. Industry relevance refers to the degree of match with the target industry type (e.g., "industrial safety"). Query relevance refers to the cosine similarity score with the search text vector.

[0108] For example, for the search text "no safety helmet on construction site", the [Safety Protection Equipment] node may obtain a high weight of 0.92 (because it is directly related to the query topic), while the [Construction Permit] node only obtains 0.35 (indirectly related).

[0109] For example, the weight calculation may adopt a weighted formula:

[0110] Semantic weight = 0.7 × query relevance + 0.3 × industry relevance, (1)

[0111] Step 402: Based on the semantic weights of the multiple nodes and a first semantic weight threshold, screen out at least one first target node whose semantic weight is greater than or equal to the first semantic weight threshold.

[0112] During the analysis of the industry knowledge graph, semantic weight is used to measure the importance of each node (i.e., entity or concept) in the current business scenario. The core goal of step 402 is to filter out key nodes based on a preset first semantic weight threshold and eliminate irrelevant or secondary nodes to improve the accuracy and efficiency of subsequent analysis.

[0113] Exemplarily, screening is performed according to a preset first semantic weight threshold (such as 0.7): nodes with weights greater than or equal to the threshold are retained (such as [safety helmet], [violation record], [safety training]), and low-weight nodes (such as [weather conditions], [construction progress]) are eliminated.

[0114] For example, the first semantic weight threshold can be dynamically adjusted: in high-risk industries (such as the chemical industry), it can be set to 0.8 to improve accuracy, and in common scenarios, it can be set to 0.6 to expand coverage. The filtered first target nodes form the starting point set for derivative relationship reasoning, ensuring that subsequent calculations focus on the core semantics.

[0115] Step 403: For each first target node, determine the maximum number of hops corresponding to the first target node according to the semantic weight of the first target node.

[0116] The maximum hop count is positively correlated with the semantic weight. For example, the maximum hop count is used in the industry graph to indicate the graph traversal depth, that is, to indicate the maximum number of associated steps when performing multi-hop reasoning in the industry graph starting from the first target node.

[0117] Exemplarily, for each first target node, the maximum number of hops is dynamically allocated according to its semantic weight:

[0118] (1) Semantic weight ≥ 0.9: 3 hops are allowed (mining deep derivative relationships).

[0119] (2) Weight ∈ [0.7, 0.9): 2 hops are allowed (balancing breadth and depth).

[0120] (3) Weight < 0.7: only 1 hop (limited divergence).

[0121] This method avoids over-expansion of secondary nodes while ensuring sufficient relationship mining of key nodes.

[0122] Step 404: Starting from the first target node, traverse the industry graph according to the maximum number of hops to determine a derivative relationship corresponding to the search text.

[0123] For example, when the semantic weight is ≥ 0.9 and 3 hops are allowed, for example: starting from [safety helmet], one can traverse to [personal protection → violation penalties → legal provisions].

[0124] In the above embodiment, a derivative relationship corresponding to the search text is determined. The following describes a detailed process of generating a target text description of a target retrieval report based on the derivative relationship.

[0125] Exemplarily, the target search report is a search report that meets preset conditions, and the target text description of the target search report generated based on the derivative relationship is as follows: Figure 5 As shown, the following steps are included:

[0126] Step 501: Generate a first text description of the target retrieval report based on the derivative relationship.

[0127] In this step, a first text description is generated based on the derived relationships obtained above. This process uses the pre-defined conditions of the target search report: key elements (such as violation clauses, disposal recommendations, and related regulations) are parsed from the derived relationships. Predefined report templates are used (for example, the security incident template includes: violation facts, disposal basis, and corrective recommendations).

[0128] Step 502: If the first text description meets the preset condition, output the first text description as the target text description.

[0129] Exemplarily, the pre-set conditions include the following three conditions: Condition 1: Completeness check, including the three elements of violation facts, basis clauses, and disposal recommendations. Condition 2: Confidence threshold, with the total weight of the derivative relationship ≥ 0.85 (to avoid low-confidence conclusions). Condition 3: Industry compliance, including compliance with industry standards.

[0130] Determine whether the first text description meets the above preset conditions.

[0131] Exemplarily, the setting methods of preset conditions can be divided into the following three situations: Exemplarily, users can define preset conditions by themselves, or choose from multiple preset options provided by the system; if the user does not actively set or select, the system will automatically use the default custom conditions of the target retrieval report as preset conditions.

[0132] If all conditions are met, the target text description is directly output and marked as "Verified Report".

[0133] If any of the preset conditions is not met, the process proceeds to step 503 .

[0134] Step 503: If the first text description does not meet the preset condition, perform an Internet search to generate a second text description, and if the second text description meets the preset condition, output the second text description as the target text description.

[0135] For example, when the first text description fails to pass the verification (such as failing to meet the requirement of completeness of condition 1 in the preset conditions, missing reference to the clause), an Internet search is performed.

[0136] For example, in an embodiment of the present application, it is also possible to determine whether an Internet search is needed by reasoning with a large model. A large reasoning model generally refers to the logical reasoning and decision-making capabilities based on a large-scale pre-trained language model. Its core goal is to enable the model to make complex logical judgments or dynamic decisions based on input information. For example, based on the first text description, it is determined whether the preset conditions are met and whether external data (such as an Internet search) is needed to assist in generating more accurate output.

[0137] For example, the missing content type (such as "penalty standards revised in 2024") can be obtained according to a preset condition. According to the type of the actual content, the missing text description content is obtained through a targeted network search.

[0138] When the second text meets the preset conditions, the target text description is the second text description.

[0139] In the case that the second text description does not meet the preset conditions, in order to generate a target text description of a target retrieval report that meets the preset conditions, more derivative relationships can be added based on the first text description to further obtain a target text description that meets the preset conditions, such as Figure 6 As shown, the method further includes the following steps:

[0140] Step 601: If the second text description does not meet the preset condition, the sub-target node corresponding to the maximum number of hops in the first text description is used as the second target node.

[0141] For example, the expansion is performed from the deepest graph node in the first text description. For example, if the semantic weight threshold of the second target node is 0.5 and the first text description is A->B->C, then the second text description is C->D->E (the weights of D and E are both greater than 0.5).

[0142] Step 602: For each second target node, determine the maximum number of hops corresponding to the second target node according to the semantic weight of the second target node.

[0143] According to the semantic weight of the second target node, the maximum number of hops corresponding to each second target node is determined. The greater the semantic weight of the second target node, the greater the corresponding maximum number of hops.

[0144] Step 603: Starting from the second target node, traverse the industry graph according to the maximum number of hops to determine a derivative relationship corresponding to the search text.

[0145] Among them, the maximum number of hops is positively correlated with the semantic weight, and the maximum number of hops is used to indicate the maximum number of associated steps when performing multi-hop reasoning in the industry graph starting from the second target node.

[0146] In this step, the depth of the derived relationship obtained from the second target node is greater than the range of the derived relationship obtained from the first target node, until the preset conditions are met and a target search report meeting the preset conditions is generated.

[0147] For example, the example process of obtaining the target search report in the embodiment of the present application is as follows: Figure 7 As shown, the following steps are included:

[0148] Step 701: Obtain the search text input by the user, Step 702: Obtain the search vector corresponding to the search text, Step 703: Perform vector retrieval in the knowledge vector database to obtain the target industry type, Step 704: Determine the derivative relationship corresponding to the search text based on the industry map corresponding to the target industry type to which the search text belongs, Step 705: Generate a target text description of the target retrieval report based on the derivative relationship, Step 706: Determine whether the preset conditions are met, Step 707: Verification is completed, and the first text description is output as the target text description, Step 708: Internet search, Step 709: Obtain the second text description, Step 710: Determine whether the preset conditions are met, Step 711: Verification is completed, and the second text description is output as the target text description, Step 712: According to the search text, obtain the target image from the surveillance video, Step 713: Generate a target retrieval report.

[0149] In the technical solution of the embodiment of the present application, steps 701 to 705 are how to generate a target text description of a target retrieval report, which has been explained in detail above and will not be repeated here.

[0150] After the target description is generated according to steps 701-702, in order to improve the comprehensiveness and accuracy of the target text description, the first text description of the target retrieval report is generated in step 705; step 706 is used to determine whether the description meets the preset conditions (such as completeness, accuracy, etc.). If the verification is completed, the first text description is directly output as the final target text description; if the conditions are not met, the system starts the Internet search process in step 708, generates a second text description after obtaining supplementary information in step 709, and performs condition verification again in step 710. If the preset conditions are not met, step 704 is executed. If the preset conditions are met, the optimized second text description is output through step 711; at the same time, in step 712, the target image that meets the conditions is retrieved from the surveillance video according to the search text, and finally in step 713, all information is integrated to generate a structured target retrieval report containing text description and target image.

[0151] The above embodiment explains in detail how to obtain the target text description content corresponding to the search text. In order to obtain the target image corresponding to the user's search text from the surveillance video and obtain a clearer retrieval report, the following describes in detail how to obtain the target image from the surveillance video based on the search text.

[0152] Exemplarily, according to the search text, a target image is obtained from the surveillance video, such as Figure 8 As shown, the following steps are included:

[0153] Step 801: Obtain key frame images in the surveillance video.

[0154] Exemplarily, obtaining the key frame image in the surveillance video includes: converting each color image frame of the surveillance video into a grayscale image, calculating the absolute difference pixel by pixel for the grayscale images of two consecutive frames, and generating a difference map; obtaining the total change between frames based on the difference map; obtaining the key frame image in the surveillance video based on the total change between frames and the preset frame change threshold, the key frame image being the image corresponding to the grayscale image whose total change between frames is greater than or equal to the preset frame change threshold.

[0155] For example, let the difference map be D(x,y), the image size be H×W, and the calculation formula of the total change S between frames be as follows:

[0156]

[0157] Step 802: performing image preprocessing on the key frame image to obtain an image vector corresponding to the key frame image, and storing the image vector in an image vector database.

[0158] Exemplarily, the key frame image is processed by a target detection algorithm to extract and crop a local target area in the image.

[0159] For example, the object detection algorithm can be from the YOLO family, such as YOLOv1, which treats detection as a single regression problem to achieve real-time detection. For example, YOLOv3 introduces multi-scale prediction and a Darknet-53 backbone network to balance speed and accuracy. For example, YOLOv5 / v8 utilize the PyTorch framework to simplify deployment and support custom training.

[0160] For example, the target detection algorithm can also be a Transformer-based end-to-end target detection network (DETR). The core idea of ​​DETR is to model the target detection task as a collective prediction problem and implement global modeling through the Transformer architecture. It should be understood that the embodiments of this application do not limit the target detection algorithm and can be flexibly selected according to the usage scenario and requirements.

[0161] Subsequently, the CLIP image encoder can be used to vectorize the entire image and the local target area respectively to obtain the global vector and local vector of each image.

[0162] For example, for each image:

[0163] The global vector (entire image embedding) is represented as For example, the local target vector is: m local targets are obtained by target detection in the image: O = {o1, o2, ..., o m}.

[0164] Among them, the vector of each target is The vectors of each target are combined into a local vector matrix:

[0165] Exemplarily, the image vector is stored in an image vector database, and the storage structure of the image database is as follows:

[0166] Table 1

[0167]

[0168] The image vector database shown in Table 1 above stores the ID of each image, the corresponding global vector and local vector, and the local target area, that is, the target category and target confidence of the local vector, etc., which are used to weight the local target area.

[0169] For example, the global vector and the local target vector may be stored separately and support separate retrieval and indexing.

[0170] Step 803: Preprocess the search text to obtain a search vector corresponding to the search text.

[0171] The user's search text is first atomized and preprocessed before being fed into the CLIP text encoder for vectorization. Atomization breaks down the complex user-entered search text into fine-grained, indivisible semantic units (atomic subqueries), each representing a separate semantic element. This process is typically implemented using natural language processing techniques, aiming to improve the accuracy and interpretability of multimodal search.

[0172] For example, “a red car runs a red light” is atomized into {red, car, runs a red light}.

[0173] For example, the following is a detailed process of vectorizing search text. The search text entered by the user is atomized to obtain n sub-queries: Q = {q1, q2, ..., q n}.

[0174] Each subquery q i After passing through the CLIP text encoder, a d-dimensional vector is generated:

[0175] Combined into a query matrix: Each row is a vector of atomic subqueries.

[0176] Step 804: Obtain the target image from the surveillance video based on the similarity between the search vector and the image vector in the image vector database.

[0177] Exemplarily, the image vector includes a global vector and a local vector, and the target image is obtained from the surveillance video according to the similarity between the search vector and the image vector in the image vector database, such as Figure 9 As shown, the following steps are included:

[0178] Step 901: Obtain local similarity based on the search vector and the local vector.

[0179] For example, the detailed steps for obtaining the local similarity are as follows:

[0180] Normalize the local vector matrix and the search vector:

[0181]

[0182] Among them, V oj,: is the local vector, V q,: is the search vector, is the normalized local vector, is the normalized search vector, i = 1, 2, 3, …, m; j = 1, 2, 3, …, n.

[0183] Next, matrix multiplication is used to calculate the local similarity matrix between the local vector and the search vector:

[0184]

[0185] in By matrix multiplication, the similarity between all local features of the search vector and the target vector is calculated to generate an n×m similarity matrix S local .

[0186] For example, if the search text has 3 search sub-query vectors (n=3) and the target image has 4 local vectors (m=4), then S local It is a 3×4 matrix, each element S local [i,j] represents the similarity between the i-th subquery and the j-th local target of the image.

[0187] Select the maximum local similarity between each subquery vector and the local vector (each subquery matches the optimal target):

[0188]

[0189] Next, the maximum local similarity of the local vectors of each local target is aggregated to obtain the local similarity of the target image:

[0190]

[0191] For example, when calculating the similarity between a local area of ​​an image and a query target, the contribution of different areas can be dynamically adjusted by weighting the local targets, thereby highlighting key areas and suppressing noise or irrelevant areas.

[0192] Exemplarily, after obtaining the local similarity, the method further includes:

[0193] Obtain target categories, target positions, and target confidences of the plurality of local targets, and weight the local similarity matrix according to the matrix of the target categories, the target positions, and the target confidences. The weighted local similarity matrix serves as the local similarity matrix between the text vector and the local vector. For example, the weighting formula is as follows:

[0194] S local_weighted =S local ⊙W, (7)

[0195] Among them, ⊙ represents element-wise multiplication, and W is the weight of one or more combinations of target category, target location, and target confidence, which can be used flexibly.

[0196] Step 902: Obtain global similarity based on the search vector and the global vector.

[0197] Exemplarily, the detailed steps for obtaining the global similarity are as follows.

[0198] The search vector With the global vector Calculate cosine similarity:

[0199] Perform vector normalization on the search vector and the global vector:

[0200]

[0201] Next, matrix multiplication is used to calculate the global similarity matrix between the global vector and the search vector:

[0202]

[0203] in,

[0204] The global similarity matrix corresponding to each subquery and the global vector of each image is summed and divided by the number of subqueries to obtain the average global similarity of the image, that is, the global similarity between the target image and the search vector:

[0205]

[0206] Step 903: Obtain the similarity between the search vector and the image vector by summing the product of the local similarity, the local feature weight corresponding to the local similarity, and the global feature weight corresponding to the global similarity.

[0207] For example, the global and local similarities are combined, and the adjustable weights global feature weight α and local feature weight β are used:

[0208] S final=α·S global_avg +β·S local_avg , (11)

[0209] Wherein, exemplarily, the sum of the local feature weight and the global feature weight is a preset value, exemplarily, α=0.4, β=0.6, and it is made to meet the normalization requirement, that is, α+β=1.

[0210] For example, the local feature weights and global feature weights can be flexibly adjusted. If local target matching is preferred, the local feature weight β is increased. If global scene matching is preferred, the global feature weight α can be increased.

[0211] Exemplarily, the target image is an image whose similarity is greater than a preset similarity threshold.

[0212] After obtaining the similarity, an image with a similarity greater than a preset similarity threshold is selected as the target image.

[0213] For example, the setting of the similarity threshold directly affects the accuracy of the search results and can be dynamically adjusted in combination with specific application scenarios.

[0214] For example, the search text entered by the user is atomized into: "red car" → q1 = "red", q2 = "car", and the search vector is generated

[0215] Local objects in the image: o1: white car; o2: red truck; o3: red car (high similarity).

[0216] For example, the local similarity calculation result is as follows:

[0217]

[0218] For the corresponding subqueries: "red" has the highest similarity with o3 (0.95), and "car" has the highest similarity with o3 (0.9). Next, we average the local similarities of the subqueries to get the local similarity:

[0219]

[0220] Assume that the global similarity is S global_avg =0.7.α=0.4,β=0.6.

[0221] According to formula (11), the similarity is S final =0.4×0.7+0.6×0.925=0.815.

[0222] Illustratively, the matrix method described above facilitates batch processing using existing vector retrieval frameworks (such as Milvus and FAISS) and can be directly extended to large-scale data scenarios.

[0223] It should be understood that although Figure 2-9 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. In addition, Figure 2-9 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.

[0224] In some possible embodiments, such as Figure 10 As shown, a multimodal retrieval system for surveillance videos is provided, including: an acquisition module 1001, a processing module 1002, and a generation module 1003, wherein:

[0225] The acquisition module 1001 is used to obtain a search text for a surveillance video input by a user; and obtain a target image from the surveillance video according to the search text.

[0226] Processing module 1002 is used to classify the search text into different industries, and determine the derivative relationship corresponding to the search text based on the industry map corresponding to the target industry type to which the search text belongs. The derivative relationship is used to generate a text description of the target retrieval report.

[0227] The generating module 1003 is configured to generate the target retrieval report based on the text description and the target image.

[0228] For further limitations on the above-mentioned apparatus, please refer to the limitations on the search report generation method above and will not be repeated here. Each module in the above-mentioned apparatus may be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above-mentioned modules may be embedded in or independent of the processor in the terminal device in hardware form, or may be stored in the memory of the terminal device in software form, so that the processor can call and execute the corresponding operations of each of the above modules.

[0229] Another embodiment provides a computer-readable storage medium for storing a computer program. This computer program includes instructions for implementing the methods described in the embodiments of this application. By installing this computer program on a computer, the computer can execute the corresponding method.

[0230] Another embodiment provides a computer program product that includes computer program code. When the computer program code is executed on a computer, it causes the computer to implement the method provided in the embodiment of the present application. In this way, a user can achieve the method by using this computer program product.

[0231] For example, Figure 11 1 is a schematic block diagram of an electronic device provided in an embodiment of the present application. The electronic device 1100 may include: a memory 1101 storing executable program code and a processor 1102 coupled to the memory 1101.

[0232] The processor 1102 calls the executable program code stored in the memory to execute any one of the methods disclosed in the embodiments of the present application. Those skilled in the art will understand that Figure 11 The electronic device structure shown in the figure does not constitute a limitation to the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0233] The processor 1102 is the control center of the electronic device. It connects the various parts of the entire electronic device using various interfaces and lines. By running or executing software programs and / or modules stored in the memory and accessing data stored in the memory, it performs various functions of the electronic device and processes data, thereby monitoring the electronic device as a whole. Optionally, the processor may include one or more processing units; preferably, the processor may integrate an application processor and a modem processor, wherein the application processor primarily processes the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into the processor.

[0234] The memory 1101 can be used to store software programs and modules. The processor executes the various functional applications and data processing of the electronic device by running the software programs and modules stored in the memory. The memory can mainly include a program storage area and a data storage area. The program storage area can store an operating system, at least one application required for a function, etc.; the data storage area can store data created based on the use of the electronic device, etc. In addition, the memory can include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state memory device.

[0235] It should be understood that in the embodiments of the present application, the processor may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0236] During implementation, each step of the above method can be completed by an integrated logic circuit of hardware in a processor or by instructions in the form of software. The steps of the method disclosed in conjunction with the embodiments of the present application can be directly embodied as being executed by a hardware processor, or can be executed by a combination of hardware and software modules in the processor. The software module can be located in a storage medium mature in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in a memory, and the processor executes the instructions in the memory, and completes the steps of the above method in conjunction with its hardware. To avoid repetition, it will not be described in detail here.

[0237] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the embodiments of this application.

[0238] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0239] In the several embodiments provided in the embodiments of the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0240] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0241] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0242] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0243] The above description is merely a specific implementation of the embodiments of the present application, but the scope of protection of the embodiments of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the embodiments of the present application should be included in the scope of protection of the embodiments of the present application. Therefore, the scope of protection of the embodiments of the present application should be based on the scope of protection of the claims.

Claims

1. A multimodal retrieval method for surveillance videos, characterized in that: include: Obtaining the search text for surveillance video input by the user; Obtaining a target image from the surveillance video according to the search text; Classifying the search text by industry, determining a derivative relationship corresponding to the search text according to an industry map corresponding to the target industry type to which the search text belongs, and generating a target text description of a target retrieval report based on the derivative relationship; The target retrieval report is generated based on the target text description and the target image.

2. The method according to claim 1, characterized in that Determining a derivative relationship corresponding to the search text according to an industry graph corresponding to the target industry type to which the search text belongs includes: Obtaining a search vector corresponding to the search text; Obtaining a plurality of keywords corresponding to a plurality of related vectors related to the search vector in a pre-stored knowledge vector database, wherein the plurality of keywords are used to characterize the target industry type; Mapping the multiple keywords to corresponding multiple nodes of the industry graph; Derivative relationships corresponding to the search text are determined based on the multiple nodes.

3. The method according to claim 2, characterized in that Determining, based on the plurality of nodes, a derivative relationship corresponding to the search text includes: Obtaining semantic weights of the multiple nodes; According to the semantic weights of the plurality of nodes and a first semantic weight threshold, screening out at least one first target node having a semantic weight greater than or equal to the first semantic weight threshold; For each first target node, the maximum number of hops corresponding to the first target node is determined according to the semantic weight of the first target node. Starting from the first target node, the industry graph is traversed according to the maximum number of hops to determine the derivative relationship corresponding to the search text, wherein the maximum number of hops is positively correlated with the semantic weight, and the maximum number of hops is used to indicate the maximum number of association steps when performing multi-hop reasoning in the industry graph starting from the first target node.

4. The method according to claim 3, characterized in that The target retrieval report is a retrieval report that meets preset conditions. The target text description of the target retrieval report generated based on the derivative relationship includes: generating a first text description of a target retrieval report based on the derived relationship; If the first text description satisfies the preset condition, outputting the first text description as the target text description; If the first text description does not meet the preset condition, an Internet search is performed to generate a second text description, and if the second text description meets the preset condition, the second text description is output as the target text description.

5. The method according to claim 4, characterized in that The method further comprises: If the second text description does not meet the preset condition, the sub-target node corresponding to the maximum number of hops in the first text description is used as the second target node; For each second target node, the maximum number of hops corresponding to the second target node is determined according to the semantic weight of the second target node. Starting from the second target node, the industry graph is traversed according to the maximum number of hops to determine the derivative relationship corresponding to the search text, wherein the maximum number of hops is positively correlated with the semantic weight, and the maximum number of hops is used to indicate the maximum number of association steps when performing multi-hop reasoning in the industry graph starting from the second target node.

6. The method according to claim 1, characterized in that Obtaining a target image from the surveillance video according to the search text includes: Obtaining key frame images in the surveillance video; Performing image preprocessing on the key frame image to obtain an image vector corresponding to the key frame image, and storing the image vector in an image vector database; Performing a preprocessing operation on the search text to obtain a search vector corresponding to the search text; The target image is obtained from the surveillance video according to the similarity between the search vector and the image vectors in the image vector database.

7. The method according to claim 6, characterized in that The image vector includes a global vector and a local vector, and obtaining the target image from the surveillance video according to the similarity between the search vector and the image vector in the image vector database includes: Obtaining local similarity according to the search vector and the local vector; Obtaining a global similarity based on the search vector and the global vector; The similarity between the search vector and the image vector is obtained by summing the product of the local similarity, the local feature weight corresponding to the local similarity, and the global feature weight corresponding to the global similarity, wherein the sum of the local feature weight and the global feature weight is a preset value, and the target image is an image whose similarity is greater than a preset similarity threshold.

8. A multimodal retrieval system for surveillance videos, characterized in that: include: An acquisition module is used to obtain the search text for the surveillance video input by the user; Obtaining a target image from the surveillance video according to the search text; a processing module configured to classify the search text by industry, and determine a derivative relationship corresponding to the search text based on an industry map corresponding to a target industry type to which the search text belongs, wherein the derivative relationship is used to generate a text description of a target retrieval report; A generating module is used to generate the target retrieval report based on the text description and the target image.

9. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that Computer instructions are stored, and when the instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Description text generation method and device, computer equipment and storage medium

    CN115221331A

  • Video searching method and device, equipment and storage medium

    CN115422399A

  • RAG question and answer method and system based on knowledge graph and medium

    CN118673126A

  • Data-driven cross-domain intelligent asset knowledge reasoning and value evaluation method and system

    CN119476499A

  • Semantic retrieval method and device of image, equipment and medium

    CN119719398A