Multi-modal retrieval enhancement generation system for mass law enforcement audio and video data
By integrating voiceprint-assisted multimodal indexing and knowledge graph database modules with speech recognition and multimodal retrieval technologies, the problem of cross-scene semantic association of massive law enforcement audio and video data was solved, achieving efficient multimodal retrieval and generation, generating structured analysis reports, and improving the efficiency and accuracy of law enforcement audio and video data processing.
Patent Information
- Application Number
- CN202511342447.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-19
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2045-09-19
AI Technical Summary
Existing technologies struggle to effectively process and understand massive amounts of law enforcement audio and video data, especially in cross-scene semantic association and multimodal information processing. The lack of systematic integration leads to the loss of contextual information and makes it difficult to meet the needs of law enforcement scenarios for cross-scene identity association and multimodal semantic modeling.
Employing a voiceprint-assisted multimodal indexing module, a multimodal retrieval and generation module, and a knowledge graph database module, this system integrates speech recognition, voiceprint analysis, and multimodal retrieval technologies. Through voiceprint-assisted knowledge graph construction and a knowledge-driven retrieval mechanism, it achieves structured indexing, semantic association, and cross-scene speaker association of law enforcement audio and video content, generating comprehensive analysis reports of related events.
It achieves cross-scene semantic association and efficient content generation, improves the recall and accuracy of multimodal retrieval, ensures real-time performance and privacy protection, outputs structured analysis reports, and supports cross-scene speaker association and event tracking.
Smart Images

Figure CN120821873A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multimodal information processing, and in particular to a multimodal retrieval enhancement generation system for massive law enforcement audio and video data. Background Art
[0002] With the rapid growth of audio and video data in law enforcement, processing and understanding this massive volume of data has become a significant challenge. This data is characterized by its length, diverse sources, and complex content. Traditional archiving methods store audio and video data as isolated fragments, resulting in a loss of context and difficulty capturing semantic relationships across scenes or speaker identities. Existing retrieval-augmented generation (RAG) technology primarily targets text data and lacks the ability to comprehensively process the multimodal information found in law enforcement audio and video.
[0003] Automatic speech recognition (ASR), speaker separation, and voiceprint analysis technologies provide new tools for audio and video analysis, but these technologies are often applied independently and lack systematic integration, making it difficult to meet the needs of law enforcement scenarios for cross-scenario identity association and multimodal semantic modeling.
[0004] Law enforcement audio and video data processing faces the following core challenges: 1. Multimodal complexity: The data includes video images, voice conversations, ambient sounds, and voiceprint features, which require comprehensive analysis to extract key information, such as event descriptions and the identities of the parties involved.
[0005] 2. Cross-scenario semantic association: Different law enforcement scenarios (such as street patrols and interrogation rooms) may involve the same personnel or events, and semantic and identity associations need to be established.
[0006] 3. Real-time and privacy protection: Law enforcement data processing requires rapid retrieval of relevant fragments while protecting sensitive information such as personal identity.
[0007] 4. Large-scale data management: The data volume is huge, and efficient indexing and retrieval mechanisms are required to support real-time queries. Summary of the Invention
[0008] The present invention aims to address the limitations of the above-mentioned traditional methods in multimodal information processing, cross-scene association, and semantic consistency. To this end, the present invention provides a multimodal retrieval enhancement generation system for massive law enforcement audio and video data, integrating speech recognition, voiceprint analysis, knowledge graph database, and multimodal retrieval technology. Through voiceprint-assisted knowledge graph construction and knowledge-driven retrieval mechanism, it realizes structured indexing, semantic association, cross-scene speaker association, efficient content generation of law enforcement audio and video content, and finally outputs a comprehensive analysis report of related events. The present invention is applicable to scenarios such as law enforcement record analysis, surveillance video retrieval, and law enforcement event tracking, and provides comprehensive event insights and archiving suggestions by generating a comprehensive analysis report of related events.
[0009] The present invention provides a multimodal retrieval enhancement generation system for massive law enforcement audio and video data, which adopts the following technical solutions: including: a multimodal indexing module based on voiceprint assistance, a multimodal retrieval and generation module, and a knowledge graph database module. The voiceprint-assisted multimodal indexing module is used to generate subtitles, transcriptions, and voiceprint vectors based on the recorder audio and video, merge the subtitles, transcriptions, and voiceprint vectors to form a multimodal representation; use a multimodal fusion encoder to process the multimodal representation and construct a knowledge graph database; The multimodal retrieval and generation module is used to receive user queries and extract multimodal keywords; query the knowledge graph database based on the multimodal keywords to obtain a text retrieval set, a visual retrieval set, and a voiceprint retrieval set, and integrate them to form a fused retrieval set; based on the fused retrieval set, use the VLM and LLM to generate a comprehensive analysis report of related events; The knowledge graph database module is used to store and manage the knowledge graph database.
[0010] Furthermore, the process of forming a multimodal representation is: Split the recorder audio and video into multiple clips; Sample visual frames from each clip, input into the visual language model, and generate captions; Generate a transcription of each segment using automatic speech recognition; Perform speaker separation on each segment to generate a voiceprint vector; Merge captions, transcriptions, and voiceprint vectors to form multimodal representations.
[0011] Furthermore, the process of sampling visual frames is as follows: the average grayscale difference between the current frame and the previous sampled frame is calculated as the rate of change. If the rate of change exceeds the rate of change threshold and the time interval exceeds the minimum interval, the current frame is used as the sampling frame. The first frame is always sampled.
[0012] Furthermore, the nodes of the knowledge graph database include speaker nodes, speech segment nodes, event nodes, and object nodes; Edges include identity-related edges, speech generation edges, behavior-related edges, and event-related edges.
[0013] Furthermore, the processing process of the multimodal fusion encoder is: Pass the subtitles and transcripts through the embedding layer to obtain the subtitle embedding and audio transcript embedding, and then pass them through the cross-attention mechanism to obtain the first layer output; The first layer output and the voiceprint vector are integrated with the identity information using the causal multi-head self-attention mechanism to obtain the second layer output; The semantic relationship between the second layer output and the knowledge graph is embedded through the graph attention network to obtain the updated node embedding and attention coefficient matrix.
[0014] Furthermore, the multimodal keywords include text keywords, visual descriptions, and latent voiceprint identifiers; Extract corresponding fragments from the knowledge graph database based on text keywords to form a text retrieval set; Convert visual descriptions into visual embedding vectors, extract corresponding fragments from the knowledge graph database, and generate a visual retrieval set; According to the latent voiceprint identifier, the voiceprint vector of the relevant speaker node is extracted from the knowledge graph database, and all associated segments of the voiceprint vector are extracted to form a voiceprint retrieval set.
[0015] Furthermore, the text retrieval set, visual retrieval set and voiceprint retrieval set are integrated, and a weighted fusion mechanism is used to remove redundant segments to form a fused retrieval set.
[0016] Furthermore, the comprehensive analysis report of the related events includes at least one of the following analysis dimensions: event chronology, entity correlation, event summary, key findings, and archiving recommendations.
[0017] Furthermore, event timing is based on the timeline sorting of edges in the knowledge graph database, which sorts events or fragments across scenarios in chronological order to form an event timeline; Entity relevance starts from the matching nodes of the fusion retrieval set, traverses the edges of the knowledge graph database, extracts relevant subgraphs, and displays the associations between entities; The event summary uses LLM to summarize the search results and subgraphs to generate a concise description of the event cause, development process, and disposal results; The key finding is to detect anomalies or patterns in events, which is achieved based on the results of subgraph extraction, text blocks and embedding similarity comparison; Archiving suggestions are based on the strength of event association and recommend whether to merge videos or events into a single archive.
[0018] Furthermore, it also includes a task management module for receiving and initializing law enforcement audio and video processing tasks and managing task status.
[0019] The above one or more technical solutions in the embodiments of the present invention have at least one of the following technical effects: 1. Voiceprint-assisted knowledge graph database: Through visual language models, speech recognition, voiceprint analysis and a multimodal fusion encoder (MEnc), the visual, audio and voiceprint information of audio and video are integrated into a knowledge graph database, preserving semantic and identity associations across scenarios. This paper designs a new multimodal fusion encoder (MEnc), which uses a hierarchical attention mechanism to fuse subtitle embeddings, audio transcription embeddings and voiceprint vectors layer by layer. Specifically, the first-layer attention module focuses on visual-audio alignment, the second layer integrates voiceprint identity information, and the third layer injects semantic relationship embeddings of the knowledge graph to achieve more accurate cross-modal interaction representation. This structure significantly improves the recall rate and precision of multimodal retrieval, reduces parameter redundancy and improves generalization ability compared to traditional concatenation fusion methods.
[0020] 2. Knowledge-driven multimodal retrieval and generation: Through the knowledge graph database, combined with text semantic matching, visual content retrieval and voiceprint association, relevant audio and video clips are accurately extracted and comprehensive responses are generated, forming a comprehensive analysis report of related events.
[0021] 3. Cross-scenario speaker association: Utilize voiceprint analysis and knowledge graph database to automatically identify and associate the same speakers in different law enforcement scenarios and build a semantic network across videos.
[0022] 4. Efficient task processing flow: A pipeline architecture built through task initialization, AI model processing, and knowledge graph database ensures system efficiency, robustness, and privacy protection, and outputs structured analysis reports.
[0023] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned by practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0025] Figure 1 It is a structural block diagram provided by the present invention.
[0026] Figure 2 This is a flow chart for constructing a knowledge graph database provided by the present invention.
[0027] Figure 3 It is a structural diagram of the multimodal fusion encoder provided by the present invention.
[0028] Figure 4 This is a multimodal retrieval and generation flow chart provided by the present invention. DETAILED DESCRIPTION
[0029] To make the purpose, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the present invention. Obviously, the embodiments described are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.
[0030] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the embodiment of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.
[0031] The following combination Figures 1 to 4 The present invention is further described in detail, which is a multimodal retrieval and enhanced generation system for massive law enforcement audio and video data. In this embodiment, Figure 1 As shown, a multimodal retrieval enhancement generation system for massive law enforcement audio and video data is provided, including: a multimodal indexing module based on voiceprint assistance, a task management module, a multimodal retrieval and generation module and a knowledge graph database module.
[0032] The voiceprint-assisted multimodal indexing module is used to generate subtitles, transcriptions, and voiceprint vectors based on the recorder audio and video, merge the subtitles, transcriptions, and voiceprint vectors to form a multimodal representation; use a multimodal fusion encoder to process the multimodal representation and construct a knowledge graph database; The multimodal retrieval and generation module is used to receive user queries and extract multimodal keywords; query the knowledge graph database based on the multimodal keywords to obtain a text retrieval set, a visual retrieval set, and a voiceprint retrieval set, and integrate them to form a fused retrieval set; based on the fused retrieval set, use the VLM and LLM to generate a comprehensive analysis report of related events; The knowledge graph database module is used to store and manage the knowledge graph database; The task management module is used to receive and initialize law enforcement audio and video processing tasks and manage task status.
[0033] The specific working process of this system is: Step S1: The Task Management Module receives a law enforcement audio and video processing task, including audio and video metadata (video_meta) and a processing directory (handle_dir). The Task Status Manager sets the task status to "Processing" and initializes the progress to 0%. The Task Status Manager then verifies the audio and video file paths. Once verified, an encrypted temporary directory (audio_asr_temp) is created to store intermediate files, ensuring the security of sensitive data.
[0034] Step S2: The multimodal indexing module based on voiceprint assistance reads the audio and video metadata, and then performs the following operations, such as Figure 2 As shown: Audio and video splitting: split the input recorder audio and video Split into multiple short segments , each segment contains visual frames and audio content, represents the jth segment, is the total number of fragments.
[0035] Visual-text semantic extraction: Sample visual frames from each segment. This embodiment adopts an intelligent frame sampling algorithm based on the inter-frame change rate and time interval, specifically including: calculating the average grayscale difference between the current frame and the previous sampled frame as the change rate. If the change rate exceeds the change rate threshold and the time interval exceeds the minimum interval, the current frame is used as the sampling frame, and the first frame is always sampled. Then, face detection and posture judgment are performed on the sampled frames, and a posture score is calculated for each detected face to form a list of each person's face. Only the photo with the highest posture score is stored to reduce storage space. The photo can be inferred by using an emotion classification model to form an emotion label. In this embodiment, the change rate threshold is 0.2 and the minimum interval is 20 frames.
[0036] All sampled frames are input into the visual language model (VLM) to generate natural language captions for the video, describing the dynamics, behaviors, and key objects of the law enforcement scene, such as "law enforcement officers questioning parties on the street, holding notebooks."
[0037] Assume that the intelligent sampling method is used to Sampling ,in, Represents the kth sample frame. Input the visual language model to get the caption of the jth segment , , Represents a visual language model.
[0038] Audio-to-text and voiceprint semantic extraction: Use automatic speech recognition (ASR) to generate time-aligned transcriptions of each segment, capturing conversational or ambient sound. ,in, represents the transcription of the jth segment, Represents an automatic speech recognition model.
[0039] Speaker separation is performed on each segment to generate speaker identifiers and voiceprint vectors. Speakers are grouped by speaker, and high-quality segments with a duration of ≥ 1.5 seconds are filtered. Aggregate voiceprints are calculated and a robust averaging strategy is applied, removing 10% of extreme vectors to obtain the voiceprint vector.
[0040] Unified multimodal representation: Merge subtitles, transcriptions, and voiceprint vectors to form a unified multimodal representation. ,in, For the recorder audio and video A multimodal representation, For the subtitles of the clip, For the The transcription of the fragments, For the The voiceprint vector of each segment.
[0041] Use Multimodal Fusion Encoder (MEnc) to process multimodal representations and build a knowledge graph database , .
[0042] Representation nodes, including speakers, speech segments, events, and objects, store multimodal representations and metadata of recorder audio and video, and are generated by the visual language model, automatic speech recognition model, and multimodal fusion encoder.
[0043] Edges represent semantic relationships between nodes, including behavioral relationships (such as "inquiry"), identity associations (such as "IS_SAME_AS"), event associations (such as "lead to"), and speech generation (such as "HAS_SPEECH"), which are generated by LLM (large language model) relationship extraction and voiceprint matching.
[0044] like Figure 3 As shown, this embodiment uses a multimodal fusion encoder to generate embeddings that capture interactive features of vision, audio, text, and voiceprints. MEnc uses a three-layer attention mechanism to fuse subtitles, transcriptions, voiceprints, and knowledge graph relationship features to generate unified entity embeddings that support subsequent node attribute definition. This design reduces parameter redundancy and ensures efficient cross-modal interaction.
[0045] Pass the subtitles and transcriptions through the embedding layer to get the subtitle embedding and audio transcription embedding ,in, , , Represented by the embedding layer.
[0046] The first layer of the multimodal fusion encoder implements caption embedding through the cross-attention mechanism Embedded with audio transcription Alignment, reducing modality bias, and calculating attention weights ,in, For visual query, , is the query weight, For audio keys, , is the key weight, is the audio value, , is the value weight, and the output of the first layer is The first layer of the multimodal fusion encoder consists of stacked layer normalization, crisscross attention layers, and addition layers.
[0047] The second layer of the multimodal fusion encoder converts the input With voiceprint vector Fusion, by injecting the voiceprint (speaker identity feature) into the pre-order alignment result ( ), to achieve content and identity association; use causal multi-head self-attention mechanism to integrate identity information, the second layer output is The second layer of the multimodal fusion encoder consists of stacked causal multi-head self-attention layers, addition layers, and layer normalization.
[0048] The third layer of the multimodal fusion encoder combines the semantic relationship of the knowledge graph to optimize the relevance of entity embedding; the input is and semantic relation embedding of knowledge graphs .in, As the initial node embedding of the third layer, it provides the basis for multimodal features for the Graph Attention Network (GAT); It is a vector representation of the relationship edge in the knowledge graph, encoding the semantic information of the semantic relationship and related metadata, such as confidence, timestamp, . The relationship is propagated through the graph attention network (GAT) to update the node embedding, and the formula is ,in, is the updated embedding of node i, is the set of neighbor nodes of node i, is the attention coefficient of node i and node j, is the linear transformation weight matrix, is the original embedding of neighbor node j. Guided Graph Attention Network (GAT) computation , to propagate the semantic relationship between nodes; at the same time, enhance , capturing cross-modal and cross-scenario semantic interactions, supporting multimodal retrieval and the generation of comprehensive analysis reports for related events. The third layer outputs updated node embeddings and attention coefficient matrices. The second layer of the multimodal fusion encoder consists of a stacked feedforward layer and a graph attention layer.
[0049] The updated node embeddings incorporate neighboring semantic information. They are no longer isolated single-node features, but rather incorporate the features of all of their neighboring nodes. The updated node embedding dimensions remain consistent with the original input node embeddings and incorporate multimodal information such as subtitles, transcriptions, and voiceprints, avoiding the one-sidedness of single-modal features. They also implement differentiated fusion with attention weights, leveraging the attention coefficient to prioritize the contributions of semantically relevant neighbors.
[0050] The attention coefficient matrix is a matrix composed of the attention coefficients of all node pairs. Its function is to supplement the importance weight of neighbor nodes to the target node. Its dimension matches the number of nodes and reflects the strength of semantic association.
[0051] The knowledge graph database is constructed using the updated node embedding and attention coefficient matrix.
[0052] In this embodiment, the nodes of the knowledge graph database include speaker nodes, speech segment nodes, event nodes, and object nodes.
[0053] The edges between nodes include identity-related edges, speech generation edges, behavior-related edges, and event-related edges, as follows: Identity association edge: connects speaker nodes and speaker nodes, relying on voiceprint vector matching + subtitle / transcription text clues.
[0054] Speech generation edge: connects the speaker node and the speech segment node, relying on voiceprint vector matching + subtitle / transcription time alignment.
[0055] Behavior-related edges: connect speaker nodes and speaker nodes, and speaker nodes and event nodes, relying on subtitle / transcribed text behavior descriptions and voiceprint vector identity confirmation.
[0056] Event-related edges: connect event nodes and event nodes, speech segment nodes and event nodes, object nodes and event nodes, relying on text causal clues and timestamp alignment of subtitles / transcriptions.
[0057] This construction phase serves as the foundation for subsequent retrieval and generation, processing the raw recorder audio and video into a structured knowledge graph database, where nodes contain multimodal features and edges record semantic associations. This offline or semi-real-time process directly supports query-driven extraction during the retrieval phase, ensuring knowledge-driven nature. The constructed multimodal representation and the layer-by-layer fusion of MEnc provide an embedding foundation for subsequent text, visual, and voiceprint matching. The encrypted temporary directory and robust averaging strategy extend to retrieval threshold adjustment and index acceleration, achieving shared optimization of privacy and efficiency.
[0058] Audio and video metadata typically includes multiple recorders, each of which forms part of a knowledge graph database. Multiple parts are connected by edges to form a complete knowledge graph database. Nodes in the knowledge graph database identify their corresponding recorders through their metadata, and each node is associated with a clip of a recorder's audio or video.
[0059] The knowledge graph database constructed by the voiceprint-assisted multimodal indexing module is stored in the knowledge graph database module.
[0060] Step S3: The multimodal retrieval and generation module queries the knowledge graph database in the knowledge graph database module according to the received user query, and generates a comprehensive analysis report of related events according to the query results. Figure 4 The specific process is as follows: Query preprocessing: Receive user queries and extract multimodal keywords from user queries, including text keywords (corresponding to , visual description (corresponding to ) and latent voiceprint identification (corresponding to This step is directly related to the multimodal representation of the construction phase, preparing the embedding vector for subsequent matching, similar to MEnc for 、 、 Layer-by-layer attention fusion.
[0061] To improve the accuracy of semantic matching, LLM can be used to reconstruct user queries and convert them into declarative sentences. For example, "the conversation between the parties in the street questioning" can be converted into "retrieve the conversation content containing the parties in the street questioning scenario", and then multimodal keywords can be extracted from the reconstructed query.
[0062] If there is no latent voiceprint identifier, the latent voiceprint identifier is inferred from the user query.
[0063] Query the knowledge graph database based on multimodal keywords to obtain text retrieval sets, visual retrieval sets, and voiceprint retrieval sets. The specific process is as follows: Text semantic matching: Based on the text keywords, the embedding similarity between them and the entity descriptions in the knowledge graph database is calculated. In this embodiment, cosine similarity or BERT-based semantic matching is used to select relevant text blocks with similarity higher than a threshold (e.g., 0.8). The corresponding audio and video clips are extracted from the selected text blocks to form a text retrieval set. If the text block contains cross-scene references, then Expand the search scope. Similarity calculation is equivalent to Semantic matching of text parts in the LLM process during the construction phase generate The entity description of provides a semantic basis for this matching. Without time-aligned transcription, Unable to capture conversation context.
[0064] Text blocks refer to structured text snippets extracted from law enforcement audio and video data, stored in a knowledge graph database, and used to describe the semantic content of the audio and video clips. They are primarily derived from subtitles and transcriptions and are associated with nodes and edges in the knowledge graph database. A text block represents a collection of text snippets that semantically match the user query during the retrieval phase. Cross-scene references refer to the semantic information contained in a text block being associated with the content of other scenes or video files through edges in the knowledge graph. In other words, cross-scene references are triggered when the semantic content of a text block (such as a conversation, event, or speaker identity) is associated with nodes or edges in different scenes (such as different times, places, or video files).
[0065] Visual content retrieval: converting visual descriptions into visual embedding vectors through MEnc , ,in, For visual description, a hierarchical index is used: first, the scene type is matched coarsely, that is, the visual embedding vector is compared with the subtitle embedding, and the embedding comparison uses the first layer of MEnc; then the details are matched finely, and the visual embedding of the sample frame corresponding to the subtitle embedding is compared (such as the cosine similarity threshold of 0.7). , is a visual embedding; if the similarity exceeds 0.7, the corresponding fragment is extracted from the knowledge graph database to generate a visual retrieval set To improve retrieval speed and accuracy, this embodiment uses the photo with the highest pose score to calculate its visual embedding, and then compares it with the visual embedding vector.
[0066] Voiceprint assisted retrieval: The voiceprint vector of the relevant speaker node is extracted from the knowledge graph database based on the potential voiceprint identifier. This embodiment uses cosine similarity for comparison query, with the threshold set to 0.75. If the similarity exceeds the threshold, all related segments of the voiceprint vector are extracted to form a voiceprint retrieval set. During retrieval, matching speaker nodes in other audio and video files are retrieved through edge expansion.
[0067] Retrieval result fusion: Integrate the text retrieval set, visual retrieval set and voiceprint retrieval set, use weighted fusion mechanism to remove redundant segments, and form a fused retrieval set. The weight is based on modality relevance. In this embodiment, the text weight is 0.4, the visual weight is 0.3, and the voiceprint weight is 0.3. Later, it can be adjusted according to actual usage (accuracy of retrieval results). The essence of building a fused retrieval set is the construction stage. The subset filtering and fusion rely on the node embedding update of the third layer GAT of MEnc. Without the multimodal representation in the construction phase, fusion cannot achieve cross-modal interaction.
[0068] Redundant fragments refer to the situation in which multiple retrieval sets ( 、 、 ) contains repeated or highly similar audio and video clips. These clips may overlap in semantic content, timestamps, scenes, or speaker identities, resulting in redundant retrieval results. For example: Include fragment (Dialogue "Please show your ID"). The same sequence is included (scene "Night Street Questioning"). Contains the same segment (based on the voiceprint of the party). If these segments point to the same (same timestamp and video ID), they are considered redundant.
[0069] A fragment deduplication method based on embedding similarity: Each fragment in the retrieval set is associated with a multimodal embedding, including subtitle embeddings, audio transcript embeddings, and voiceprint embeddings (voiceprint vectors). The subtitle embeddings, audio transcript embeddings, and voiceprint embeddings are initially weighted and then input into MEnc. The third layer of MEnc outputs the updated embeddings of the nodes corresponding to each fragment. The cosine similarity is used to calculate the similarity of the updated embeddings. If the similarity exceeds a threshold, such as 0.9, the fragment is considered redundant and only one is retained. Generally, the fragment with the highest confidence is retained. For example: Include fragment (Dialogue "Please show your ID"), Contains the same segment (scene "Nighttime Street Interrogation"). Calculate and compare the updated embeddings of the two segments. If the similarity is greater than 0.9, it is confirmed that the two segments point to the same segment, and one is retained.
[0070] The construction and retrieval of the knowledge graph database use the same similarity threshold. According to the accuracy of the retrieval results, the threshold of the retrieval stage is adjusted, and the threshold of the construction stage is updated synchronously.
[0071] Based on the fusion search set, VLM and LLM are used to generate a comprehensive analysis report of related events.
[0072] Extract key information from the fusion search set, including the conversation content (from Extracted from ), scene description (from ), speaker identity (from Extracted) and sentiment analysis (optionally, using an additional sentiment classification model). Use VLM and LLM to integrate multimodal information and generate a comprehensive response: ,in, It is an extended form of the knowledge graph database. is the search function, For user queries, The LLM first constructs a prompt template, injects search results (fused search sets), and then generates a natural language response to obtain a comprehensive analysis report of related events.
[0073] The report contains at least one of the following analysis dimensions: event chronology, entity correlation, event summary, key findings, and archiving recommendations. Specifically: (1) Event timing is a timeline sorting based on the edges of the knowledge graph database. It sorts events or fragments across scenes in chronological order to form an event timeline and marks the transition, such as from street patrol to interrogation room.
[0074] Data source: metadata (such as timestamps and video metadata) of edges (such as "leads to", "IS_SAME_AS", and "participated in events") and nodes in the knowledge graph database; each audio and video clip is associated with a timestamp (obtained from video metadata or time-aligned transcription).
[0075] Sorting process: Use a graph database (Neo4j) to query and traverse the knowledge graph database, starting from the matching speaker node or event node, and collect related segments along the edge, such as connecting the same party across videos through the "IS_SAME_AS" edge.
[0076] Apply a timeline algorithm (such as timestamp-based topological sorting) to the collected clips or event nodes to generate an ordered sequence. For example, first sort the clips within the same video, then connect cross-video events through logical edges. LLM injects prompt templates based on the sorting results to generate a natural language timeline.
[0077] Cross-scene processing: Extend to other videos (such as street patrol timestamp T1 to interrogation room timestamp T2) through the "IS_SAME_AS" edge, and mark the transition (such as "the time interval is about 20 minutes"). Example: From The “interrogation” events (T1: street patrol) and “questioning” events (T2: interrogation room) are extracted and sorted by timestamps to generate “first law enforcement action: [time 1], follow-up: [time 2]”.
[0078] Efficiency optimization: Graph database index traversal is accelerated, with an average response time of less than 1 second.
[0079] (2) Entity relevance is based on the extraction of subgraphs from the knowledge graph database. Starting from the matching nodes of the fusion retrieval set, it traverses the edges of the knowledge graph database, extracts relevant subgraphs, and displays the associations between entities (such as speakers, events, and objects), such as links between people, places, and events.
[0080] Data source: nodes and edges in the knowledge graph database.
[0081] Subgraph extraction and processing: Use graph database query to start from the matching node (such as the speaker node) of the fusion search set, traverse the edges, and extract related subgraphs. For example, extract cross-scene related entities from the "party" node along the "IS_SAME_AS" edge. Specifically, use depth-first search (DFS) or breadth-first search (BFS) to traverse the neighbors of the node, limiting the depth (such as 3 layers) to avoid overly large subgraphs. Using GAT Filter out low-relevance entities. LLM processes subgraph data and generates natural language descriptions, such as "person association: the same party and multiple law enforcement officers; location association: from street A to square B."
[0082] Cross-scene processing: Connect entities in different videos through the “IS_SAME_AS” edge (e.g., “party on the street - IS_SAME_AS -> party on the square”), and extract subgraphs to highlight continuity.
[0083] Example: Extract a subgraph from the knowledge graph database, including the "party" node, the "law enforcement officer" node, the "IS_SAME_AS" edge, and the "surround" edge, to generate "entity association: the same party (confirmed by voiceprint matching) and multiple law enforcement officers; event association: subsequent follow-up verification due to failure to cooperate with inquiries."
[0084] Efficiency optimization: index-accelerated subgraph extraction and privacy protection (such as encrypted voiceprint-related edges).
[0085] (3) Event summary is to summarize the search results and subgraphs using LLM to generate a concise description of the cause, development process and disposal results of the event.
[0086] Data source: Fusion of key information in the retrieval set and results of subgraph extraction.
[0087] Summarization process: LLM first builds a prompt template (such as "Generate an event summary based on the following event chain and text blocks: cause, development, result"), injects retrieval results (such as text blocks, subgraph nodes and edges), and uses LLM to generate a natural language summary.
[0088] Cross-scene processing: summaries integrate cross-video events, such as connecting street and interrogation room descriptions via “IS_SAME_AS” edges.
[0089] Example: LLM input subgraph ("Interrogation->Interrogation") and text block ( = "Relative A is missing"), generates "Cause: Dispute caused by relative A's missing; Development process: Persuasion and mediation by law enforcement personnel; Disposal result: Resolved through legal channels".
[0090] Efficiency optimization: LLM prompts template pre-optimization to reduce calculation iterations.
[0091] (4) Key findings are detecting anomalies or patterns in events, such as contradictions (inconsistencies in dialogue) and repetitions (e.g., the same party is involved in multiple cases).
[0092] Data sources: results of subgraph extraction, text blocks, and embedding similarity comparison.
[0093] Anomaly detection process: 1. Contradictory point detection: Compare the semantic consistency between text blocks and use LLM or embedding similarity to identify inconsistencies, such as "deny violation" in street transcription and "admit violation" in interrogation room. 2. Pattern recognition: Through GAT Scan the subgraph to detect repetitive patterns, such as multiple "IS_SAME_AS" edges indicating repeated involvement of a party. 3. LLM Summary Anomalies: Inject prompt templates (e.g., "Detect inconsistencies and repetitive patterns in the following subgraphs") to generate key findings.
[0094] Cross-scenario processing: Use the "IS_SAME_AS" edge to detect cross-video anomalies, such as inconsistent behavior of the same party.
[0095] Example: Detect duplicate edges in the subgraph "Party - Violation -> Traffic Rules" (via "IS_SAME_AS") and generate "Key Finding: The party has multiple similar violation records (cross-video correlation discovery), recommending enhanced monitoring."
[0096] Efficiency optimization: automatic filtering based on thresholds (e.g., similarity < 0.7 indicates a contradiction), and GAT accelerated mode scanning.
[0097] (5) Archiving recommendations are based on the strength of association and recommendation of merged archiving. Based on the strength of event association, it is recommended whether to merge videos or events into a single archive and mark high-risk entities.
[0098] Data source: results of subgraph extraction and attention coefficient of GAT.
[0099] Association strength calculation: The attention coefficient is used to quantify association strength. For example, a high attention coefficient for the "IS_SAME_AS" edge indicates a strong association. If the attention coefficient is greater than a threshold (e.g., 0.7), a merge recommendation is made (e.g., combining street and interrogation room videos into the same event). LLM generates recommendations: It injects a prompt template (e.g., "Based on association strength, recommend archiving methods") and outputs natural language recommendations.
[0100] Cross-scene processing: Calculate cross-video strength through the "IS_SAME_AS" edge. If the strength is high, it is recommended to "merge into a single case file."
[0101] For example, the attention coefficient of the subgraph "Party in the street - IS_SAME_AS -> Party in the square" is 0.8, which generates the "Archive suggestion: merge the street and square videos into a single case file and mark high-risk persons."
[0102] Efficiency optimization: threshold automation, privacy protection (such as encryption of high-risk entities).
[0103] This embodiment realizes the association between retrieval generation and voiceprint-assisted knowledge graph database construction: the retrieval generation stage directly relies on the output of the construction stage, performs multi-channel retrieval of user queries with the knowledge graph database as the core data source, and integrates the results to generate responses. This query triggering process and the connection between the construction stage support end-to-end pipeline: the output of the construction is directly fed into the retrieval function; the cross-scene voiceprint matching in the construction stage strengthens the expansion of the retrieval (such as the voiceprint retrieval set covering multiple videos), which is reflected in the report as key findings and archiving suggestions; MEnc's multi-layer attention reduces parameter redundancy, improves generalization, and achieves a recall rate greater than 95%. In the embodiment, symbols are constructed (such as "Law enforcement officers questioning people on the streets at night") is directly used for query processing and supports automatic archiving.
[0104] This embodiment realizes cross-scene semantics and speaker association: Speech Segment Node Management: Create a speech segment node for each speech segment, storing the text transcription, timestamp, emotion label (inferred using the sentiment analysis model, such as "angry" or "calm"), and voiceprint vector. Create a "generate" edge (HAS_SPEECH) from the speaker node to the speech segment node, and inject metadata such as the scene type ("street patrol").
[0105] Cross-scenario voiceprint matching: This system traverses speaker nodes from different law enforcement scenarios (based on video metadata classification) and calculates the cosine similarity between the aggregated voiceprints. If the similarity exceeds a threshold, a "similar" edge (IS_SAME_AS) is created to automatically associate the speaker with the same speaker. Node attributes are also updated to record the confidence level of the association. To handle noise, a robust matching strategy is used: the threshold is adjusted based on context (e.g., location and temporal proximity). If the similarity is between 0.7 and 0.75, a manual verification prompt is triggered.
[0106] Semantic association expansion: semantic relationship edges through knowledge graph , linking events and entities across scenarios. For example, linking the "interrogation" behavior of the same person in a street patrol video with the "confession" behavior in an interrogation room video. Constructing a cross-scenario subgraph: Starting from the matching speaker node, traverse the connected entities and relationships to form an event chain. Integrate this into a comprehensive analysis report of linked events. For example, annotate cross-scenario transitions in the event timeline, list cross-video links in entity associations, and highlight recurring patterns in key findings, such as "the same person was involved in multiple cases."
[0107] Privacy and efficiency optimization: Voiceprint data is encrypted during the association process and decrypted only during authorized queries; cross-scenario queries are accelerated using graph database indexes, with an average response time of less than 1 second.
[0108] Step S4: The task management module updates the task status to "completed", records the progress 100% and result statistics, including the generated comprehensive analysis report of related events.
[0109] If an error occurs, the task status is set to "failed", the error information is recorded, and sensitive data is cleaned.
[0110] The following two specific examples illustrate the implementation process of this system.
[0111] Example 1: Law enforcement record analysis Input: A set of body camera videos, totaling 50 hours, including street patrol and interrogation scenes.
[0112] Processing flow: Task initialization: Receive video metadata and create an encrypted temporary directory.
[0113] Knowledge graph database construction: Segment audio and video into multiple segments, sample visual frames using an intelligent sampling algorithm based on the inter-frame change rate and time interval, then determine facial posture and score them to form a list of each person's face. Only the photo with the highest posture score is stored, and subtitles, audio transcriptions, and voiceprint vectors are generated to build a knowledge graph database (entities: law enforcement officers, parties involved, relationships: interrogation).
[0114] Voiceprint analysis and association: Identify eight speakers, generate aggregated voiceprint vectors, create speaker and voice segment nodes, and associate the same person in street patrols and interrogations through voiceprint similarity.
[0115] Query processing: When a user queries "the content of the conversation between the parties during street questioning", the system retrieves relevant fragments through the knowledge graph database, generates a comprehensive response, and outputs a comprehensive analysis report of related events.
[0116] Example of a comprehensive analysis report of related events: Event chronology: First law enforcement action: [time 1], law enforcement officers handled a family dispute; subsequent law enforcement action: [time 2], handled an internal family conflict, with a time interval of approximately 20 minutes.
[0117] Entity relevance: Personnel relevance: involving the same family members such as "relative A" and the person who called the police; location relevance: [location A]; event relevance: continuous stages of internal family conflicts.
[0118] Event summary: Cause: Dispute arising from the disappearance of relative A; Development: Persuasion and mediation by law enforcement officers; Result: Resolved through legal channels without forced intervention.
[0119] Key findings: The incidents were continuous conflicts within the same family, and the person who called the police was the key instigator.
[0120] Archiving suggestion: Merge archives into the same event to provide complete context.
[0121] Results: The system achieved high accuracy on 500 queries, significantly outperforming the traditional RAG method.
[0122] Example 2: Surveillance Video Retrieval Input: A set of surveillance videos with a total length of 30 hours, including scenes of urban streets and public places, such as nighttime street patrols and incident responses in public squares.
[0123] Processing flow: Task initialization: Receive video metadata, including timestamp and location information, and create an encrypted temporary directory to store intermediate audio files and voiceprint data.
[0124] Knowledge graph database construction: Videos are segmented and visual frames are sampled using an intelligent sampling algorithm based on inter-frame change rates and time intervals (for example, sampling frames of flashing lights from nighttime street footage). Facial poses are then determined and scored to form a list of each person's face, with only the photo with the highest pose score stored (for example, a clear frontal shot of the person's face). Captions are generated (for example, "Law enforcement vehicles flash lights on a nighttime street, and law enforcement officers surround the person"), audio transcriptions (capturing conversations such as "Please show your ID card"), and voiceprint vectors are generated to build a knowledge graph database (entities: person, law enforcement officer, vehicle; relationships: surround, question).
[0125] Voiceprint Analysis and Association: Six speakers (including repeated speakers) were identified, an aggregated voiceprint vector was generated, and speaker and speech segment nodes were created. Cosine similarity (threshold 0.75) was used to correlate the voiceprints of individuals in videos of urban streets with those of the same individuals in videos of public squares, enabling cross-scene identity linking (for example, linking a "failure to cooperate with questioning" on the street with a "follow-up verification" in a public square).
[0126] Query processing: For a user query about "behavior and conversations of people on the street at night," a multimodal search is performed on the knowledge graph database: text matching extracts conversation transcripts, visual search matches flashing light scenes, and voiceprint search correlates the identities of the people involved. A comprehensive response is generated, including key fragment links, and a comprehensive analysis report of the related events is output.
[0127] Example of a comprehensive analysis report of related events: Event timeline: Initial surveillance: [Time 1], nighttime street patrol discovered the individual; subsequent response: [Time 2], follow-up verification action in the public square, approximately 15 minutes apart.
[0128] Entity Relevance: Personnel Relevance: The same party (confirmed by voiceprint matching) and multiple law enforcement officers; Location Relevance: Moved from [Street A] to [Square B]; Event Relevance: Subsequent follow-up verification resulting from failure to cooperate with inquiries.
[0129] Event summary: Cause: The person involved was stopped and checked for violating traffic regulations; Development: The person involved fled and was tracked to the square by law enforcement officers; Disposition: The person involved cooperated in verification and record-keeping.
[0130] Key findings: The person involved has multiple similar violation records (discovered through cross-video correlation), and it is recommended to strengthen monitoring.
[0131] Archiving suggestion: Combine street and square videos into a single case file and label high-risk individuals.
[0132] Results: The system accurately retrieves key segments and generates comprehensive responses and reports containing dialogue, scene description, and speaker identity, achieving a recall rate of 95% in 300 queries, outperforming the baseline method.
[0133] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A multimodal retrieval and enhanced generation system for massive law enforcement audio and video data, characterized by: include: Based on voiceprint-assisted multimodal indexing module, multimodal retrieval and generation module and knowledge graph database module, The voiceprint-assisted multimodal indexing module is used to generate subtitles, transcriptions, and voiceprint vectors based on the recorder audio and video, merge the subtitles, transcriptions, and voiceprint vectors to form a multimodal representation; use a multimodal fusion encoder to process the multimodal representation and construct a knowledge graph database; The multimodal retrieval and generation module is used to receive user queries and extract multimodal keywords; query the knowledge graph database based on the multimodal keywords to obtain a text retrieval set, a visual retrieval set, and a voiceprint retrieval set, and integrate them to form a fusion retrieval set; Based on the fusion search set, VLM and LLM are used to generate a comprehensive analysis report of related events; The knowledge graph database module is used to store and manage the knowledge graph database.
2. A multimodal retrieval enhancement generation system for massive law enforcement audio and video data according to claim 1, characterized in that: The process of forming a multimodal representation is: Split the recorder audio and video into multiple clips; Sample visual frames from each clip, input into the visual language model, and generate captions; Generate a transcription of each segment using automatic speech recognition; Perform speaker separation on each segment to generate a voiceprint vector; Merge captions, transcriptions, and voiceprint vectors to form multimodal representations.
3. A multimodal retrieval enhancement generation system for massive law enforcement audio and video data according to claim 2, characterized in that: The process of sampling visual frames is as follows: the average grayscale difference between the current frame and the previous sampled frame is calculated as the rate of change. If the rate of change exceeds the rate of change threshold and the time interval exceeds the minimum interval, the current frame is used as the sampling frame. The first frame is always sampled.
4. A multimodal retrieval enhancement generation system for massive law enforcement audio and video data according to claim 1, characterized in that: The nodes of the knowledge graph database include speaker nodes, speech segment nodes, event nodes and object nodes; Edges include identity-related edges, speech generation edges, behavior-related edges, and event-related edges.
5. A multimodal retrieval enhancement generation system for massive law enforcement audio and video data according to claim 1 or 4, characterized in that: The processing process of the multimodal fusion encoder is: Pass the subtitles and transcripts through the embedding layer to obtain the subtitle embedding and audio transcript embedding, and then pass them through the cross-attention mechanism to obtain the first layer output; The first layer output and the voiceprint vector are integrated with the identity information using the causal multi-head self-attention mechanism to obtain the second layer output; The semantic relationship between the second layer output and the knowledge graph is embedded through the graph attention network to obtain the updated node embedding and attention coefficient matrix.
6. A multimodal retrieval enhancement generation system for massive law enforcement audio and video data according to claim 1, characterized in that: The multimodal keywords include text keywords, visual descriptions, and latent voiceprint identifiers; Extract corresponding fragments from the knowledge graph database based on text keywords to form a text retrieval set; Convert visual descriptions into visual embedding vectors, extract corresponding fragments from the knowledge graph database, and generate a visual retrieval set; According to the latent voiceprint identifier, the voiceprint vector of the relevant speaker node is extracted from the knowledge graph database, and all associated segments of the voiceprint vector are extracted to form a voiceprint retrieval set.
7. A multimodal retrieval enhancement generation system for massive law enforcement audio and video data according to claim 1 or 6, characterized in that: Integrate the text retrieval set, visual retrieval set and voiceprint retrieval set, use the weighted fusion mechanism to remove redundant segments, and form a fused retrieval set.
8. A multimodal retrieval enhancement generation system for massive law enforcement audio and video data according to claim 1, characterized in that: The comprehensive analysis report of related events includes at least one of the following analysis dimensions: event chronology, entity correlation, event summary, key findings, and archiving recommendations.
9. A multimodal retrieval enhancement generation system for massive law enforcement audio and video data according to claim 8, characterized in that: Event timing is based on the timeline sorting of edges in the knowledge graph database, sorting events or fragments across scenarios in chronological order to form an event timeline; Entity relevance starts from the matching nodes of the fusion retrieval set, traverses the edges of the knowledge graph database, extracts relevant subgraphs, and displays the associations between entities; The event summary uses LLM to summarize the search results and subgraphs to generate a concise description of the event cause, development process, and disposal results; The key finding is to detect anomalies or patterns in events, which is achieved based on the results of subgraph extraction, text blocks and embedding similarity comparison; Archiving suggestions are based on the strength of event association and recommend whether to merge videos or events into a single archive.
10. A multimodal retrieval enhancement generation system for massive law enforcement audio and video data according to claim 1, characterized in that: It also includes a task management module for receiving and initializing law enforcement audio and video processing tasks and managing task status.
Citation Information
Patent Citations
Law enforcement recording method and system
CN111355912A
Apparatus and method for automatically generating and updating knowledge graph from multi-modal source
CN114270339A
Intelligent audio-visual interaction system based on large language model
CN117093744A
Multi-modal data distributed retrieval method and system based on mapping knowledge domain and vector matching
CN118551086A
Extraction system and method based on highlight video in automobile field
CN119418241A