A general scene retrieval analysis method and system based on multi-modal feature fusion

By constructing an offline feature library and using multimodal feature fusion technology, the problems of low video processing efficiency and inaccurate analysis results in existing technologies have been solved. This has enabled efficient cross-modal retrieval and dynamic knowledge-enhanced analysis, thereby improving the application capabilities of video analytics in security monitoring and smart city fields.

CN120821872BActive Publication Date: 2025-12-26SHENZHEN KAOLA YOURAN TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511325356.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-17
Publication Date
2025-12-26
Estimated Expiration
2045-09-17

AI Technical Summary

Technical Problem

Existing technologies suffer from low video processing efficiency in offline scenarios, difficulty in cross-modal feature fusion, and insufficient support for dynamic knowledge bases in multimodal feature contexts. This results in insufficient real-time performance and accuracy of analysis results, making it difficult to meet the needs of in-depth applications in fields such as security monitoring and smart cities.

Method used

By constructing an offline feature library and employing multimodal feature fusion technology and dynamic knowledge enhancement mechanisms, cross-modal video retrieval and interactive enhanced analysis are achieved, including video preprocessing, cross-modal retrieval, dynamic knowledge-enhanced question answering, and interactive enhanced analysis, thereby improving video processing efficiency and the authority of analysis results.

Benefits of technology

It effectively improves the efficiency of offline video processing, enables cross-modal feature fusion retrieval, enhances the authority of analysis results and the ability to issue early warnings of anomalies, supports the rapid location of the activity trajectory of target objects at different monitoring points, and automatically generates analysis reports that conform to industry standards.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120821872B_ABST
    Figure CN120821872B_ABST
Patent Text Reader

Abstract

The application discloses a kind of general scene retrieval analysis method and system based on multimodal feature fusion, the method includes video analysis step and application service step, the application service step includes receiving user input request, based on the video summary description and multidimensional standardization video label, can carry out cross-modal video retrieval, dynamic knowledge enhancement question and answer and interactive enhancement analysis step, the present application realizes efficient video pre-processing by constructing offline feature library, uses cross-modal feature fusion technology to improve retrieval precision, combines dynamic knowledge base to enhance analysis authority, and supports interactive enhancement analysis to realize abnormal early warning, with the advantages of improving offline video processing efficiency, realizing cross-modal feature fusion retrieval, enhancing the authority of analysis result.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of video analysis and intelligent retrieval, in particular to a general scene retrieval analysis method and system based on multi-modal feature fusion. BACKGROUND

[0002] In the multi-modal feature context, general scene retrieval analysis usually has the following problems: first, the existing technology excessively relies on online databases, resulting in low video processing efficiency in offline scenarios and failing to meet real-time requirements; second, traditional methods lack effective cross-modal feature fusion mechanisms, making it difficult to implement cross-video file correlation analysis, especially for scattered correlation event positioning in different time periods; in addition, existing systems generally lack dynamic knowledge base support, resulting in a lack of authoritative basis for video event interpretation and affecting the accuracy of analysis results. These problems seriously restrict the deep application of video analysis technology in the fields of security monitoring, smart cities, etc. The existing technical solutions usually use single modal features for retrieval, which cannot effectively fuse multi-dimensional information such as spatio-temporal parameters and semantic labels, and lack intelligent linkage mechanisms with industry standards in abnormal event early warning. In view of the above problems, the existing technology needs to be improved. SUMMARY

[0003] One of the purposes of the present application is to provide a general scene retrieval analysis method and system based on multi-modal feature fusion, which has the advantages of improving offline video processing efficiency, implementing cross-modal feature fusion retrieval, and enhancing the authority of analysis results.

[0004] One of the purposes of the present application is achieved by the following technical solutions:

[0005] The general scene retrieval analysis method based on multi-modal feature fusion comprises the following steps:

[0006] A general scene retrieval analysis method based on multi-modal feature fusion, the method comprising a video analysis step and an application service step; wherein,

[0007] The video analysis step comprises preprocessing offline video to generate structured video description information; the preprocessing comprises:

[0008] Frame extraction is performed on the video to obtain multiple frames of images;

[0009] Image understanding is performed on the multiple frames of images to generate frame-level understanding description and corresponding time stamp of each frame of image;

[0010] Based on the frame-level understanding description and time stamp of each frame of image, summary processing is performed to generate a video summary description of the entire video content;

[0011] The multi-frame images are subjected to semantic label extraction to obtain an initial label set, and a normalization cleaning algorithm is used to process the initial label set to eliminate redundancy and form a multi-dimensional standardized video label covering scene types, components and environmental parameters.

[0012] The application service step includes receiving a user input request, and based on the video summary description and the multi-dimensional standardized video label, performing at least one of the following operations:

[0013] Cross-modal video retrieval: analyzing the semantic content of the user query, matching the query semantics with the video summary description and the multi-dimensional standardized video label based on the Euclidean distance algorithm, and outputting the associated video clips and timestamps;

[0014] Dynamic knowledge-enhanced question answering: retrieving the video description information related to the user question through RAG technology, and associating the specification information in the external knowledge base to generate an enhanced answer that integrates internal and external information;

[0015] After the operation of the application service step is performed, an optional interactive enhanced analysis step is further included, which responds to the circled instruction initiated by the user in the result of the application service step, and calls a visual question answering model to perform image understanding and question answering on the specified target area, and returns the question answering result to the user.

[0016] As an optional technical solution, the cross-modal video retrieval specifically includes:

[0017] According to the user prompt word, the spatio-temporal parameters and semantic content of the user query are analyzed;

[0018] According to the time range and location label in the query, candidate frames are screened in the feature library;

[0019] The query semantic vector and the candidate frame feature vector are subjected to Euclidean distance calculation;

[0020] According to the similarity threshold, a video clip list is outputted, and the confidence and timestamp are labeled.

[0021] As an optional technical solution, the dynamic knowledge-enhanced question answering includes:

[0022] The video analysis result is associated with the external knowledge base through RAG technology;

[0023] The specification articles and historical records associated with the query event in the knowledge base are retrieved;

[0024] The video analysis result and the external specification are integrated to generate an enhanced answer.

[0025] As an optional technical solution, the interactive enhanced analysis includes:

[0026] retrieving all segments in which the target appears in the whole video library based on feature similarity;

[0027] outputting the tracking result across videos in time sequence.

[0028] The second object of the present application is to provide a general scene retrieval analysis system based on multi-modal feature fusion, which comprises:

[0029] a video analysis module for preprocessing offline videos to generate structured video description information;

[0030] an application service module for receiving user input requests and performing operations based on the structured information generated by the video analysis module, the application service module comprising:

[0031] a cross-modal video retrieval sub-module for analyzing the semantic content of a user query and matching the query semantics with the video summary description and multi-dimensional standardized video labels based on the Euclidean distance algorithm, and outputting associated video segments and timestamps;

[0032] a dynamic knowledge enhancement sub-module for retrieving video description information related to the user's question through RAG technology and associating normative information in an external knowledge base to generate an enhanced answer that integrates internal and external information;

[0033] an interactive enhancement analysis sub-module for responding to a user-initiated circled instruction and calling a visual question answering model to perform image understanding and question answering on the target area and returning the question answering result to the user.

[0034] As an alternative technical solution, the video analysis module comprises:

[0035] a frame extraction unit for extracting multiple frames of images from a video;

[0036] an image understanding unit for understanding the multiple frames of images, generating frame-level understanding descriptions and corresponding timestamps for each frame of image, and performing summary processing to generate a video summary description;

[0037] a label processing unit for performing semantic label extraction on the multiple frames of images and processing the initial label set using a normalization cleaning algorithm to form multi-dimensional standardized video labels covering scene types, components and environmental parameters.

[0038] As an alternative technical solution, the image understanding unit comprises:

[0039] a metadata analysis sub-unit for analyzing and extracting the timestamp corresponding to each frame of image from the video stream;

[0040] a feature extraction subunit configured to extract a deep feature vector of the frame image using a CLIP model;

[0041] a semantic generation subunit configured to generate the frame-level understanding description and the video summary description using a visual language model.

[0042] As an alternative technical solution, the dynamic knowledge enhancement sub-module performs the following when generating an answer:

[0043] The linkage retrieves and queries the specifications and event records related to the query.

[0044] As an alternative technical solution, the interactive enhancement analysis sub-module includes the following:

[0045] A target circumscription tool triggers enhancement analysis in response to user labeling;

[0046] An intelligent reply unit configured to invoke a visual question and answer model to perform image understanding and question and answer on the target area and return the question and answer result to the user.

[0047] A warning pushing unit configured to automatically push a warning to a management terminal when detecting a target abnormal behavior.

[0048] As an alternative technical solution, the enhancement analysis further includes the following:

[0049] A related segment collation unit configured to retrieve related segments in the full library of videos based on the target feature and output all video segments in which the target appears in time sequence.

[0050] The present application has the following advantages:

[0051] The present application realizes efficient video preprocessing by constructing an offline feature library, improves retrieval accuracy by using cross-modal feature fusion technology, enhances the authority of the analysis by combining a dynamic knowledge base, and supports interactive enhancement analysis to realize abnormal warning, thereby improving the efficiency of offline video processing, realizing cross-modal feature fusion retrieval, and enhancing the authority of the analysis result.

[0052] Other advantages, objects, and features of the present application will be apparent to those skilled in the art in view of the following detailed description, and in some aspects will be apparent from the context of the following detailed description, or can be learned from practice of the present application. BRIEF DESCRIPTION OF DRAWINGS

[0053] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings, in which:

[0054] Figure 1 The present application is a method flowchart;

[0055] Figure 2 Logic diagram for overall function implementation of the present application;

[0056] Figure 3 Logic diagram for video retrieval of the present application;

[0057] Figure 4 Logic diagram for question and answer of the present application;

[0058] Figure 5 Logic diagram for video analysis of the present application. DETAILED DESCRIPTION

[0059] The preferred embodiments of the present application will be described in detail below with reference to the accompanying drawings. It should be understood that the preferred embodiments are only for illustrating the present application, and are not intended to limit the protection scope of the present application.

[0060] In the prior art, general scene retrieval analysis has long been faced with problems of low efficiency of offline scene processing, difficulty in positioning cross-period events, and insufficient accuracy of event explanation. The traditional method relies on online databases for real-time processing, which is difficult to cope with unstable networks or large amounts of data; video retrieval mostly uses single-modal feature matching, which cannot effectively associate related events scattered in different time periods; event explanation lacks professional and normative support, resulting in limited credibility of analysis results. For example, in the security monitoring scene, when investigators need to trace the activity track of a target object in multi-day monitoring videos, the prior art cannot quickly complete offline video analysis, accurately associate cross-camera segments, and automatically reference related external norms to assist event research and judgment.

[0061] To solve the above problems, first, a scheme of constructing a localized feature library is proposed, which effectively solves the contradiction between feature extraction and storage methods; for the problem of cross-video retrieval, it is found that multi-modal feature fusion can break through the limitations of single-modal retrieval; for the problem of explanation accuracy, it is realized that a dynamic knowledge association mechanism can improve professional judgment ability. By organically combining offline feature library construction, cross-modal retrieval algorithm, and knowledge enhancement mechanism, a complete offline analysis and intelligent interaction system is formed.

[0062] Therefore, the present application proposes a general scene retrieval analysis method based on multi-modal feature fusion, as shown in Figure 1 and Figure 2 The method comprises the following steps:

[0063] The method comprises a video analysis step and an application service step; wherein,

[0064] The video analysis step comprises preprocessing offline videos to generate structured video description information; the preprocessing comprises:

[0065] Frame extraction is performed on the video to obtain multiple frames of images;

[0066] image understanding is performed on the multiple frames of images to generate frame-level understanding descriptions and corresponding timestamps for each frame of image;

[0067] Based on the frame-level understanding descriptions and timestamps of each frame of image, a summary processing is performed to generate a video summary description for the entire video content;

[0068] Semantic label extraction is performed on the multiple frames of images to obtain an initial label set, and a normalization cleaning algorithm is used to process the initial label set to eliminate redundancy and form a multi-dimensional standardized video label covering scene type, component and environmental parameter;

[0069] The application service step includes receiving a user input request, and based on the video summary description and multi-dimensional standardized video label, at least one of the following operations is performed:

[0070] Cross-modal video retrieval: analyzing the semantic content of the user query, matching the query semantics with the video summary description and multi-dimensional standardized video label based on the Euclidean distance algorithm, and outputting the associated video segment and timestamp;

[0071] Dynamic knowledge enhanced question answering: retrieving the video description information related to the user question through RAG technology, and associating the specification information in the external knowledge base to generate an enhanced answer that integrates internal and external information;

[0072] After the operation of the application service step is performed, an optional interactive enhanced analysis step is further included, which responds to the circled instruction initiated by the user in the result of the application service step, and calls a visual question answering model to perform image understanding and question answering on the specified target area, and returns the question answering result to the user.

[0073] Among them, the frame extraction in pre-processing adopts preset interval frame extraction, that is, the frame extraction frequency is dynamically adjusted according to the video content complexity, which can adopt fixed time interval or dynamic frame extraction algorithm based on scene change to ensure the balance between key frame coverage rate and processing efficiency; environmental parameters can be obtained by analyzing video metadata through OpenCV library; RAG technology integrates knowledge base retrieval results into question answering generation process through retrieval enhancement generation mechanism.

[0074] Specifically, when constructing an offline feature library (which can be understood as Figure 1The video library stores original video files, and the video analysis stores relevant feature data obtained after analysis. The information integrity and storage overhead are balanced by a preset frame extraction strategy. Image visual features, environmental parameters, and semantic labels are synchronously extracted to form a structured feature index. In the cross-modal retrieval stage, after the retrieval range is narrowed by time and space parameters, the semantic content and multi-modal features are jointly matched by the Euclidean distance algorithm to output cross-video associated segments. In the dynamic knowledge enhanced question and answer process, the RAG technology retrieves relevant regulations related to the event in real time, and the regulations are fused with the video analysis results to generate a professional report. In the interactive enhanced analysis, the user can encircle suspicious targets, cross-video tracking is achieved through feature matching, and abnormal behavior detection is performed in combination with a preset rule library.

[0075] Compared with the prior art, the traditional method relies on an online database, resulting in processing delay. The present solution realizes offline efficient processing through a local feature library. The existing single-modal retrieval cannot associate cross-period events. The present solution improves retrieval accuracy through multi-modal feature fusion. The existing question and answer system lacks specification support. The present solution realizes dynamic knowledge enhancement through RAG technology. In particular, in the cross-camera tracking scenario, the prior art requires manual comparison of multiple video streams. The present solution can automatically output all associated segments and time sequences of the target.

[0076] Through the above technical solutions, the present application effectively improves the offline scene processing efficiency, realizes multi-modal accurate retrieval across videos, and enhances the professionalism and credibility of event explanation. In the security investigation scene, the investigator can quickly locate the activity track of the target object at different monitoring points and automatically generate an analysis report that meets the industry specifications. In the equipment inspection scene, the maintenance personnel can retrieve historical similar events by inputting abnormal components and quickly judge the fault level in combination with the knowledge base specification.

[0077] For the specific implementation of each step of the foregoing method, we will further illustrate it through the following embodiments.

[0078] Embodiment One

[0079] As shown in Figure 3 , in this embodiment, the temporal and spatial parameters and semantic content of the user query are further parsed according to the user prompt words. The candidate frames are filtered in the feature library according to the time range and location label in the query. The query semantic vector and the candidate frame feature vector are calculated by the Euclidean distance. The video segment list is output according to the similarity threshold, and the confidence and timestamp are labeled.

[0080] The user prompt word analysis refers to extracting the space-time parameters and semantic content in the query through natural language processing technology. Specifically, the BERT model can be used to realize text feature vectorization, thereby accurately capturing the multi-dimensional requirements of the user query. The time range and location label screening refers to quickly locating the candidate data in the feature library according to the query conditions. Specifically, the database index technology can be used to realize the joint retrieval of space-time parameters, thereby reducing the subsequent calculation complexity. The similarity threshold filtering refers to setting a dynamic threshold to screen high-relevant results. Specifically, the statistical distribution analysis can be used to determine the optimal cutoff value to avoid the interference of low-quality fragments on the output results.

[0081] Specifically, the text query input by the user is first converted into a vector representation containing space-time parameters and semantic features. The video frame data stored in the feature library is indexed by space-time labels, so that the candidate frame screening process can quickly locate the subset that meets the time range and location conditions. In the semantic matching stage, the query vector and the candidate frame feature vector are calculated by the Euclidean distance algorithm to obtain the correlation score. The scoring results are filtered by a threshold to generate a ranking list. The final output video segment list not only contains timestamp information, but also carries a confidence index to evaluate the reliability of the retrieval results.

[0082] Compared with the prior art, the traditional video retrieval method usually only relies on a single modal feature or a fixed time window for matching, resulting in low efficiency of cross-period event positioning. The similarity calculation based on full library traversal in the prior art produces a large amount of redundant operations, and lacks a joint screening mechanism of space-time labels, making it difficult to meet the retrieval needs of large-scale video data. The present scheme significantly reduces the computational complexity by using a hierarchical retrieval strategy, first narrowing down the candidate range based on space-time labels, and then performing semantic matching calculation. At the same time, the dynamic threshold mechanism can adapt to the data distribution of different scenarios, avoiding the retrieval deviation caused by manually setting a fixed threshold.

[0083] Through the above technical solutions, the present application solves the problems of large consumption of computing resources and low relevance of results in the cross-video event positioning process. The joint screening of space-time labels effectively reduces the size of the candidate data, making the subsequent semantic matching calculation more efficient. The Euclidean distance algorithm combined with dynamic threshold filtering ensures that the output video segments not only meet the space-time constraints but also have high semantic relevance. The confidence labeling mechanism provides a quantitative basis for the reliability of the results, assisting users in quickly judging the effectiveness of the retrieval. The multi-modal feature fusion strategy realizes the event correlation analysis across video segments, meeting the precise positioning needs in complex scenarios.

[0084] Embodiment Two

[0085] As Figure 4As shown, in this embodiment, a dynamic knowledge enhanced question and answer implementation method is further proposed, including establishing an association mechanism between video analysis results and external knowledge base through retrieval enhancement generation technology, retrieving relevant regulations and historical records in the knowledge base related to the query event, and fusing the video analysis results and external regulations to generate a question and answer report.

[0086] Among them, the retrieval enhancement generation technology refers to an artificial intelligence technology combining semantic retrieval and text generation, which can be implemented by combining a pre-trained language model with a vector database, extracting relevant information from the knowledge base through semantic matching and generating structured text.

[0087] Among them, the external knowledge base refers to a standardized document storage system independent of the video feature library, which can be implemented by a hybrid architecture of relational database and document database, and is used to store industry standards, historical event records and other structured and unstructured data.

[0088] Among them, the regulations and historical records refer to industry standard files and similar event processing cases related to the query event, which can be filtered by keyword matching and semantic similarity calculation to ensure the accuracy and timeliness of knowledge reference.

[0089] Among them, the question and answer report refers to a composite output document containing video analysis conclusions and regulatory basis, which can be implemented by template filling and natural language generation technology, and the objectivity and compliance of the content are guaranteed by a double verification mechanism.

[0090] Specifically, when the user initiates an event query, the system first parses the key elements in the video analysis results, dynamically retrieves the current regulations and historical processing records related to the event from the external knowledge base through the retrieval enhancement generation technology. The retrieval process uses a semantic vector matching algorithm to calculate the similarity between the query content and the knowledge base entries, and selects the regulation texts with a relevance higher than a preset threshold. Then, the system cross- verifies the objective facts obtained by video analysis with the retrieved regulatory requirements, and integrates the verification results into a question and answer report containing fact description, regulation reference and processing suggestion through a text generation model. In this process, the update status of the knowledge base affects the retrieval results in real time, ensuring that the referenced regulation version is always the latest effective standard.

[0091] Compared with the prior art, the traditional method relies on a static knowledge base, resulting in a lag in regulation reference, and cannot dynamically associate video analysis results with external knowledge. This scheme realizes the dynamic calling of the knowledge base through the retrieval enhancement generation technology, makes the regulation text retrieval range accurately correspond to the video analysis content, and introduces historical records as auxiliary decision basis, effectively eliminating the explanation errors caused by knowledge update delay or scene understanding deviation.

[0092] By the technical solution, the application solves the problem of explanation accuracy caused by the lack of dynamic knowledge support, and uses a double verification mechanism to compare the video analysis result with external specifications in real time, ensuring that the event explanation meets both objective facts and industry standards. By dynamically associating the latest specification version, the compliance risk caused by the lagging of the knowledge base update is avoided, and the professionalism and credibility of the question and answer report are improved.

[0093] Embodiment Three

[0094] As Figure 5 shown, the figure discloses the specific steps of video analysis. Through parallel analysis tasks for each image frame, the system extracts rich semantics and information, including:

[0095] (1) Video frame extraction: the system extracts frames from the input video file at a preset time interval (such as 1 frame per second), and discretizes the continuous video stream into a series of continuous image frames.

[0096] (2) Frame-level image understanding: use large-scale visual language models (such as GPT-Vision, CLIP, BLIP-2, etc.) to perform deep semantic understanding on single-frame images, generating natural language descriptions of the frame (i.e. "frame-level image understanding description"), such as "a man in a blue shirt is approaching a white door".

[0097] (3) Environment parameter analysis: use computer vision libraries (such as OpenCV) to analyze the timestamp embedded in the video frame or metadata to accurately obtain the shooting time (date and time point) of each frame. In some embodiments, the model can also analyze the first frame or key frame to infer comprehensive environmental information (such as weather "sunny", lighting "night", location "intersection"), forming a "video environment" summary.

[0098] Semantic label annotation: use advanced visual recognition models (such as ResNet, ViT, VLM) or semantic segmentation models to identify objects, scenes and activities in the image, and generate a series of preliminary semantic labels for each frame (such as

pedestrian, blue shirt, white door, walking

[0099] (4) Video-level information integration and extraction: based on frame-level analysis, higher-level integration is performed to form a macro description of the entire video file. Including:

[0100] (4.1) Video description generation: based on the understanding description of all frames, use a time series model or a large language model for summary processing to generate a coherent "video description" that summarizes the entire video content.

[0101] (4.2) Label system construction and cleaning:

[0102] a. Tag extraction: Collect the preliminary semantic tags of all frames to form an initial tag set and its frame index.

[0103] b. Tag cleaning and normalization: The system uses a dedicated normalization cleaning algorithm to process the initial tag set. The purpose is to:

[0104] Eliminate redundancy: Merge synonyms (such as "car" and "sedan"), near-synonyms.

[0105] Hierarchical organization: Build a structured tag system that covers at least the following dimensions:

[0106] Scene type (Scene Type): such as "traffic intersection", "indoor lobby", "perimeter defense".

[0107] Object parts (Object / Parts): such as "vehicle", "face", "license plate", "tire".

[0108] Environmental parameters (Environmental Parameters): such as "visibility: low", "temperature: high temperature", "weather: rain and snow".

[0109] (4.3) Ensure accuracy and interpretability: After cleaning and normalization, the final output is an accurate and structured "video tag" set, which can greatly improve the accuracy and efficiency of subsequent retrieval.

[0110] (5) Feature library storage: Finally, all extracted and generated structured information, including image feature vectors, frame understanding descriptions, timestamps, video descriptions, video environments, cleaned video tag sets, and their corresponding relationships with frames, are stored in the local offline feature library.

[0111] Video analysis is the basis of the overall scheme of the invention, aiming to convert unstructured raw video data into a structured, machine-understandable and searchable multi-modal feature information library.

[0112] Example Four

[0113] This embodiment further proposes to receive user's circle instruction for target area in video frame, use DetectWithAttributes algorithm to extract features of the circled area, retrieve all segments where the target appears in the full library video based on feature similarity, and output the cross-video tracking result in time sequence.

[0114] The receiving of the user's circled instruction for the target region in the video frame refers to obtaining the coordinate information of the user's attention region through an interactive labeling tool. Specifically, the rectangular frame selection or free drawing method can be used to achieve this. This feature enables the analysis process to focus on a specific target region. The use of the DetectWithAttributes algorithm to extract the features of the circled region refers to simultaneously extracting the visual features and semantic attributes of the target. Specifically, the joint output of the target detection model and the attribute classification model can be used to achieve this. This feature enhances the target representation capability by generating a composite feature vector. The retrieval of all segments where the target appears in the entire library of videos based on feature similarity refers to matching the extracted feature vector with all video frames in the offline feature library based on similarity. Specifically, the Euclidean distance algorithm threshold screening method can be used to achieve this. This feature breaks through the limitations of single-video analysis and enables cross-file retrieval. The output of the cross-video tracking results in chronological order refers to sorting and reorganizing the retrieved discrete video segments based on timestamps. Specifically, the time axis alignment algorithm can be used to achieve this. This feature forms a continuous behavior trajectory to restore the complete event chain.

[0115] Specifically, after the user circulates the target region in the video frame, the coordinate information of this region is transmitted to the feature extraction module. For example, the DetectWithAttributes algorithm can be used to extract features such as color, shape, and motion state of the target based on target detection, generating a feature vector containing multi-dimensional information. This vector is input into the offline feature library for full comparison. By calculating the similarity scores with each video frame feature, candidate segments with similarity scores exceeding a predetermined threshold are selected. All eligible segments are sorted by original video timestamps, and a cross-video time series tracking report is finally generated.

[0116] Compared with the prior art, the traditional method usually relies on fixed algorithms to automatically identify targets and cannot handle user-defined attention regions. When performing cross-video retrieval, relying on only a single visual feature leads to high false detection rates. This scheme realizes human-computer collaborative analysis by introducing user circled instructions, generates a composite feature vector using the DetectWithAttributes algorithm to improve retrieval accuracy, and solves the problem of broken cross-period event association through full library retrieval and time series reorganization techniques.

[0117] Through the above technical solutions, the present application can accurately lock the user-specified target, quickly locate all spatiotemporal segments where the target appears in the offline video library, and integrate the scattered retrieval results into a continuous behavior trajectory. This scheme effectively improves the cross-video event association analysis capability and provides complete spatiotemporal evidence chain for abnormal behavior tracking.

[0118] Embodiment Five

[0119] Based on the design idea of the foregoing method, the embodiment further proposes a system comprising a video analysis module and an application service module, wherein the application service module comprises a cross-modal video retrieval sub-module, a dynamic knowledge enhancement sub-module, and an interactive enhancement analysis sub-module. The video analysis module is configured to preprocess offline videos to generate structured video description information. The application service module is configured to receive user input requests and perform operations based on the structured information generated by the video analysis module. The functions of the sub-modules included in the application service module are as follows:

[0120] The cross-modal video retrieval sub-module is configured to analyze the semantic content of a user query and match the query semantics with the video summary description and multi-dimensional standardized video labels based on the Euclidean distance algorithm, and output the associated video segments and timestamps.

[0121] The dynamic knowledge enhancement sub-module is configured to retrieve video description information related to a user question through RAG technology and associate the information with the specification information in an external knowledge base to generate an enhanced answer that integrates internal and external information.

[0122] The interactive enhancement analysis sub-module is configured to respond to a circle instruction initiated by a user and call a visual question and answer model to perform image understanding and question answering on the target area, and return the question and answer results to the user.

[0123] In the embodiment, the video analysis module refers to a unit for structuring the original video. Specifically, the video analysis module can be implemented by extracting image features using a preset time interval frame extraction combined with a CLIP model, analyzing timestamps using OpenCV to obtain environmental parameters, and generating semantic labels using a VLM model. The video analysis module is configured to build an offline feature library to eliminate the dependence on an online database. The video analysis module supports custom configuration, and the default frame extraction interval can be set. In a specific implementation, a normalization cleaning algorithm can be used to implement data standardization by uniformly labeling scene types, device components, and temperature and humidity parameters, so as to eliminate the feature differences of multi-source data.

[0124] The cross-modal video retrieval sub-module refers to a core algorithm module for cross-video retrieval. Specifically, the cross-modal video retrieval sub-module can be implemented by using a spatiotemporal feature fusion technology combined with Euclidean distance calculation, and is configured to solve the problem of cross-period event positioning difficulty.

[0125] The dynamic knowledge enhancement sub-module refers to a knowledge enhancement processing unit. Specifically, the dynamic knowledge enhancement sub-module can be implemented by using RAG technology to establish a vector association between video features and an external knowledge base, and is configured to improve the accuracy of event explanation by jointly retrieving video analysis results and specification provisions.

[0126] The interactive enhanced analysis submodule includes a man-machine cooperative analysis interface, supports multi-screen interaction, supports multi-terminal adaptation such as PC / tablet, includes providing a floating search box, dynamically loading a list and other conventional interactive components, and more importantly, providing a target delineation tool, which can be implemented by using the DetectWithAttributes algorithm to extract user-labeled region features and performing cross-video tracking, for triggering enhanced analysis and pushing early warning information.

[0127] Specifically, the video analysis module extracts video frames at fixed intervals, for example, one frame per second or five seconds, extracts image feature vectors through the CLIP model, analyzes GPS coordinates and timestamps in the video metadata as environmental parameters, and generates semantic description labels using the BERT model; the cross-modal video retrieval submodule receives user input of spatiotemporal conditions and semantic keywords, filters candidate frames that meet the time range and location labels in the feature library, concatenates the query semantic vector and the candidate frame feature vector, and calculates the Euclidean distance, and outputs the associated segment when the threshold is exceeded; the dynamic knowledge enhancement submodule synchronously retrieves video analysis results and external knowledge bases during the question and answer process, for example, associates and matches detected device abnormal states with industry safety specification clauses, and generates an analysis report containing specification references; the multi-dimensional label processing module cleans the original labels, for example, converts expressions such as "high temperature" and "hot" into standardized labels such as "temperature ≥ 35℃"; the intelligent interaction module responds to the target area delineated by the user on the video frame, extracts the color histogram and texture features of the area, and retrieves the same target in all videos based on feature similarity.

[0128] Through the above technical solutions, the present application realizes efficient processing capability of offline video, supports feature extraction and retrieval tasks in an environment without network connection; enhances the positioning accuracy of cross-period events, and can accurately identify associated video segments scattered in different time periods; establishes a dynamic knowledge association mechanism, automatically references the latest specification clauses during event analysis; unifies the label standards of multi-source data, eliminating retrieval errors caused by expression differences; provides man-machine cooperative analysis function, realizes closed-loop processing of abnormal events through target feature tracking.

[0129] Embodiment six

[0130] The embodiment is an extension and expansion of the technical content of Embodiment Five. The video processing module includes a frame extraction unit, an image understanding unit, and a label processing unit. The frame extraction unit is used to extract multiple frames of images from the video. The image understanding unit is used to understand the multiple frames of images, generate frame-level understanding descriptions and corresponding timestamps for each frame of image, and perform summary processing to generate a video summary description. The label processing unit is used to perform semantic label extraction on the multiple frames of images, and uses a normalization cleaning algorithm to process the initial label set to form a multi-dimensional standardized video label covering scene types, components, and environmental parameters.

[0131] Further, in the embodiment, the image understanding unit includes a metadata analysis subunit, a feature extraction subunit, and a semantic generation subunit. The metadata analysis subunit is used to analyze and extract the timestamp corresponding to each frame of image from the video stream, such as extracting the timestamp metadata from the video file using a computer vision library (such as OpenCV). The feature extraction subunit is used to extract the deep feature vector of the frame image using a model (such as CLIP). The semantic generation subunit is used to generate frame-level understanding descriptions and video summary descriptions using a visual language model (such as VLM).

[0132] The frame extraction unit uses a preset time interval for frame extraction, which is a frame extraction frequency dynamically adjusted according to the video content. Specifically, it can use a fixed interval or a dynamic interval strategy based on scene changes to balance data coverage and processing efficiency. The CLIP model is used to extract high-dimensional visual features of the image. OpenCV analyzes the timestamp, which means that the timestamp is associated with environmental parameters through the time information in the video metadata, to establish the mapping relationship between the space-time dimension and the physical environment. The label processing unit can use the BERT model to generate semantic labels, which is a semantic understanding model based on natural language processing. Specifically, it automatically generates standardized labels based on text descriptions to eliminate subjective bias in manual labeling.

[0133] Specifically, the frame extraction unit reduces the number of redundant frames of the video through a preset interval strategy to provide structured input for subsequent processing. In the feature extraction unit, the CLIP model extracts deep features of the image through a residual structure to solve the problem of insufficient representation of complex scenes by traditional models. OpenCV analyzes the timestamp to associate environmental parameters such as light and temperature, forming a multi-dimensional description of space-time physical properties. The BERT model performs semantic understanding on the video content to generate text labels in a unified format. The three work together to convert video data into a multi-modal feature vector containing visual features, environmental parameters, and semantic labels, and construct a standardized offline feature library.

[0134] Through the technical solution, the application realizes automatic feature extraction and standardized processing of offline video data, and solves the problem of low processing efficiency caused by dependence on an online database. A structured feature library is constructed through multi-modal feature fusion, providing a unified data basis for subsequent cross-modal retrieval, while eliminating subjective errors caused by manual annotation and improving the scalability of offline scene analysis.

[0135] Embodiment Seven

[0136] The embodiment further proposes a working mode of the cross-modal video retrieval sub-module. First, the time range and location information of the user query are preliminarily filtered in the feature library, the Euclidean distance algorithm is used to match the query semantics and video frame features, and a video segment list and corresponding time stamp sorted by confidence are returned.

[0137] Among them, the cross-modal video retrieval sub-module refers to a multi-dimensional matching system integrating text semantics and image features, which can be implemented by a multi-modal feature fusion framework, and is used for simultaneous processing of correlation analysis of different data types. The time range and location information of the user query refers to the search conditions containing the time interval and geographic coordinates, which can be implemented by a timestamp analysis module and a location label matching algorithm, and is used for quickly filtering irrelevant data in the feature library. Confidence sorting refers to prioritizing the results according to the matching similarity, which can be implemented by a threshold filtering and sorting algorithm, and is used to ensure the controllability of the quality of the output results.

[0138] Specifically, when the user inputs the query conditions containing the time interval and geographic coordinates, the system first performs preliminary filtering in the offline feature library in the time and space dimensions, such as filtering out video frame data within a specific date range and located in a specified area. Then, the text semantic vector input by the user is calculated with the image feature vector of the candidate video frame in cross-modal similarity, wherein the text semantic vector can be generated by a natural language processing model, and the image feature vector is extracted by a convolutional neural network. The vector space distance between the two is calculated by the Euclidean distance algorithm to obtain the matching score of each candidate frame. The system filters out low-score results according to the preset confidence threshold, and ranks the results that meet the conditions in descending order of score, while associating the timestamp information of the original video, forming a spatiotemporal association event chain containing multiple video segments.

[0139] Through the above technical solution, the application can effectively solve the problem of cross-video event positioning difficulty, reduce the data processing amount through the spatiotemporal filtering mechanism, and improve the accuracy of fuzzy queries by using cross-modal feature matching, finally outputting an association event chain with time sequence labeling. This scheme enables users to quickly obtain associated event segments scattered in different video files, providing reliable data support for subsequent behavior analysis and decision-making.

[0140] Embodiment Eight

[0141] The embodiment further proposes that the dynamic knowledge base module performs linkage retrieval of relevant regulations and event records in connection with the query when generating the answer.

[0142] The linkage retrieval refers to active invocation of external knowledge base resources through multi-dimensional retrieval strategies, which can be implemented by using a retrieval algorithm based on semantic correlation degree, so as to ensure that the retrieval range covers industry regulation provisions and historical event records by simultaneously matching keywords and event context correlation.

[0143] Further, a structured analysis report can be generated by combining video analysis results and knowledge base information. The structured analysis report generation refers to cross-modal alignment of video features and external knowledge, which can be implemented by using a knowledge fusion algorithm to establish the association between video behaviors and regulation provisions by spatial mapping of video feature vectors and knowledge base text vectors.

[0144] Specifically, after receiving a user query request, the dynamic knowledge base module first extracts the core elements in the query, such as event type, time range, spatial location, and other parameters, through semantic analysis. Based on these parameters, the module starts the linkage retrieval mechanism to retrieve industry regulation provisions matching the event type in the external knowledge base and record cases with similar spatio-temporal attributes in the historical database in parallel. During the retrieval process, a multi-dimensional matching strategy is adopted, for example, for the query of “wearing safety helmet at construction site”, not only the keywords such as “safety helmet” and “construction” are matched, but also the extended concepts such as “high-altitude operation” and “labor protection” are associated. After completing the retrieval, the module fuses the feature data obtained by video analysis with the retrieval results, for example, the coordinates of the person detected in the video without wearing a safety helmet are associated with Article 3.2.1 of “Technical Code for Safety in Construction”, and a structured report containing violation clause references and historical similar case comparisons is generated by using a knowledge fusion algorithm.

[0145] Compared with the prior art, the traditional video analysis system only relies on the internal feature database for event interpretation, and lacks external knowledge support, resulting in a lack of basis for the conclusion. The present scheme expands the retrieval range by combining the event context through a multi-dimensional retrieval strategy, for example, when retrieving safety helmet related regulations, it is simultaneously associated with the inspection procedures for labor protection equipment; in the report generation link, the prior art usually adopts a simple result splicing method, while the present scheme realizes the deep association between video detection data and text clauses by using cross-modal vector alignment technology, for example, the frame sequence of the person without wearing a safety helmet is mapped to the specific clause in the regulation.

[0146] By the technical solution, the application solves the problem of limited event explanation accuracy due to lack of dynamic knowledge base support, and realizes effective association of video analysis conclusion and industry specifications. Specifically, it can automatically match video behavior with relevant regulations, generate analysis conclusions with clear specification basis, and avoid omissions or misjudgments that may occur during manual comparison. At the same time, through the association analysis of historical cases, a reference disposal example is provided for event explanation, improving the credibility and practicality of the conclusion.

[0147] Embodiment Nine

[0148] The embodiment further proposes that the intelligent interaction module includes a target delineation tool, an intelligent reply unit, and a warning push unit. The target delineation tool is used to trigger enhanced analysis in response to user labeling; the intelligent reply unit is used to call a visual question and answer model to perform image understanding and question and answer on the target area, and return the question and answer result to the user; and the warning push unit is used to automatically push a warning to a management terminal when a target abnormal behavior is detected.

[0149] The trigger condition for enhanced analysis is that the user suspects some content (such as finding a suspicious target or event) after performing retrieval or viewing the question and answer result, and initiates a request through an interactive interface (such as the target delineation tool).

[0150] The execution process includes:

[0151] The system extracts the visual features of the specific target or area delineated by the user. Then, instead of simply performing similarity retrieval in the entire database, the target and its context information (image of the frame, existing labels, etc.) are submitted to a large-scale visual language model for specialized "image understanding and question and answer". The large model will perform deep reasoning and analysis based on the specific target (which may also involve the system calling the model to extract the visual features and semantic attributes of the area, generating a feature vector containing color distribution, shape contour, and object category, etc. Specifically, the vector features of the delineated area can be extracted using a specific model, and the feature extraction method of the foregoing embodiments can be referred to), answering the user's specific questions (for example: "What is this person doing?", "What is the model of this vehicle?", "Is this behavior abnormal?"). Output and response: the system returns the question and answer result of the large model to the user. At the same time, the system background will continuously monitor the results of such analysis, and if it is judged to meet the preset abnormal rules (such as identifying "illegal intrusion"), it will automatically push warning information to the management terminal.

[0152] The scheme triggers enhanced analysis through the target delineation tool triggered by the user's delineation instruction, combines manual experience with algorithm feature extraction, effectively improves the relevance of target features, and timely discovers and pushes warning signals through the warning push unit, ensuring the timeliness of monitoring.

[0153] Embodiment Ten

[0154] The embodiment further proposes that the enhanced analysis further includes an associated segment collating unit for retrieving associated segments in the full library video based on the target feature of the circled region, and outputting all video segments where the target appears in time sequence.

[0155] The full library video retrieval refers to cross-file matching based on a pre-constructed multi-modal feature library. Specifically, the feature vector comparison can be achieved by using the Euclidean distance algorithm, and the cross-video file association analysis can be achieved by using the spatiotemporal tags of the offline feature library.

[0156] The time sequence output refers to time sequence reorganization of the retrieval results. Specifically, the time stamp alignment technology can be used to achieve this. By establishing a spatiotemporal coordinate chain of the target appearance time, a complete trajectory tracking record is formed. The continuity of the behavior trajectory is strengthened through time sequence output, and structured time sequence data is provided for anomaly detection.

[0157] Specifically, when the user circulates the target region in the video frame through the interactive interface, the system calls the model to extract the visual features and semantic attributes of the region, and generates a feature vector containing information such as color distribution, shape contour, and object category (the vector features of the circled region can be extracted using the model, and the feature extraction method can refer to the aforementioned embodiments). The feature vector is input into the pre-constructed multi-modal feature library, and similarity matching is performed with the feature data of all offline video frames. In the matching process, the Euclidean distance algorithm is used to calculate the Euclidean distance between the feature vectors. The smaller the distance, the more similar it is. When the distance exceeds a predetermined threshold, the system records the timestamp and position information of the corresponding video segment. All retrieval results that meet the conditions are sorted and reorganized according to the timestamp. The target appearance time scattered in different videos is connected through the time axis alignment technology, and finally a cross-video spatiotemporal trajectory map is generated.

[0158] Compared with the prior art, the traditional video analysis system only supports single file retrieval and relies on online feature extraction, and cannot realize cross-video association analysis in offline scenarios. The present scheme completes cross-file retrieval without the need for real-time connection to the database through the pre-constructed multi-modal feature library and offline similarity matching mechanism. At the same time, interactive feature extraction is used instead of full-frame calculation, and the calculation resources are focused on the user's attention area, so that the target tracking efficiency is improved by several times. The time stamp alignment technology breaks through the traditional file-based storage retrieval mode and realizes true cross-spatiotemporal event association.

[0159] By the technical solution, the application effectively solves the problem of time-space fragmentation in cross-video event tracking, and can quickly locate the appearance record of the target object in different time periods in an offline environment. The interactive feature extraction mechanism reduces the calculation load while improving the search accuracy, and the time sequence output mode directly presents the target moving track, providing complete time-space data chain for cross-period event analysis. Through the time sequence output mechanism, the discrete search results are converted into structured behavior records, improving the integrity and explainability of abnormal event analysis.

[0160] Finally, it should be pointed out that the above description is only a preferred embodiment of the present application, and does not limit the present application in any form. Although the present application has been disclosed as above with a preferred embodiment, it is not intended to limit the present application. Any person skilled in the art can make some changes or modifications to the above disclosed technical content to obtain equivalent embodiments with equivalent changes, without departing from the scope of the technical solution of the present application. Any modification, equivalent change and modification of the above embodiments made in accordance with the technical essence of the present application still fall within the scope of the technical solution of the present application.

Claims

1.A general scene retrieval analysis method based on multi-modal feature fusion, characterized in that: The method comprises a video analysis step and an application service step; wherein The video analysis step comprises preprocessing offline videos to generate structured video description information; the preprocessing comprises: frame extraction of the video to obtain a plurality of frames of images; image understanding of the plurality of frames of images to generate frame-level understanding descriptions and corresponding time stamps of the frames of images; summary processing based on the frame-level understanding descriptions and time stamps of the frames of images to generate a video summary description of the entire video content; semantic label extraction of the plurality of frames of images to obtain an initial label set, and processing of the initial label set by a normalization cleaning algorithm to eliminate redundancy and form multi-dimensional standardized video labels covering scene types, components and environmental parameters; The application service step comprises receiving a user input request, and based on the video summary description and multi-dimensional standardized video labels, performing at least one of the following operations: cross-modal video retrieval: analyzing the semantic content of the user query, matching the query semantics with the video summary description and multi-dimensional standardized video labels based on the Euclidean distance algorithm, and outputting the associated video segments and time stamps; dynamic knowledge enhanced question answering: retrieving the video description information related to the user question through RAG technology, and associating the specification information in the external knowledge base to generate an enhanced answer integrating internal and external information; After the operation of the application service step is performed, an optional interactive enhanced analysis step is further included, which responds to the circled instruction initiated by the user in the result of the application service step, and calls a visual question answering model to perform image understanding and question answering on the specified target area, and returns the question answering result to the user. 2.The general scene retrieval analysis method based on multi-modal feature fusion according to claim 1, characterized in that: The cross-modal video retrieval specifically comprises: analyzing the spatio-temporal parameters and semantic content of the user query according to the user prompt words; filtering candidate frames in the feature library according to the time range and location label in the query; calculating the Euclidean distance between the query semantic vector and the candidate frame feature vector; outputting a video segment list according to the similarity threshold, and labeling the confidence and time stamp. 3.The general scene retrieval analysis method based on multi-modal feature fusion according to claim 1, characterized in that: The dynamic knowledge enhanced question answering comprises: associating the video analysis result with the external knowledge base through RAG technology; retrieving specification articles and historical records associated with the query event in the knowledge base; fusing the video analysis result with the external specification to generate an enhanced answer. 4.The general scene retrieval analysis method based on multi-modal feature fusion according to claim 1, characterized in that: The interactive enhanced analysis further comprises: retrieving all segments where the target appears in the entire video library based on feature similarity; outputting the cross-video tracking result in time sequence. 5.A general scene retrieval analysis system based on multi-modal feature fusion, characterized in that, The system comprises: a video analysis module for preprocessing offline videos to generate structured video description information; the video analysis module comprises: a frame extraction unit for frame extraction of the video to obtain a plurality of frames of images; an image understanding unit for understanding the plurality of frames of images to generate frame-level understanding descriptions and corresponding time stamps of the frames of images, and performing summary processing to generate a video summary description; a label processing unit for semantic label extraction of the plurality of frames of images, and processing of an initial label set by a normalization cleaning algorithm to form multi-dimensional standardized video labels covering scene types, components and environmental parameters; An application service module is configured to receive a user input request and perform an operation based on structured information generated by the video analysis module, and the application service module includes: A cross-modal video retrieval submodule is configured to parse semantic content of a user query and match query semantics with video summary descriptions and multi-dimensional standardized video tags based on an Euclidean distance algorithm, and output associated video clips and timestamps; A dynamic knowledge enhancement submodule is configured to retrieve video description information related to a user question through RAG technology and associate normative information in an external knowledge base to generate an enhanced answer that integrates internal and external information; An interactive enhancement analysis submodule is configured to respond to a user-initiated circled instruction and call a visual question and answer model to perform image understanding and question answering on a specified target area and return a question and answer result to the user. 6.The general scene retrieval analysis system based on multi-modal feature fusion according to claim 5, characterized in that: The image understanding unit includes: A metadata analysis subunit is configured to parse and extract a timestamp corresponding to each frame of an image from a video stream; A feature extraction subunit is configured to extract a deep feature vector of a frame of an image using a CLIP model; A semantic generation subunit is configured to generate a frame-level understanding description and a video summary description using a visual language model. 7.The universal scene retrieval analysis system based on multi-modal feature fusion of claim 5, wherein: The dynamic knowledge enhancement submodule performs the following when generating an answer: Retrieval and query of related norms and event records. 8.The universal scene retrieval analysis system based on multi-modal feature fusion of claim 5, wherein: The interactive enhancement analysis submodule includes: A target circled tool that triggers enhancement analysis in response to user labeling; An intelligent reply unit configured to call a visual question and answer model to perform image understanding and question answering on the target area and return a question and answer result to the user; A warning push unit that automatically pushes a warning to a management terminal when a target abnormal behavior is detected. 9.The universal scene retrieval analysis system based on multi-modal feature fusion of claim 8, wherein: The enhancement analysis further includes: An associated clip collation unit configured to retrieve associated clips in a full library of videos based on a target feature and output all video clips in which the target appears in chronological order.

Citation Information

Patent Citations

  • Knowledge base construction method based on video content reading analysis

    CN118966329A

  • Communication method and device based on multi-modal data, equipment and storage medium

    CN120455540A