Event detection method fusing dual-scale ROI semantic calibration and timing gating
Patent Information
- Application Number
- CN202611066228.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-17
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2046-07-17
AI Technical Summary
[0003]然而,目前的开放词汇视觉事件检测方法在实际监控场景中仍存在以下技术问题,首先,单帧检测中候选框区域容易出现局部相似但语义不同的误报
[0008] The event detection method fused with dual-scale ROI semantic calibration and temporal gating described above employs a dual-scale ROI mechanism that divides processing between tightly clipped ROIs and extended context ROIs. Tightly clipped ROIs retain target details consistent with the scale of the pre-labeled sample library for accurate cosine similarity matching. Extended context ROIs incorporate scene context information surrounding the candidate target and combine positive and negative concepts to complete semantic discrimination. By leveraging the negative scene semantics identified by the extended ROI (such as cutting boards, kitchens, delivery boxes, and tools), false alarms can be suppressed for scenarios with similar appearances, such as holding a knife and cutting vegetables, unpacking packages, or using tools in the kitchen. This solves the problem of misjudgment due to similar local targets but different scene semantics from the feature input level. Simultaneously, semantic uncertainty is defined based on positive and negative concept matching scores. Differentiated positive calibration and negative suppression strengths are adaptively generated from the uncertainty and positive/negative concept scores. This allows semantically ambiguous candidate ROIs to receive stronger semantic calibration, while semantically clear candidate ROIs avoid over-calibration. Furthermore, higher negative concept scores result in stronger negative suppression, replacing the traditional manual fixed threshold calibration mode and solving the problem of fixed parameters in different events. This paper addresses the issues of fluctuating calibration effects and unstable negative sample suppression capabilities under different event types and scene conditions. In the multi-frame event confirmation stage, a main frame candidate target reuse strategy is adopted. Local ROI dynamic scoring captures the temporal changes of the candidate target region itself, combined with adaptive injection of global temporal context to supplement the dynamic background information of the entire scene. Then, a gated fusion unit automatically allocates fusion weights based on the reliability of three types of evidence: static sample matching, dynamic temporal verification, and conceptual semantic discrimination. This avoids the high computational overhead of full-segment video temporal modeling and the interference of global background changes, and also compensates for the deficiency of simply counting alarms in a single frame, which cannot perceive the temporal changes of the target region. It achieves unified adaptive fusion of multi-dimensional evidence, effectively distinguishing between real dynamic events and static similar objects, without requiring manual configuration of fixed fusion weights for different event types. Furthermore, this application is built on a frozen open-vocabulary graph-text encoder architecture. Adding new event types only requires supplementing the corresponding text prompts, positive and negative concept texts, and a small number of positive and negative sample features to complete the expansion, without retraining the complete detection model. It possesses excellent engineering deployment convenience and event expansion capabilities.
Smart Images

Figure CN122574747B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of image processing and video event detection technology, and in particular to an event detection method that combines dual-scale ROI semantic calibration with temporal gating. Background Technology
[0002] With the development of computer vision technology, open-vocabulary visual event detection has been widely used in surveillance scenarios, such as the possession of dangerous equipment, the holding of sticks, the raising of banners or signs, the distribution of leaflets, and the setting off of fireworks and firecrackers—visual event scenarios that require alarms to be triggered by images or multiple frames of images. Existing visual event detection systems typically include steps such as object detection, visual feature extraction, sample library matching, or classifier decision-making. With the development of visual language models, image-text pre-trained models such as CLIP can map images and text to a unified semantic space, while open-vocabulary detection models such as Grounding DINO can detect candidate objects in images based on text cues. These models reduce the cost of training a detector separately for each event.
[0003] However, current open-vocabulary visual event detection methods still suffer from the following technical problems in real-world monitoring scenarios. First, in single-frame detection, candidate bounding box regions are prone to false positives due to local similarities but semantic differences. For example, in detecting events involving the possession of dangerous tools, scenes such as chopping vegetables in the kitchen, using a utility knife to open a package, and tool use may have local targets with similar appearances to the target event. Similarly, in detecting fireworks, streetlights, car lights, and welding sparks may resemble the target event in terms of local bright spot shapes, making them even more difficult to distinguish due to rain and fog. Simply matching the feature vectors of the bounding boxes is insufficient to effectively differentiate them using the surrounding context. Second, in existing methods, positive event text is typically only used for enhancement or classification, while negative, easily confused concepts are often handled with manual thresholds or simple rules, failing to automatically adjust the calibration intensity based on the semantic similarity between positive and negative for each candidate ROI. This can lead to over-calibration of clear samples, insufficient calibration of blurred samples, or unstable negative sample suppression.
[0004] In multi-frame video event confirmation, some methods directly perform temporal modeling on the entire video segment, resulting in high computational costs and susceptibility to changes in the global background. Other methods only count the number of alarms per frame, failing to determine whether the candidate target region itself has undergone event-related temporal changes. For events such as fireworks displays, both the semantic changes of local ROIs and the global inter-frame context are valuable, but existing methods typically do not integrate these two aspects with static sample matching and conceptual semantic judgment in a unified and adaptive manner.
[0005] Therefore, there is an urgent need for a visual event detection method that can reduce false alarms due to local similarity, improve the stability of scene discrimination with confusion between positive and negative concepts, and effectively integrate multi-frame temporal information. Summary of the Invention
[0006] Based on this, it is necessary to provide an event detection method that combines dual-scale ROI semantic calibration and temporal gating to address the aforementioned technical problems. This method can reduce false alarms due to local similarity, improve the stability of scene discrimination between positive and negative concepts, and enhance the accuracy of visual event detection.
[0007] An event detection method that fuses dual-scale ROI semantic calibration and temporal gating, the method comprising: Acquire multiple consecutive images to be detected, a set of positive concept texts corresponding to the target event, a set of negative concept texts corresponding to the target event, and a pre-labeled positive and negative sample feature library; select one frame from the multiple consecutive images as the main frame, input the main frame and the text prompt words corresponding to the target event into the text-guided target detection model, and output at least one candidate target box and the corresponding detection confidence; For each candidate bounding box, two types of regions of interest at different scales are generated: tightly clipped ROI and extended contextual ROI. A frozen image encoder is used to extract visual features corresponding to tightly cropped ROIs and visual features corresponding to extended context ROIs; a frozen text encoder is used to extract positive concept text feature sets corresponding to positive concept text sets and negative concept text feature sets corresponding to negative concept text sets. The visual features of the extended context ROI are matched with the positive concept text feature set and the negative concept text feature set respectively to obtain the corresponding positive concept score and negative concept score. The semantic uncertainty of the ROI is calculated based on the positive concept score and the negative concept score. Based on the semantic uncertainty, the positive concept score and the negative concept score, positive calibration strength and negative suppression strength are adaptively generated. The visual features of the extended context ROI are semantically calibrated based on the positive calibration strength and the negative suppression strength to obtain the calibrated extended ROI features and output the concept flow score. The visual features of the tightly clipped ROI are matched with the features in the pre-labeled positive and negative sample feature library using cosine similarity to calculate the static flow score. For reference frames other than the main frame in a multi-frame continuous image, the positions of the candidate target boxes in the main frame are reused to extract the corresponding reference ROI images, and the reference ROI features are extracted by the image encoder. The average semantic difference between the ROI features of the main frame and the ROI features of each reference frame is calculated, and the local ROI dynamic score is obtained by mapping. At the same time, the whole image features of all frames are temporally encoded to obtain global temporal context features. The context injection strength is adaptively adjusted according to the inter-frame dynamic degree of all frames, and the global temporal context features are injected into the visual features of the tightly cropped ROI in the main frame to obtain temporally enhanced ROI features. The global dynamic flow score is calculated by performing cosine similarity matching between the temporally enhanced ROI features and the features in the pre-labeled positive and negative sample feature library. The static flow score, local ROI dynamic score, global dynamic flow score, and conceptual flow score are input into the gating fusion unit. The weighting weights corresponding to each score are adaptively calculated and weighted fusion is performed to obtain the final fused alarm score. When the fused alarm score meets the preset judgment conditions, the corresponding event alarm result is output.
[0008] The event detection method fused with dual-scale ROI semantic calibration and temporal gating described above employs a dual-scale ROI mechanism that divides processing between tightly clipped ROIs and extended context ROIs. Tightly clipped ROIs retain target details consistent with the scale of the pre-labeled sample library for accurate cosine similarity matching. Extended context ROIs incorporate scene context information surrounding the candidate target and combine positive and negative concepts to complete semantic discrimination. By leveraging the negative scene semantics identified by the extended ROI (such as cutting boards, kitchens, delivery boxes, and tools), false alarms can be suppressed for scenarios with similar appearances, such as holding a knife and cutting vegetables, unpacking packages, or using tools in the kitchen. This solves the problem of misjudgment due to similar local targets but different scene semantics from the feature input level. Simultaneously, semantic uncertainty is defined based on positive and negative concept matching scores. Differentiated positive calibration and negative suppression strengths are adaptively generated from the uncertainty and positive / negative concept scores. This allows semantically ambiguous candidate ROIs to receive stronger semantic calibration, while semantically clear candidate ROIs avoid over-calibration. Furthermore, higher negative concept scores result in stronger negative suppression, replacing the traditional manual fixed threshold calibration mode and solving the problem of fixed parameters in different events. This paper addresses the issues of fluctuating calibration effects and unstable negative sample suppression capabilities under different event types and scene conditions. In the multi-frame event confirmation stage, a main frame candidate target reuse strategy is adopted. Local ROI dynamic scoring captures the temporal changes of the candidate target region itself, combined with adaptive injection of global temporal context to supplement the dynamic background information of the entire scene. Then, a gated fusion unit automatically allocates fusion weights based on the reliability of three types of evidence: static sample matching, dynamic temporal verification, and conceptual semantic discrimination. This avoids the high computational overhead of full-segment video temporal modeling and the interference of global background changes, and also compensates for the deficiency of simply counting alarms in a single frame, which cannot perceive the temporal changes of the target region. It achieves unified adaptive fusion of multi-dimensional evidence, effectively distinguishing between real dynamic events and static similar objects, without requiring manual configuration of fixed fusion weights for different event types. Furthermore, this application is built on a frozen open-vocabulary graph-text encoder architecture. Adding new event types only requires supplementing the corresponding text prompts, positive and negative concept texts, and a small number of positive and negative sample features to complete the expansion, without retraining the complete detection model. It possesses excellent engineering deployment convenience and event expansion capabilities. Attached Figure Description
[0009] Figure 1 This is a flowchart illustrating an event detection method that combines dual-scale ROI semantic calibration with temporal gating in one embodiment. Figure 2 This is a schematic diagram of a single-frame detection process in one embodiment; Figure 3 This is a schematic diagram of a multi-frame process in one embodiment. Detailed Implementation
[0010] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0011] In one embodiment, such as Figure 1 As shown, an event detection method integrating dual-scale ROI semantic calibration and temporal gating is provided, including the following steps: Step 1: Obtain the multi-frame continuous images to be detected, the set of positive concept texts corresponding to the target event, the set of negative concept texts, and the pre-labeled positive and negative sample feature library; select one frame from the multi-frame continuous images as the main frame, input the text prompt words corresponding to the target event into the text-guided target detection model, and output at least one candidate target box and the corresponding detection confidence.
[0012] The text prompts include event descriptions in Chinese or English, such as "person holding an aknife," "a person holding a dangerous instrument," and "fireworks sparks." The text-guided target detection model can be GroundingDINO, OWL-ViT, GLIP, or other models with open-vocabulary localization capabilities. A pre-labeled positive and negative sample feature library stores feature vectors extracted from high-quality pre-labeled positive and negative sample images using a frozen image encoder. By inputting the main frame into the detection model, candidate bounding boxes and their detection confidence scores are obtained, providing a foundation for subsequent ROI extraction and event determination.
[0013] Step 2: For each candidate bounding box, generate two types of regions of interest at different scales: tightly clipped ROI and extended contextual ROI.
[0014] The compactly cropped ROI is obtained by adding a preset small margin to the boundary of the candidate target box, which is used to maintain size alignment with the pre-annotated positive and negative sample feature library. The extended context ROI is obtained by extending the boundary of the candidate target box outward according to a preset ratio, which is used to preserve the scene context information around the candidate target. By dividing the work into two scales of ROIs, both the target details consistent with the sample library are preserved, and semantic judgment is performed using the surrounding scene, which helps to reduce false positives of local similarity. The expansion ratio can be a fixed ratio, or it can be adaptively determined according to the area of the candidate target box, the target category, the image resolution, or the detection confidence.
[0015] Step 3: Use a frozen image encoder to extract the visual features corresponding to the tightly cropped ROI and the extended context ROI respectively; use a frozen text encoder to extract the positive concept text feature set corresponding to the positive concept text set and the negative concept text feature set corresponding to the negative concept text set respectively.
[0016] Image and text encoders can employ CLIPViT-L / 14, or alternatively, SigLIP, EVA-CLIP, OpenCLIP, Qwen-VL-Embedding, or other image-text aligned visual encoders. Using a frozen encoder ensures consistency in the feature space, reducing deployment computational overhead and training costs. Positive concept texts describe the event itself or key components of the event, while negative concept texts describe easily confused security scenarios, similar objects, or similar lighting phenomena. By extracting features through a frozen encoder, visual and textual features are ensured to reside in the same aligned semantic space.
[0017] Step 4: Perform similarity matching between the visual features of the extended context ROI and the positive concept text feature set and the negative concept text feature set, respectively, to obtain the corresponding positive concept score and negative concept score; calculate the semantic uncertainty of the ROI based on the positive concept score and negative concept score; adaptively generate positive calibration strength and negative suppression strength based on the semantic uncertainty, positive concept score and negative concept score; perform semantic calibration on the visual features of the extended context ROI based on the positive calibration strength and negative suppression strength to obtain the calibrated extended ROI features, and output the concept flow score.
[0018] The cosine similarity between the visual features of the extended context ROI and the textual features of all positive and negative concepts is calculated separately, and the maximum value is taken as the positive concept score and negative concept score of the ROI. Semantic uncertainty is used to quantify the semantic ambiguity of the current ROI. The closer the positive and negative scores are, the higher the uncertainty, corresponding to a stronger calibration intensity; conversely, the lower the uncertainty, the less over-calibration is needed. The semantic direction of the visual features is adjusted by adaptively generated calibration intensity, so that ambiguous but positive ROIs are moderately closer to the event semantic space, and ROIs that are obviously biased towards negative and confusing scenarios are moved away from the event semantic space, avoiding over-calibration of already clear ROIs. The output concept flow score is used for subsequent multi-stream fusion decision.
[0019] Step 5: Perform cosine similarity matching between the visual features of the tightly clipped ROI and the features in the pre-labeled positive and negative sample feature library to calculate the static flow score.
[0020] The pre-annotated positive and negative sample feature library is composed of high-quality positive and negative sample images with manual annotation, whose features are extracted using the same image encoder. The tightly cropped ROIs use the same cropping scale as the sample library, ensuring the accuracy of similarity matching. The static flow score reflects the similarity between candidate targets and known event samples in a single frame, and is one of the core criteria for single-frame detection. When the input is a single-frame image, the single-frame event detection results can be directly output by combining the concept safety rejection results, including candidate boxes, detection confidence, positive and negative sample similarity, positive and negative concept scores, uncertainty, calibration strength, alarm status, and alarm reason.
[0021] Step 6: For the reference frames other than the main frame in a multi-frame continuous image, the positions of the candidate target boxes in the main frame are reused to extract the corresponding reference ROI images, and the reference ROI features are extracted by the image encoder; the average semantic difference between the main frame ROI features and the ROI features of each reference frame is calculated, and the local ROI dynamic score is obtained by mapping; at the same time, the whole image features of all frames are temporally encoded to obtain global temporal context features, and the context injection strength is adaptively adjusted according to the inter-frame dynamic degree of all frames. The global temporal context features are injected into the visual features of the tightly cropped ROI in the main frame to obtain temporally enhanced ROI features.
[0022] This step reuses the candidate bounding box positions from the main frame to reduce the overhead of repeated detection across multiple frames and maintain consistency in the identity of candidate targets. The average semantic difference can be calculated using the mean of cosine distances; a larger difference indicates more drastic changes in the target region and a higher local ROI dynamic score. When temporally encoding the full-image features of all frames, the inter-frame cosine similarity matrix is first calculated, and a relative position penalty term is introduced to attenuate attention in distant frames. Temperature-scaled softmax is used to obtain self-attention weights, which are then weighted and aggregated, and the context vector corresponding to the main frame is taken as the global temporal context feature. Reference frames can be used to extract ROIs at the same location or after light position correction, eliminating the need for repeated target detection and significantly reducing multi-frame computational overhead. The local ROI dynamic score captures the temporal changes within the candidate target region itself, reflecting the degree of semantic fluctuation in the target region. The global temporal context feature is obtained through temporal encoding of the entire image sequence, carrying the scene dynamic information of the entire video. The injection strength adaptively adjusts with the inter-frame dynamics; the more drastic the scene changes, the higher the injection strength, ensuring that the injected ROI features possess both local discriminative power and temporal consistency.
[0023] Step 7: Perform cosine similarity matching between the temporal enhanced ROI features and the features in the pre-labeled positive and negative sample feature library to calculate the global dynamic flow score.
[0024] The temporal enhanced ROI feature integrates global temporal information from multiple frames. The global dynamic stream score obtained based on this feature can reflect the degree of event matching in the temporal dimension, complementing the static stream score and used to distinguish between real dynamic events and static similar objects.
[0025] Step 8: Input the static flow score, local ROI dynamic score, global dynamic flow score, and conceptual flow score into the gating fusion unit, adaptively calculate the weighting weights corresponding to each score, and perform weighted fusion to obtain the final fused alarm score; when the fused alarm score meets the preset judgment conditions, output the corresponding event alarm result.
[0026] The gated fusion unit automatically assigns weights based on the scores of each stream. Streams with higher scores generally indicate more reliable evidence, and their corresponding weights increase accordingly. Compared to fixed static, dynamic, and conceptual weights, this method can automatically adjust the evidence sources based on different event types and scenarios. For example, dynamic streams may have higher weights in fireworks and firecracker events, while static sample matching and conceptual context may have higher weights in scenarios involving dangerous equipment. When the number of input frames, the number of static candidates, and the fusion alarm score all meet preset conditions, a multi-frame event alarm is output; otherwise, the timing confirmation condition is not met.
[0027] In the above method, step 3 generates compactly clipped ROIs and extended contextual ROIs, which are then processed separately. The compactly clipped ROIs are used in step 8 for similarity matching with the pre-labeled sample library, preserving target details consistent with the sample library. The extended contextual ROIs are used in step 6 for similarity matching with the positive and negative concept text feature sets to obtain positive concept scores. and negative concept scores This allows for semantic judgment using the surrounding context of the target. Based on this, a one-way safety veto is implemented according to the negative concept score and the difference between positive and negative scores: when the extended context ROI clearly points to a negative safety scenario—the negative concept score is higher than a preset threshold and the difference between positive and negative scores is lower than a preset interval threshold—the static judgment result is directly set to non-alarm. This effectively suppresses candidate boxes that are only similar in appearance (e.g., holding a dangerous tool) but are in a harmless context (e.g., chopping vegetables in a kitchen). This solves the problem that simply matching the feature vector similarity of the target box makes it difficult to effectively distinguish candidate boxes using the surrounding context, thus achieving the beneficial effect of reducing false positives due to local similarity.
[0028] Furthermore, an uncertainty-based adaptive calibration mechanism is introduced in step 7, firstly based on the positive concept score obtained in step 6. and negative concept scores The semantic uncertainty is calculated, which quantitatively characterizes the degree of ambiguity between the positive and negative semantics of the current ROI; then, based on the uncertainty... , and Adaptive generation of positive calibration and negative suppression intensities ensures stronger positive calibration for blurred ROIs, stronger negative suppression for ROIs with high negative concept scores, and weaker calibration for sharp ROIs. This avoids the failure of manually fixed parameters under different events and scene conditions. The calibrated extended ROI features are used for subsequent judgment. This solves the problems of over-calibration of sharp samples, insufficient calibration of blurred samples, and unstable negative sample suppression, achieving the beneficial effect of improving the stability of positive and negative concept confusion scene discrimination.
[0029] Furthermore, by using positive and negative concept texts as input, step 8 utilizes a pre-labeled positive and negative sample feature library for matching, eliminating the need to retrain the object detection model or image encoder for each event. When a new event type needs to be added, only the corresponding text prompts, positive and negative concept texts, and a small number of positive and negative sample features need to be added to complete the expansion, achieving the beneficial effects of reducing deployment costs and improving expansion flexibility.
[0030] In one embodiment, the tight-clipping ROI is obtained by adding a preset small margin to the boundary of the candidate target box, which is used to maintain size alignment with the pre-labeled positive and negative sample feature library; the extended context ROI is obtained by extending the boundary of the candidate target box outward according to a preset ratio, which is used to preserve the scene context information around the candidate target.
[0031] Specifically, such as Figure 2 As shown, in the single-frame detection process, a compact cropped ROI and an extended context ROI are generated for each candidate bounding box. The compact cropped ROI has a smaller padding value, ensuring that the target object occupies the main area of the image, facilitating alignment and comparison with features in the pre-labeled sample library. The extended context ROI can be expanded by 1.5 times or 2 times, which can be adaptively adjusted according to the area of the candidate bounding box. Through dual-scale division of labor, the compact cropped ROI is used for sample matching, and the extended context ROI is used for semantic concept judgment, avoiding the dilemma of a single ROI needing to both preserve target details and include context, effectively reducing false positives due to local similarity.
[0032] In one embodiment, given a set of positive concept texts and negative concept text set The corresponding feature matrix is obtained by freezing the text encoder and then performing L2 normalization. and For the detected target ROI, extended cropping is used to extract scene context features, which are then normalized using L2 to obtain... Calculate the cosine similarity between the ROI and all positive and negative concept text features, and take the maximum value as the concept matching score for that ROI. and The calibration intensity is automatically calculated based on the positive and negative scores of each ROI. The semantic uncertainty is then calculated as follows:
[0033] in, This represents semantic uncertainty, with a value range of [0,1]. The positive concept score is the maximum cosine similarity between the extended contextual ROI visual features and the set of positive concept text features. The negative concept score is the maximum cosine similarity between the extended context ROI visual features and the negative concept text feature set; m is the positive and negative concept boundary, satisfying... .
[0034] Specifically, define uncertainty measure ∈[0,1], when positive and negative fractions are close ≈1 indicates that the current ROI semantics are ambiguous and require strong correction; when the difference is large... ≈0 indicates that the current ROI has clear semantic distinction and requires no correction. Here, m represents the boundary between positive and negative concepts; when m>0, the visual features are closer to a positive concept, and vice versa. This quantification method can accurately measure the semantic confusion level of each ROI, providing a quantitative basis for subsequent adaptive calibration.
[0035] In one embodiment, the calculation processes for the positive calibration intensity and the negative suppression intensity are as follows:
[0036]
[0037] in, The negative suppression intensity For positive calibration intensity, To calibrate the total strength coefficient, For semantic uncertainty, For positive concept fractions, The score is a negative concept score.
[0038] Specifically, The overall calibration strength is determined by preset global calibration coefficients. In confusing scenarios, negative concept scores are high, so negative calibration is automatically strengthened to suppress false detections; in ambiguous scenarios, uncertainty is high, so positive calibration is automatically strengthened to widen the positive-negative boundary. When the positive concept score itself is zero, the positive calibration strength is naturally zero, avoiding over-calibration of irrelevant ROIs. The two sets of calibration coefficients are adapted to the different logics of negative suppression and positive enhancement, respectively, to achieve differentiated calibration per ROI.
[0039] In one embodiment, the visual features of the extended context ROI are semantically calibrated based on the positive calibration intensity and the negative suppression intensity to obtain calibrated extended ROI features, including: The mean vectors of the positive concept text feature set and the negative concept text feature set are calculated separately. After L2 normalization, the positive concept anchor vector and the negative concept anchor vector are obtained. Feature calibration is completed by superimposing the positive concept anchor direction and suppressing the negative concept anchor direction. The ROI features are then expanded as follows:
[0040] in, To extend ROI features after calibration, To expand the original visual features of the contextual ROI, For positive calibration intensity, The negative suppression intensity For positive concept anchor vectors, For negative concept anchor vectors, , .
[0041] Specifically, the positive concept anchor vector is obtained by L2 normalizing the mean of the positive concept text feature set, and the negative concept anchor vector is obtained by L2 normalizing the mean of the negative concept text feature set, representing the semantic centers of positive events and negative confusion scenarios, respectively. The calibration process is achieved by superimposing features in the positive anchor direction and subtracting features in the negative anchor direction on the original visual features, causing the feature vectors to shift towards positive events and away from negative confusion scenarios in the semantic space, while avoiding over-calibration of already clear ROIs.
[0042] In one embodiment, temporal encoding is performed on the whole-frame graph features to obtain global temporal context features, including: For the whole image features of all frames, calculate the inter-frame cosine similarity matrix, introduce a relative position penalty term to attenuate and correct the similarity matrix, and obtain the inter-frame self-attention weights through temperature scaling softmax operation. The whole image features of each frame are weighted and aggregated through the self-attention weights to obtain the global temporal context features. The process of introducing a relative position penalty term to attenuate and correct the similarity matrix is as follows:
[0043] In the formula, This is the corrected similarity matrix. This is the original inter-frame cosine similarity matrix. This is the relative position penalty coefficient. Let be the inter-frame distance matrix, and let the first frame be the distance matrix. Line number The element value of the column is the first element. Frame and the The absolute value of the difference between the frame sequence numbers of the frames.
[0044] like Figure 3 As shown, given continuous T Frame Image Extracting each frame using a frozen image encoder D 3D global features and L2 normalization constitute Only for the main frame Obtained through target detection N Each candidate region is used to extract local features and normalize them to form a... The goal of temporal context injection is to fuse global temporal dynamic information from multiple frames into... V In this way, the injected features possess both local discriminative power and temporal consistency.
[0045] A multi-frame global feature sequence is constructed, and the inter-frame cosine similarity matrix S is calculated. Introducing a relative position penalty applies a temporal locality prior, attenuating the attention weights of distant frames, which aligns with the objective law of temporal locality in video. The corrected similarity matrix is then subjected to temperature-scaled softmax to obtain self-attention weights. Weighted aggregation yields the temporal features corresponding to each frame. The temporal vector corresponding to the main frame is taken as the global temporal context vector, where the temperature parameter can adopt the standard temperature of the CLIP model. The specific process is as follows: Then, a temporal locality prior is introduced, and a relative position penalty is applied to reduce attention in far frames:
[0046] Temperature-scaled softmax is applied along the rows to obtain self-attention weights, which are then weighted and aggregated.
[0047] in This is the standard temperature for the CLIP model. (Take...) As a global time-series context vector.
[0048] In one embodiment, the process of injecting global temporal context features into the visual features of the tightly cropped ROI in the main frame includes: The cross-attention weights between the visual features of each tightly cropped ROI in the main frame and the global temporal context features are calculated to generate the corresponding context feature matrix. A global dynamic score is obtained based on the original inter-frame cosine similarity matrix. The effective injection weights are adaptively adjusted according to the temporal dynamic score. The context feature matrix is multiplied by the effective injection weights and then superimposed onto the original tightly cropped ROI visual features. After normalization, the temporally enhanced ROI features are obtained.
[0049]
[0050] in, To effectively inject weights, Inject weights based on the base. This is a time-series dynamic fraction, with a value range of [0,1]. To enhance the ROI feature matrix for time series, The original visual feature matrix of the ROI tightly cropped in the main frame. For context feature matrix, For global dynamic scores, The total number of frames. Indicates the first i Frame and the j Cosine similarity of global image features in a frame.
[0051] Specifically, the cross-attention between the main frame ROI target features and the temporal context is calculated, and the context matrix is generated through outer product: .
[0052] Global dynamic score This reflects the drastic changes in the overall video: when all frames are highly similar, Approaching 0; when the inter-frame difference is large, Approaching 1. Videos with drastic changes require stronger temporal context injection; therefore, the effective injection weight is increased by 0.15+ on top of the base weight. The scaling factor is 0.15, which is a fixed bias constant to ensure that there is a minimum basic injection strength even when the global dynamic score is 0. Finally, the context features are superimposed on the original ROI features with effective weights and normalized to obtain the temporally enhanced ROI features, so that the subsequent dynamic flow score can perceive the global temporal consistency.
[0053] In one embodiment, the adaptive weighted fusion process of the gated fusion unit is as follows:
[0054] in, For the first i The weighted weights corresponding to the path fractions. For the first i The input score for the path, For gated temperature parameters, The final alarm score is determined by merging the alarms.
[0055] Specifically, The gating temperature parameter controls the sharpness of the weight distribution. The three scores are normalized using softmax to obtain the corresponding weights. The weights are then multiplied by their corresponding scores and summed to obtain the final fused score. This gating mechanism automatically identifies the most reliable evidence source in the current scenario and dynamically adjusts the contribution ratio of each stream, eliminating the need to manually configure fixed fusion weights for each type of event and adapting to the different characteristics of various events.
[0056] In one embodiment, the method further includes: When the negative concept score corresponding to the calibrated extended ROI feature is higher than the preset threshold and the difference between the positive concept score and the negative concept score is lower than the preset interval threshold, the static judgment result of the corresponding candidate target is directly set to a non-alarm state; when the positive concept score is dominant, the concept judgment result is not used as an alarm basis alone, but participates in the subsequent fusion judgment together with the static flow score.
[0057] Specifically, this mechanism is a one-way security veto mechanism, which is triggered only when the extended context ROI clearly points to a negative security scenario and the interval between positive and negative concepts is insufficient, directly excluding false alarms with high confidence. The positive concept judgment result is not used as the sole basis for alarm confirmation, but must be combined with the sample matching result for joint judgment, which can strictly control the false alarm rate and adapt to the high reliability requirements of security alarm scenarios.
[0058] In one embodiment, the cosine similarity matching process includes: Calculate the cosine similarity between the feature to be matched and all features in the pre-labeled positive sample library, and select the features with the highest similarity. k The mean of each sample is used as the positive sample matching score; the cosine similarity between the feature to be matched and all features in the pre-labeled negative sample library is calculated, and the top features with the highest similarity are selected. k The mean of each sample is used as the negative sample matching score; the judgment score of the corresponding stream is output based on the positive sample matching score and the negative sample matching score.
[0059] Specifically, the judgment score can be output using threshold, margin, softmax, or weighted margin modes. Using the top-k mean method can reduce the impact of single-sample bias on the matching results and improve the stability of the matching results; combining the matching scores of positive and negative samples can further distinguish the target from easily confused samples and improve detection accuracy.
[0060] In one alternative implementation, the text-guided object detection model is not limited to GroundingDINO, but may also employ other open vocabulary detection, phrase localization, or visual grounding models; the image encoder and text encoder are not limited to CLIPViT-L / 14, but may also employ other image-text alignment models, and the feature dimension is not limited to 768 dimensions; positive and negative concept text may be generated by manual configuration, sample statistics, knowledge base, or large language model.
[0061] It should be understood that, although Figure 1 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figure 1 At least some of the steps in the process may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.
[0062] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0063] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these modifications and improvements all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. An event detection method that fuses dual-scale ROI semantic calibration and temporal gating, characterized in that, The method includes: Acquire multiple consecutive images to be detected, a set of positive concept texts corresponding to the target event, a set of negative concept texts corresponding to the target event, and a pre-labeled positive and negative sample feature library; select one frame from the multiple consecutive images as the main frame, input the main frame and the text prompt words corresponding to the target event into the text-guided target detection model, and output at least one candidate target box and the corresponding detection confidence; For each candidate bounding box, two types of regions of interest at different scales are generated: tightly clipped ROI and extended contextual ROI. A frozen image encoder is used to extract the visual features corresponding to the tightly cropped ROI and the visual features corresponding to the extended context ROI, respectively; a frozen text encoder is used to extract the positive concept text feature set corresponding to the positive concept text set and the negative concept text feature set corresponding to the negative concept text set, respectively. The visual features of the extended context ROI are matched with the positive concept text feature set and the negative concept text feature set for similarity to obtain corresponding positive concept scores and negative concept scores. The semantic uncertainty of the ROI is calculated based on the positive concept scores and the negative concept scores. Positive calibration strength and negative suppression strength are adaptively generated based on the semantic uncertainty, the positive concept scores and the negative concept scores. The visual features of the extended context ROI are semantically calibrated based on the positive calibration strength and the negative suppression strength to obtain calibrated extended ROI features, and the concept flow score is output. The visual features of the tightly clipped ROI are matched with the features in the pre-labeled positive and negative sample feature library using cosine similarity to calculate the static flow score. For the reference frames other than the main frame in the multi-frame continuous images, the positions of the candidate target boxes in the main frame are reused to extract the corresponding reference ROI images, and the reference ROI features are extracted by the image encoder; the average semantic difference between the main frame ROI features and the ROI features of each reference frame is calculated, and the local ROI dynamic score is obtained by mapping; at the same time, the whole image features of all frames are temporally encoded to obtain global temporal context features, and the context injection strength is adaptively adjusted according to the inter-frame dynamic degree of all frames. The global temporal context features are injected into the visual features of the tightly cropped ROI in the main frame to obtain temporally enhanced ROI features. The global dynamic flow score is calculated by performing cosine similarity matching between the temporally enhanced ROI features and the features in the pre-labeled positive and negative sample feature library. The static flow score, the local ROI dynamic score, the global dynamic flow score, and the conceptual flow score are input into the gating fusion unit. The weighting weights corresponding to each score are adaptively calculated and weighted fusion is performed to obtain the final fused alarm score. When the fused alarm score meets the preset judgment conditions, the corresponding event alarm result is output.
2. The method according to claim 1, characterized in that, The tightly clipped ROI is obtained by adding a preset small margin to the boundary of the candidate target box, which is used to maintain size alignment with the pre-labeled positive and negative sample feature library; The extended context ROI is obtained by extending and cropping the candidate target box outwards according to a preset ratio based on the boundary of the candidate target, and is used to preserve the scene context information around the candidate target.
3. The method according to claim 1, characterized in that, The calculation process for the semantic uncertainty is as follows: in, This represents semantic uncertainty, with a value range of [0,1]. The positive concept score is the maximum cosine similarity between the extended contextual ROI visual features and the set of positive concept text features. The negative concept score is the maximum cosine similarity between the extended context ROI visual features and the negative concept text feature set; m is the positive and negative concept boundary, satisfying... .
4. The method according to claim 1, characterized in that, The calculation processes for the positive calibration intensity and the negative suppression intensity are as follows: in, The negative suppression intensity For positive calibration intensity, To calibrate the total strength coefficient, For semantic uncertainty, For positive concept fractions, The score is a negative concept score.
5. The method according to claim 1, characterized in that, Based on the positive calibration intensity and the negative suppression intensity, the visual features of the extended context ROI are semantically calibrated to obtain calibrated extended ROI features, including: The mean vectors of the positive concept text feature set and the negative concept text feature set are calculated separately. After L2 normalization, the positive concept anchor vector and the negative concept anchor vector are obtained. Feature calibration is completed by superimposing the positive concept anchor direction and suppressing the negative concept anchor direction. The ROI features are then expanded as follows: in, To extend ROI features after calibration, To expand the original visual features of the contextual ROI, For positive calibration intensity, The negative suppression intensity For positive concept anchor vectors, It is a negative concept anchor vector.
6. The method according to claim 1, characterized in that, Temporal encoding is performed on the whole-frame features of all frames to obtain global temporal context features, including: For the whole image features of all frames, calculate the inter-frame cosine similarity matrix, introduce a relative position penalty term to attenuate and correct the similarity matrix, and obtain the inter-frame self-attention weights through temperature scaling softmax operation. The whole image features of each frame are weighted and aggregated through the self-attention weights to obtain the global temporal context features. The process of introducing a relative position penalty term to attenuate and correct the similarity matrix is as follows: In the formula, This is the corrected similarity matrix. This is the original inter-frame cosine similarity matrix. This is the relative position penalty coefficient. Let be the inter-frame distance matrix, and let the first frame be the distance matrix. Line 1 The element value of the column is the first element. Frame and the The absolute value of the difference between the frame sequence numbers of the frames.
7. The method according to claim 6, characterized in that, The process of injecting global temporal context features into the visual features of the tightly clipped ROI in the main frame includes: The cross-attention weights between the visual features of each tightly cropped ROI in the main frame and the global temporal context features are calculated to generate the corresponding context feature matrix. A global dynamic score is obtained based on the original inter-frame cosine similarity matrix. The effective injection weights are adaptively adjusted according to the temporal dynamic score. The context feature matrix is multiplied by the effective injection weights and then superimposed onto the original tightly cropped ROI visual features. After normalization, the temporally enhanced ROI features are obtained. in, To effectively inject weights, Inject weights based on the base. This is a time-series dynamic fraction, with a value range of [0,1]. To enhance the ROI feature matrix for time series, The original visual feature matrix of the ROI tightly clipped in the main frame. For context feature matrix, For global dynamic scores, The total number of frames. Indicates the first i Frame and the j Cosine similarity of global image features in a frame.
8. The method according to claim 1, characterized in that, The adaptive weighted fusion process of the gated fusion unit is as follows: in, For the first i The weighted weights corresponding to the path fractions. For the first i The input score for the path, For gated temperature parameters, The final alarm score is determined by merging the alarms.
9. The method according to claim 1, characterized in that, The method further includes: When the negative concept score corresponding to the calibrated extended ROI feature is higher than the preset threshold and the difference between the positive concept score and the negative concept score is lower than the preset interval threshold, the static judgment result of the corresponding candidate target is directly set to a non-alarm state; when the positive concept score is dominant, the concept judgment result is not used as an alarm basis alone, but participates in the subsequent fusion judgment together with the static flow score.
10. The method according to claim 1, characterized in that, The cosine similarity matching process includes: Calculate the cosine similarity between the feature to be matched and all features in the pre-labeled positive sample library, and select the features with the highest similarity. k The mean of each sample is used as the positive sample matching score; the cosine similarity between the feature to be matched and all features in the pre-labeled negative sample library is calculated, and the top features with the highest similarity are selected. k The mean of each sample is used as the negative sample matching score; the judgment score of the corresponding stream is output based on the positive sample matching score and the negative sample matching score.
Citation Information
Patent Citations
Feature enhancement and fusion-based weak supervision video anomaly detection method and system
CN118470608A
Method and apparatus for context-embedding and region-based object detection
US20210383166A1