Video cross-view alignment retrieval method based on high-dynamic scene

By using a scene-dynamic enhancement frame extraction mechanism and a dual-encoder collaborative alignment system, the problems of alignment accuracy and generalization in cross-view video processing are solved, enabling efficient retrieval of fine-grained actions, which is applicable to fields such as intelligent monitoring and action teaching.

CN121833994APending Publication Date: 2026-04-10UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
UNIV OF ELECTRONICS SCI & TECH OF CHINA
Filing Date
2025-12-23
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing cross-view video processing methods struggle to balance alignment accuracy and generalization when dealing with non-temporally synchronized video, fine-grained motion differentiation, and complex visual environments. They suffer from issues such as time drift and local misalignment, and lack the ability to capture frame-level information for high-dynamic motion scenes.

Method used

By adopting a scene-dynamic enhancement frame extraction mechanism, a bidirectional benchmark testing framework, and a dual encoder collaborative alignment system, key frames are retained through double-density sampling and inter-frame feature similarity screening. A multi-positive example aggregation module, a bidirectional sorting consistency target, and intramodal cohesion loss are designed to achieve accurate matching and generalization of cross-view features.

Benefits of technology

It significantly improves the accuracy and generalization ability of cross-view video retrieval, and can achieve efficient alignment at the fine-grained action level in real-world scenarios, making it suitable for practical applications such as intelligent monitoring and action teaching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121833994A_ABST
    Figure CN121833994A_ABST
Patent Text Reader

Abstract

The invention discloses a video cross-view alignment retrieval method based on a high-dynamic scene, belongs to the field of image processing, and particularly relates to a cross-view alignment retrieval method for solving the problem of two-way action migration from third person demonstration to first person execution in a real scene. The method comprises the following steps: sampling an input video through a scene dynamic enhancement frame extraction mechanism to extract double-density effective frames, comparing and screening inter-frame visual feature similarity, abandoning redundant similar frames, retaining key frames with remarkable dynamic change, and accurately capturing high-dynamic action features; and learning uniform significant feature representation by combining a dual-encoder architecture and a parameter efficient adaptation technology, and solving the core problem that common features are difficult to extract in cross-view retrieval through complementary strategies such as multi-positive-example aggregation, bidirectional sorting consistency constraint and intra-modal cohesion, so that the method is suitable for the cross-view retrieval of the multi-view retrieval. Visual representation, language anchoring and cross-view alignment performance are synchronously optimized, finally fine-grained efficient retrieval is achieved in a real scene, and lightweight deployment cross-view bidirectional action migration and retrieval are supported.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image processing, and in particular, it is a cross-view alignment retrieval method for solving the bidirectional action transfer problem from third-person demonstration to first-person execution in real-world scenarios. Background Technology

[0002] In recent years, with the rapid development of human-computer interaction and intelligent assistance technologies, cross-perspective video association technology based on first- and third-person perspectives has been increasingly widely applied in areas such as action imitation learning, process guidance alignment, and process quality assessment. Human cognition naturally supports rapid switching between two perspectives; that is, people can observe and imitate the actions of others from different angles. This ability relies on perspective switching, action schemas, and perspective signal mapping mechanisms. Artificial intelligence agents with such cross-perspective association capabilities can achieve generalized reuse of cross-perspective demonstrations, effectively integrate key knowledge from both perspectives, reduce supervision requirements, and maintain stability when the visual environment changes, providing core support for a range of downstream applications. However, as application scenarios demand fine-grained action alignment, existing cross-perspective video processing methods exhibit significant limitations in addressing the fundamental differences between the two perspectives. Especially when dealing with non-temporally synchronized video, fine-grained action differentiation, and complex visual environments, existing models struggle to balance alignment accuracy and generalization, leading to problems such as time drift and local misalignment. Therefore, how to achieve efficient cross-perspective video retrieval at the fine-grained action level in real-world scenarios has become a pressing technical challenge in this field.

[0003] The core challenge of the aforementioned problems lies in the dual constraints of the inherent differences in perspective and adaptation to complex scenes. First, the inherent characteristics of perspective present significant challenges. First-person cameras capture narrow fields of view, large hand movements, and frequent occlusion, while third-person cameras provide broader background information and drastically different scene compositions; the visual features of the two perspectives are fundamentally different. Furthermore, even within the same process, adjacent fine-grained actions are visually highly similar but semantically completely different, further increasing the difficulty of differentiation and making it difficult for models to accurately capture core semantic relationships. Second, insufficient data adaptation and generalization capabilities become key bottlenecks. Some existing methods rely on synchronized first- and third-person video pairs to explore perspective consistency or employ restrictive alignment methods, lacking generality for unstructured data; other methods extend unpaired data through generative alignment, but suffer from high training and deployment costs, making it difficult to meet practical application needs. Furthermore, existing models lack dedicated optimization objectives for cross-view retrieval and suffer from insufficient frame-level information capture in highly dynamic action scenes. Traditional fixed frame extraction methods tend to miss key action details or retain a large number of redundant frames, making it difficult for the model to handle feature ambiguity caused by rapid action changes. This results in insufficient ability to address issues such as temporal jitter, sub-action variation, and intramodal drift, further limiting their performance in real-world scenarios. Therefore, existing cross-view video association technologies face severe performance bottlenecks when dealing with fine-grained action alignment, asynchronous data adaptation, highly dynamic motion capture, and real-world scene interference. Summary of the Invention

[0004] The cross-view video retrieval method of the present invention effectively solves the key problems existing in the prior art, such as inaccurate high-dynamic motion feature capture, poor cross-view data adaptability, fine-grained motion alignment deviation, and insufficient retrieval generalization ability, through scene dynamic enhancement frame extraction mechanism, bidirectional benchmark testing framework, and dual encoder collaborative alignment system. The scene dynamic enhancement frame extraction mechanism accurately preserves key frames of high-dynamic actions and removes redundant information through double-density sampling and inter-frame feature similarity screening, solving the problems of lost action details or data redundancy caused by traditional fixed frame extraction, and providing high-quality data support for feature learning. The bidirectional benchmark testing framework solves the problems of lack of real scene interference adaptation and low data utilization efficiency of existing benchmarks through structured candidate set design and training mode compatible with strong and weak supervision, and achieves accurate matching of real needs for cross-view retrieval. The dual encoder architecture is equipped with three complementary alignment objectives. Among them, the multi-positive example aggregation module absorbs temporal jitter and sub-action mutation, the bidirectional sorting consistency objective fixes the action order through hard negative sample constraints, and the intramodal consistency objective strengthens the clustering of features in the same view. The three work together to solve the problems of cross-view feature shift, temporal drift and local misalignment, significantly improving the accuracy and generalization ability of cross-view retrieval at the fine-grained action level.

[0005] The technical solution of this invention is a video cross-view alignment retrieval method based on high dynamic scenes, the method comprising:

[0006] Step 1: Define the video cross-view alignment retrieval task and clarify the core objectives and data organization form of cross-view video retrieval in real-world scenarios;

[0007] The video cross-viewpoint alignment retrieval task is to support both "first-person view → third-person view" retrieval and "third-person view → first-person view" retrieval, requiring one-to-one matching at the fine-grained action level; and defining each complete set of procedural action records as a session. , conversation A collection of all sessions Each session is segmented into fine-grained actions. A fragment, a conversation from a single perspective. The set of fragments is represented as Define a complete set of procedural actions as a session. Each session is segmented into multiple fragments based on fine-grained actions; these fragments are identified by pairing tags. Cross-view segments that are related to the same fine-grained action. This refers to a conversation from a certain perspective. Time indexing of medium- to fine-grained action segments In other words, from another perspective, and The time index of the corresponding fine-grained action is the same. Corresponding to the same action; each segment contains coarse-grained action tags. Scene tags Fine-grained action description text ;

[0008] During retrieval, given a first-person perspective query fragment The candidate pool contains third-person perspective segments from the same conversation and cross-conversation similar action distractors. The model needs to retrieve a unique matching segment that satisfies the constraint that the similarity of the matching segment is higher than that of all distractors. ;

[0009] in, For feature similarity function; To query fragments High-dimensional feature vectors; To match segments, For interference segments, Represents a query fragment In its own session Time index; Represents a query fragment Cross-view pairing identifiers ;

[0010] Step 2: First, a scene dynamic enhancement frame extraction mechanism is used to accurately sample and extract high-value keyframes from the input video, avoiding redundant information interference, while accurately capturing high-dynamic motion features. This mechanism first adopts a double-density sampling strategy, extracting video frames at twice the original frame rate to ensure that no rapidly changing dynamic motion details are missed. Then, the visual feature similarity between adjacent frames is calculated, and keyframes are filtered through a preset similarity threshold: when the similarity between two frames is higher than the threshold, they are judged as redundant similar frames and discarded; when the similarity is lower than the threshold, they are judged as keyframes with significant dynamic changes and retained.

[0011] Step 3: Establish dual encoders: a video encoder and a text encoder;

[0012] Step 4: Design two training modes: cross-view strong supervision and single-view weak supervision. The strong supervision mode is used for scenarios with cross-view paired data, directly using the cross-view paired identifiers as labels, and defining the label function:

[0013] ;

[0014] That is, all sharing the same perspective across different viewpoints and within the same viewpoint. All segments with values ​​are positive samples;

[0015] Weakly supervised mode is suitable for scenarios where only single-view data is available. It constructs coarse feature signatures as labels based on scene labels and coarse-grained action labels, and defines a label function:

[0016] ;

[0017] Simultaneously, a time-tolerable set is introduced. :

[0018] ;

[0019] in The time tolerance threshold is used to accommodate minor errors in action segmentation. Segments within the tolerance set and segments with the same signature are considered positive samples.

[0020] Both supervision modes share the positive sample set. and negative sample set ;

[0021] Step 5: Design a multi-positive-example aggregation loss function to aggregate and optimize the similarity scores of all positive samples;

[0022] Step 6: Design a bidirectional sorting consistency loss function and introduce marginal parameters. Control the score gap between positive samples and hard negative samples;

[0023] For each anchor point sample Filter out samples with anchor points Highest similarity One negative sample is used as Top- Difficult sample set Select anchor point samples The positive sample with the highest similarity is selected as the optimal positive sample. Optimize the sorting relationship in both directions: ;

[0024] in, For anchor point samples With the best positive sample Similarity score, For anchor point samples With difficult samples Similarity score, and Similarity score for reverse retrieval; For the hinge function, only when Losses may occur at that time; and Used to normalize loss values, avoiding batch size or Impact on loss scale; Indicates the anchor point sample Highest similarity A set of difficult negative samples consisting of 10 negative samples;

[0025] This loss ensures that the scoring relationship in both directions is constrained. and ;

[0026] Step 7: Determine the intramodal cohesion loss, and use supervised contrastive loss to constrain the feature distributions of both video and text modalities; for batch sizes of... The sample set, respectively for the video feature matrix and text feature matrix The loss is calculated as the average of the losses of the two modes. The loss function is defined as follows:

[0027] ;

[0028] in For matrix The Row eigenvectors, matrices For video or text feature matrices, This is a temperature coefficient used to adjust the smoothness of feature similarity. For the first The set of positive samples is defined as the anchor sample and the feature dot product of all positive samples. The scaled exponential sum, with the denominator being the exponential sum of the feature dot products of the anchor sample and all other samples; by minimizing this loss, the feature distance of samples with the same label within the modality is compressed, while the feature distance of samples with different labels is widened.

[0029] Step 8: Apply symmetric loss The three alignment losses mentioned above are weighted and fused to form the final training loss function. ;

[0030] Step 9: Based on the model trained in the above steps, input a query segment from a single perspective during the inference phase. The encoder generates feature embeddings, calculates similarity with the segment embeddings in the candidate pool, sorts them by score, and returns the optimal matching result. Load the trained optimal dual encoder, freeze all parameters, and perform separate operations on the query segment. and candidate pool Each segment Encode and generate normalized video embeddings. and and text features and The similarity score between the query embedding and each candidate fragment embedding is calculated using the scaled cosine similarity formula:

[0031]

[0032]

[0033]

[0034]

[0035] Candidate segments are sorted from highest to lowest score, and the segment ranked first is returned as the search result.

[0036] Furthermore, the specific method of step 3 is as follows: the video encoder is built based on Timesformer-B and uses a frozen-in-time attention mechanism to extract spatiotemporal features of the video; the text encoder is built based on a 12-layer Transformer, initialized from DistilBERT, and used to encode text semantic features; the original output features of the video encoder are processed... conduct Normalization yields video clips Normalization characteristics :

[0037] ;

[0038] in, These are learnable parameters for the video encoder; for Norm, express 3D real space;

[0039] Similarly, the raw output features of the text encoder Normalization yields the text description. Normalization characteristics :

[0040] ;

[0041] in, These are learnable parameters for the text encoder;

[0042] for Batch size of video clip collection Construct batch video feature matrices and text feature matrix This leads to the video-text cross-modal similarity matrix. Cross-view video - video similarity ; This represents the i-th video segment. This represents the trainable temperature parameter used to scale the inner product of the feature vectors.

[0043] Furthermore, the specific method for step 5 is as follows: for batch sizes of... The sample set, traversing each anchor point sample Calculate its relationship with all positive samples Similarity score The values ​​are transformed into non-negative values ​​using an exponential function and then summed to obtain the numerator; the anchor sample is compared with all other samples. The similarity score index is used as the denominator; the ratio of the numerator to the denominator is converted into a loss value through logarithmic operation, and the overall loss function is defined as:

[0044] ;

[0045] in For indicator functions, when The value is 1 if the condition is met, and 0 otherwise, and is used to filter positive samples. This loss is used to normalize the loss value of a batch of samples; by minimizing this loss, the overall similarity of the anchor sample with all positive samples is significantly higher than its similarity with negative samples.

[0046] Furthermore, the final loss function in step 8 is:

[0047] ;

[0048] in, These are the interpolation coefficients, used to adjust the weight ratio of the symmetry loss and the three-term alignment loss; , , These are the weighting coefficients for multi-positive-example aggregation loss, bidirectional ranking consistency loss, and intramodal cohesion loss, respectively.

[0049] The core function of symmetric loss is to provide a stable semantic benchmark for cross-viewpoint retrieval through video-text cross-modal alignment, avoiding the model learning viewpoint-specific surface features. It is defined as follows:

[0050] ;

[0051] in, It is the identity matrix. The cross-entropy function is used to calculate the cross-modal alignment loss in both the video-to-text and text-to-video directions, and the average value is taken as the symmetry loss.

[0052] The cross-view video retrieval method proposed in this invention is specifically designed for fine-grained bidirectional motion retrieval in first- and third-view videos in real-world scenarios. It effectively establishes a one-to-one anti-interference alignment relationship at the action level across viewpoints, making it suitable for various practical applications such as intelligent monitoring, motion teaching, and virtual reality. The main innovation of this invention lies in its efficient dual-encoder feature alignment scheme based on an architectural design, and the construction of a multi-constraint collaborative training mechanism. Through four core innovative strategies, it effectively addresses key challenges in cross-view retrieval, including the difficulty in extracting common features across viewpoints, temporal jitter interference, insufficient consistency in bidirectional retrieval, and viewpoint-specific drift. First, the scene-dynamic enhancement frame extraction mechanism accurately captures the core features of highly dynamic actions and eliminates redundant information through double-density sampling and inter-frame similarity filtering. Second, the style dual-encoder architecture, combined with efficient parameter adaptation technology, achieves unified alignment of feature semantics across modalities and viewpoints. Third, three complementary loss strategies—multiple positive example aggregation, bidirectional ranking consistency constraints, and intramodal cohesion—address issues such as temporal drift, bidirectional retrieval inconsistency, and viewpoint-specific drift, respectively. Finally, a flexible training mode compatible with both strong and weak supervision reduces dependence on cross-viewpoint pairing data and improves the model's generalization ability. These innovative designs work synergistically to support lightweight deployment, comprehensively improving the accuracy, robustness, and efficiency of cross-viewpoint video retrieval in real-world complex scenarios. Attached Figure Description

[0053] Figure 1 This is a task description diagram designed for the present invention.

[0054] Figure 2 The flowchart is designed for the method of this invention. Detailed Implementation

[0055] This invention proposes a real-world cross-view video retrieval method and benchmark framework E2EFollow based on fine-grained motion. This method aims to achieve accurate association retrieval between first- and third-view videos by learning viewpoint-invariant, language-anchored feature representations, improving the accuracy and generalization of cross-view alignment. This invention addresses various challenges arising from viewpoint differences, high-dynamic motion capture, and scene complexity through core technology design and innovative framework construction. Specifically, it includes the following key technical solutions: First, an innovative scene dynamic enhancement frame extraction mechanism is designed. A double-density sampling strategy is used to extract effective frames from the input video. An inter-frame visual feature similarity comparison algorithm automatically discards redundant similar frames and retains key frames with significant dynamic changes. This accurately captures the core features of high-dynamic motion while compressing the data volume, laying a high-quality data foundation for subsequent feature learning. Second, a bidirectional benchmark framework, E2EFollow, is constructed. A structured candidate set is designed using a large dataset collected with laboratory eye-tracking equipment. Each query corresponds to a single correct answer and multiple distractor options, perfectly meeting the needs of real-world application scenarios. It also supports mixed training of single-view and cross-view data, is compatible with strongly supervised and weakly supervised labels, and significantly improves data utilization efficiency. Finally, a dual-encoder architecture is designed to learn viewpoint-invariant and language-anchored feature representations. Accurate retrieval is achieved through three complementary alignment objectives: a multi-positive-example aggregation module effectively absorbs the effects of temporal jitter and sub-action variation; a bidirectional ranking consistency objective uses hard negative samples to fix the action order in both the first-view to third-view and third-view to first-view directions, avoiding ranking errors; and an intramodal consistency objective strengthens feature clustering within the same viewpoint, suppressing viewpoint-specific drift. Through the synergistic effect of these innovative designs, this invention significantly improves the accuracy of cross-viewpoint retrieval at the fine-grained action level, efficiently addressing the performance shortcomings of existing methods in viewpoint difference adaptation, asynchronous data processing, high-dynamic motion capture, and real-world scene interference. It provides an efficient technical solution for cross-viewpoint video association technology and has broad application prospects in motion guidance, process evaluation, and decision support.

[0056] To address the core challenges in cross-view video retrieval in real-world scenarios, including a lack of structured scene adaptation, chaotic data associations, inconsistencies in cross-modal and cross-view feature semantics, and sensitivity to temporal jitter and viewpoint specificity, we constructed a learning framework for precise one-to-one action matching between first- and third-view perspectives. As a challenging task, this invention needs to meet the requirements of adapting to real-world scene interference, unifying cross-modal semantic alignment, enabling flexible training in data-scarce scenarios, and ensuring stable and reliable retrieval ranking, thereby guaranteeing the accuracy, generalization, and efficiency of bidirectional cross-view retrieval.

[0057] To address the issues of excessive redundancy in original video frames and the easy obscuring of high-dynamic motion features, the network needs to reduce data processing costs while preserving core motion information. Therefore, we designed a scene dynamic enhancement frame extraction mechanism. A double-density sampling strategy is employed to extract video frames, avoiding the omission of rapidly changing dynamic motion details; through computation... and Keyframes are selected based on visual feature similarity between adjacent frames, redundant similar frames are removed, and high-dynamic motion features are accurately captured. Similarity calculation is performed.

[0058]

[0059] Redundant frames are filtered by setting a preset threshold, ultimately resulting in a keyframe sequence. .

[0060] To achieve feature semantic unification across modalities (video-text) and perspectives (first- and third-person), and to address the issue of indirect feature comparison between different modalities and perspectives, a dual-encoder architecture is constructed. The video encoder is based on TimeSformer-B and employs a frozen-in-time attention mechanism to extract video segments. Spatiotemporal characteristics of keyframe sequences:

[0061]

[0062] The text encoder is based on a 12-layer Transformer (initialized from DistilBERT) and encodes action description text. The semantic features of the two encoder outputs are analyzed. Normalization, mapped to Dimensional Shared Space:

[0063]

[0064] During batch processing, a video feature matrix is ​​constructed. With text feature matrix A temperature-coefficient scaled cosine similarity metric was used to quantify feature association calculations for cross-modal similarity. Similarity with cross-view videos .

[0065] To adapt to real-world scenarios where cross-view pairing data is scarce and only single-view data is available, and to reduce data dependency costs, we provide effective supervision signals to the model through a unified label definition. We design a strong-weak dual-supervision training mechanism. In the strong supervision mode, the cross-view pairing identifier is used as the label, directly using the uniquely corresponding pairing identifier across the cross-view as the label:

[0066]

[0067] The weakly supervised mode uses scene tags plus coarse-grained action tags as joint tags:

[0068]

[0069] Simultaneously, a time-tolerable set is introduced. Compatible with action segmentation errors. Both modes share the same positive and negative sample sets. and Provide stable supervision signals for the model.

[0070] To address the many-to-one correspondence and local temporal drift issues in cross-view retrieval, we improve the model's tolerance to temporal jitter and sub-action variation by aggregating the feature responses of all positive samples. We design a multi-positive-example aggregation loss, which aggregates the feature responses of all positive samples to enhance the model's tolerance to temporal jitter and sub-action variation. This loss function solves the problem that existing single-positive-example losses cannot adapt to scenarios such as multiple segments of the same action and minor deviations in action segmentation, leading to the model's sensitivity to temporal jitter. For each anchor sample... Calculate the similarity score between the anchor sample and all positive samples, sum them using an exponential function, and use the sum as the numerator; calculate the similarity score between the anchor sample and all other samples. The similarity score index is used as the denominator. Finally, the ratio is converted into a loss value through logarithmic calculation:

[0071]

[0072] By minimizing this loss, the overall similarity between the anchor sample and all positive samples is significantly higher than that of the negative samples, thus absorbing temporal interference. This enables the model to automatically aggregate the feature responses of all positive samples, effectively absorbing interference from temporal jitter and sub-action mutations. It can still maintain accurate cross-view alignment in many-to-one correspondence scenarios, improving the robustness of the model.

[0073] To ensure bidirectional consistency in cross-perspective retrieval, enhance the model's ability to distinguish similar distractors, and force true positive samples to outperform hard negative samples in both the "query → candidate" and "candidate → query" directions, thereby improving the stability and reliability of retrieval ranking, we designed a bidirectional ranking consistency loss. This loss function addresses the problem that existing losses only optimize retrieval ranking in one direction, leading to bidirectional inconsistency and weak ability to distinguish similar distractors. Marginal parameters are introduced. By controlling the score difference between positive samples and hard negative samples, the ranking relationships in both the "query → candidate" and "candidate → query" directions are constrained.

[0074]

[0075] This ensures that positive samples outperform hard negative samples in both bidirectional retrieval, improving ranking stability. It achieves bidirectional consistency across different perspectives, significantly enhancing the model's ability to distinguish similar interference items, resulting in more stable and reliable retrieval ranking results that meet the accuracy requirements of real-world scenarios.

[0076] To suppress viewpoint-specific drift, enhance feature clustering of samples with the same label across video and text modalities, and ensure the distinguishability of similar semantic samples, we design an intramodal cohesion loss. This loss function addresses the problem that features of the same action can easily be misclassified as different categories due to differences in shooting perspectives. Supervised contrastive loss is used to constrain the feature distributions of video and text modalities.

[0077]

[0078] By compressing the feature distance of samples with the same label and widening the distance of samples with different labels, the distinguishability of similar semantic samples is ensured. This effectively suppresses feature drift caused by viewpoint specificity, strengthens feature aggregation of samples with the same semantic meaning within the modality, and ensures that similar action samples can still maintain good distinguishability after cross-viewpoint alignment, thereby improving the classification and retrieval accuracy of the model.

[0079] Finally, to balance semantic anchoring, cross-view alignment, ranking stability, and intramodal consistency, we weighted and fused the symmetric loss with the three alignment losses mentioned above to form the total loss function:

[0080]

[0081] Symmetric loss provides a semantic benchmark through video-text cross-modal alignment:

[0082]

[0083] Through multi-loss collaborative optimization, the model achieves comprehensive improvements in cross-modal semantic alignment, cross-view matching accuracy, retrieval ranking stability, and feature consistency, ultimately enabling fine-grained and efficient retrieval in real-world scenarios and supporting lightweight training and deployment. Specifically, the multi-loss collaborative optimization enhances the model's overall retrieval performance, ensuring stable handling of complex cross-view retrieval needs in real-world scenarios.

Claims

1. A video cross-view alignment retrieval method based on high dynamic range scenes, the method comprising: Step 1: Define the video cross-view alignment retrieval task and clarify the core objectives and data organization form of cross-view video retrieval in real-world scenarios; The video cross-viewpoint alignment retrieval task is to support both "first-person view → third-person view" retrieval and "third-person view → first-person view" retrieval, requiring one-to-one matching at the fine-grained action level; and defining each complete set of procedural action records as a session. , conversation A collection of all sessions Each session is segmented into fine-grained actions. A fragment, a conversation from a single perspective. The set of fragments is represented as Define a complete set of procedural actions as a session. Each session is segmented into multiple fragments based on fine-grained actions; these fragments are identified by pairing tags. Cross-view segments that are related to the same fine-grained action. This refers to a conversation from a certain perspective. Time indexing of medium- to fine-grained action segments In other words, from another perspective, and The time index of the corresponding fine-grained action is the same. Corresponding to the same action; each segment contains coarse-grained action tags. Scene tags Fine-grained action description text ; During retrieval, given a first-person perspective query fragment The candidate pool contains third-person perspective segments from the same conversation and cross-conversation similar action distractors. The model needs to retrieve a unique matching segment that satisfies the constraint that the similarity of the matching segment is higher than that of all distractors. ; in, For feature similarity function; To query fragments High-dimensional feature vectors; To match segments, For interference segments, Represents a query fragment In its own session Time index; Represents a query fragment Cross-view pairing identifiers ; Step 2: First, a scene dynamic enhancement frame extraction mechanism is used to accurately sample and extract high-value keyframes from the input video, avoiding redundant information interference, while accurately capturing high-dynamic motion features. This mechanism first adopts a double-density sampling strategy, extracting video frames at twice the original frame rate to ensure that no rapidly changing dynamic motion details are missed. Then, the visual feature similarity between adjacent frames is calculated, and keyframes are filtered through a preset similarity threshold: when the similarity between two frames is higher than the threshold, they are judged as redundant similar frames and discarded; when the similarity is lower than the threshold, they are judged as keyframes with significant dynamic changes and retained. Step 3: Establish dual encoders: a video encoder and a text encoder; Step 4: Design two training modes: cross-view strong supervision and single-view weak supervision. The strong supervision mode is used for scenarios with cross-view paired data, directly using the cross-view paired identifiers as labels, and defining the label function: ; That is, all sharing the same perspective across different viewpoints and within the same viewpoint. All segments with values ​​are positive samples; Weakly supervised mode is suitable for scenarios where only single-view data is available. It constructs coarse feature signatures as labels based on scene labels and coarse-grained action labels, and defines a label function: ; Simultaneously, a time-tolerable set is introduced. : ; in The time tolerance threshold is used to accommodate minor errors in action segmentation. Segments within the tolerance set and segments with the same signature are considered positive samples. Both supervision modes share the positive sample set. and negative sample set ; Step 5: Design a multi-positive-example aggregation loss function to aggregate and optimize the similarity scores of all positive samples; Step 6: Design a bidirectional sorting consistency loss function and introduce marginal parameters. Control the score gap between positive samples and hard negative samples; For each anchor point sample Filter out samples with anchor points Highest similarity One negative sample is used as Top- Difficult sample set Select anchor point samples The positive sample with the highest similarity is selected as the optimal positive sample. Optimize the sorting relationship in both directions: ; in, For anchor point samples With the best positive sample Similarity score, For anchor point samples With difficult samples Similarity score, and Similarity score for reverse retrieval; For the hinge function, only when Losses may occur at that time; and Used to normalize loss values, avoiding batch size or Impact on loss scale; Indicates the anchor point sample Highest similarity A set of difficult negative samples consisting of 10 negative samples; This loss ensures that the scoring relationship in both directions is constrained. and ; Step 7: Determine the intramodal cohesion loss, and use supervised contrastive loss to constrain the feature distributions of both video and text modalities; for batch sizes of... The sample set, respectively for the video feature matrix and text feature matrix The loss is calculated as the average of the losses of the two modes. The loss function is defined as follows: ; in For matrix The Row eigenvectors, matrices For video or text feature matrices, This is a temperature coefficient used to adjust the smoothness of feature similarity. For the first The set of positive samples is defined as the anchor sample and the feature dot product of all positive samples. The scaled exponential sum, with the denominator being the exponential sum of the feature dot products of the anchor sample and all other samples; by minimizing this loss, the feature distance of samples with the same label within the modality is compressed, while the feature distance of samples with different labels is widened. Step 8: Apply symmetric loss The three alignment losses mentioned above are weighted and fused to form the final training loss function. ; Step 9: Based on the model trained in the above steps, input a query segment from a single perspective during the inference phase. The encoder generates feature embeddings, calculates similarity with the segment embeddings in the candidate pool, sorts them by score, and returns the optimal matching result. Load the trained optimal dual encoder, freeze all parameters, and perform separate operations on the query segment. and candidate pool Each segment Encode and generate normalized video embeddings. and and text features and The similarity score between the query embedding and each candidate fragment embedding is calculated using the scaled cosine similarity formula: ; ; ; ; Candidate segments are sorted from highest to lowest score, and the segment ranked first is returned as the search result.

2. The video cross-view alignment retrieval method based on high dynamic scenes as described in claim 1, characterized in that, The specific method of step 3 is as follows: the video encoder is built based on Timesformer-B and uses a frozen-in-time attention mechanism to extract spatiotemporal features of the video; the text encoder is built based on a 12-layer Transformer, initialized from DistilBERT, and used to encode text semantic features; the original output features of the video encoder are processed... conduct Normalization yields video clips Normalization characteristics : ; in, These are learnable parameters for the video encoder; for Norm, express 3D real space; Similarly, the raw output features of the text encoder Normalization yields the text description. Normalization characteristics : ; in, These are learnable parameters for the text encoder; for Batch size of video clip collection Construct batch video feature matrices and text feature matrix This leads to the video-text cross-modal similarity matrix. Cross-view video - video similarity ; This represents the i-th video segment. This represents the trainable temperature parameter used to scale the inner product of the feature vectors.

3. The video cross-view alignment retrieval method based on high dynamic scenes as described in claim 1, characterized in that, The specific method for step 5 is as follows: For batch sizes of... The sample set, traversing each anchor point sample Calculate its relationship with all positive samples Similarity score The summation of these values, after transforming them into non-negative values ​​using an exponential function, is used as the numerator. Calculate anchor point samples and all other samples The similarity score index is used as the denominator; the ratio of the numerator to the denominator is converted into a loss value through logarithmic operation, and the overall loss function is defined as: ; in For indicator functions, when The value is 1 if the condition is met, and 0 otherwise, and is used to filter positive samples. This loss is used to normalize the loss value of a batch of samples; by minimizing this loss, the overall similarity of the anchor sample with all positive samples is significantly higher than its similarity with negative samples.

4. The video cross-view alignment retrieval method based on high dynamic scenes as described in claim 1, characterized in that, The final loss function in step 8 is: ; in, These are the interpolation coefficients, used to adjust the weight ratio of the symmetry loss and the three-term alignment loss; , , These are the weighting coefficients for multi-positive-example aggregation loss, bidirectional ranking consistency loss, and intramodal cohesion loss, respectively. The core function of symmetric loss is to provide a stable semantic benchmark for cross-viewpoint retrieval through video-text cross-modal alignment, avoiding the model learning viewpoint-specific surface features. It is defined as follows: ; in, It is the identity matrix. The cross-entropy function is used to calculate the cross-modal alignment loss in both the video-to-text and text-to-video directions, and the average value is taken as the symmetry loss.