Multilevel semantic interactive cross-modal tracking method based on unreal engine

By generating a virtual simulation environment in Unreal Engine and building an end-to-end multi-level bootstrap framework, the problems of limited availability of data sets and insufficient fusion of semantics and text in the prior art are solved, high-quality semantic target tracking data sets are achieved and the tracking accuracy and recall of the model are improved.

CN119941790AActive Publication Date: 2025-05-06BEIJING UNIV OF POSTS & TELECOMM
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510001144.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-05-06
Estimated Expiration
2045-01-02

AI Technical Summary

Technical Problem

In the prior art, semantic target tracking methods have problems such as limited availability of data sets and insufficient fusion of semantic and text, resulting in low recall and difficult tracking of complex dynamic targets.

Method used

Using a multi-level semantic interactive cross-modal tracking method based on Unreal Engine, we generate a highly realistic virtual simulation environment through Unreal Engine, reduce dependence on manual annotation, and build an end-to-end multi-level guidance framework to integrate semantic information to improve model performance.

Benefits of technology

High-quality and low-cost semantic tracking dataset generation is achieved, which significantly improves the recall and tracking accuracy of the model's semantic targets, and improves the overall performance and recall of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941790A_ABST
    Figure CN119941790A_ABST
Patent Text Reader

Abstract

The invention provides a multi-level semantic interactive cross-modal tracking method based on an unreal engine. The method comprises the following steps: constructing a virtual simulation world by using an unreal engine 5; constructing pedestrian and vehicle virtual multi-target tracking data; constructing a text-track matching pair to generate multi-modal tracking data; constructing a multi-target tracking model fusing the multi-modal semantic features layer by layer; enhancing perceptual query features by using text features; mapping the decoding perception feature to a semantic space by using a linear layer, and calculating the similarity between the decoding perception feature and a coded text feature; updating the target track information by utilizing the perception query result; according to the method, the problem that a track semantic data set is missing is solved, and the accuracy and recall rate of semantic target tracking of the model in a complex dynamic environment are remarkably improved by combining a layer-by-layer semantic interaction module with the cross-modal alignment capability of CLIP.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of signal processing technology and deep learning, and specifically to a multi-level semantic interactive cross-modal tracking method based on Unreal Engine. Background Art

[0002] In recent years, the field of visual tasks is experiencing a trend of integration with natural language descriptions, giving rise to a series of innovative technologies. The semantic target tracking task is one of them, which aims to open a new chapter in human-computer interaction: users communicate directly with the tracking system through natural language instructions, accurately identifying and tracking the target objects specified in the image or video. This task poses a dual challenge: on the one hand, the system must have the ability to deeply understand complex natural languages; on the other hand, it also needs to be equipped with efficient image processing capabilities. Therefore, a visual target tracking method that can fully understand semantic instructions is required, which is very challenging.

[0003] Despite the innovation in this field, the availability of relevant datasets is relatively limited. Typically, existing benchmark datasets are re-annotated on top of public multi-target tracking benchmarks. As a result of this approach, the accuracy of the resulting new benchmark is affected by both the inherent limitations of the original benchmark and the subjective judgment bias in the manual annotation process, which in turn affects the performance of downstream tasks. To address this problem, the present invention proposes a method for multi-level semantic interactive cross-modal tracking based on Unreal Engine. The method overcomes the shortcomings of existing datasets by constructing a high-precision, low-cost benchmark dataset. Specifically, by using Unreal Engine to generate a highly realistic virtual simulation environment, the reliance on manual annotation is eliminated, and large-scale production of datasets is achieved. In addition, in response to the recall problem caused by insufficient fusion of semantics and text in the prior art, the present invention also introduces an end-to-end multi-level guided framework. The framework ensures that semantic information can be effectively integrated between each level from the autoencoder stage to the prediction head, significantly improving the performance of the model in practical applications. Summary of the invention

[0004] The present invention provides a multi-level semantic interactive cross-modal tracking method based on Unreal Engine, characterized in that the method comprises:

[0005] Step 1: Build a virtual simulation world through Unreal Engine 5, collect the trajectories of pedestrians and vehicles in the world, and generate multi-target tracking data;

[0006] Step 2: Clean, filter, describe, mark, and evaluate the target trajectory to ensure the accuracy and availability of the data;

[0007] Step 3: Combine the appearance description and motion description of the marker according to certain rules with the help of a large language model, and associate them with the target trajectory segment, finally obtaining a high-quality and low-cost semantic tracking dataset;

[0008] Step 4: construct a multi-target tracking model that integrates multi-modal semantic features layer by layer;

[0009] Step 5: The video frames are cut and sent to the model. Two backbone networks are used to extract the features of the image and text respectively, and the extended mode encoder is used to achieve feature fusion.

[0010] Step 6: Initialize a set of detection queries and combine them with the previously output tracking queries to form perception queries. Then apply text features to guide the perception queries and input them into the decoder together with the fusion features to capture the semantic targets in the video.

[0011] Step 7: Use the linear layer to map the decoded query features into the semantic space, and compare the similarity between the decoded query features and the encoded text features to select the query results with higher similarity;

[0012] Step 8: Use a multi-layer perceptron to perform category classification and coordinate regression processing on the filtered query results.

[0013] Step 9, update the trajectory tracking result, keep the perception query until the next frame and turn it into a tracking query, and repeat steps 5 to 9 until all image sequences in the video are tracked;

[0014] Specifically, in step 1, the resources disclosed in Unreal Engine 5 are used, and the appearance (color and size) of 3D modeling is changed to generate vehicle and crowd instances in the system. Then they are placed in the virtual world, and the internal vehicle and pedestrian movement rules are used to simulate the urban traffic system. Finally, a camera is placed in the world, and the video screen is recorded, and the coordinate information and appearance information of the target in the screen can be directly obtained. By using this method, the 3D modeling of the object and the camera position are continuously changed, and the data of multi-target tracking is generated in batches;

[0015] Specifically, in step 2, the multi-target tracking data obtained in step 1 is cleaned, filtered, described, labeled, and quality evaluated, including: first, filtering out the trajectories that are severely blocked by buildings and scenes with fewer targets collected in step 1; second, describing and recording the motion of the target trajectory fragments. Finally, the quality of the processed label data is evaluated to filter out descriptions with longer duration and obvious behavior;

[0016] Specifically, in step 3, the motion description annotated in step 2 is first combined with the appearance description automatically generated by 3D modeling. Then, a large language model is used to derive multiple groups of synonyms from the combined words, and some common behavior description phrases are manually screened out. Finally, the corresponding description phrases are associated with the trajectory segments, the start and end times of the trajectory are recorded, and a text-trajectory matching pair is generated. This method greatly reduces the cost of manual annotation and ensures the accuracy and objectivity of the data set;

[0017] Specifically, in step 4, the model includes an image encoder, a frozen text encoder, a fusion encoder, a semantic guidance module, a tracking decoder, and a semantic relevance branch prediction module. Among them, the text encoder selects RoBERTa and CLIP, and the fusion encoder and fusion decoder both adopt the DETR structure, which contains multiple layers of Transformer. The input of the semantic guidance module sets a learnable query for perceiving different semantic features, with a size of M×N dim , where M is the number of learnable queries;

[0018] Specifically, in step 5, first, the video generated in step 3 is cut into frames to obtain an image sequence, and the image at time t is feature extracted using a CNN-based backbone network to obtain an image pyramid feature I t Next, the frozen text encoder is used to encode the text instruction into a vector S. Then, I is transformed into t Each layer of is fused with S to generate cross-modal features E t ;

[0019] Specifically, in step 6, a set of detection queries are initialized with learnable variables, denoted as And combined with the tracking query from the previous frame output Generate a perceptual query, denoted as Q t . Use the semantic guidance module to interact with the perceptual query and the semantic instruction: t Added to its position encoding, it is fed into the self-attention module to integrate Q t The deep features of the text features are linearly mapped to the same latitude, and the calculation formula is: S emb =Linear(S); then use cross attention to Q t Perform text guidance to generate perceptual queries with text prompts. The calculation formula is: Finally, the fusion decoder is used to allow the perceptual query to capture the high-dimensional information of the target in the fusion feature, and the output is recorded as

[0020] Specifically, in step 7, the linear layer is used to transform the Mapped to the semantic feature space, the calculation formula is: Use CLIP text encoder to re-encode the text instructions. The calculation formula is: S clip =Encoder clip (S). And with the powerful cross-modal alignment capability of CLIP, calculation With S clip The similarity between the two modalities finally obtains the relevance between each query and the text instruction, denoted as R cos , the calculation formula is: Further, the query results with similarity greater than 0.5 are filtered out;

[0021] Specifically, in step 8, a multi-layer perceptron is used to perform category classification and coordinate regression processing on the filtered query results to obtain category information and coordinate information of the semantic target in the current frame;

[0022] Specifically, in step 9, the trajectory queue is updated using the result. If the target is Capture, then create a new track, if by If captured, the corresponding old trajectory is updated. Record and update the tracking query for the next frame Repeat steps 5 to 9 until all image sequences in the video are tracked;

[0023] The present invention aims at the fact that multimodal tracking is difficult to use in actual scenes due to the lack of semantic description of trajectories and the existing semantic tracking methods only rely on a single cross-modal interaction layer to fuse text and visual features, resulting in insufficient perception of different semantic targets and difficulty in tracking complex dynamic targets. A multi-level semantic interactive cross-modal tracking method based on Unreal Engine is proposed. By using the 3D modeling capability of Unreal Engine 5, a simulated world is generated. Through semi-automatic annotation tools, trajectory-semantic pairs can be generated at low cost and high quality, filling the gap of data sets in this field. At the same time, a semantic guidance module is used to promote the interaction between perceptual queries and text, and with the powerful cross-modal alignment capability of CLIP, the fusion features of perceptual queries are aligned with text features. Through this layer-by-layer semantic fusion framework, the accuracy of the model's recall and tracking of semantic targets is greatly improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.

[0025] Figure 1 A flowchart of the steps of a multi-level semantic interactive cross-modal tracking method based on Unreal Engine; DETAILED DESCRIPTION

[0026] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0027] Although innovative research in the field of semantic object tracking is endless, the availability of datasets and the limitations of existing technologies are still important challenges facing this field. Traditionally, researchers create new benchmark datasets by adding annotations to existing public benchmarks, but the effectiveness of this method is often limited by the limitations of the original dataset and the subjective errors in manual annotation, which affects the improvement of model performance. At the same time, existing technologies often only use a single interaction layer to integrate text and visual features, resulting in insufficient perception of semantic targets and difficulty in effectively tracking dynamic and complex scenes. The present invention proposes a multi-level semantic interactive cross-modal tracking method based on Unreal Engine. First, the Unreal Engine is used to generate a highly accurate and cost-effective benchmark dataset, and an end-to-end multi-level guided framework is proposed. The framework significantly improves the model's recognition and prediction capabilities for actual scenes by effectively integrating semantic information between each level from the autoencoder stage to the prediction head, thereby improving the overall performance and recall of the system. The specific steps are as follows:

[0028] S101, generating a simulated world using Unreal Engine 5;

[0029] Specifically, various 3D models are selected as basic materials in the public resource library of Unreal Engine 5. Then, these models are personalized and their appearance attributes are changed, such as adjusting the color and size of vehicles and crowds, so as to generate instances that meet the system requirements. The vehicle models are configured in a diversified manner, with different body colors and vehicle specifications set to match the types of vehicles in the real world; the clothing color and body size of the character models are changed to generate diversified character targets. Then, these customized vehicle and crowd models are placed in the virtual city environment built by Unreal Engine 5 to generate a virtual simulation world.

[0030] S102, collecting pedestrian and vehicle trajectories in the world and generating multi-target tracking data;

[0031] Specifically, select key observation points in the virtual world and place cameras there, ensuring that the cameras cover key traffic nodes and hot spots to capture sufficient dynamic information. Start the video recording function to collect the coordinates and appearance data of the target as it moves in the picture. Continue to change the placement of the camera, such as raising or lowering its shooting angle, or moving the camera to a new viewing point. Repeat the above steps to expand the recording range, increase the diversity of scenarios, and systematically collect different trajectory data for multi-target tracking;

[0032] S103, 3D modeling generates text description and manually calibrates trajectory motion description;

[0033] Specifically, for the multi-target tracking trajectory labeling description text obtained in S102, a preliminary appearance text description containing basic information such as the target type, color, size, etc. is first directly constructed based on the target 3D model, and then the movement mode of each target in the trajectory is manually annotated, such as straight-line movement, turning, stopping, accelerating, etc., to generate a preliminary dynamic text description.

[0034] S104, cleaning, filtering, quality assessment, etc. are performed on the text-trajectory matching pairs;

[0035] Specifically, the text-trajectory matching data obtained in S103 is cleaned, filtered, and evaluated for quality, including: first, screening and removing the video data with incomplete trajectory information due to building occlusion, secondly eliminating scenes with a small number of targets or unclear behavior patterns, and then evaluating the quality of the remaining data, eliminating trajectories with unclear descriptions or short durations, and only retaining data with longer durations and clear behavior patterns in rich contexts;

[0036] S105, combining the descriptions with the help of a large language model and associating them with the track segments to obtain semantic tracking data;

[0037] Specifically, the appearance text description annotated in step S103 is first randomly combined with the dynamic text description. For example, "red" is selected and matched with "straight" to obtain the description of "red straight car". Then, a large language model is used to generate a series of synonyms for these combinations, and common behavior description phrases are manually screened out. Next, the selected descriptive phrases are associated with the trajectory segments filtered by S104, and the start and end time of the trajectory are recorded to ensure the generation of accurate and diverse text-trajectory matching pairs, and finally generate semantic tracking data. This method greatly reduces the cost of manual annotation and ensures the accuracy and objectivity of the data set.

[0038] S106, extracting features using two backbone networks and fusing image and text features using a cross-modal encoder;

[0039] Specifically, in the application, the semantic instructions and image sequences are input into the model to track specific semantic targets in the image sequence. First, the image at time t is read in the image sequence, and the CNN-based backbone network is used to extract the features of the image at time t to obtain the image pyramid feature I t Next, the frozen text encoder is used to encode the text instruction into a vector S. Then, I is transformed into t Each layer of is fused with S to generate cross-modal features E t ;

[0040] S107, uses text to guide detection queries and tracking queries, and semantic-aware queries capture semantic targets in fused features through decoders;

[0041] Specifically, we first initialize a set of detection queries with learnable variables, denoted as And combined with the tracking query from the previous frame output Generate a perceptual query, denoted as Q t . Use the semantic guidance module to interact with the perceptual query and the semantic instruction: t Added to its position encoding, it is fed into the self-attention module to integrate Q t The deep features of the text features are linearly mapped to the same latitude, and the calculation formula is: S emb =Linear(S); then use cross attention to Q t Perform text guidance to generate perceptual queries with text prompts. The calculation formula is: Finally, the fusion decoder is used to allow the perceptual query to capture the high-dimensional information of the target in the fusion feature, and the output is recorded as

[0042] S108, using a linear layer to map the decoded features to a semantic space and calculate similarity with the text features;

[0043] Specifically, the S107 obtained by the linear layer Mapped to the semantic feature space, the calculation formula is: Use CLIP text encoder to re-encode the text instructions. The calculation formula is: S clip =Encoder clip (S). And with the powerful cross-modal alignment capability of CLIP, calculation With S clip The similarity between the two modalities finally obtains the relevance between each query and the text instruction, denoted as R cos , the calculation formula is: Further, the query results with similarity greater than 0.5 are filtered out;

[0044] S109, for decoding categories and coordinates with high similarity, detection query generates new tracks, and tracking query updates old tracks;

[0045] Specifically, a multi-layer perceptron is used to classify and regress the search results to obtain the category information and coordinate information of the semantic target in the current frame. The results are used to update the trajectory queue. Capture, then create a new track, if by If captured, the corresponding old trajectory is updated. Record and update the tracking query for the next frame

[0046] S110, tracking until the sequence ends and recording the position information of all tracks;

[0047] Specifically, steps S106 to S110 are repeated until all image sequences in the video are tracked; the system will eventually record the trajectory information of all semantic targets in the video, including their movement paths and position changes over time;

[0048] The multi-level semantic interactive cross-modal tracking method based on Unreal Engine proposed in the present invention uses Unreal Engine 5 to build a highly realistic simulation environment, and uses semi-automatic annotation tools to create accurate trajectory-semantic data pairs at low cost and high efficiency, effectively making up for the shortcomings of the existing technology in terms of data sets; secondly, semantic features are used to guide queries, which, as a core hub, promotes deep interaction between perceptual queries and texts; and combined with the excellent ability of the CLIP model in cross-modal alignment, the fusion features extracted from the perceptual queries and the text features can achieve high consistency. This unique layer-by-layer semantic interaction mechanism significantly enhances the system's detection sensitivity and recall rate for various semantic targets, while also improving the overall tracking accuracy, enabling the model to more accurately complete target tracking tasks in complex dynamic scenes.

Claims

1. A multi-level semantic interactive cross-modal tracking method based on Unreal Engine, characterized in that: The method comprises the following steps: Step 1: Build a virtual simulation world through Unreal Engine 5, collect the trajectories of pedestrians and vehicles in the world, and generate multi-target tracking data; Step 2: Clean, filter, describe, mark, and evaluate the target trajectory to ensure the accuracy and availability of the data; Step 3: Combine the appearance description and motion description of the marker according to certain rules with the help of a large language model, and associate them with the target trajectory segment, finally obtaining a high-quality and low-cost semantic tracking dataset; Step 4: construct a multi-target tracking model that integrates multi-modal semantic features layer by layer; Step 5: The video frames are cut and sent to the model. Two backbone networks are used to extract the features of the image and text respectively, and the extended mode encoder is used to achieve feature fusion. Step 6: Initialize a set of detection queries and combine them with the previously output tracking queries to form perception queries. Then apply text features to guide the perception queries and input them into the decoder together with the fusion features to capture the semantic targets in the video. Step 7: Use the linear layer to map the decoded query features into the semantic space, and compare the similarity between the decoded query features and the encoded text features to select the query results with higher similarity; Step 8, using a multi-layer perceptron to perform category classification and coordinate regression processing on the screened query results; Step 9: Update the trajectory tracking result, keep the perception query until the next frame and turn it into a tracking query, and repeat steps 5 to 9 until all image sequences in the video are tracked.

2. The multi-level semantic interactive cross-modal tracking method based on Unreal Engine according to claim 1, characterized in that: In step 1, the virtual simulation world is built through Unreal Engine 5, the pedestrian and vehicle trajectories in the world are collected, and multi-target tracking data is generated. Specifically, the resources disclosed in Unreal Engine 5 are used, and the appearance (color and size) of 3D modeling is changed to generate vehicle and crowd instances in the system. Then they are placed in the virtual world, and the internal vehicle and pedestrian movement rules are used to simulate the urban traffic system. Finally, a camera is placed in the world to record the video screen, and the coordinate information and appearance information of the target in the screen can be directly obtained. By using this method, the 3D modeling of objects and the camera position are continuously changed to generate data for multi-target tracking in batches.

3. The multi-level semantic interactive cross-modal tracking method based on Unreal Engine according to claim 1, characterized in that: In step 2, the target trajectory is cleaned, filtered, described, labeled, and quality evaluated to ensure the accuracy and availability of the data. Specifically, first, the trajectories collected in step 1 that are severely blocked by buildings and scenes with fewer targets are filtered out; second, the movement of the target trajectory fragments is described and recorded. Finally, the quality of the processed label data is evaluated to filter out descriptions with longer duration and obvious behavior.

4. The multi-level semantic interactive cross-modal tracking method based on Unreal Engine according to claim 1, characterized in that: In step 3, the motion description annotated in step 2 is combined with the appearance description automatically generated by 3D modeling. The large language model is then used to derive multiple groups of synonyms for the combined words, and some common behavior description phrases are manually screened. Finally, the corresponding description phrases are associated with the trajectory segments, the start and end times of the trajectory are recorded, and a text-trajectory matching pair is generated. This method greatly reduces the cost of manual annotation and ensures the accuracy and objectivity of the data set.

5. The multi-level semantic interactive cross-modal tracking method based on Unreal Engine according to claim 1, characterized in that: In step 4, a multi-target tracking model that integrates multimodal semantic features layer by layer is constructed, which includes an image encoder, a frozen text encoder, a fusion encoder, a semantic guidance module, a tracking decoder, and a semantic relevance branch prediction module. Among them, the text encoder uses RoBERTa and CLIP, and the fusion encoder and fusion decoder both adopt the DETR structure, which contains multiple layers of Transformer. The input of the semantic guidance module sets a learnable query for perceiving different semantic features, with a size of M×N dim , where M is the number of learnable queries.

6. The multi-level semantic interactive cross-modal tracking method based on Unreal Engine according to claim 1, characterized in that: In step 5, the video generated in step 3 is first cut into frames to obtain an image sequence, and the image at time t is feature extracted using a CNN-based backbone network to obtain an image pyramid feature I t Next, the frozen text encoder is used to encode the text instruction into a vector S. Then, I is transformed into t Each layer of is fused with S to generate cross-modal features E t .

7. The multi-level semantic interactive cross-modal tracking method based on Unreal Engine according to claim 1, characterized in that: In step 6, a set of detection queries are initialized with learnable variables, denoted as And combined with the tracking query from the previous frame output Generate a perceptual query, denoted as Q t . Use the semantic guidance module to interact with the perceptual query and the semantic instruction: t Added to its position encoding, it is fed into the self-attention module to integrate Q t The deep features of the text features are linearly mapped to the same latitude, and the calculation formula is: S emb =Linear(S); then use cross attention to Q t Perform text guidance to generate perceptual queries with text prompts. The calculation formula is: Finally, the fusion decoder is used to allow the perceptual query to capture the high-dimensional information of the target in the fusion feature, and the output is recorded as 8. The multi-level semantic interactive cross-modal tracking method based on Unreal Engine as claimed in claim 1, characterized in that: In step 7, the obtained value in step 6 is transformed into Mapped to the semantic feature space, the calculation formula is: Use CLIP text encoder to re-encode the text instructions. The calculation formula is: S clip =Encoder clip (S). And with the powerful cross-modal alignment capability of CLIP, calculation With S clip The similarity between the two modalities finally obtains the relevance between each query and the text instruction, denoted as R cos , the calculation formula is: Further, the query results with similarity greater than 0.5 are filtered out.

9. The multi-level semantic interactive cross-modal tracking method based on Unreal Engine as claimed in claim 1, characterized in that: In step 8, a multi-layer perceptron is used to perform category classification and coordinate regression processing on the filtered query results to obtain the category information and coordinate information of the semantic target in the current frame.

10. The multi-level semantic interactive cross-modal tracking method based on Unreal Engine according to claim 1, characterized in that: In step 9, the trajectory tracking result is updated, and the perception query is retained until the next frame and becomes a tracking query. Steps 5 to 9 are repeated until all image sequences of the video are tracked. Specifically, the trajectory queue is updated using the result. If the target is Capture, then create a new track, if by If captured, the corresponding old trajectory is updated. Record and update the tracking query for the next frame Repeat steps 5 to 9 until all image sequences in the video are tracked.

Citation Information

Patent Citations

  • Depth representation learning and fusion method based on multi-modal trajectory

    CN116956224A

  • Short-time natural language target tracking method based on visual language large model

    CN117746024A

  • Traffic dynamic shielding tracking method and system based on multi-modal fusion knowledge graph

    CN118537835A

  • Multi-source multi-modal activity recognition in aerial video surveillance

    US20170024899A1