Zero-training vehicle re-identification method based on visual language large model

Through the zero-trained vehicle re-identification method based on visual language big model, dynamic multi-grained text generation and adaptive feature fusion are used to solve the problem of vehicle re-identification methods relying on supervised training and single-modal limitations in the prior art, and efficient vehicle re-identification in an open environment is achieved.

CN120220085AActive Publication Date: 2025-06-27HUAZHONG NORMAL UNIV

Patent Information

Application Number
CN202510185338.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-06-27
Estimated Expiration
2045-02-19

AI Technical Summary

Technical Problem

The existing vehicle recognition method relies on supervision training and is difficult to adapt to new scenarios, and the single-modal visual features are susceptible to light changes and viewing angle differences.

Method used

The zero-trained vehicle re-identification method based on visual language big model is adopted to realize vehicle visual feature analysis and similarity sorting through dynamic multi-grained text generation, adaptive feature fusion and combined contrast reasoning.

Benefits of technology

It improves the robustness and generalization ability of vehicle re-identification, and can be applied without tuning in an open environment across cities and cameras, significantly reducing model tuning costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220085A_ABST
    Figure CN120220085A_ABST
Patent Text Reader

Abstract

The invention relates to the field of computer vision and intelligent transportation, and provides a zero-training vehicle re-identification method based on a visual language large model, which comprises the following steps of: generating a dynamic multi-granularity text, analyzing visual features of a vehicle by using the visual language large model to generate structured hierarchical description, and constructing a hierarchical generation framework. Generating basic semantic tags of vehicle types, colors and visual angles and key detail area descriptions guided by local semantics in a layered manner, and dynamically adjusting description levels according to confidence; adaptive feature fusion: realizing adaptive fusion of visual-text features for rough sorting of vehicle similarity; and performing combined comparison reasoning, dividing the TopN images of the vision-text rough sorting list into N / 2 comparison groups, and performing multi-image conjoint analysis by using a visual language large model to realize fine sorting. According to the method, the multi-level fine-grained text description of the vehicle image is generated, and vehicle re-identification in an open scene is realized without training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and intelligent transportation, and particularly to a vehicle re-identification method without training. Background Art

[0002] Vehicle re-identification refers to judging the identity of a vehicle based on the appearance features of the vehicle, which is an important means to make up for the limitations of license plate recognition verification. The existing vehicle re-identification methods mainly have the following limitations: relying on supervised training, the existing methods require a large amount of labeled data for model training and are difficult to adapt to new scenarios; having single-modal limitations, only using visual features is vulnerable to factors such as illumination changes and perspective differences. The existing vision-language large models can effectively generate image descriptions. Therefore, it is urgent to explore a zero-training vehicle re-identification method based on vision-language large models. Summary of the Invention

[0003] In view of the above problems in the prior art, the present invention proposes a zero-training vehicle re-identification method based on a vision-language large model, which effectively improves the robustness and generalization ability of vehicle re-identification.

[0004] To achieve the above object, the technical solution adopted by the present invention is as follows.

[0005] The present invention provides a zero-training vehicle re-identification method based on a vision-language large model, including the following steps:

[0006] (1) Dynamic multi-granularity text generation, using a vision-language large model to analyze the visual features of a vehicle to generate a structured hierarchical description, by constructing a hierarchical generation framework, hierarchically generating basic semantic labels of vehicle type, color, and perspective and key detail area descriptions guided by local semantics, and dynamically adjusting the description level according to the confidence level;

[0007] (2) Adaptive feature fusion, realizing the adaptive fusion of visual-text features for rough sorting of vehicle similarity, including: calculating the weight of visual features guided by image quality, allocating the weight of text features driven by semantic confidence, and rough sorting of visual-text fusion;

[0008] (3) Combinatorial contrast reasoning, dividing the TopN images in the visual-text rough sorting list into N / 2 comparison groups, and using a vision-language large model for multi-image joint analysis to achieve fine sorting, including: automatic image stitching, group contrast based on prompt engineering, and sorting integration based on a probabilistic graph model.

[0009] In the above technical solution, the specific implementation steps of the dynamic multi-granularity text generation are as follows:

[0010] (1-1) Construct a hierarchical generation framework. The first level generates basic descriptions of vehicle models and colors, as well as visible perspectives (front, rear, side, front-side, rear-side). The second level generates detailed descriptions of the roof area. The third level guides the vision-language large model to focus on specific areas according to the perspective to generate detailed descriptions. For example, for the front perspective, it focuses on the description of windows, grilles, headlights, and logos.

[0011] (1-2) Confidence threshold control. Use a multi-modal model such as the CLIP model as a feature detector. When the local feature detection confidence C_d < 0.6, the corresponding detailed description is masked. When the global feature matching degree C_t < 0.4, the re-detection process is triggered.

[0012] In the above technical solution, the specific implementation steps of the adaptive feature fusion are as follows:

[0013] (2-1) Calculate the visual feature weight guided by image quality. Based on the Tenengrad gradient function and entropy function, calculate the clarity Q_v of the image, Q_v ∈ [0,1], and calculate the visual feature weight according to the image clarity.

[0014] (2-2) Allocate the text feature weight driven by semantic confidence. Calculate the text feature weight according to the local feature detection confidence C_d and the global feature matching degree C_t.

[0015] (2-3) Visual-text fusion rough ranking. Normalize the image weight and text weight. Use a multi-modal model that supports long texts, such as LongCLIP, to extract the image features of all query images and test images, and the multi-level text features generated by the multi-level generation framework. Calculate the visual similarity S_v between the query image and all images, and the text similarities at different levels (S_t, S_r, S_d). Calculate the final similarity by weighted summation of the visual similarity and text similarities according to the visual feature weight and text feature weights at different levels. According to the final similarity, obtain the rough ranking list of all query vehicle images.

[0016] In the above technical solution, the specific implementation steps of the combined contrast reasoning are as follows:

[0017] (3-1) Automatic image stitching. Construct a four-image stitching template "[query image][candidate image A][candidate image B][candidate image C]". After scaling the input images to a unified size of 224×224, perform vertical stitching and add learnable spatial position encoding at the stitching boundary.

[0018] (3-2) Based on group comparison with prompt engineering, construct a structured prompt template: "As a vehicle recognition expert, compare the similarities from the following dimensions: 1. Degree of vehicle model contour matching (weight 40%); 2. Consistency of color spectrum (weight 30%); 3. Consistency of detailed features, such as sunroof, logo, grille, windows, etc. (weight 30%). Output format: Optimal match ID > Second-best match ID > Least match ID", and establish an abnormal response filtering mechanism to automatically discard the result when the model output does not conform to the preset format;

[0019] (3-3) Based on the sorting and integration of the probabilistic graphical model, establish a Markov chain transition matrix among the top N (usually taken between 50 and 100) images in the rough ranking list, calculate the steady-state distribution probability through the random walk algorithm, and then perform confidence-weighted output to form the final sorted list.

[0020] Compared with the prior art, the beneficial effects of the zero-training vehicle re-identification method based on the vision-language large model of the present invention are as follows:

[0021] 1. Break through the dependence of traditional vehicle re-identification on closed-scene data, and can be directly applied to open environments across cities and cameras without fine-tuning, significantly reducing the model tuning cost;

[0022] 2. Construct a dynamic multi-granularity text generation module to generate multi-granularity semantic descriptions including vehicle global attributes (vehicle model, color) and local details (logo, decoration, damage), and can support two-way queries of "searching images by text" and "searching text by image" at the same time. Brief Description of the Drawings

[0023] Figure 1 It is the implementation flowchart of the embodiment of the present invention. Detailed Embodiment

[0024] The following will describe the present invention in detail with reference to the drawings and specific embodiments.

[0025] As Figure 1 shown, a zero-training vehicle re-identification method based on the vision-language large model provided by the embodiment of the present invention includes the following steps:

[0026] Step 1, dynamic multi-granularity text generation, use the vision-language large model to perform visual feature analysis on the vehicle to generate a structured hierarchical description, and by constructing a hierarchical generation framework, hierarchically generate basic semantic labels of vehicle model, color, and perspective and descriptions of key detail areas guided by local semantics, and dynamically adjust the description level according to the confidence level;

[0027] Step 2, Adaptive Feature Fusion, to achieve the adaptive fusion of visual-text features for the rough ranking of vehicle similarity, including: calculating the weights of visual features guided by image quality, allocating the weights of text features driven by semantic confidence, and rough ranking of visual-text fusion;

[0028] Step 3, Combinatorial Contrast Reasoning, divides the TopN images in the visual-text rough ranking list into N / 2 comparison groups, and uses a vision-language large model for multi-image joint analysis to achieve fine ranking, including: automatic image stitching, group contrast based on prompt engineering, and sorting integration based on probabilistic graphical models.

[0029] Next, this method will be described in detail.

[0030] The above Step 1 specifically includes the following steps:

[0031] (1-1) Construct a hierarchical generation framework. The first layer generates basic descriptions of vehicle models, colors, and visible perspectives (front, rear, side, front-side, rear-side). The second layer generates detailed descriptions of the roof area. The third layer guides the vision-language large model to focus on specific areas according to the perspective to generate detailed descriptions, such as focusing on the descriptions of windows, grilles, headlights, and logos for the front perspective.

[0032] (1-2) Confidence Threshold Control. Use a multi-modal model such as the CLIP model as a feature detector. When the local feature detection confidence C_d < 0.6, the corresponding detailed description is masked. When the global feature matching degree C_t < 0.4, the re-detection process is triggered.

[0033] The above Step 2 specifically includes the following steps:

[0034] (2-1) Calculating the weights of visual features guided by image quality. Based on the Tenengrad gradient function and entropy function, calculate the clarity Q_v of the image, Q_v ∈ [0,1]. The formula for calculating the weights of visual features is as follows:

[0035] α = sigmoid(Q_v × 0.5) × 0.7 (Formula 1)

[0036] (2-2) Allocating the weights of text features driven by semantic confidence. The formula for calculating the weights of text features is as follows:

[0037] β = 0.25 × C_t + 1.0 × C_d (Formula 2)

[0038] (2-3) Rough ranking of visual-text fusion. Normalize the image weights and text weights. The calculation formula is as follows:

[0039]

[0040] Using a multi-modal model that supports long texts, such as LongCLIP, extract the image features of all query images and test images, as well as the multi-level text features generated by a multi-level generation framework. Calculate the visual similarity \(S_v\) between the query image and all images, and the text similarities at different levels (\(S_t\), \(S_r\), \(S_d\)). The formula for the final similarity is as follows:

[0041]

[0042] According to the final similarity, obtain the rough ranking list of all query vehicle images.

[0043] The above step 3 specifically includes the following steps:

[0044] (3-1) Automatically splice images to construct a four-image splicing template "[query image][candidate image A][candidate image B][candidate image C]". After scaling the input images to a unified size of 224×224, perform vertical splicing and add learnable spatial position encoding at the splicing boundary.

[0045] (3-2) Group comparison based on prompt engineering to construct a structured prompt template: "As a vehicle recognition expert, please compare the similarities from the following dimensions: 1. Model contour matching degree (weight 40%) 2. Color spectrum consistency (weight 30%) 3. Consistency of detail features, such as sunroof, logo, grille, windows, etc. (weight 30%). Output format: optimal matching ID > second-best matching ID > least matching ID", and establish an abnormal response filtering mechanism to automatically discard the result when the model output does not conform to the preset format.

[0046] (3-3) Ranking integration based on a probabilistic graph model. Establish a Markov chain transition matrix between the first N (usually taken between 50 and 100) images in the rough ranking list. Calculate the steady-state distribution probability through the random walk algorithm, and then perform confidence-weighted output to form the final ranking list.

[0047] The above step (3-3) specifically includes the following steps:

[0048] (3-3-1) Construct a Markov transition matrix. First, define the state space, and use the set \(G=\{g1,g2,...,g\}\) of the first N images in the rough ranking list as the state nodes of the Markov chain. Each state \(s\) corresponds to the image \(g\). N} as the state nodes of the Markov chain. Each state \(s\) i , corresponding to the image \(g\) iPerform similarity score conversion based on the output result of (3-2), numerically process each group of comparison inference results (such as "g1>g2>g3"). Assign a score of score1 = 1.0 to the image ranked 1, a score of score2 = 0.7 to the image ranked 2, and a score of score3 = 0.4 to the image ranked 3. Then perform normalization processing on each group of scores, and the calculation formula is as follows:

[0049]

[0050] where j and k are other images in the comparison group. Then, calculate the transition probability and construct the state transition matrix T∈R N×N . Among them, the element T ij represents the probability of transitioning from state s i to s j , and the specific calculation formula is as follows:

[0051]

[0052] Among them, the diagonal element T ii is set to the self-loop probability ∈ = 0.05 to ensure the ergodicity of the Markov chain.

[0053] (3-3-2) Random walk to calculate the steady-state distribution, initialize the probability distribution Then perform damped random walk iteration, set the damping coefficient d = 0.85, and perform iterative update. The calculation formula is as follows:

[0054] π (t+1) = d·T T π (t) +(1 - d)·π 0 (Formula 7)

[0055] Obtain the final steady-state probability π * indicating the similarity of each test vehicle image.

[0056] (3-3-3) Confidence-weighted output, perform weighted calculation according to the number of occurrences C i of the image in the comparison group. The calculation formula is as follows:

[0057] S i = π i ×log(1 + C i ) (Formula 8)

[0058] Use the weighted similarity as the final similarity to obtain the final sorted list.

[0059] The content not described in detail in this specification belongs to the prior art known to those skilled in the art.

[0060] Although the present invention has been described in detail above and embodiments have been given, the present invention and the applicable embodiments are not limited thereto. Those skilled in the art of the present technology can make various modifications according to the principles of the present invention, and some of the methods in the present invention can also be applied to other systems. Therefore, any modifications made according to the principles of the present invention should be understood to fall within the protection scope of the present invention.

Claims

1. A zero-training vehicle re-identification method based on a large visual language model, characterized in that: The method comprises the following steps: (1) Dynamic multi-granularity text generation: Use the large visual language model to analyze the visual features of the vehicle and generate a structured hierarchical description. By building a hierarchical generation framework, basic semantic labels for vehicle type, color, and perspective, as well as descriptions of key detail areas guided by local semantics, are generated in layers, and the description level is dynamically adjusted based on the confidence level. (2) Adaptive feature fusion, which realizes the adaptive fusion of visual and textual features for rough sorting of vehicle similarity, including: visual feature weight calculation guided by image quality, text feature weight allocation driven by semantic confidence, and visual and textual fusion rough sorting; (3) Combinatorial contrastive reasoning: divide the top N images in the visual-textual rough ranking list into N / 2 contrastive groups, and use the visual language large model to perform multi-image joint analysis to achieve precise ranking, including: automatic image stitching, group comparison based on prompt engineering, and ranking integration based on probabilistic graph models.

2. The zero-training vehicle re-identification method based on a large visual language model according to claim 1 is characterized in that The dynamic multi-granularity text generation described in step (1) is specifically implemented as follows: (1-1) Construct a hierarchical generation framework. The first level generates basic descriptions of vehicle models and colors and visible perspectives, including front, rear, side, front-side, and rear-side. The second level generates detailed descriptions of the roof area. The third level generates detailed descriptions by guiding the visual language model to focus on specific areas based on the perspective. (1-2) Confidence threshold control, using the multimodal model as a feature detector, when the local feature detection confidence C_d < 0.6, the corresponding detailed description is shielded, and when the global feature matching degree C_t < 0.4, the re-detection process is triggered.

3. The zero-training vehicle re-identification method based on a large visual language model according to claim 1 is characterized in that The specific implementation steps of the adaptive feature fusion in step (2) are as follows: (2-1) Image quality guided visual feature weight calculation, based on Tenengrad gradient function and entropy function, calculate the image clarity Q_v, Q_v∈[0,1], and calculate the visual feature weight according to the image clarity; (2-2) Semantic confidence-driven text feature weight allocation, which calculates the text feature weight based on the local feature detection confidence C_d and the global feature matching degree C_t; (2-3) Visual-text fusion rough sorting, normalize image weights and text weights, use a multimodal model that supports long texts to extract image features of all query images and test images and multi-level text features generated by a multi-level generation framework, calculate the visual similarity S_v and text similarities of different levels (S_t, S_r, S_d) between the query image and all images respectively, and calculate the final similarity by weighted summation of visual similarity and text similarity based on the visual feature weights and text feature weights of different levels. Based on the final similarity, a rough sorted list of all query vehicle images is obtained.

4. The zero-training vehicle re-identification method based on a large visual language model according to claim 1 is characterized in that The combined contrast reasoning described in step (3) is specifically implemented as follows: (3-1) Automatic image stitching: construct a four-image stitching template "[query image][candidate image A][candidate image B][candidate image C]", scale the input images to a uniform size of 224×224, and then stitch them vertically, adding a learnable spatial position code to the stitching boundary; (3-2) Based on the group comparison of the prompt engineering, a structured prompt template is constructed: "As a vehicle recognition expert, please compare the similarity from the following dimensions:

1. Model outline matching (weight 40%) 2. Color spectrum consistency (weight 30%) 3. Detail feature consistency, such as sunroof, car logo, grid, window, etc. (weight 30%). Output format: best matching ID> second best matching ID> least matching ID", and an abnormal response filtering mechanism is established to automatically discard the result when the model output does not conform to the preset format; (3-3) Based on the sorting integration of the probabilistic graphical model, a Markov chain transfer matrix is ​​established between the first N images in the rough sorting list. The steady-state distribution probability is calculated by the random walk algorithm, and the confidence-weighted output is performed to form the final sorting list, with N ranging from 50 to 100.

5. The zero-training vehicle re-identification method based on a large visual language model according to claim 4 is characterized in that The sorting integration based on the probabilistic graphical model specifically includes: (3-3-1) Construct the Markov transition matrix. First, define the state space and convert the first N images in the rough sorting list into a set G = {g1, g2, ..., g N } as the state node of the Markov chain, each state s i , corresponding to image g i , perform similarity score conversion according to the output results of step (3-2), perform numerical processing on each set of comparative reasoning results, and then standardize each set of scores, calculate the transition probability based on the standardized scores, and construct the state transfer matrix T∈R N×N , where the diagonal elements T ii The self-loop probability is set to ∈=0.05 to ensure the ergodicity of the Markov chain; (3-3-2) Random walk calculation of steady-state distribution and initialization of probability distribution Perform damped random walk iteration, set the damping coefficient d = 0.85, perform iterative update, and obtain the final steady-state probability π * Indicates the similarity of each test vehicle image; (3-3-3) Confidence weighted output, based on the number of times the image appears in the comparison group C i Calculate the weights, take the weighted similarity as the final similarity, and get the final sorted list.

Citation Information

Patent Citations

  • Vehicle multi-target detection and trajectory tracking method based on re-identification

    CN111914664A

  • Fine-grained video-text retrieval method based on context Transform network

    CN114282060A

  • Visual intention analysis method and system based on cross-modal pyramid alignment

    CN116434255A

  • Multi-granularity image-text matching method and system based on deep fusion

    CN117093692A

  • Image-text cross-modal vehicle retrieval model training method in vehicle dense scene

    CN118968516A

Cited By

  • Maintenance guidance method and device for household appliances

    CN120670561A

  • Method and apparatus for maintenance guidance of a household appliance

    CN120670561B