Knowledge distillation method for text and graph cross-modal pedestrian retrieval model

By employing a three-stage knowledge distillation method to progressively transfer knowledge from the image and text encoders of the student model, the problem of large parameters and high computational cost in cross-modal learning models is solved, achieving lightweight design and performance improvement. This method is suitable for cross-modal pedestrian retrieval tasks involving text and images.

CN120930720APending Publication Date: 2025-11-11UESTC (SHENZHEN) ADVANCED RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510963970.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing cross-modal learning models have a large number of parameters and high computational costs, making it difficult to meet the needs of practical applications.

Method used

A three-stage knowledge distillation method is adopted to perform stepwise knowledge transfer on the image encoder and text encoder of the student model, and the cross-modal matching ability is improved through joint optimization. Supervised training is carried out by combining task loss and distillation loss to obtain a lightweight model.

Benefits of technology

It improves the model's ability to capture fine-grained semantic information and its cross-modal feature alignment performance, reduces model complexity and computational cost, and enhances the overall performance and robustness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120930720A_ABST
    Figure CN120930720A_ABST
Patent Text Reader

Abstract

The invention discloses a knowledge distillation method for a text and graph cross-modal pedestrian retrieval model, which comprises the following steps: constructing a teacher model and a student model, and initializing the student model, both the teacher model and the student model being provided with a text encoder and an image encoder; performing three-stage knowledge distillation on the student model, in the first stage, performing knowledge distillation on an image encoder of the student model through the teacher model, and in the second stage, performing knowledge distillation on a text encoder of the student model through the teacher model; in the third stage, knowledge distillation is carried out on a text encoder and an image encoder of the student model at the same time through the teacher model; and training the student model until convergence according to the task loss and the distillation loss of each stage to obtain a lightweight student model. According to the invention, association of the student model and the teacher model in cross-modal task learning can be enhanced, and the ability of the student model to obtain text image multi-modal feature height alignment is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of machine learning technology, and in particular to a knowledge distillation method for cross-modal pedestrian retrieval models for text and images. Background Technology

[0002] Cross-modal pedestrian retrieval based on text and images is an interdisciplinary research area between text image retrieval and pedestrian re-identification. Its core task is to retrieve semantically matching target images from a pedestrian image database based on given text descriptions. This technology effectively overcomes the limitations of traditional pedestrian re-identification techniques in scenarios lacking query images and has broad application prospects in real-world scenarios. For example, in intelligent security, it assists police in quickly identifying criminal suspects based on witness descriptions, improving case-solving efficiency; in personnel screening, it quickly identifies suspicious individuals, preventing potential security risks; and in intelligent photo albums, it quickly locates target tourist images based on natural language descriptions, improving image management efficiency and user experience. The development of cross-modal pedestrian retrieval technology based on text and images will provide strong support for building a more intelligent, efficient, and safe social environment.

[0003] Although cross-modal person retrieval based on text and images is an emerging research area in recent years, it has already attracted widespread attention from the academic community. Early research methods were mainly based on a two-stream encoder architecture, using convolutional neural networks and recurrent neural networks for image and text feature extraction, respectively, and improving retrieval performance by designing complex cross-modal alignment mechanisms. Among them, the GAN-RNN (Generative Adversarial Network-Recurrent Neural Network) method uses VGG (Visual Geometry Group) and LSTM (Long Short-Term Memory) as image and text encoders, respectively. By dynamically calculating the association weights between each word and visual feature units, it highlights the contribution of key semantic information to the matching results and achieves word-level feature alignment. The CMPM (Cross-Modal Projection Matching) method uses MobileNet (a lightweight convolutional neural network) and Bi-LSTM (Bidirectional LSTM) as image and text encoders, respectively. It proposes cross-modal projection matching loss functions and cross-modal projection classification loss functions, directly training projection matching in the global feature space, achieving significant performance improvements on benchmark datasets. Furthermore, to mitigate the challenges posed by the modal differences between text and images, some studies have explored using the same network structure to extract cross-modal features, i.e., using convolutional network structures to encode both text and image information, thus mitigating the differences between different information in text and images through the consistency of network architecture. In summary, early research methods primarily improved model performance by introducing specific network modules or complex preprocessing procedures. While these methods effectively improved retrieval accuracy, they also increased model complexity and computational cost, posing challenges for practical applications.

[0004] In recent years, the breakthrough development of Transformer has provided a new paradigm for cross-modal learning. Based on large-scale text-image pair pre-trained visual language models, a unified mapping of cross-modal semantic spaces has been achieved through a contrastive learning framework. In the field of text-image cross-modal person retrieval, researchers have proposed a series of innovative methods based on large-scale visual language pre-trained models. For example, the IRRA (Image-Text Retrieval with Relevance Aggregation) method uses the CLIP (Contrastive Language-Image Pre-training) pre-trained model for initialization and designs an IRRA module, enhancing the model's ability to model implicit text-image relationships through a text masking learning mechanism. The RaSa (Relation-aware Semantic Alignment) method introduces relation-aware and sensitivity-aware learning branches based on the ALBEF (Align before Fuse) model, effectively improving the model's ability to discriminate weak positive sample pairs under noise interference, further promoting the development of text-image cross-modal person retrieval methods designed based on visual language pre-trained models. Although cross-modal pedestrian retrieval methods based on visual language pre-trained models have achieved remarkable results in academia, their large parameter scale and high computational cost make it difficult to meet the needs of practical applications. Summary of the Invention

[0005] The purpose of this application is to provide a knowledge distillation method for cross-modal pedestrian retrieval models based on text and images, in order to solve the technical problems of existing cross-modal learning methods, such as large scale of parameters and high computational cost. The various technical effects of the preferred solutions among the many technical solutions provided in this application are detailed below.

[0006] To achieve the above objectives, this application provides the following technical solutions:

[0007] Firstly, this application provides a knowledge distillation method for a cross-modal pedestrian retrieval model based on text and images, comprising: constructing a teacher model and a student model; initializing the student model, wherein both the teacher model and the student model have a text encoder and an image encoder; performing three-stage knowledge distillation on the student model, wherein in the first stage, knowledge distillation is performed on the image encoder of the student model through the teacher model; in the second stage, knowledge distillation is performed on the text encoder of the student model through the teacher model; and in the third stage, knowledge distillation is performed simultaneously on the text encoder and image encoder of the student model through the teacher model; and training the student model until convergence based on the task loss and distillation loss of each stage, thereby obtaining a lightweight student model.

[0008] In some embodiments, the knowledge distillation of the image encoder of the student model through the teacher model in the first stage includes:

[0009] The student model and the teacher model respectively extract features from the first image training data to obtain a first output result and a second output result. The teacher model extracts features from the first text training data to obtain a third output result. The first task loss and the first distillation loss are calculated based on the first output result, the second output result and the third output result. The parameters of the image encoder of the student model are adjusted by minimizing the first task loss and the first distillation loss.

[0010] In some embodiments, the first task loss includes a first contrastive learning loss, and the first distillation loss includes a first cross-entropy loss, wherein the first contrastive learning loss is the contrastive learning loss between the first output and the third output, and the first cross-entropy loss is the cross-entropy loss between the matching probability distribution of the first output and the second output and the true label.

[0011] In some embodiments, the knowledge distillation of the text encoder of the student model through the teacher model in the second stage includes: extracting features from the second image training data through the student model and the teacher model respectively to obtain a fourth output result and a fifth output result; extracting features from the second text training data through the student model and the teacher model respectively to obtain a sixth output result and a seventh output result; calculating a second task loss and a second distillation loss based on the fourth output result, the fifth output result, the sixth output result, and the seventh output result; and adjusting the parameters of the text encoder of the student model by minimizing the second task loss and the second distillation loss.

[0012] In some embodiments, the second task loss includes a second contrastive learning loss and a third contrastive learning loss, wherein the second contrastive learning loss is the contrastive learning loss between the fifth output result and the sixth output result, and the third contrastive learning loss is the contrastive learning loss between the fourth output result and the sixth output result.

[0013] In some embodiments, the second distillation loss includes a second cross-entropy loss and a first KL divergence loss, wherein the second cross-entropy loss is the cross-entropy loss between the matching probability distribution of the sixth output and the seventh output and the true label, and the first KL divergence loss is the KL divergence loss of fitting the matching output probability distribution of the fourth output and the sixth output to the matching output probability distribution of the fifth output and the seventh output.

[0014] In some embodiments, the step of simultaneously performing knowledge distillation on the text encoder and image encoder of the student model through the teacher model in the third stage includes: extracting features from the third image training data through the student model and the teacher model to obtain an eighth output result and a ninth output result; extracting features from the third text training data through the student model and the teacher model to obtain a tenth output result and an eleventh output result; calculating a third task loss and a third distillation loss based on the eighth output result, the ninth output result, the tenth output result, and the eleventh output result; and adjusting the parameters of the text encoder and the image encoder of the student model by minimizing the third task loss and the third distillation loss.

[0015] In some embodiments, the third task loss includes a fourth contrastive learning loss and a fifth contrastive learning loss, wherein the fourth contrastive learning loss is the contrastive learning loss between the eighth output result and the eleventh output result, and the fifth contrastive learning loss is the contrastive learning loss between the ninth output result and the tenth output result.

[0016] In some embodiments, the third distillation loss includes a third cross-entropy loss, a second KL divergence loss, a feature-wise distance distillation loss, and a similarity-wise distance distillation loss; wherein, the third cross-entropy loss is the cross-entropy loss between the matching probability distributions of the tenth and eleventh output results and the true labels; the second KL divergence loss is the KL divergence loss fitted by the matching output probability distributions of the eighth and tenth output results to the matching output probability distributions of the ninth and eleventh output results; the feature-wise distance distillation loss includes a first feature-wise distance distillation loss between the eighth and ninth output results, and a second feature-wise distance distillation loss between the tenth and eleventh output results; the similarity-wise distance distillation loss includes a first similarity-wise distance distillation loss based on the similarity between the tenth and eleventh output results, and a second similarity-wise distance distillation loss based on the similarity between the eighth and ninth output results.

[0017] Secondly, this application provides a computer program product stored on a data carrier and designed to perform the knowledge distillation method for a text-image cross-modal pedestrian retrieval model as described above.

[0018] Implementing one of the technical solutions described in this application has the following advantages or beneficial effects: In this application, the student model undergoes a three-stage knowledge distillation process. The first and second stages focus on single-modal knowledge transfer in images and text, respectively, ensuring that the student model fully absorbs the knowledge from the teacher model within a single modality. In the third stage, by jointly optimizing the student model's text encoder and image encoder, the learning space of the student model for the teacher model's cross-modal matching ability is expanded, enhancing the model's ability to capture fine-grained semantic information, thereby improving the model's performance in fine-grained feature alignment. Furthermore, this application trains the student model based on the task loss function and distillation loss function at each stage. By designing a complementary supervision strategy for the task loss function and distillation loss function, the association between the student model and the teacher model in cross-modal task learning is strengthened. Appropriate guidance from the teacher model in text-image matching knowledge distillation is also introduced, further enhancing the student model's ability to achieve high-level alignment of multimodal features in text and images, thereby improving the overall performance and robustness of the model. Attached Figure Description

[0019] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings:

[0020] Figure 1 This is a flowchart illustrating the knowledge distillation method for a text-image cross-modal pedestrian retrieval model according to an embodiment of this application;

[0021] Figure 2 This is a schematic diagram of the knowledge distillation framework of an embodiment of this application;

[0022] Figure 3 This is a schematic diagram of the complementary supervision strategy in the third stage of an embodiment of this application;

[0023] Figure 4 This is a schematic diagram of the compression process of the three-stage progressive knowledge distillation model in this application embodiment;

[0024] Figure 5 This is a structural block diagram of the processing device according to an embodiment of this application.

[0025] In the diagram: 1. Processing device; 10. Memory; 11. Processor. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of this application clearer, various exemplary embodiments described below will be referenced to the accompanying drawings, which form part of the exemplary embodiments and depict various exemplary embodiments that may be adopted to implement this application. Unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. It should be understood that they are merely examples of processes, methods, and apparatuses consistent with some aspects of this application disclosed as detailed in the appended claims, and other embodiments may be used, or structural and functional modifications may be made to the embodiments listed herein without departing from the scope and spirit of this application.

[0027] In the description of this application, it should be understood that the terms "center," "longitudinal," "lateral," etc., indicate the orientation or positional relationship based on the accompanying drawings, and are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the referred element must have a specific orientation, or be constructed and operated in a specific orientation. The terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. The term "multiple" means two or more. The terms "connected" and "linked" should be interpreted broadly, for example, they can be fixed connections, detachable connections, integral connections, mechanical connections, electrical connections, communication connections, direct connections, indirect connections through an intermediate medium, and can be the internal connection of two elements or the interaction relationship between two elements. The term "and / or" includes any and all combinations of one or more of the related listed items. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.

[0028] To illustrate the technical solutions described in this application, specific embodiments are provided below, showing only the parts related to the embodiments of this application.

[0029] This application relates to a knowledge distillation method that can be applied to a text-image cross-modal pedestrian retrieval model. This model is generally used to retrieve matching target images from a pedestrian image database or to obtain corresponding text descriptions from images. It includes an image encoder, a text encoder, and a cross-modal alignment module. The image encoder can be used to extract pedestrian image features (e.g., pose, clothing color), and the text encoder can be used to extract semantic features of the text description (e.g., objects, spatial relationships). A teacher model can be used to guide the training of student models. The trained student models, as lightweight models, can be used for practical deployment.

[0030] The knowledge distillation proposed in this application is a three-stage progressive knowledge distillation framework.

[0031] like Figures 1 to 4 As shown, this application provides a knowledge distillation method for a text-image cross-modal pedestrian retrieval model, including the following steps (steps S1 to S3):

[0032] S1. Construct teacher and student models, and initialize the student model. Both the teacher and student models have text encoders and image encoders.

[0033] Specifically, the student model can be initialized using pre-trained weights, that is, the student model can be initialized using a publicly available pre-trained model. In other embodiments, the student model can also be initialized through random initialization, teacher-guided initialization, or other methods.

[0034] In some embodiments, the knowledge distillation method for a text-image cross-modal pedestrian retrieval model may further include fine-tuning the teacher model. Specifically, the teacher model may be fine-tuned with all parameters, with partial layer adjustments, or with task-adaptive adjustments. This enables the teacher model to adapt to new tasks and improves output quality.

[0035] S2. Perform three-stage knowledge distillation on the student model. In the first stage, the image encoder of the student model is distilled using the teacher model. In the second stage, the text encoder of the student model is distilled using the teacher model. In the third stage, the text encoder and image encoder of the student model are distilled simultaneously using the teacher model.

[0036] In some embodiments, the knowledge distillation of the image encoder of the student model by the teacher model in the first stage may include: extracting features from the first image training data by the student model and the teacher model respectively to obtain a first output result and a second output result; extracting features from the first text training data by the teacher model to obtain a third output result; calculating a first task loss and a first distillation loss based on the first output result, the second output result and the third output result; and adjusting the parameters of the image encoder of the student model by minimizing the first task loss and the first distillation loss.

[0037] Specifically, the first image training data can be feature extracted by the image encoder of the student model to obtain the first output result, the first image training data can be feature extracted by the image encoder of the teacher model to obtain the second output result, and the first text training data can be feature extracted by the text encoder of the teacher model to obtain the third output result.

[0038] In some embodiments, calculating the first task loss and the first distillation loss based on the first output result, the second output result, and the third output result may include: calculating the first task loss based on the first output result and the third output result; and calculating the first distillation loss based on the first output result and the second output result.

[0039] In some embodiments, the first task loss may include a first contrastive learning loss, and the first distillation loss may include a first cross-entropy loss. Specifically, the first contrastive learning loss may be a contrastive learning loss between a first output and a third output, and the first cross-entropy loss may be a cross-entropy loss between the matching probability distribution of the first and second outputs and the true label.

[0040] In the first stage, the first output can be the output features of the image encoder of the student model, the second output can be the output features of the image encoder of the teacher model, and the third output can be the output features of the text encoder of the teacher model. All output features are L2 norm normalized.

[0041] For ease of description, F is defined as follows. a and F b ∈R N×d Define the features for the output features of the two encoders after L2 norm normalization. and characteristics The matching output probability distribution is the target probability distribution, and the feature F a and feature F b The matching output probability distribution is the predicted probability distribution; let f denote the encoder function, and F denote the output features after L2 norm normalization following encoder processing. Define the input batch size as N, d as the feature dimension, and the text encoder as f. t The image encoder is defined as f v The teacher model is labeled T, and the student model is labeled S.

[0042] In some embodiments, feature F a To feature F b The cross-entropy loss between the matching output probability and the true label can be calculated using the following formula:

[0043]

[0044] in,

[0045]

[0046] In some embodiments, feature F a With feature F b The contrast loss function can be calculated using the following formula:

[0047]

[0048] In some embodiments, the KL divergence of the predicted probability distribution fitted to the target probability distribution can be calculated using the following formula:

[0049]

[0050] In some embodiments, feature F a and feature F b The feature-wise distance loss function can be calculated using the following formula:

[0051]

[0052] In some embodiments, feature F a To feature F b Similarity and features and characteristics The per-similarity distance distillation loss function between similarities can be calculated using the following formula:

[0053]

[0054] In some embodiments, the first task loss can be represented by the following formula:

[0055]

[0056] in, The output features of the text encoder representing the teacher model, The output features of the image encoder for the student model, This represents the loss function used to compare the first and third output results.

[0057] In some embodiments, the first distillation loss can be expressed by the following formula:

[0058]

[0059] in, It could refer to the cross-entropy loss function between the first and third output results.

[0060] In some embodiments, the knowledge distillation of the text encoder of the student model by the teacher model in the second stage may include: extracting features from the second image training data by the student model and the teacher model respectively to obtain a fourth output result and a fifth output result; extracting features from the second text training data by the student model and the teacher model respectively to obtain a sixth output result and a seventh output result; calculating the second task loss and the second distillation loss based on the fourth output result, the fifth output result, the sixth output result and the seventh output result; and adjusting the parameters of the text encoder of the student model by minimizing the second task loss and the second distillation loss.

[0061] Specifically, the image encoder of the student model can be used to extract features from the second image training data to obtain the fourth output result; the image encoder of the teacher model can be used to extract features from the second image training data to obtain the fifth output result; the text encoder of the student model can be used to extract features from the second text training data to obtain the sixth output result; and the text encoder of the teacher model can be used to extract features from the second text training data to obtain the seventh output result.

[0062] In some embodiments, calculating the second task loss and the second distillation loss based on the fourth output result, the fifth output result, the sixth output result, and the seventh output result may include: calculating the second task loss based on the fourth output result, the fifth output result, and the sixth output result, and calculating the second distillation loss based on the fourth output result, the fifth output result, the sixth output result, and the seventh output result.

[0063] In some embodiments, the second task loss may include a second contrastive learning loss and a third contrastive learning loss, wherein the second contrastive learning loss may be a contrastive learning loss between the fifth output and the sixth output, and the third contrastive learning loss may be a contrastive learning loss between the fourth output and the sixth output.

[0064] In some embodiments, the second distillation loss may include a second cross-entropy loss and a first KL divergence loss, wherein the second cross-entropy loss may be the cross-entropy loss between the matching probability distribution of the sixth and seventh output results and the true label, and the first KL divergence loss may be the KL divergence loss that fits the matching output probability distribution of the fourth and sixth output results to the matching output probability distribution of the fifth and seventh output results.

[0065] In the second stage, the fourth output can be the output features of the image encoder of the student model, the fifth output can be the output features of the image encoder of the teacher model, the sixth output can be the output features of the text encoder of the student model, and the seventh output can be the output features of the text encoder of the teacher model. All output features are L2 norm normalized.

[0066] In some embodiments, the loss of the second task can be expressed by the following formula:

[0067]

[0068] in, The contrastive learning loss is used between the student model's text encoder and the teacher model's image encoder. The contrastive learning loss is used between the text and image encoders within the student model.

[0069] In some embodiments, the second distillation loss can be expressed by the following formula:

[0070]

[0071] in, For the second cross-entropy loss, This is the first KL divergence loss.

[0072] In some embodiments, the first KL divergence loss can be calculated using the following formula:

[0073]

[0074] In some embodiments, the third stage of knowledge distillation of the text encoder and image encoder of the student model by the teacher model may include: extracting features from the third image training data by the student model and the teacher model to obtain the eighth and ninth output results; extracting features from the third text training data by the student model and the teacher model to obtain the tenth and eleventh output results; calculating the third task loss and the third distillation loss based on the eighth, ninth, tenth, and eleventh output results; and adjusting the parameters of the text encoder and the image encoder of the student model by minimizing the third task loss and the third distillation loss.

[0075] Specifically, the third image training data can be feature extracted by the image encoder of the student model to obtain the eighth output result, the third image training data can be feature extracted by the image encoder of the teacher model to obtain the ninth output result, the third text training data can be feature extracted by the text encoder of the student model to obtain the tenth output result, and the third text training data can be feature extracted by the text encoder of the teacher model to obtain the eleventh output result.

[0076] In some embodiments, the third task loss may include a fourth contrastive learning loss and a fifth contrastive learning loss, wherein the fourth contrastive learning loss may be a contrastive learning loss between the eighth output and the eleventh output, and the fifth contrastive learning loss may be a contrastive learning loss between the ninth output and the tenth output.

[0077] In some embodiments, the third distillation loss may include a third cross-entropy loss, a second KL divergence loss, a feature-wise distance distillation loss, and a similarity-wise distance distillation loss. Specifically, the third cross-entropy loss may be the cross-entropy loss between the matching probability distributions of the tenth and eleventh output results and the true labels; the second KL divergence loss may be the KL divergence loss obtained by fitting the matching output probability distributions of the eighth and tenth output results to the matching output probability distributions of the ninth and eleventh output results; the feature-wise distance distillation loss may include a first feature-wise distance distillation loss between the eighth and ninth output results, and a second feature-wise distance distillation loss between the tenth and eleventh output results; and the similarity-wise distance distillation loss may include a first similarity-wise distance distillation loss based on the similarity between the tenth and eleventh output results, and a second similarity-wise distance distillation loss based on the similarity between the eighth and ninth output results.

[0078] In the third stage, the eighth output can be the output features of the image encoder of the student model, the ninth output can be the output features of the image encoder of the teacher model, the tenth output can be the output features of the text encoder of the student model, and the eleventh output can be the output features of the text encoder of the teacher model. All output features are normalized to L2 norm.

[0079] In some embodiments, the loss of the third task can be calculated using the following formula:

[0080]

[0081] in, This represents the fourth contrastive learning loss. This represents the learning loss in the fifth contrastive comparison.

[0082] In some embodiments, the third distillation loss can be calculated using the following formula:

[0083]

[0084] in, This represents the third cross-entropy loss. This represents the second KL divergence loss. This represents the distillation loss per characteristic distance. This represents the distance distillation loss per similarity.

[0085] In some embodiments, the feature-per-feature distance distillation loss can be calculated using the following formula:

[0086]

[0087] in, This represents the first characteristic distance distillation loss. This represents the second characteristic distance distillation loss.

[0088] In some embodiments, the feature-per-feature distance distillation loss can be calculated using the following formula:

[0089]

[0090] in, This represents the first characteristic distance distillation loss. This represents the second characteristic distance distillation loss.

[0091] In the third stage, this application designs a complementary supervision strategy to address the potential competition between multiple task losses and distillation losses. Specifically, in the third task loss, this application only uses comparative learning of the student model and the teacher model's cross-modal approach, which can effectively alleviate the competition between the text and image encoders within the student model for task learning and the cross-modal encoders of the student and teacher models for task learning.

[0092] Meanwhile, the third distillation loss in this application combines the different encoder structures of the student model and the teacher model to construct a cross-entropy distillation loss function, namely the third cross-entropy loss; a KL divergence distillation loss function, namely the second KL divergence loss; a feature-wise distance distillation loss function; and a similarity-wise distance distillation loss function. Specifically, the third cross-entropy loss can be used to fit the predicted probability distribution of the teacher model between single modalities; the KL divergence distillation loss function can be used to fit the predicted probability distribution of the teacher model between cross-modal matches; the feature-wise distance distillation loss function can be used to narrow the gap in feature extraction between the student model and the teacher model; and the similarity-wise distance distillation loss function can be used to improve the student model's ability to extract fine-grained features within a single modality.

[0093] S3. Based on the task loss and distillation loss at each stage, train the student model until convergence to obtain a lightweight student model.

[0094] In some embodiments, the overall loss at each stage is calculated by weighted summation as follows:

[0095]

[0096] in, For mission losses, The value represents the distillation loss, and α is the weighting coefficient.

[0097] The following details the data utilized by the knowledge distillation method for the text-image cross-modal pedestrian retrieval model of this application, as well as the settings of the training parameters for the text-image cross-modal pedestrian retrieval model of this application:

[0098] Regarding datasets, this application can utilize three commonly used public datasets in the field of cross-modal pedestrian retrieval: the CUHK-PEDES dataset, the ICFG-PEDES dataset, and the RSTPReid dataset. That is, the image training data and text training data in this application can be derived from these datasets. The teacher model in this application can use the IRRA algorithm, and the student model can use the xsmall model from TinyCLIP. While the IRRA algorithm can train an excellent teacher model, it has a large number of parameters and incurs significant memory and latency overhead during inference. Initializing the student model using the pre-trained weights of the xsmall model can accelerate the convergence speed of the student model training.

[0099] For input data preprocessing, text input can have "<|startoftext|>" and "<|endoftext|>" markers added before and after the sequence to clearly define text boundaries, ultimately truncating or padding to a length of 77 tokens. Image data input can first be scaled to 384×128, then horizontally flipped with a 50% probability, followed by 10-pixel boundary padding and random cropping to 384×128, and finally, random erasure data augmentation with a region ratio of 0.02 to 0.4 is applied. During training, each batch of data is processed simultaneously through a student model and a teacher model with fixed weights.

[0100] Regarding training parameter settings, this application uses the Adam optimizer for gradient calculation and backpropagation, with smoothing constants β1 = 0.9 and β2 = 0.999, an initial learning rate of 1e-5, and a weight decay coefficient of 4e-5. Each training phase can last for 60 epochs, with a fixed batch size of 64. The learning rate scheduling strategy can combine linear warm-up with cosine decay, performing linear warm-up starting from 1e-6 for the first 5 epochs, followed by cosine decay. The temperature coefficient of the task loss can be set to τ = 0.02, and the temperature coefficient of the distillation loss can be set to τ = 0.5. The weighting coefficient α of the loss function adopts a phased decay strategy. In the early training stage (first 10 epochs), a high distillation weight (α = 0.9) is maintained, allowing the student model to quickly absorb the discriminative knowledge from the teacher model. In the middle training stage, α decreases from 0.6 (10-20 epochs) to 0.3 (20-40 epochs), enhancing the model's autonomous optimization ability. In the final stage (40-60 epochs), α is further reduced to 0.1, fully releasing the representational potential of the student model. This dynamic balancing mechanism effectively alleviates the tension between knowledge transfer and model adaptation. It should be noted that the above is only one specific embodiment of this application, and the selection of datasets, model selection, and parameter values ​​involved do not constitute any limitation on this application. In fact, this application can arbitrarily adjust the above embodiment according to actual needs.

[0101] Simulation experiments and results:

[0102] To verify the effectiveness of the knowledge distillation method proposed in this application, a comparative experiment is conducted with existing methods applied to lightweight text-image cross-modal pedestrian retrieval models (including the InfoNCE-based method, Cross-model KD method, MOTIS method, and ConaCLIP method). The InfoNCE-based method uses the xsmall model for initialization and the InfoNCE contrastive loss for fine-tuning; the Cross-model KD method, proposed by Hinton et al., is a knowledge distillation method for unimodal applications and is directly applied to text-image cross-modal pedestrian retrieval tasks; the MOTIS and ConaCLIP methods are both used in lightweight text-image retrieval models and have been transferred to text-image cross-modal pedestrian retrieval tasks. Since knowledge distillation methods specifically designed for lightweight text-image cross-modal pedestrian retrieval models are still in the early stages of research, the comparison methods in this experiment are all based on replication.

[0103] Table 1 shows the comparison results on three public datasets: CUHK-PEDES, ICFG-PEDES, and RSTPReid. As can be seen from the table, the knowledge distillation method proposed in this application significantly outperforms other comparative methods across all metrics. Specifically, R1 (Recall@1) represents the proportion of correct results among the first results; R5 (Recall@5) represents the proportion of at least one correct result among the first five results; R10 (Recall@10) represents the proportion of at least one correct result among the first ten results; and mAP (Mean Average Precision) represents the overall accuracy.

[0104]

[0105] Table 1

[0106] In this application, a three-stage knowledge distillation process is employed for the student model. The first and second stages focus on single-modal knowledge transfer in images and text, respectively, ensuring the student model fully absorbs the teacher model's knowledge within each single modality. In the third stage, by jointly optimizing the student model's text encoder and image encoder, the learning space for the student model's cross-modal matching ability of the teacher model is expanded, enhancing the model's ability to capture fine-grained semantic information and thus improving its performance in fine-grained feature alignment. Furthermore, this application trains the student model based on the task loss function and distillation loss function at each stage. By designing a complementary supervision strategy for the task loss function and distillation loss function, the correlation between the student model and the teacher model in cross-modal task learning is strengthened. Appropriate guidance from the teacher model in text-image matching knowledge distillation is also introduced, further enhancing the student model's ability to achieve high-level alignment of multimodal features between text and images, thereby improving the overall performance and robustness of the model.

[0107] Those skilled in the art will understand that all or part of the features / steps of the above-described method embodiments can be implemented by methods, data processing systems, or computer programs. These features may be implemented without hardware, entirely in software, or in a combination of hardware and software. The aforementioned computer program may be stored in one or more computer-readable storage media. When the computer program is executed (e.g., by a processor), it performs the steps of the knowledge distillation method embodiments of the text-image cross-modal pedestrian retrieval model described above.

[0108] The aforementioned storage media capable of storing program code include: static hard disks, solid-state hard disks, random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), optical storage devices, magnetic storage devices, flash memory, magnetic disks or optical disks and / or combinations thereof, that is, they can be implemented by any type of volatile or non-volatile storage devices or combinations thereof.

[0109] like Figure 5 As shown, this application also provides an embodiment of a processing device 1, including one or more processors 11 and a memory 10; wherein, the memory 10 is used to store one or more computer programs, and the one or more processors 11 are used to execute one or more computer programs stored in the memory 10, so that the processors 11 execute the features / steps of the knowledge distillation method embodiment for the text-image cross-modal pedestrian retrieval model described above.

[0110] This application also provides a computer program product, which is stored on a data carrier and designed to execute the knowledge distillation method for a text-image-oriented cross-modal pedestrian retrieval model as described above. Therefore, the computer program product according to this application produces the same advantages as those described in the detailed description of the device according to this application. The computer program product can be executed as computer-readable instruction code using any suitable programming language such as JAVA, C++, etc. Furthermore, the computer program product can be provided on a network, such as the Internet, or downloaded from a network, such as the Internet, by a network, such as the Internet, when needed. The computer program product can be implemented using a computer program, i.e., software, or by one or more dedicated electronic circuits, i.e., hardware, or in any mixed form, i.e., by using software components and hardware components, or a software, hardware, or a hybrid of software and hardware.

[0111] The above description is merely a preferred embodiment of this application. Those skilled in the art will understand that various changes or equivalent substitutions can be made to these features and embodiments without departing from the spirit and scope of this application. Furthermore, under the teachings of this application, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of this application. Therefore, this application is not limited to the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application are within the protection scope of this application.

Claims

1. A knowledge distillation method for a cross-modal pedestrian retrieval model oriented towards text and images, characterized in that, include: Construct a teacher model and a student model, and initialize the student model, wherein both the teacher model and the student model have a text encoder and an image encoder; The student model is subjected to three-stage knowledge distillation. In the first stage, the image encoder of the student model is distilled using the teacher model. In the second stage, the text encoder of the student model is distilled using the teacher model. In the third stage, the text encoder and image encoder of the student model are distilled simultaneously using the teacher model. The student model is trained until convergence based on the task loss and distillation loss at each stage, resulting in a lightweight student model.

2. The knowledge distillation method for a text-image cross-modal pedestrian retrieval model according to claim 1, characterized in that, The first stage of knowledge distillation of the image encoder of the student model using the teacher model includes: The student model and the teacher model respectively extract features from the first image training data to obtain a first output result and a second output result. The teacher model extracts features from the first text training data to obtain a third output result. The first task loss and the first distillation loss are calculated based on the first output result, the second output result and the third output result. The parameters of the image encoder of the student model are adjusted by minimizing the first task loss and the first distillation loss.

3. The knowledge distillation method for a cross-modal pedestrian retrieval model oriented towards text and images according to claim 2, characterized in that, The first task loss includes a first contrastive learning loss, and the first distillation loss includes a first cross-entropy loss, wherein the first contrastive learning loss is the contrastive learning loss between the first output result and the third output result, and the first cross-entropy loss is the cross-entropy loss between the matching probability distribution of the first output result and the second output result and the true label.

4. The knowledge distillation method for a text-image cross-modal pedestrian retrieval model according to claim 1, characterized in that, The second stage of knowledge distillation of the text encoder of the student model by the teacher model includes: extracting features from the second image training data by the student model and the teacher model respectively to obtain a fourth output result and a fifth output result; extracting features from the second text training data by the student model and the teacher model respectively to obtain a sixth output result and a seventh output result; calculating a second task loss and a second distillation loss based on the fourth output result, the fifth output result, the sixth output result, and the seventh output result; and adjusting the parameters of the text encoder of the student model by minimizing the second task loss and the second distillation loss.

5. The knowledge distillation method for a text-image cross-modal pedestrian retrieval model according to claim 4, characterized in that, The second task loss includes a second contrastive learning loss and a third contrastive learning loss, wherein the second contrastive learning loss is the contrastive learning loss between the fifth output result and the sixth output result, and the third contrastive learning loss is the contrastive learning loss between the fourth output result and the sixth output result.

6. The knowledge distillation method for a text-image cross-modal pedestrian retrieval model according to claim 4, characterized in that, The second distillation loss includes a second cross-entropy loss and a first KL divergence loss, wherein the second cross-entropy loss is the cross-entropy loss between the matching probability distribution of the sixth output result and the seventh output result and the true label, and the first KL divergence loss is the KL divergence loss that fits the matching output probability distribution of the fourth output result and the sixth output result to the matching output probability distribution of the fifth output result and the seventh output result.

7. The knowledge distillation method for a text-image cross-modal pedestrian retrieval model according to claim 1, characterized in that, The third stage, which involves simultaneously performing knowledge distillation on the text encoder and image encoder of the student model using the teacher model, includes: extracting features from the third image training data using the student model and the teacher model to obtain an eighth and a ninth output result; extracting features from the third text training data using the student model and the teacher model to obtain a tenth and an eleventh output result; calculating a third task loss and a third distillation loss based on the eighth, ninth, tenth, and eleventh output results; and adjusting the parameters of the text encoder and the image encoder of the student model by minimizing the third task loss and the third distillation loss.

8. The knowledge distillation method for a text-image cross-modal pedestrian retrieval model according to claim 7, characterized in that, The third task loss includes a fourth contrastive learning loss and a fifth contrastive learning loss, wherein the fourth contrastive learning loss is the contrastive learning loss between the eighth output result and the eleventh output result, and the fifth contrastive learning loss is the contrastive learning loss between the ninth output result and the tenth output result.

9. The knowledge distillation method for a text-image cross-modal pedestrian retrieval model according to claim 7, characterized in that, The third distillation loss includes a third cross-entropy loss, a second KL divergence loss, a feature-wise distance distillation loss, and a similarity-wise distance distillation loss; wherein, the third cross-entropy loss is the cross-entropy loss between the matching probability distribution of the tenth output result and the eleventh output result and the true label; the second KL divergence loss is the KL divergence loss fitted by the matching output probability distribution of the eighth output result and the tenth output result to the matching output probability distribution of the ninth output result and the eleventh output result; the feature-wise distance distillation loss includes a first feature-wise distance distillation loss between the eighth output result and the ninth output result, and a second feature-wise distance distillation loss between the tenth output result and the eleventh output result; the similarity-wise distance distillation loss includes a first similarity-wise distance distillation loss based on the similarity between the tenth output result and the eleventh output result, and a second similarity-wise distance distillation loss based on the similarity between the eighth output result and the ninth output result.

10. A computer program product, characterized in that, The computer program product is stored on a data carrier and is designed to execute the knowledge distillation method for a text-image cross-modal pedestrian retrieval model as described in any one of claims 1-9.