Training method and device of text pedestrian retrieval model based on self-feedback learning, and medium

The text-based pedestrian retrieval model training method based on self-feedback learning solves the problem of inaccurate alignment between text and images, achieves stable retrieval in complex scenarios, and improves retrieval accuracy and robustness.

CN122364495APending Publication Date: 2026-07-10JIANGNAN UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JIANGNAN UNIV
Filing Date
2026-04-01
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

In existing technologies, text pedestrian retrieval suffers from low accuracy due to poor alignment between text and images. Furthermore, cross-modal semantic alignment is unstable due to factors such as changes in viewing angle, lighting conditions, and occlusion.

Method used

A text-based pedestrian retrieval model training method based on self-feedback learning is adopted. Through self-supervised semantic self-feedback learning and cross-modal answer mutual supervision learning, answer features of images and text are generated. Then, through comparative learning by attention network, the model parameters are optimized to achieve accurate alignment.

Benefits of technology

It improves the accuracy and robustness of text-based pedestrian retrieval, maintains stable retrieval ranking performance in complex scenarios, reduces false detections caused by missing single attributes or noise, and enhances the accuracy and reliability of cross-modal matching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122364495A_ABST
    Figure CN122364495A_ABST
Patent Text Reader

Abstract

This application discloses a training method, device, and medium for a text pedestrian retrieval model based on self-feedback learning, relating to the field of cross-modal text pedestrian retrieval technology. The method includes: determining image features, text features, and cue features; inputting the image features, text features, and cue features into a text pedestrian retrieval model; performing self-supervised semantic self-feedback learning through the text pedestrian retrieval model to obtain a self-feedback loss; performing cross-modal answer mutual supervision learning to obtain a mutual supervision loss; performing contrastive learning to obtain KL divergence loss and contrastive loss; and iteratively optimizing the model parameters of the text pedestrian retrieval model based on the above losses. This application aims to solve the problem of low retrieval accuracy in existing technologies due to poor alignment accuracy between text and images during text pedestrian retrieval, thereby achieving accurate alignment of text and images during text pedestrian retrieval to improve retrieval accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of cross-modal text pedestrian retrieval technology, and in particular to a training method, device and medium for a text pedestrian retrieval model based on self-feedback learning. Background Technology

[0002] With the development of applications such as intelligent security, public safety, and smart cities, text-based pedestrian retrieval technology has been widely used. Text-based pedestrian retrieval refers to retrieving target pedestrian images from a large-scale pedestrian image database based on natural language descriptive text. Its core lies in achieving fine alignment between textual semantics and visual semantics.

[0003] Text descriptions and pedestrian images differ in semantic granularity. Text often describes pedestrian appearance in a relatively coarse-grained manner, while images contain a large number of fine-grained visual details. Inconsistencies can easily arise during alignment, leading to cross-modal semantic dispersion. This semantic dispersion manifests in actual retrieval as different attributes in the text description failing to consistently map to the correct regions in the image, resulting in unstable alignment and reduced retrieval accuracy.

[0004] Existing techniques typically mitigate the aforementioned alignment instability problem by performing local-level matching through image region segmentation or text phrase extraction. However, these methods rely on predefined region segmentation or phrase segmentation rules, which can easily disrupt the coherence of the overall semantic representation of pedestrians and weaken the ability to comprehensively model contextual information. This makes it difficult to obtain stable and consistent cross-modal semantic representations, resulting in poor alignment accuracy.

[0005] In addition, the same pedestrian can appear significantly different in different images due to factors such as changes in perspective, lighting conditions, and occlusion. Text descriptions also have great flexibility in sentence structure, word choice, and descriptive focus. These visual and linguistic variations further increase the difficulty of cross-modal alignment, ultimately leading to lower retrieval accuracy. Summary of the Invention

[0006] To address the aforementioned problems and technical requirements, this applicant proposes a training method, device, and medium for a text pedestrian retrieval model based on self-feedback learning. This aims to solve the problem of low retrieval accuracy in existing technologies due to poor alignment between text and images, thereby achieving accurate alignment of text and images during text pedestrian retrieval and improving retrieval accuracy.

[0007] This application provides a training method for a text-based pedestrian retrieval model based on self-feedback learning, the method comprising:

[0008] Acquire training sample data, which includes: image samples, text samples, and prompt information. Image samples and text samples describing image samples are matched to obtain sample pairs. N image features corresponding to an image sample are determined, and M text features and K prompt features corresponding to a text sample are determined based on prompt information. The K prompt features correspond to K parallel subspaces. The image features include local image features and global image features, and the text features include local text features and global text features. N image features, M text features, and K cue features are input into a text-based pedestrian retrieval model. The model uses the cue features as cue constraints to generate image and text answer features for each subspace. Self-supervised semantic self-feedback learning is performed based on global image features, image answer features, global text features, and text answer features to obtain a self-feedback loss. Cross-modal mutual supervision learning is performed on the image and text answer features for each subspace to obtain a mutual supervision loss. KL divergence loss and contrast loss are obtained by comparing the target local text features (selected through an attention network), global text features, and target local image features (selected through an attention network), and global image features. Finally, the model parameters are iteratively optimized based on the self-feedback loss, mutual supervision loss, contrast loss, and KL divergence loss until a preset number of iterations is reached, indicating that the text-based pedestrian retrieval model training is complete. Among them, the self-feedback loss is the cross-modal consistency assessment of the same sample in different subspaces.

[0009] According to the training method of the text pedestrian retrieval model based on self-feedback learning provided in the embodiments of this application, self-supervised semantic self-feedback learning is performed based on global image features, image answer features, global text features, and text answer features to obtain self-feedback loss, including: Calculate the adaptive weights of the image answer feature generated by the i-th image feature under the constraint of the k-th prompt feature and the j-th global text feature, and calculate the adaptive weights of the text answer feature generated by the i-th text feature under the constraint of the k-th prompt feature and the j-th global image feature. The adaptive weights of all subspaces are weighted and summed to obtain the feedback score; The self-feedback loss is derived from the feedback score.

[0010] According to the training method of the text pedestrian retrieval model based on self-feedback learning provided in the embodiments of this application, the adaptive weights of the image answer feature generated by the i-th image feature under the constraint of the k-th prompt feature and the j-th global text feature are calculated, as well as the adaptive weights of the text answer feature generated by the i-th text feature under the constraint of the k-th prompt feature and the j-th global image feature are calculated, including: The image answer feature generated by the i-th image feature under the k-th prompt feature as a prompt constraint and the j-th global text feature are input into the preset first adaptive weight calculation formula, and the text answer feature generated by the i-th text feature under the k-th prompt feature as a prompt constraint and the j-th global image feature are input into the second adaptive weight calculation formula to obtain the adaptive weight output by the adaptive weight calculation formula. The formula for calculating the first adaptive weight includes: ; The formula for calculating the second adaptive weight includes: ; in, This represents the adaptive weights corresponding to the image. Represents cosine similarity. This represents the image answer feature generated under the constraint of the k-th prompt feature, where the i-th image feature is used as the prompt feature. Represents the j-th global text feature. This represents the current subspace among K subspaces. This represents the image answer feature generated by the i-th image feature in the current subspace. This represents the adaptive weight corresponding to the text. Let i represent the text answer feature generated under the constraint of the k-th prompt feature, where the i-th text feature is used as the prompt feature. Represents the j-th global image feature. Let i represent the text answer feature generated by the i-th text feature in the current subspace.

[0011] According to the training method of the text pedestrian retrieval model based on self-feedback learning provided in the embodiments of this application, the adaptive weights of all subspaces are weighted and summed to obtain a feedback score, including: The adaptive weights of each subspace are input into a preset feedback score calculation formula to obtain the feedback score output by the formula: The formula for calculating the feedback score includes: ; ; in, This represents the cross-modal semantic alignment score corresponding to the image side. This represents the cross-modal semantic alignment score corresponding to the text side. This represents the temperature hyperparameter, where, in The time is a non-matching image-text pair.

[0012] According to the training method of the text pedestrian retrieval model based on self-feedback learning provided in the embodiments of this application, the self-feedback loss is obtained based on the feedback score, including: Input the feedback score into the preset self-feedback loss calculation formula to obtain the self-feedback loss output by the self-feedback loss calculation formula; The formula for calculating self-feedback loss includes: ; in, Indicates self-feedback loss. This represents the semantic alignment score of the matched text and image in the image orientation. This represents the semantic alignment score of the matched image and text in the text direction, where the matched image and text are the information indicating a successful match.

[0013] According to the training method of the text pedestrian retrieval model based on self-feedback learning provided in the embodiments of this application, cross-modal mutual supervision learning is performed on the image answer features and text answer features corresponding to each subspace to obtain mutual supervision loss, including: Input the image answer features and text answer features corresponding to each subspace into the preset mutual supervision loss calculation formula, and sum the mutual supervision loss output by the mutual supervision loss calculation formula to obtain the final mutual supervision loss; The formula for calculating mutual supervision loss includes: ; in, , ; in, Indicates mutual supervision losses, This indicates temperature hyperparameters. This represents the similarity score between image answer features generated under the same cue constraints within the k-th subspace. This represents the similarity score between text answer features generated under the same cue constraints within the k-th subspace. This represents the current subspace among K subspaces. This represents the image answer feature generated under the constraint of the k-th prompt feature, where the i-th image feature is used as the prompt feature. Let represent the text answer feature generated by the i-th text feature in the current subspace. Let i represent the text answer feature generated under the constraint of the k-th prompt feature, where the i-th text feature is used as the prompt feature. Let i represent the image answer feature generated by the i-th image feature in the current subspace.

[0014] According to the training method of the text pedestrian retrieval model based on self-feedback learning provided in the embodiments of this application, the KL divergence loss and contrast loss are obtained by comparing the target local text features and global text features selected by the attention network, and the target local image features and global image features selected by the attention network, including: The local text features and local image features of the target are normalized to construct a local cross-modal similarity matrix; Similarity calculations are performed on global text features and global image features to obtain a global cross-modal similarity matrix; The global cross-modal similarity matrix is ​​normalized to obtain the global cross-modal similarity probability distribution; The difference between the global cross-modal similarity probability distribution and the preset target distribution is calculated to obtain the KL loss, which is then used to optimize the global cross-modal similarity probability distribution. Positive and negative sample pairs are constructed based on the local cross-modal similarity matrix, and the contrastive loss is calculated based on the positive and negative sample pairs.

[0015] According to the training method of the text pedestrian retrieval model based on self-feedback learning provided in the embodiments of this application, N image features corresponding to image samples are determined, and M text features and K prompt features corresponding to text samples are determined based on prompt information, including: Input the image sample into the preset image encoder to obtain N image features output by the image encoder; Input the prompt information and text sample into the preset text encoder to obtain M text features and K prompt features output by the text encoder.

[0016] This application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps of the training method for a text pedestrian retrieval model based on self-feedback learning as described above.

[0017] This application also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the training method for a text pedestrian retrieval model based on self-feedback learning as described above.

[0018] The training method, device, and medium for a text pedestrian retrieval model based on self-feedback learning provided in this application embodiment determine the image features, text features, and prompt features of the training sample data; input the image features, text features, and prompt features into the text pedestrian retrieval model; and generate image answer features and text answer features corresponding to each subspace through the text pedestrian retrieval model using prompt features as prompt constraints. Self-supervised semantic self-feedback learning is performed based on global image features, image answer features, global text features, and text answer features to obtain a self-feedback loss. This application achieves semantic weighted alignment between answer features and image and text features by calculating the self-feedback loss, solving the problem of alignment instability caused by inconsistent semantic granularity and polymorphic changes. Cross-modal answer mutual supervision learning is performed on the image answer features and text answer features corresponding to each subspace to obtain a mutual supervision loss. This application utilizes cross-modal answer mutual supervision learning to eliminate semantic ambiguity that may be caused by the adaptive knowledge alignment in the previous step. The consistency constraints are mutually supervised to ensure that the image answer features and text answer features remain consistent, ensuring the consistency of fine-grained attributes between the two modalities. Based on the target local text features and global text features selected by the attention network, and the target local image features and global image features selected by the attention network, KL divergence loss and contrast loss are obtained through comparative learning. This application optimizes the global modal similarity distribution based on KL divergence and enhances the discriminative power based on contrast loss, forming a spatial topology structure with the image as the core and the text surrounding it, thus ensuring semantic alignment. Based on self-feedback loss, mutual supervision loss, contrast loss, and KL divergence loss, the model parameters of the text pedestrian retrieval model are iteratively optimized until the number of iterations reaches the preset number, at which point the text pedestrian retrieval model is considered to have completed training. This enables the trained text pedestrian retrieval model to achieve accurate alignment of text and image when performing pedestrian retrieval, improving the accuracy of the retrieval results. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is one of the flowcharts illustrating the training method for a text pedestrian retrieval model based on self-feedback learning provided in this application embodiment; Figure 2 This is the second flowchart illustrating the training method of the text pedestrian retrieval model based on self-feedback learning provided in the embodiments of this application; Figure 3This is a framework diagram of the training method for a text pedestrian retrieval model based on self-feedback learning provided in the embodiments of this application; Figure 4 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0022] This application provides a training method for a text-based pedestrian retrieval model based on self-feedback learning. This method can be applied to smart terminals and servers. This application uses the application of this method in a server as an example for illustration, and some other descriptions in the embodiments are illustrative and not intended to limit the scope of protection of this application, and will not be described in detail thereafter. The specific implementation of the method is as follows... Figure 1 As shown: Step 101: Obtain training sample data.

[0023] The training sample data includes: image samples, text samples, and prompts. Image samples and text samples describing the image samples are matched to obtain sample pairs.

[0024] Step 102: Determine N image features corresponding to the image sample, and determine M text features and K prompt features corresponding to the text sample based on the prompt information.

[0025] Among them, the K cue features correspond to K parallel subspaces. The image features include local image features and global image features, and the text features include local text features and global text features.

[0026] The prompt information is used to guide subspace modeling, and is not simply an input of manual annotation, prior rules, or external instructions. This subspace is a semantic subspace, meaning that multiple semantic subspaces are generated through the prompt information.

[0027] Step 103: Input N image features, M text features, and K cue features into the text pedestrian retrieval model. Using the cue features as cue constraints, the text pedestrian retrieval model generates image and text answer features for each subspace. Self-supervised semantic self-feedback learning is performed based on global image features, image answer features, global text features, and text answer features to obtain a self-feedback loss. Cross-modal mutual supervision learning is performed on the image and text answer features corresponding to each subspace to obtain a mutual supervision loss. Comparative learning is performed on the target local text features selected through an attention network, global text features, and target local and global image features selected through an attention network to obtain KL divergence loss and contrast loss. Based on the self-feedback loss, mutual supervision loss, contrast loss, and KL divergence loss, the model parameters of the text pedestrian retrieval model are iteratively optimized until the preset number of iterations is reached, indicating that the text pedestrian retrieval model training is complete.

[0028] Among them, the self-feedback loss is a cross-modal consistency evaluation of the same sample in different subspaces. It is used to adaptively weight the semantic subspace, forming a closed-loop optimization process of "semantic evaluation-semantic weighting-semantic alignment", which effectively distinguishes it from the existing supervision methods based on external annotation.

[0029] The training method for a text pedestrian retrieval model based on self-feedback learning provided in this application embodiment determines the image features, text features, and prompt features of the training sample data; inputs the image features, text features, and prompt features into the text pedestrian retrieval model, and generates image answer features and text answer features corresponding to each subspace by using prompt features as prompt constraints; performs self-supervised semantic self-feedback learning based on global image features, image answer features, global text features, and text answer features to obtain a self-feedback loss. This application achieves semantic weighted alignment between answer features and image and text features by calculating the self-feedback loss, solving the problem of alignment instability caused by inconsistent semantic granularity and polymorphic changes; performs cross-modal answer mutual supervision learning on the image answer features and text answer features corresponding to each subspace to obtain a mutual supervision loss. This application utilizes the consistency of cross-modal answers to eliminate the semantic ambiguity that may be caused by the adaptive knowledge alignment in the previous step. Mutual supervision of sexual constraints ensures consistency between image and text answer features, guaranteeing the consistency of fine-grained attributes across the two modalities. Based on comparative learning of target local text features (selected through an attention network), global text features, target local image features (selected through an attention network), and global image features, KL divergence loss and contrast loss are obtained. This application optimizes the global modal similarity distribution based on KL divergence and enhances discriminative power based on contrast loss, forming a spatial topology structure with the image as the core and text surrounding it, ensuring semantic alignment. Based on self-feedback loss, mutual supervision loss, contrast loss, and KL divergence loss, the model parameters of the text pedestrian retrieval model are iteratively optimized until the preset number of iterations is reached, indicating that the text pedestrian retrieval model training is complete. This enables the trained text pedestrian retrieval model to achieve accurate alignment of text and images during pedestrian retrieval, improving the accuracy of retrieval results.

[0030] In one specific embodiment, the specific implementation of determining N image features corresponding to an image sample, and determining M text features and K prompt features corresponding to a text sample based on prompt information, includes: Input the image sample into the preset image encoder to obtain N image features output by the image encoder; input the prompt information and text sample into the preset text encoder to obtain M text features and K prompt features output by the text encoder.

[0031] Among them, image features are used express, It includes N image features, and text features are used. express, It includes M text features.

[0032] Image features include local image features and global image features. Text features include local text features and global text features. .

[0033] This includes pre-building a set of structured prompt questions (prompt information).

[0034] Design a set of structured prompt questions. This is used to query fields with different semantic attributes.

[0035] The prompt information includes basic information (such as gender), features and colors of body parts (such as head, upper body, lower body and feet), and K semantic dimensions (subspaces) such as accessories and overall clothing.

[0036] For example, the categories and colors of tops, bottoms, and shoes; the presence of long garments and their colors; personal items and accessories, etc.

[0037] Furthermore, multiple fillable semantic slots (MASK) are set in the prompt message to induce semantic expression in the subspace.

[0038] In one specific embodiment, the specific implementation of generating image answer features and text answer features corresponding to each subspace through a text pedestrian retrieval model using prompt features as prompt constraints includes: The answer extractor of the text pedestrian retrieval model pairs the cue features with image features and text features respectively, and extracts the field-level semantic answer from the corresponding modal features under the cue constraints of the cue features.

[0039] The answer extractor is composed of a cross-modal attention layer and a Transformer network.

[0040] Specifically, through the formula , Represents an image. Representing the text, we obtain image answer features and text answer features respectively.

[0041] Among them, the image answer features include global image answer features, and the text answer features include global text answer features.

[0042] This application decomposes a high-dimensional feature space into multiple independent semantic subspaces by using cue features as cue constraints.

[0043] In one specific embodiment, the self-supervised semantic self-feedback learning based on global image features, image answer features, global text features, and text answer features, and the specific implementation of obtaining the self-feedback loss includes: Calculate the adaptive weights of the image answer feature generated by the i-th image feature under the constraint of the k-th prompt feature and the j-th global text feature, and calculate the adaptive weights of the text answer feature generated by the i-th text feature under the constraint of the k-th prompt feature and the j-th global image feature; perform weighted summation on the adaptive weights of all subspaces to obtain the feedback score; obtain the self-feedback loss based on the feedback score.

[0044] This application requires multiple iterative loops. Within each iteration, adaptive weights are calculated for the image answer features and global text features, as well as adaptive weights for the text answer features and global image features. The adaptive weights across all subspaces are weighted and summed to obtain a feedback score; a self-feedback loss is then derived based on the feedback score.

[0045] In one specific embodiment, the specific implementation of calculating the adaptive weights of the image answer feature generated by the i-th image feature under the constraint of the k-th prompt feature and the j-th global text feature, and the adaptive weights of the text answer feature generated by the i-th text feature under the constraint of the k-th prompt feature and the j-th global image feature, includes: The image answer feature generated by the i-th image feature under the k-th prompt feature as a prompt constraint and the j-th global text feature are input into the preset first adaptive weight calculation formula, and the text answer feature generated by the i-th text feature under the k-th prompt feature as a prompt constraint and the j-th global image feature are input into the second adaptive weight calculation formula to obtain the adaptive weight output by the adaptive weight calculation formula.

[0046] The formula for calculating the first adaptive weight is shown in formula (1.1): ... (1.1) The formula for calculating the second adaptive weight is shown in formula (1.2): ... (1.2) in, This represents the adaptive weights corresponding to the image. Represents cosine similarity. This represents the image answer feature generated under the constraint of the k-th prompt feature, where the i-th image feature is used as the prompt feature. Represents the j-th global text feature. This represents the current subspace among K subspaces. This represents the image answer feature generated by the i-th image feature in the current subspace. This represents the adaptive weight corresponding to the text. Let i represent the text answer feature generated under the constraint of the k-th prompt feature, where the i-th text feature is used as the prompt feature. Represents the j-th global image feature. Let i represent the text answer feature generated by the i-th text feature in the current subspace.

[0047] In one specific embodiment, the weighted summation of the adaptive weights across all subspaces to obtain the feedback score includes the following specific implementation: The adaptive weights of each subspace are input into the preset feedback score calculation formula to obtain the feedback score output by the feedback score calculation formula.

[0048] The formulas for calculating the feedback score are shown in formulas (2.1) and (2.2): ... (2.1) ... (2.2) in, This represents the cross-modal semantic alignment score corresponding to the image side. This represents the cross-modal semantic alignment score corresponding to the text side. This represents the temperature hyperparameter, where, in The time is a non-matching image-text pair.

[0049] The feedback score is derived from the consistency evaluation of the unified sample across different semantic subspaces. It is used to perform adaptive weighting on the semantic subspace, thereby forming a closed-loop optimization process of "semantic evaluation - semantic weighting - semantic alignment".

[0050] In one specific embodiment, the specific implementation of obtaining the self-feedback loss based on the feedback score includes: Input the feedback score into the preset self-feedback loss calculation formula to obtain the self-feedback loss output by the self-feedback loss calculation formula.

[0051] The formula for calculating the self-feedback loss is shown in formula (3): …(3) in, Indicates self-feedback loss. This represents the semantic alignment score of the matched text and image in the image orientation. This represents the semantic alignment score of the matched image and text in the text direction, where the matched image and text are the information indicating a successful match.

[0052] Specifically, the feedback score is a weighted similarity. This application constructs a contrastive learning objective based on the weighted similarity, forming a bidirectional self-feedback loss, which enables the model to automatically emphasize the semantic subspace related to matching and suppress the noise subspace, thus solving the problem of alignment instability caused by inconsistent semantic granularity and polymorphic changes.

[0053] This application calculates the weight of each subspace by calculating the similarity between the answer features and the sample features, performs self-supervised semantic feedback, automatically suppresses noise channels and strengthens effective attribute channels throughout the process, and achieves the purpose of adaptive alignment according to semantic channels.

[0054] Furthermore, this application provides more stable alignment for occlusion, small objects, and weak signal attributes (because the weights will naturally decrease and will not be forcibly aligned), thereby alleviating the problem of unstable alignment caused by visual and textual changes.

[0055] This application does not rely on any semantic or attribute-level manual annotation. It only uses existing image-text sample pairs as supervision signals and achieves semantic credibility and self-evaluation and self-feedback optimization through cross-modal consistency within the sample pairs, thereby realizing closed-loop training of "semantic self-evaluation - semantic self-weighting - semantic self-alignment".

[0056] In one specific embodiment, cross-modal mutual supervision learning is performed on the image answer features and text answer features corresponding to each subspace to obtain the mutual supervision loss. The specific implementation includes: The image and text answer features corresponding to each subspace are input into a preset mutual supervision loss calculation formula. The mutual supervision loss output by the mutual supervision loss calculation formula is summed to obtain the final mutual supervision loss.

[0057] The formula for calculating the mutual supervision loss is shown in formula (4): …………(4) in, , .

[0058] in, Indicates mutual supervision losses, This indicates temperature hyperparameters. This represents the similarity score between image answer features generated under the same cue constraints within the k-th subspace. This represents the similarity score between text answer features generated under the same cue constraints within the k-th subspace. This represents the current subspace among K subspaces. This represents the image answer feature generated under the constraint of the k-th prompt feature, where the i-th image feature is used as the prompt feature. Let represent the text answer feature generated by the i-th text feature in the current subspace. Let i represent the text answer feature generated under the constraint of the k-th prompt feature, where the i-th text feature is used as the prompt feature. This represents the image answer feature generated by the i-th image feature in the current subspace. To eliminate semantic ambiguity that might arise from adaptive alignment, the image answer feature and the text answer feature are forced to maintain consistency. Specifically, the consistency of fine-grained attributes (such as color and accessories) between the two modalities is determined by calculating the point-to-point mutual supervision loss between the image and text answer features in each subspace.

[0059] While adaptive alignment ensures high global semantic similarity, it may lead to local semantic contradictions (e.g., the text says "backpack," but the image response channel says "no backpack"). Therefore, this application constrains the consistency between image and text responses, forming point-to-point mutual supervision to reduce cross-modal semantic conflicts and improve fine-grained consistency and interpretability. Ultimately, it achieves semantic consistency at both the global and local levels.

[0060] In one specific embodiment, the specific implementation of obtaining KL divergence loss and contrastive loss by comparing and learning the target local text features (selected through an attention network), global text features, target local image features (selected through an attention network), and global image features, and obtaining KL divergence loss and contrastive loss includes: The target local text features and target local image features are normalized to construct a local cross-modal similarity matrix; the similarity of global text features and global image features is calculated to obtain a global cross-modal similarity matrix; the global cross-modal similarity matrix is ​​normalized to obtain a global cross-modal similarity probability distribution; the difference between the global cross-modal similarity probability distribution and the preset target distribution is calculated to obtain the KL loss, which is used to optimize the global cross-modal similarity probability distribution; positive and negative sample pairs are constructed based on the local cross-modal similarity matrix, and the contrastive loss is calculated based on the positive and negative sample pairs; the KL loss and the contrastive loss are weighted and summed to obtain the joint loss.

[0061] The formula for calculating the comparative loss is shown in formula (5): ………………(5) in, .

[0062] in, This represents the contrast loss in the i2t direction. This represents the set of text features that share the same identity as image feature i. This represents the similarity score for images with the same identity as image feature i. This represents all samples within the current loop. This indicates the sample index.

[0063] Among them, the same identity means that the image features belong to the same ID.

[0064] The contrast loss corresponding to the t2i direction is represented by the formula above, which is also used to calculate the loss when the x and y coordinates of the matrix are interchanged in the it2 and t2i directions.

[0065] Similarly, the KL loss is the sum of the losses calculated in both directions.

[0066] This application combines global features and local features filtered by an attention grid, and employs a hybrid loss function ( Optimize the topology of the shared embedded space.

[0067] Among them, KL divergence loss is used to optimize the global cross-modal similarity distribution, and contrast loss is used to enhance the discriminative power at the instance level. The two are combined to form a feature space topology structure with "image as the core and text surrounding".

[0068] This application combines KL divergence loss and contrastive loss to obtain a structured and interpretable distribution where image features form a compact core. Conversely, text features form a ring-like topology around them. This topology anchors the identity center to more stable image features while allowing more diverse text features to flexibly arrange themselves around them. This topology is particularly desirable in text-image retrieval tasks because it preserves a clear semantic center for each identity while accommodating modality-specific variations. Furthermore, the clear separation between identity clusters indicates that the joint objective effectively avoids identity overlap, even in dense regions of the feature space. This application combines KL divergence to optimize the global cross-modal similarity distribution and many-to-many contrastive loss to capture semantic variations and adapt to cross-modal semantic diversity.

[0069] The resulting feature space exhibits a stable topological structure with image features at its core and text features surrounding it, effectively improving the robustness of retrieval.

[0070] In one specific embodiment, the optimization of the text pedestrian retrieval model based on self-feedback loss, mutual supervision loss, contrastive loss, and KL divergence loss includes the following specific implementations: After calculating the self-feedback loss, mutual supervision loss, contrastive loss, and KL divergence loss, the final loss is obtained, and the network is trained end-to-end based on the final loss.

[0071] The total loss function is given by formula (6): …………(6) in, Indicates the final loss. This indicates the preset coefficient.

[0072] In one specific embodiment, after the text pedestrian retrieval model is trained, the image to be retrieved and the text to be retrieved are obtained. The image to be retrieved and the text to be retrieved are input into the text pedestrian retrieval model to obtain the target image corresponding to the text to be retrieved, which is output by the text pedestrian retrieval model.

[0073] This application can simultaneously consider multiple semantic clues during the retrieval process. Even if some attributes are occluded in the image or only present local features, effective matching can still be achieved through other semantic information, thereby reducing false detections caused by missing single attributes or noise. It achieves stable retrieval ranking performance under complex natural language descriptions and real-world pedestrian images, improving the practicality and robustness of text pedestrian retrieval.

[0074] This application enables stable recognition using local semantic information without relying on the explicit visibility of the complete target. Through a semantically guided attention mechanism, the model focuses more on key regions related to the text description during the retrieval process, thereby improving the accuracy and reliability of cross-modal matching.

[0075] Below, through Figure 2 This application will be described in detail as follows: Step 201: Construct image features, text features, and prompt features.

[0076] Step 202: Input the image features, text features, and prompt features into the answer extractor to obtain the image answer features and text answer features under prompt constraints.

[0077] Step 203: Based on image features, text features, image answer features, and text answer features, perform self-supervised semantic self-feedback learning, cross-modal answer consistency constraint processing, and fine-grained pedestrian-level semantic alignment processing to obtain self-feedback loss, mutual supervision loss, KL divergence loss, contrast loss, and final loss.

[0078] Step 204: Iteratively optimize the model parameters of the text pedestrian retrieval model based on self-feedback loss, mutual supervision loss, KL divergence loss, contrastive loss, and final loss until the preset number of iterations is reached, and then determine that the text pedestrian retrieval model training is complete.

[0079] Below, through Figure 3 This application is illustrated in the form of a framework diagram. Figure 3 CMAC in PLSA is a processing module, corresponding to the processing logic of the text description section.

[0080] Among them, Figure 3 The text descriptions in the illustrations also correspond to the descriptions above.

[0081] This application constructs structured prompt questions with multiple semantic fields (prompt-driven questions) to guide the model to self-organize the extraction of multi-subspace semantic representations in a high-dimensional feature space. Images and text are treated as independent latent knowledge bases, and a self-supervised semantic feedback mechanism enables adaptive correction of cross-modal semantics. Simultaneously, cross-modal answer consistency constraints are introduced to strengthen point-to-point supervision between image and text answers at the subspace level. Furthermore, an individual-level semantic alignment strategy is combined to maintain identity-level topological consistency in the shared embedding space. This method requires no additional manual annotation or external knowledge support, effectively alleviating the problems of cross-modal semantic discrepancies and semantic polymorphism, and improving the accuracy and robustness of text-based person retrieval. Moreover, each answer is refined through self-supervised feedback, thus transforming semantic alignment into a reasoning process that ensures global semantic continuity, rather than static feature matching.

[0082] This application adaptively decomposes features into interpretable semantic subspaces using prompts, avoiding semantic fragmentation caused by hard segmentation. Through a self-supervised semantic feedback learning mechanism, the global features of the samples themselves are used as a knowledge base, dynamically suppressing noise interference from occlusion or viewpoint without relying on expensive external annotations or predefined attribute sets, thus adapting to semantic polymorphism. By forcing mutual supervision between image and text answers at the subspace level, ambiguity in cross-modal alignment is effectively reduced, improving the matching accuracy of subtle attributes (such as color and accessories). Through a pedestrian-level semantic alignment strategy, combined with KL divergence and contrastive loss, a more robust feature space topology is constructed, preserving intra-class diversity while enhancing inter-class discriminability. Ultimately, this improves the accuracy of model retrieval.

[0083] Figure 4 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 4 As shown, the electronic device may include a processor 401, a communications interface 402, a memory 403, and a communication bus 404. The processor 401, communications interface 402, and memory 403 communicate with each other via the communication bus 404. The processor 401 can call logical instructions in the memory 403 to execute a training method for a text-based pedestrian retrieval model based on self-feedback learning.

[0084] Furthermore, the logical instructions in the aforementioned memory 403 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0085] On the other hand, the present invention also provides a computer program product, the computer program product including a computer program stored on a non-transitory computer-readable storage medium, the computer program including program instructions, and when the program instructions are executed by a computer, the computer is able to execute the training method of the text pedestrian retrieval model based on self-feedback learning provided by the above methods.

[0086] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the training method for the text pedestrian retrieval model based on self-feedback learning provided in the above embodiments.

[0087] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0088] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0089] Finally, it should be noted that the above descriptions are merely preferred embodiments of this application, and this application is not limited to the above embodiments. It is understood that other improvements and variations directly derived or conceived by those skilled in the art without departing from the spirit and concept of this application should be considered to be included within the protection scope of this application.

Claims

1. A training method for a text-based pedestrian retrieval model based on self-feedback learning, characterized in that, The method includes: Acquire training sample data, which includes: image samples, text samples, and prompt information. Image samples and text samples describing image samples are matched to obtain sample pairs. N image features corresponding to an image sample are determined, and M text features and K prompt features corresponding to a text sample are determined based on prompt information. The K prompt features correspond to K parallel subspaces. The image features include local image features and global image features, and the text features include local text features and global text features. N image features, M text features, and K cue features are input into a text-based pedestrian retrieval model. The model uses the cue features as cue constraints to generate image and text answer features for each subspace. Self-supervised semantic self-feedback learning is performed based on global image features, image answer features, global text features, and text answer features to obtain a self-feedback loss. Cross-modal mutual supervision learning is performed on the image and text answer features for each subspace to obtain a mutual supervision loss. KL divergence loss and contrast loss are obtained by comparing the target local text features (selected through an attention network), global text features, and target local image features (selected through an attention network), and global image features. Finally, the model parameters are iteratively optimized based on the self-feedback loss, mutual supervision loss, contrast loss, and KL divergence loss until a preset number of iterations is reached, indicating that the text-based pedestrian retrieval model training is complete. Among them, the self-feedback loss is the cross-modal consistency assessment of the same sample in different subspaces.

2. The training method for the text pedestrian retrieval model based on self-feedback learning according to claim 1, characterized in that, Self-supervised semantic self-feedback learning is performed based on global image features, image-answer features, global text features, and text-answer features to obtain a self-feedback loss, including: Calculate the adaptive weights of the image answer feature generated by the i-th image feature under the constraint of the k-th prompt feature and the j-th global text feature, and calculate the adaptive weights of the text answer feature generated by the i-th text feature under the constraint of the k-th prompt feature and the j-th global image feature. The adaptive weights of all subspaces are weighted and summed to obtain the feedback score; The self-feedback loss is derived from the feedback score.

3. The training method for the text pedestrian retrieval model based on self-feedback learning according to claim 2, characterized in that, Calculate the adaptive weights of the image answer feature generated by the i-th image feature under the constraint of the k-th prompt feature and the j-th global text feature, and calculate the adaptive weights of the text answer feature generated by the i-th text feature under the constraint of the k-th prompt feature and the j-th global image feature, including: The image answer feature generated by the i-th image feature under the k-th prompt feature as a prompt constraint and the j-th global text feature are input into the preset first adaptive weight calculation formula, and the text answer feature generated by the i-th text feature under the k-th prompt feature as a prompt constraint and the j-th global image feature are input into the second adaptive weight calculation formula to obtain the adaptive weight output by the adaptive weight calculation formula. The formula for calculating the first adaptive weight includes: ; The formula for calculating the second adaptive weight includes: ; in, This represents the adaptive weights corresponding to the image. Represents cosine similarity. This represents the image answer feature generated under the constraint of the k-th prompt feature, where the i-th image feature is used as the prompt feature. Represents the j-th global text feature. This represents the current subspace among K subspaces. This represents the image answer feature generated by the i-th image feature in the current subspace. This represents the adaptive weight corresponding to the text. Let i represent the text answer feature generated under the constraint of the k-th prompt feature, where the i-th text feature is used as the prompt feature. Represents the j-th global image feature. Let i represent the text answer feature generated by the i-th text feature in the current subspace.

4. The training method for the text pedestrian retrieval model based on self-feedback learning according to claim 3, characterized in that, The feedback scores include: the cross-modal semantic alignment score corresponding to the image side and the cross-modal semantic alignment score corresponding to the text side; The adaptive weights of all subspaces are weighted and summed to obtain the feedback score, which includes: The adaptive weights of each subspace are input into a preset feedback score calculation formula to obtain the feedback score output by the formula: The formula for calculating the feedback score includes: ; ; in, This represents the cross-modal semantic alignment score corresponding to the image side. This represents the cross-modal semantic alignment score corresponding to the text side. This represents the temperature hyperparameter, where, in The time is a non-matching image-text pair.

5. The training method for the text pedestrian retrieval model based on self-feedback learning according to claim 4, characterized in that, The self-feedback loss is derived from the feedback score, including: Input the feedback score into the preset self-feedback loss calculation formula to obtain the self-feedback loss output by the self-feedback loss calculation formula; The formula for calculating self-feedback loss includes: ; in, Indicates self-feedback loss. This represents the semantic alignment score of the matched text and image in the image orientation. This represents the semantic alignment score of the matched image and text in the text direction, where the matched image and text are the information indicating a successful match.

6. The training method for the text pedestrian retrieval model based on self-feedback learning according to any one of claims 1-5, characterized in that, Cross-modal mutual supervision learning is performed on the image answer features and text answer features corresponding to each subspace to obtain the mutual supervision loss, including: Input the image answer features and text answer features corresponding to each subspace into the preset mutual supervision loss calculation formula, and sum the mutual supervision loss output by the mutual supervision loss calculation formula to obtain the final mutual supervision loss; The formula for calculating mutual supervision loss includes: ; in, , ; in, Indicates mutual supervision losses, This indicates temperature hyperparameters. This represents the similarity score between image answer features generated under the same cue constraints within the k-th subspace. This represents the similarity score between text answer features generated under the same cue constraints within the k-th subspace. This represents the current subspace among K subspaces. This represents the image answer feature generated under the constraint of the k-th prompt feature, where the i-th image feature is used as the prompt feature. Let represent the text answer feature generated by the i-th text feature in the current subspace. Let i represent the text answer feature generated under the constraint of the k-th prompt feature, where the i-th text feature is used as the prompt feature. Let i represent the image answer feature generated by the i-th image feature in the current subspace.

7. The training method for the text pedestrian retrieval model based on self-feedback learning according to any one of claims 1-5, characterized in that, Based on comparative learning of target local text features (selected through an attention network), global text features, target local image features (selected through an attention network), and global image features, KL divergence loss and contrastive loss are obtained, including: The local text features and local image features of the target are normalized to construct a local cross-modal similarity matrix; Similarity calculations are performed on global text features and global image features to obtain a global cross-modal similarity matrix; The global cross-modal similarity matrix is ​​normalized to obtain the global cross-modal similarity probability distribution; The difference between the global cross-modal similarity probability distribution and the preset target distribution is calculated to obtain the KL loss, which is then used to optimize the global cross-modal similarity probability distribution. Positive and negative sample pairs are constructed based on the local cross-modal similarity matrix, and the contrastive loss is calculated based on the positive and negative sample pairs.

8. The training method for the text pedestrian retrieval model based on self-feedback learning according to any one of claims 1-5, characterized in that, Determine N image features corresponding to the image samples, and determine M text features and K prompt features corresponding to the text samples based on the prompt information, including: Input the image sample into the preset image encoder to obtain N image features output by the image encoder; Input the prompt information and text sample into the preset text encoder to obtain M text features and K prompt features output by the text encoder.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the training method for the text pedestrian retrieval model based on self-feedback learning as described in any one of claims 1 to 8.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the training method for a text pedestrian retrieval model based on self-feedback learning as described in any one of claims 1 to 8.