Text image pet re-identification method based on noise learning
By adopting a noise-based learning method in pet identification task, using Gaussian hybrid model and consensus division strategy, the problem of low recognition accuracy under the influence of noise data is solved, and higher robustness and matching accuracy are achieved, which is suitable for pet image recognition in complex backgrounds.
Patent Information
- Application Number
- CN202510450272.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art is difficult to effectively process noise data in pet recognition scenarios, resulting in low recognition accuracy and high computational complexity of traditional methods, making it difficult to adapt to pet image recognition of complex backgrounds.
Using a noise learning-based approach, through Gaussian hybrid model (GMM) and consensus division strategy, the loss function is optimized using smoothness to enhance the robustness of noise data, and supports fine-grained feature learning, effectively filtering noise data and calibrating labels.
It significantly improves the robustness and matching accuracy of noise data in cross-modal tasks, provides higher flexibility and theoretical guarantees, and can handle pet images under different lighting, backgrounds and postures, improving generalization capabilities in practical applications.
Smart Images

Figure CN119992113A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of text image re-identification, and more specifically, relates to a text image pet re-identification method based on noise learning. Background Art
[0002] The current methods of finding lost pets are not very effective, resulting in a low success rate of finding them. Therefore, improving the efficiency of pet retrieval and shortening the search time have become key issues that the industry urgently needs to break through, and are also an important demand for the application of new technologies.
[0003] When searching for lost pets, text image re-identification technology mainly adopts cross-modal retrieval. The main technical routes can be divided into two categories: one is global matching method and the other is local matching method. Global matching method realizes cross-modal feature mapping in a common latent space by integrating text and visual networks and combining matching loss function. This method pays too much attention to the correspondence of overall features and ignores the more detailed local feature interaction between text and image, so there is a bottleneck in its performance improvement. Local matching method focuses on exploring feature interactions at the detail level. This type of method attempts to accurately correspond specific body parts in the image with entities in the text description to achieve more accurate semantic matching. However, due to the need to process complex local associations, this type of method often requires more computing resources. With the rapid progress of visual language pre-training models, some studies have begun to use the rich correspondence knowledge accumulated in these models to enhance the matching effect of local and global features. This type of method has shown excellent performance in many cross-modal tasks, but they are generally based on an ideal assumption: that the text and image in the training data are accurately corresponding. However, in practical applications, noisy data is often encountered, making this assumption difficult to establish.
[0004] The so-called noisy data includes mislabeling negative samples as positive samples or incorrectly labeling the positive sample relationship. There are two main solutions to this problem: one is the sample screening method, and the other is the robust loss function method. The sample screening method uses the memory characteristics of deep neural networks to identify and exclude noise data through cyclic training, thereby focusing on the learning of high-quality samples; the robust loss function method is committed to designing a loss calculation mechanism that can accommodate noise and improve the training stability of the model in a noisy environment. Although these methods perform well in many fields, they still face unique challenges in the task of pet text image re-identification. First, the diversity of pet appearance and the ambiguity of text descriptions make it difficult for global matching methods to capture fine-grained associations; second, although local matching methods can handle details, changes in pet posture, lighting, etc. are easy to interfere with local feature extraction, and the computational complexity is high. In addition, existing noise processing methods fail to fully deal with the complex noise types in pet data, so there is still room for improvement in performance in pet scenes. Summary of the invention
[0005] In view of the deficiencies in the above-mentioned technologies, the present invention provides a text image pet re-identification method based on noise learning, which solves the problem of the influence of noise data in the complex background of pet identification and the incompatibility of traditional methods in pet identification scenarios.
[0006] In order to solve the above technical problems, the present invention is implemented in the following ways: A method for pet re-identification from text images based on noise learning, comprising the following steps: S1. Collect pet image-description text pair dataset and normalize it. S2, vectorize the images and description texts in the normalized dataset respectively; S3, performing global and local similarity calculations on the vectorized data set; S4, cross-modal retrieval model training; S5. Model reasoning and retrieval.
[0007] Furthermore, the step S1 specifically includes the following sub-steps: S11, collecting a pet image-description text pair dataset, wherein the dataset includes a pet image dataset and a description text dataset corresponding thereto, wherein the image includes images and description texts of different pet breeds under different background environments, shooting angles, and lighting conditions; S12, preprocessing the images collected in step S11, including target labeling, uniformly adjusting the image size to 384*128 pixels, and color normalization; and selecting the length range of the description text.
[0008] Furthermore, the step S2 specifically includes the following sub-steps: S21, vectorizing the data set after the normalization process in step S12, and extracting the global features and local features of the image and text respectively through the visual encoder and text encoder of the pre-trained CLIP model; S22, dividing the data set after vectorization processing in step S21 into input training data, verification data and test data.
[0009] Furthermore, the step S3 specifically includes the following sub-steps: S31, using the global features obtained in step S21, using cosine similarity to calculate the image text pair The global similarity of is expressed as follows: in, Represents image-text pairs The global similarity of and Respectively represent the global features of the i-th image and the j-th text obtained after global embedding; S32. Extract the self-attention maps of image and text data from the last layer Transformer block of the CLIP model , and then extract the correlation weight between the global Token and the local Token. The specific expression is as follows: in, N Indicates the number of local Tokens in the image ,M Indicates the number of local tokens in the text. Indicates The self-attention map of sample image data, Indicates Self-attention map of sample text data, Indicates The correlation weight between the global Token and all local Tokens in the sample image data, Indicates The correlation weights of the global Token and all local Tokens in the sample text data; select the image with the highest correlation weight with the text K local Token, denoted as and ,in represents the K local tokens with the highest correlation weights in the i-th image data, represents the K local tokens with the highest relevance weights in the i-th text data, v j Represents the jth Token of the current image, Represents the jth Token of the current text; S33, the selected step S32 K The local tokens of the image and text are L2 normalized respectively. The specific expressions are as follows: Then perform embedding transformation on them respectively, and the expressions are as follows: in, It represents the result of L2 normalization of K local Tokens in the i-th image data. It represents the result of L2 normalization of K local tokens in the i-th text data. express The result after embedding transformation is express The result after embedding transformation, MLP(·) represents multi-layer perceptron, FC(·) represents linear layer, and MaxPool(·) represents maximum pooling function; S34, obtained according to step S33 and , the cosine similarity is used to calculate the local similarity between the image and the text. The specific expression is as follows: in, Represents image-text pairs The local similarity of .
[0010] Furthermore, the step S4 specifically includes the following sub-steps: S41. For each training round, in a small batch Input image text pairs , and calculate the loss based on the function, the specific expression is as follows: in, Represents image-text pairs The loss function is Indicates images, T i Indicates text, represents the ReLU function used to ensure that the loss value is non-negative, represents the edge parameter, Representing images The weighted average similarity with the positive sample text pair, Represents text T i The weighted average similarity with the positive sample image pair, Indicates The image and j The global or local similarity of texts, Indicates j The image and i The global or local similarity of texts, t represents the temperature coefficient, Representing images and text The corresponding label, Representing images and text T i The corresponding label, B Represents small batch The number of samples in , Representing images The weight of the weighted average similarity with the positive sample text pair, Represents text T i The weight of the weighted average similarity with the positive sample image pair is expressed as follows:
[0011] S42. Use a two-component Gaussian mixture model to fit the sample loss distribution and calculate the posterior probability, which is expressed as follows: in, Represents image-text pairs is the probability of a clean sample, Represents image-text pairs is the probability of the noise sample; the threshold is set to 0.5, and the sample pairs are divided into a clean set and a noise set according to the posterior probability. The expression is as follows: in, Represents a clean set, represents the noise set; S43, through the consistency of global and local partitioning results, the final clean set, noise set and uncertain set are obtained. in, represents the final clean set, represents the final noise set, represents the final uncertain set, represents the clean set obtained using global similarity, represents the clean set obtained using local similarity, represents the noise set obtained using global similarity, Represents the noise set obtained by using local similarity; then recalibrate the labels of the sample pairs according to the division results, the expression is as follows: in, Represents the image and text sample pair after calibration The label of (1 for clean, 0 for noisy); S44. Use the calibrated labels to calculate the final matching loss, the specific expression of which is as follows: in, represents the final matching loss, represents the loss function calculated based on the global similarity in step S41, represents the loss function calculated based on the local similarity in step S41; S45, then use the optimizer to update the model parameters to minimize the matching loss; S46. After each training round, the model performance is evaluated using the verification data divided in step S21 to adjust the hyperparameters and monitor overfitting, and the optimized model parameters are output.
[0012] Furthermore, the step S5 specifically includes the following sub-steps: S51, using the test data divided in step S21, extracting global features and local features of the pet images and corresponding description texts in the test data through the model trained in step 46; S52: Calculate the global similarity of the image-text pair according to step S31 and step S34 S g Similarity with local S l ; S53: For each image-text pair , calculate the final similarity S The expression is as follows: in, S Represents the final similarity of the image-text pair, S g Represents the global similarity of the image-text pair, S l Indicates the local similarity of the image-text pair; S54. Sort the image-text pairs according to the final similarity, and return the most matching pet image result.
[0013] Compared with the prior art, the present invention has the following beneficial effects: The present invention uses a Gaussian mixture model (GMM) and a consensus partitioning strategy, performs smooth optimization on the loss function, enhances the robustness to noise data, supports fine-grained feature learning, and focuses on positive samples to effectively filter noisy data and calibrate labels, significantly improving the robustness and matching accuracy to noise data in cross-modal tasks, and providing higher flexibility and theoretical guarantees. At the same time, through cross-modal feature extraction and noise robustness design, the method is suitable for a variety of scenes with complex backgrounds (such as indoors, outdoors, at night), and can process pet images under different lighting, backgrounds, and postures, improving the generalization ability in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 Schematic diagram of the process of the re-identification method of the present invention; Figure 2 Schematic diagram of global and local feature embedding of the re-identification method of the present invention. DETAILED DESCRIPTION
[0015] The specific implementation of the present invention is further described in detail below with reference to the accompanying drawings and specific examples.
[0016] like Figure 1~2 As shown, a text image pet re-identification method based on noise learning includes the following steps: S1. Collect a dataset of pet images and description text pairs and perform normalization processing, which includes the following steps: S11, collecting a pet image-description text pair dataset, wherein the dataset includes a pet image dataset and a description text dataset corresponding thereto, wherein the image includes images and description texts of different pet breeds under different background environments, shooting angles, and lighting conditions; S12, preprocessing the images collected in step S11, including target labeling, uniformly adjusting the image size to 384*128 pixels, and color normalization; selecting the description text within a reasonable length range (23≤number of words≤77).
[0017] S2. Vectorize the images and description texts in the normalized dataset separately, including the following steps: S21, vectorizing the data set after the normalization process in step S12, and extracting the global features and local features of the image and text respectively through the visual encoder and text encoder of the pre-trained CLIP model; S22, dividing the data set after vectorization processing in step S21 into input training data, verification data and test data.
[0018] S3, performing global and local similarity calculations on the vectorized data set, specifically including the following sub-steps: S31, using the global features obtained in step S21, using cosine similarity to calculate the image text pair The global similarity of is expressed as follows: in, Represents image-text pairs The global similarity of and Respectively represent the global features of the i-th image and the j-th text obtained after global embedding; S32. Extract the self-attention maps of image and text data from the last layer Transformer block of the CLIP model , and then extract the correlation weight between the global Token and the local Token. The specific expression is as follows: in, N Indicates the number of local Tokens in the image ,M Indicates the number of local tokens in the text. Indicates The self-attention map of sample image data, Indicates Self-attention map of sample text data, Indicates The correlation weight between the global Token and all local Tokens in the sample image data, Indicates The correlation weight between the global token and all local tokens in the sample text data; Select the image and text with the highest relevance weight respectively K local Token, denoted as and ,in represents the K local tokens with the highest correlation weights in the i-th image data, represents the K local tokens with the highest relevance weights in the i-th text data, v j Represents the jth Token of the current image, Represents the jth Token of the current text; S33, the selected step S32 K The local tokens of images and texts are and L2 normalization is performed respectively, and the specific expressions are as follows: Then perform embedding transformation on them respectively, and the expressions are as follows: in, Represents the K local Tokens in the i-th image data, namely The result after L2 normalization is: Represents the K local Tokens in the i-th text data, namely The result after L2 normalization is: express The result after embedding transformation is express The result obtained after embedding transformation, MLP(·) represents multi-layer perceptron, FC(·) represents linear layer, and MaxPool(·) represents maximum pooling function; S34, obtained according to step S33 and , the cosine similarity is used to calculate the local similarity between the image and the text. The specific expression is as follows: in, Represents image-text pairs The local similarity of .
[0019] S4. Cross-modal retrieval model training includes the following steps: S41. For each training round, in a small batch Input image text pairs , and calculate the loss based on the function, the specific expression is as follows: in, Represents image-text pairs The loss function is Indicates images, Indicates text, represents the ReLU function used to ensure that the loss value is non-negative, represents the edge parameter, Representing images The weighted average similarity with the positive sample text pair, Represents text The weighted average similarity with the positive sample image pair, Indicates The image and j The global or local similarity of texts, Indicates j The image and The global or local similarity of texts, t represents the temperature coefficient, Representing images and text The corresponding label, Representing images and text The corresponding label, B Represents small batch The number of samples in , Representing images The weight of the weighted average similarity with the positive sample text pair, Represents text The weight of the weighted average similarity with the positive sample image pair is expressed as follows:
[0020] S42. Use a two-component Gaussian mixture model (GMM) to fit the sample loss distribution and calculate the posterior probability, which is expressed as follows: in, Represents image-text pairs is the probability of a clean sample, Represents image-text pairs is the probability of a noise sample; Set the threshold to 0.5, and divide the sample pairs into clean sets and noise sets according to the posterior probability. The expression is as follows: in, Represents a clean set, represents the noise set; S43, through the consistency of global and local partitioning results, the final clean set, noise set and uncertain set are obtained. in, represents the final clean set, represents the final noise set, represents the final uncertain set, represents the clean set obtained using global similarity, represents the clean set obtained using local similarity, represents the noise set obtained using global similarity, Represents the noise set obtained by using local similarity; then recalibrate the labels of the sample pairs according to the division results, the expression is as follows: in, Represents the image and text sample pair after calibration The label of (1 for clean, 0 for noisy); S44. Use the calibrated labels to calculate the final matching loss, the specific expression of which is as follows: in, represents the final matching loss, represents the loss function calculated based on the global similarity in step S41, represents the loss function calculated based on the local similarity in step S41; S45, then use the optimizer to update the model parameters to minimize the matching loss; S46. After each training round, the model performance is evaluated using the verification data divided in step S21 to adjust the hyperparameters and monitor overfitting, and the optimized model parameters are output.
[0021] S5, model reasoning and retrieval, specifically includes the following steps: S51, using the test data divided in step S21, extracting global features and local features of the pet images and corresponding description texts in the test data through the model trained in step 46; S52: Calculate the global similarity of the image-text pair according to step S31 and step S34 Similarity with local ; S53: For each image-text pair , calculate the final similarity S The expression is as follows: in, S Represents the final similarity of the image-text pair, Represents the global similarity of the image-text pair, Indicates the local similarity of the image-text pair; S54. Sort the image-text pairs according to the final similarity, and return the most matching pet image result.
[0022] The above description is only an implementation mode of the present invention. It is stated again that for ordinary technicians in this technical field, several improvements can be made to the present invention without departing from the principle of the present invention. These improvements are also included in the protection scope of the claims of the present invention.
Claims
1. A method for pet re-identification from text images based on noise learning, characterized by: The following steps are involved: S1. Collect pet image-description text pair dataset and normalize it. S2, vectorize the images and description texts in the normalized dataset respectively; S3, performing global and local similarity calculations on the vectorized data set; S4, cross-modal retrieval model training; S5. Model reasoning and retrieval.
2. The method for pet re-identification from text image based on noise learning as claimed in claim 1, characterized in that: The step S1 specifically includes the following sub-steps: S11, collecting a pet image-description text pair dataset, wherein the dataset includes a pet image dataset and a description text dataset corresponding thereto, wherein the image includes images and description texts of different pet breeds under different background environments, shooting angles, and lighting conditions; S12, preprocessing the images collected in step S11, including target labeling, uniformly adjusting the image size to 384*128 pixels, and color normalization; and selecting the length range of the description text.
3. A method for pet re-identification from text image based on noise learning as claimed in claim 2, characterized in that: The step S2 specifically includes the following sub-steps: S21, vectorizing the data set after the normalization process in step S12, and extracting the global features and local features of the image and text respectively through the visual encoder and text encoder of the pre-trained CLIP model; S22, dividing the data set after vectorization processing in step S21 into input training data, verification data and test data.
4. The method for pet re-identification from text image based on noise learning as claimed in claim 3, characterized in that: The step S3 specifically includes the following sub-steps: S31, using the global features obtained in step S21, using cosine similarity to calculate the image text pair The global similarity of is expressed as follows: in, Represents image-text pairs The global similarity of and Respectively represent the global features of the i-th image and the j-th text obtained after global embedding; S32. Extract the self-attention maps of image and text data from the last layer Transformer block of the CLIP model and , and then extract the correlation weight between the global Token and the local Token. The specific expression is as follows: in, N Indicates the number of local Tokens in the image ,M Indicates the number of local tokens in the text. Indicates The self-attention map of sample image data, Indicates Self-attention map of sample text data, Indicates The correlation weight between the global Token and all local Tokens in the sample image data, Indicates The correlation weight between the global token and all local tokens in the sample text data; Select the image and text with the highest relevance weight respectively K local Token, denoted as and ,in represents the K local tokens with the highest correlation weights in the i-th image data, represents the K local tokens with the highest relevance weights in the i-th text data, Represents the jth Token of the current image, Represents the jth Token of the current text; S33. For the selected K The local tokens of the image and text are L2 normalized respectively. The specific expressions are as follows: Then perform embedding transformation on them respectively, and the expressions are as follows: in, It represents the result of L2 normalization of K local Tokens in the i-th image data. It represents the result of L2 normalization of K local tokens in the i-th text data. express The result after embedding transformation is express The result after embedding transformation, MLP(·) represents multi-layer perceptron, FC(·) represents linear layer, and MaxPool(·) represents maximum pooling function; S34, obtained according to step S33 and , the cosine similarity is used to calculate the local similarity between the image and the text. The specific expression is as follows: in, Represents image-text pairs The local similarity of .
5. A method for pet re-identification from text image based on noise learning as claimed in claim 4, characterized in that: The step S4 specifically includes the following sub-steps: S41. For each training round, in a small batch Input image text pairs , and calculate the loss based on the function, the specific expression is as follows: in, Represents image-text pairs The loss function is Indicates images, Indicates text, represents the ReLU function used to ensure that the loss value is non-negative, represents the edge parameter, Representing images The weighted average similarity with the positive sample text pair, Represents text The weighted average similarity with the positive sample image pair, Indicates The image and The global or local similarity of texts, Indicates The image and The global or local similarity of texts, t represents the temperature coefficient, Representing images and text The corresponding label, Representing images and text The corresponding label, B Represents small batch The number of samples in , , Representing images The weight of the weighted average similarity with the positive sample text pair, Represents text The weight of the weighted average similarity with the positive sample image pair is expressed as follows: ; S42. Use a two-component Gaussian mixture model to fit the sample loss distribution and calculate the posterior probability, which is expressed as follows: in, Represents image-text pairs is the probability of a clean sample, Represents image-text pairs is the probability of a noise sample; Set the threshold to 0.5, and divide the sample pairs into clean sets and noise sets according to the posterior probability. The expression is as follows: in, Represents a clean set, represents the noise set; S43, through the consistency of global and local partitioning results, the final clean set, noise set and uncertain set are obtained. in, represents the final clean set, represents the final noise set, represents the final uncertain set, represents the clean set obtained using global similarity, represents the clean set obtained using local similarity, represents the noise set obtained using global similarity, represents the noise set obtained using local similarity; Then recalibrate the labels of the sample pairs according to the division results. The expression is as follows: in, Represents the image and text sample pair after calibration Labels; S44. Use the calibrated labels to calculate the final matching loss, the specific expression of which is as follows: in, represents the final matching loss, represents the loss function calculated based on the global similarity in step S41, represents the loss function calculated based on the local similarity in step S41; S45, then use the optimizer to update the model parameters to minimize the matching loss; S46. After each training round, the model performance is evaluated using validation data to adjust hyperparameters and monitor overfitting, and the optimized model parameters are output.
6. A method for pet re-identification from text image based on noise learning as claimed in claim 5, characterized in that: The step S5 specifically includes the following sub-steps: S51, using the test data, extracting global features and local features of the pet images and corresponding description texts in the test data through the trained model; S52: Calculate the global similarity of the image-text pair according to step S31 and step S34 Similarity with local ; S53: For each image-text pair , calculate the final similarity S The expression is as follows: in, S Represents the final similarity of the image-text pair, S g Represents the global similarity of the image-text pair, S l Indicates the local similarity of the image-text pair; S54. Sort the image-text pairs according to the final similarity, and return the most matching pet image result.
Citation Information
Patent Citations
Cross-modal image-text retrieval method and system based on image-text semantic similarity optimization
CN118484545A
Cited By
Multi-modal data hybrid retrieval method and device, equipment and storage medium
CN121478992A
Rumor detection method based on multi-modal data conversion network
CN121744248A
Text image re-identification method based on phrase-level mask and large language model
CN121962783A