Method and apparatus for training human body recognition model, device, and medium

By training the human body recognition model, and using the adaptive encoder and the image encoder of the image-text model to extract features, the problem of low recognition accuracy in the self-supervised re-recognition method is solved, and higher recognition accuracy is achieved.

WO2025139388A1PCT designated stage expired Publication Date: 2025-07-03SHENZHEN INTELLIFUSION TECHNOLOGIES CO LTD

Patent Information

Application Number
PCT/CN2024/130476
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-29
Filing Date
2024-11-07
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

The existing self-supervised re-identification methods have low recognition accuracy in human body recognition and are difficult to meet the practical application needs.

Method used

By obtaining the second input human image and image tag, the human body recognition model is trained, and feature extraction is performed using a preset adaptive encoder and the image encoder in the trained image-text model to improve feature distinction.

Benefits of technology

The training accuracy and recognition accuracy of the human body recognition model are improved, and the model's distinction ability in human body image recognition is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024130476_03072025_PF_FP_ABST
    Figure CN2024130476_03072025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of artificial intelligence, and relates in particular to a method and an apparatus for training a human body recognition model, a device, and a medium. The method comprises: training a human body recognition model by using second input human body images and image labels of the second input human body images, so as to obtain a trained human body recognition model, wherein the trained human body recognition model is used for performing feature extraction on a human body image, and the human body recognition model comprises a preset adaptive encoder, and an image encoder in a trained image-text model. In the present application, the human body recognition model, which comprises the preset adaptive encoder and the image encoder in the trained image-text model, is trained, and during training, the preset adaptive encoder and the image encoder in the trained image-text model are used to separately perform feature extraction, so that features extracted by the adaptive encoder are fused into image features, enhancing feature discriminability and thereby improving the training accuracy of the human body recognition model.
Need to check novelty before this filing date? Find Prior Art

Description

A training method, device, equipment and medium for human body recognition model Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a training method, device, equipment and medium for a human body recognition model.

[0002] This application claims priority to the Chinese patent application filed with the China Patent Office on December 29, 2023, with application number 202311871584.5 and invention name “A training method, device, equipment and medium for a human body recognition model”, the entire contents of which are incorporated by reference into this application. Background Art

[0003] With the advancement and development of artificial intelligence technology and the growing needs of corporate management and public safety, human body recognition technology has been widely used in all aspects of social life because of its ability to track, match and identify target people across time and space. It is also one of the research hotspots in the field of computer vision in recent years.

[0004] The principle of supervised learning-based re-identification methods is to use human images as the input of the re-identification model and manually annotated human identity labels as the expected output of the model, thereby training the model to extract the identity features of the human image and classify the human identity. Since supervised learning methods require manual annotation of a large number of paired data labels, in practical applications, it is costly to annotate large-scale datasets for each application scenario, and this method is greatly limited in practical applications. To solve this problem, some self-supervised re-identification methods have been developed in recent years, mainly by clustering unlabeled data or transferring knowledge from the labeled source data domain to the target data domain. However, the model performance of existing self-supervised re-identification methods is not satisfactory. Compared with supervised algorithms, the recognition accuracy of the models obtained based on self-supervised re-identification is significantly reduced. Therefore, in the process of self-supervised re-identification learning, how to improve the recognition accuracy of the re-identification model has become an urgent problem to be solved. Technical issues

[0005] In view of this, embodiments of the present invention provide a method, apparatus, device, and medium for training a human recognition model to solve the problem of low re-identification accuracy of the re-identification model during self-supervised re-identification learning.

[0006] In a first aspect, an embodiment of the present invention provides a method for training a human body recognition model, the method comprising:

[0007] Obtaining a second input human image and an image label of the second input human image;

[0008] The human body recognition model is trained using the second input human body image and the image label of the second input human body image to obtain a trained human body recognition model. The trained human body recognition model is used to extract features from the human body image. The human body recognition model includes a preset adaptive encoder and an image encoder in a trained image-text model.

[0009] In a second aspect, an embodiment of the present invention provides a training device for a human body recognition model, the training device comprising:

[0010] an acquisition module, configured to acquire a second input human body image and an image label of the second input human body image;

[0011] A training module is used to train the human body recognition model using the second input human body image and the image label of the second input human body image to obtain a trained human body recognition model. The trained human body recognition model is used to extract features from the human body image. The human body recognition model includes a preset adaptive encoder and an image encoder in a trained image-text model.

[0012] In a third aspect, an embodiment of the present invention provides a computer device, comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor implements the training method described in the first aspect when executing the computer program.

[0013] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the training method as described in the first aspect is implemented.

[0014] Compared with the prior art, the present invention has the following beneficial effects:

[0015] A second input human body image and an image label of the second input human body image are obtained, and the second input human body image and the image label of the second input human body image are used to train a human body recognition model to obtain a trained human body recognition model. The trained human body recognition model is used to extract features from the human body image. The human body recognition model includes a preset adaptive encoder and an image encoder in a trained image-text model. In the present application, a human body recognition model including a preset adaptive encoder and an image encoder in a trained image-text model is trained. During the training process, the preset adaptive encoder and the image encoder in the trained image-text model are used to extract features respectively, so as to facilitate the fusion of the features extracted by the adaptive encoder into the image features, improve the discrimination of the features, and thus improve the training accuracy of the human body recognition model. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] FIG1 is a schematic diagram of an application environment of a training method for a human body recognition model provided by an embodiment of the present invention;

[0017] FIG2 is a flow chart of a method for training a human body recognition model according to an embodiment of the present invention;

[0018] FIG3 is a schematic structural diagram of a training device for a human body recognition model provided by one embodiment of the present invention;

[0019] FIG4 is a schematic diagram of the structure of a computer device provided by an embodiment of the present invention. Modes for Carrying Out the Invention

[0020] In order to illustrate the technical solution of the present invention, specific embodiments are provided below.

[0021] A method for training a human recognition model provided by one embodiment of the present invention can be applied in an application environment such as that shown in Figure 1 , wherein a client communicates with a server. The client includes, but is not limited to, PDAs, desktop computers, laptops, ultra-mobile personal computers (UMPCs), netbooks, cloud computing devices, personal digital assistants (PDAs), and other computer devices. The server can be a standalone server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0022] Refer to Figure 2, which is a flow chart of a training method for a human body recognition model provided by an embodiment of the present invention. The above-mentioned training method for a human body recognition model can be applied to the server in Figure 1. As shown in Figure 2, the training method for the human body recognition model can include the following steps.

[0023] S201: Acquire a second input human body image and an image label of the second input human body image.

[0024] In step S201 , the second input human body image is an image with an image tag of a human body identifier.

[0025] In this embodiment, a second input human image and an image label of the second input human image are obtained. The second input human image can be an original image captured by an image acquisition device such as a camera, or can be an image obtained by preprocessing the original image. The preprocessing operation can include a denoising operation such as filtering. The image label of the second input human image is an identification label of the human body in the corresponding image, such as an identity label.

[0026] Optionally, before training the human body recognition model using the second input human body image and the image label of the second input human body image, the method further includes:

[0027] Performing pseudo-label annotation on the first input human body image using a pre-trained image-text model and preset text prompt words to obtain a pseudo-label of the first input human body image;

[0028] According to the first input human image and the pseudo label, supervised training is performed on the pre-trained image-text model to update the parameters of the pre-trained image-text model.

[0029] In this embodiment, a pre-trained image-text model is obtained, wherein the pre-trained image-text model is a model for multimodal tasks of processing images and text, and includes an image encoder and a text encoder. The image encoder is used to extract image features of a human image, and the text encoder is used to extract text features of text used to describe the human image. The first input human image is an unlabeled image. The first input human image can be an original image captured by an image acquisition device such as a camera, or it can be an image obtained by pre-processing the original image.

[0030] It should be noted that when using an image encoder to extract image features, the image embedding features and the image position encoding features can be extracted; when using a text encoder to extract text features, the word embedding features and the text position encoding features can be extracted.

[0031] In another embodiment, the pre-trained image-text model may be a pre-trained CLIP (Contrastive Language-Image Pre-Training) model. The CLIP model is obtained by training on 400 million image-text pairs, and has high accuracy and strong generalization ability.

[0032] A pre-trained image-text model is used to annotate the first input human body image with pseudo labels, wherein the pseudo labels are text prompt words with the most similar features found in the search for the first input human body image, and the text prompt words with the most similar features to the first input human body image are used as the labels of the first input human body image.

[0033] In this embodiment, when using a pre-trained image-text model to pseudo-label the first input human image, it is necessary to annotate the first input human image with a corresponding text label. Therefore, a preset text prompt word is required as an alternative text label so that the text prompt word that is most similar to the features of the first input human image can be searched from multiple preset text prompt words.

[0034] An image encoder in a pre-trained image-text model is used to extract image features of a first input human image to obtain image features. A text encoder in the pre-trained image-text model is used to extract features of preset text prompt words to obtain text features. Based on the similarity between the text features and the image features, a text prompt word that is most similar to the features of the first input human image is determined as a pseudo label for the first input human image, wherein the pseudo label can be a single preset text prompt word or a combination of multiple preset text prompt words.

[0035] In this embodiment, by pseudo-labeling the first input human image, the pre-trained image-text model can be fine-tuned according to the pseudo-label of the first input human image to obtain a trained image-text model, so that the trained image-text model can adapt to feature extraction of human images, so that the trained image-text model can better adapt to downstream tasks.

[0036] In another embodiment, if the first input human body image is a plurality of images, when a pre-trained image-text model is used to pseudo-label the first input human body image, pseudo-label labeling can also be performed by clustering. For example, a plurality of first input human body images and classification dimensions of the plurality of first input human body images are obtained. For example, the plurality of first input human body images include human body images of 5 people, then the classification dimension of the plurality of first input human body images is 5, and the image encoder in the pre-trained image-text model is used to feature encode the first input human body image to obtain the image coding feature of each first input human body image, and the image coding features of all the first input human body images are clustered to obtain clustering results of 5 cluster centers, and one of the image coding features of each cluster center is extracted, and the similarity between the image coding feature and the text feature of the preset text prompt word is calculated to determine the corresponding text prompt word of each cluster center, and the corresponding text prompt word is used as the pseudo label of the first input human body image contained in each cluster center.

[0037] According to the first input human image and the pseudo label, supervised training is performed on the pre-trained image-text model to obtain a trained image-text model, wherein during the supervised training, parameters of the image encoder and the text encoder in the pre-trained image-text model are adjusted.

[0038] In this embodiment, a first input human image and a pseudo-label are respectively input to an image encoder and a text encoder in a pre-trained image-text model, which output image encoding features and text encoding features. Based on the image encoding features and the text encoding features, a similarity loss between the image encoding features and the text encoding features is calculated. The pre-trained image-text model is trained based on the similarity loss to obtain a trained image-text model. During training, the parameters of the image encoder and the text encoder are continuously adjusted to converge the similarity loss, thereby obtaining a trained image-text model. Alternatively, training is terminated when a preset number of training cycles is reached, thereby obtaining a trained image-text model.

[0039] In this embodiment, supervised training is performed on a pre-trained image-text model based on a first input human image and pseudo-labels to obtain a trained image-text model. The pseudo-labels are obtained based on the pre-trained image-text model. The pseudo-labels are used to perform supervised training on the pre-trained image-text model and adjust parameters of the image encoder and text encoder in the pre-trained image-text model, thereby improving training efficiency.

[0040] Optionally, obtain a pre-trained image-text model, including:

[0041] Obtain an initial image-text model, an image dataset, and text labels for the image dataset. The initial image-text model includes an initial image encoder based on a self-attention mechanism structure and an initial text encoder based on a self-attention mechanism structure.

[0042] Use the initial image encoder to encode the features of the image dataset to obtain the initial image features, and use the initial text encoder to encode the features of the text labels to obtain the initial text features;

[0043] Calculate the initial similarity loss based on the initial image features and the initial text features;

[0044] The initial image-text model is pre-trained according to the initial similarity loss to obtain a pre-trained image-text model.

[0045] In this embodiment, an initial image-text model, an image dataset, and text labels for the image dataset are obtained. The initial image-text model includes an initial image encoder based on a self-attention mechanism structure and an initial text encoder based on a self-attention mechanism structure. The image dataset can be a general dataset, such as LAION-5B. The LAION-5B dataset is a multimodal image-text dataset that can be used to pre-train the initial image-text model. The image dataset and the text labels for the image dataset are image-text pairs, i.e., the text labels are used to represent the textual representation of the corresponding image dataset. For example, if the image dataset is a human image with human identifier ID1, the text label is ID1.

[0046] The image dataset is feature-encoded using the initial image encoder to obtain initial image features. When the image dataset is feature-encoded using the initial image encoder, the images in the image dataset can be segmented to obtain corresponding image blocks, where each image block has an equal size. For example, the image size in the image dataset is adjusted to an equal size. For example, the image size in the image dataset is adjusted to If the image size in the image dataset is smaller than , then fill the images in the image dataset. If the image size in the image dataset is larger than , then compress the images in the image dataset. After adjusting to a fixed size, the image is segmented to obtain image blocks of equal size. When segmenting, the image can be segmented into image blocks containing fixed rows and fixed columns according to the image size. For example, The image is divided into 24 rows and 8 columns, totaling 192 image blocks. The image blocks are feature-encoded using an initial image encoder, and embedded coding features and position coding features of each image are extracted. The embedded coding features and position coding features are fused to obtain initial image features. When fusion is performed, the embedded coding features and position coding features can be fused by feature addition, or other fusion methods can be used, which are not limited in this embodiment.

[0047] The initial text encoder is used to perform feature encoding on the text label to obtain the initial text feature. When the initial text encoder is used to perform feature encoding on the text label, the text label is segmented, and the initial text encoder is used to extract the word embedding encoding feature and position encoding feature of each segmented word. The word embedding encoding feature and the position encoding feature are feature fused to obtain the initial text feature. When fusing features, the word embedding encoding feature and the position encoding feature can be feature added and fused, or other fusion methods can be used for fusion, which is not limited in this embodiment.

[0048] When calculating the initial similarity loss based on the initial image features and the initial text features, the similarity loss function is used for calculation. The initial similarity loss calculation formula is as follows:

[0049] in, is the initial similarity loss, is the initial image feature, is the initial text feature.

[0050] In another embodiment, when calculating the initial similarity loss based on the initial image features and the initial text features, since the obtained initial image features are the image features of each image block, the initial image features can be averaged through a multi-layer stacked encoder to obtain the average features of each image block, and the average features can be output as the final initial image features. The multi-layer stacked encoder is composed of a plurality of encoders connected together, and the output features of the previous encoder are used as the input features of the next encoder. The initial text features can be averaged through a multi-layer stacked encoder to obtain the average features of each word, and the average features of each word can be output as the final initial text features. Therefore, when calculating the initial similarity loss, the similarity loss between the average features of each image block and the average features of each word can be calculated.

[0051] The initial image-text model is pre-trained according to the initial similarity loss to obtain a pre-trained image-text model. During pre-training, the initial similarity loss is converged by continuously adjusting the parameters in the image encoder and the text encoder to obtain a pre-trained image-text model. Alternatively, when the number of pre-training reaches a preset number, the pre-training is stopped to obtain a pre-trained image-text model.

[0052] In this embodiment, the initial image-text model is pre-trained to obtain a pre-trained image-text model. The image encoder and text encoder in the pre-trained image-text model can extract image features and text features in the image-text pair, so that the image features and text features in the image-text pair are similar, so that the pre-trained image-text model can be used to search for multiple text prompt words with the closest features for the unlabeled image, and the text pseudo-label of the unlabeled image is obtained by combining multiple text prompt words.

[0053] Optionally, pseudo-labeling the first input human image using a pre-trained image-text model and preset text prompt words to obtain a pseudo label for the first input human image includes:

[0054] Use the image encoder in the pre-trained image-text model to perform feature encoding on the first input human image to obtain image encoding features;

[0055] Use the text encoder in the pre-trained image-text model to perform feature encoding on the preset text prompt words to obtain text encoding features;

[0056] According to the image encoding features and the text encoding features, the target text prompt word is matched to the first input human image, and the target text prompt word is determined as a pseudo label of the first input human image.

[0057] In this embodiment, a pre-trained image-text model is used to perform feature encoding on a first input human image to obtain image encoding features. When performing feature encoding on the first input human image, the first input human image is segmented to obtain image blocks of equal size. Embedded coding features and position coding features are extracted from each image, and the embedded coding features and position coding features are fused to obtain image coding features. When fusion is performed, the embedded coding features and position coding features can be fused by feature addition, or other fusion methods can be used, which are not limited in this embodiment.

[0058] Use the text encoder in the pre-trained image-text model to perform feature encoding on the preset text prompt words to obtain text encoding features, wherein the preset text prompt words are pre-set human identification information, such as identity information, etc. When the preset text prompt words are feature encoded, the preset text prompt words are segmented, and the word embedding encoding features and position encoding features of each segmented word are extracted. The word embedding encoding features and the position encoding features are feature fused to obtain text encoding features. When fusing features, the word embedding encoding features and the position encoding features can be feature added and fused, or other fusion methods can be used for fusion, which is not limited in this embodiment.

[0059] Based on the image coding features and the text coding features, a target text prompt word is matched to the first input human image, and the target text prompt word is determined as a pseudo label for the first input human image. When matching the target text prompt word to the first input human image, the target text prompt word can be determined by calculating the distance between the image coding features and the text coding features, and the preset text prompt word corresponding to the text coding feature with the smallest distance between the image coding features and the text coding features is used as the target text prompt word.

[0060] In this embodiment, the image encoder in the pre-trained image-text model is directly used to perform feature encoding on the first input human image to obtain image encoding features, and the text encoder in the pre-trained image-text model is used to perform feature encoding on the preset text prompt words to obtain text encoding features. According to the image encoding features and the text encoding features, the target text prompt words are matched to the first input human image. There is no need to manually annotate the first input human image, and the pseudo label of the first input human image can be directly obtained through the model, so that the corresponding pseudo label can be used for supervised training to improve training efficiency.

[0061] Optionally, matching the target text prompt word to the first input human body image according to the image encoding feature and the text encoding feature includes:

[0062] Calculate the similarity between the image encoding feature and each text encoding feature to obtain a similarity value;

[0063] According to the similarity value, the first input human body image is matched with the target text prompt word.

[0064] In this embodiment, the cosine formula can be used to calculate the similarity between the image coding feature and each text coding feature to obtain the similarity value between the image coding feature and each text coding feature, and the preset text prompt word with the maximum similarity value is selected as the target text prompt word.

[0065] S202: Use the second input human body image and the image label of the second input human body image to train the human body recognition model to obtain a trained human body recognition model. The trained human body recognition model is used to extract features from the human body image. The human body recognition model includes a preset adaptive encoder and an image encoder in the trained image-text model.

[0066] In step S202, the human body recognition model includes a preset adaptive encoder and an image encoder in a trained image-text model, wherein the image encoder in the trained image-text model is a trained encoder, so when training the human body recognition model, only the parameters in the preset adaptive encoder are adjusted to obtain a trained human body recognition model, and the trained human body recognition model is used for feature extraction.

[0067] In this embodiment, a human body recognition model is trained using a second input human body image and an image label of the second input human body image. During training, supervised training is used. A loss is calculated based on the image label of the second input human body image and a classification result output by the human body recognition model. When calculating the loss, a cross-entropy loss function can be used. Parameters in a preset adaptive encoder are adjusted based on the loss to obtain a trained human body recognition model.

[0068] It should be noted that the human body recognition model includes a preset adaptive encoder and an image encoder in a trained image-text model, wherein the image encoder in the trained image-text model is a trained image encoder, and the image encoder in the trained image-text model can extract high-precision image features of human body images. The preset adaptive encoder is a low-dimensional encoder composed of a fully connected layer. The preset adaptive encoder is connected in parallel with the image encoder in the trained image-text model, and the image encoder and the preset adaptive encoder are used respectively to perform feature encoding on the human body image.

[0069] It should be noted that the preset adaptive encoder can be composed of two fully connected layers with the same structure, connected in series, and the feature dimensions of the fully connected layers are equal to the dimensions of the encoded features output by the image encoder in the trained image-text model. This facilitates feature fusion of the encoded features output by the preset adaptive encoder and the encoded features output by the image encoder in the trained image-text model.

[0070] It should be noted that when training the human body recognition model, whether to stop training can be determined based on whether the number of training times reaches a preset number of training times. If the number of training times reaches the preset number of training times, training of the human body recognition model is stopped. Whether to stop training can also be determined based on whether the loss converges. If the calculated loss converges, training of the human body recognition model is stopped. This embodiment does not limit the training process of the human body recognition model.

[0071] In this embodiment, the human body recognition model is trained by supervised training, thereby improving the training accuracy. Thus, when the trained human body recognition model is used to perform human body recognition on human images, the recognition accuracy of human body recognition can be improved.

[0072] After obtaining the trained human body recognition model, the trained human body recognition model is used to recognize the human body image to be recognized to obtain the recognition result.

[0073] Optionally, using the second input human image and the image label of the second input human image to train a human body recognition model to obtain a trained human body recognition model includes:

[0074] Using the human body recognition model to extract features from the second input human body image to obtain image features;

[0075] Performing human body identification classification on the image features to obtain human body identification classification results;

[0076] Calculate the classification loss based on the human body identification classification results and image labels;

[0077] According to the classification loss, the human body recognition model is supervised and trained, and the parameters of the adaptive encoder preset in the human body recognition model are adjusted to obtain a trained human body recognition model.

[0078] In this embodiment, when supervised training is performed on a human body recognition model, the human body recognition model is used to extract features from a second input human body image to obtain image features, where the image features include features extracted by the image encoder in the human body recognition model and features extracted by a preset adaptive encoder. The image features are then classified by human body identification to obtain a human body identification classification result. A classification loss is calculated based on the human body identification classification result and the image label. The classification loss can be calculated using a cross-entropy loss function. Based on the classification loss, supervised training is performed on the human body recognition model, and parameters of the preset adaptive encoder in the human body recognition model are adjusted to obtain a trained human body recognition model.

[0079] It should be noted that when the human body recognition model is supervised and trained, the parameters of the image encoder in the human body recognition model are fixed.

[0080] In this embodiment, the second input human body image and the image label of the second input human body image are used to perform supervised training on the human body recognition model. During supervised training, the parameters of the image encoder in the human body recognition model are fixed, and only the parameters in the preset adaptive encoder are adjusted, thereby reducing the adjusted parameters and improving the training efficiency. Since the image encoder in the human body recognition model is the encoder of the trained image-text model, the image encoder can be used to extract high-precision image features. Therefore, in the initial stage of training the human body recognition model, the image features extracted by the preset human body recognition model have low accuracy. By fixing the parameters of the image encoder in the human body recognition model, high-precision image features can be obtained, thereby improving the accuracy of the image features extracted by the human body recognition model for the second input human body image, so as to facilitate more accurate adjustment of the parameters in the preset adaptive encoder, thereby improving the training efficiency of the human body recognition model.

[0081] Optionally, a human body recognition model is used to perform feature extraction on the second input human body image to obtain image features, including:

[0082] Performing feature extraction on the second input human image using the image encoder in the human body recognition model to obtain first image features;

[0083] Using a preset adaptive encoder in the human body recognition model to perform feature extraction on the second input human body image to obtain second image features;

[0084] The first image feature is fused with the second image feature to obtain an image feature.

[0085] In this embodiment, the image encoder in the human body recognition model is used to extract features of the second input human body image to obtain a first image feature, wherein the first image feature is the image block feature after the second input human body image is divided into blocks, and the preset adaptive encoder in the human body recognition model is used to extract features of the second input human body image to obtain a second image feature, wherein the feature dimension of the second image feature is equal to that of the first image feature, and the first image feature and the second image feature are feature fused to obtain an image feature, wherein when the features are fused, the first image feature and the second image feature can be added to perform fusion.

[0086] In this embodiment, by performing feature fusion between features output by the image encoder and the preset adaptive encoder, and using the fused features as image features, feature accuracy is improved, thereby improving the calculation accuracy of the classification loss when calculating the classification loss.

[0087] In another embodiment, when performing feature fusion on the first image feature and the second image feature, the first image feature and the second image feature can also be weighted and summed to obtain the image feature, and dynamic weights are set for the first image feature and the second image feature. At the initial stage of training, the image encoder is a trained encoder, and the accuracy of the output first image feature is relatively high. The preset adaptive encoder is the initial encoder, and the accuracy of the output second image feature is relatively low. At the initial stage of training, a higher first weight value can be set for the first image feature, and a lower second weight value can be set for the second image feature, and the sum of the first weight value and the second weight value is 1. It should be noted that when setting the initial values ​​of the first weight value and the second weight value, the first weight value is set to be greater than one-half, and the second weight value is set to be less than one-half.

[0088] As the number of training times increases, the first weight value and the second weight value are adjusted, the first weight value is gradually reduced, and the second weight value is gradually increased. When adjusting the first weight value and the second weight value, they can be adjusted step by step according to the number of training times. For example, when the number of training times increases by 10 times, the first weight value is reduced by one tenth, and the second weight value is increased by one tenth, until the first weight value is reduced to one half, and the adjustment of the values ​​of the first weight value and the second weight value is stopped.

[0089] Optionally, using the second input human image and the image label of the second input human image to train a human body recognition model to obtain a trained human body recognition model further includes:

[0090] Get positive sample images with the same label as the image, and negative sample images with different labels from the image;

[0091] Use the human body recognition model to extract features from the positive sample image to obtain the features of the positive sample image, and use the human body recognition model to extract features from the negative sample image to obtain the features of the negative sample image;

[0092] Calculate triplet loss based on image features, positive sample image features, and negative sample image features;

[0093] According to the classification loss and triplet loss, the parameters of the preset adaptive encoder in the human recognition model are adjusted to obtain a trained human recognition model.

[0094] In this embodiment, a positive sample image with the same label as the image and a negative sample image with a different label are obtained, wherein the human body in the positive sample image is the same person as the human body in the second input human body image, and the human body in the negative sample image is a different person from the human body in the second input human body image. A human body recognition model is used to extract features from the positive sample image to obtain positive sample image features, and a human body recognition model is used to extract features from the negative sample image to obtain negative sample image features. The positive sample image features should be similar to the image features of the second input human body image, and the negative sample image features should be distant from the image features of the second input human body image. A triplet loss is calculated based on the image features, the positive sample image features, and the negative sample image features.

[0095] Among them, the triplet loss refers to setting three samples in a feature space: the second input human image, the positive sample image and the negative sample image. The second input human image and the positive sample image are different samples of the same type, and the second input human image and the negative sample image are different samples. In the feature space, the distance between the second input human image and the positive sample image of the same category is closer than the distance between the second input human image and the negative sample image of different categories, and the distance between the second input human image and the negative sample image is much larger than the distance between the second input human image and the positive sample image.

[0096] Based on the classification loss and the triplet loss, the parameters of the adaptive encoder preset in the human recognition model are adjusted to obtain a trained human recognition model. In this embodiment, the classification loss and the triplet loss are added to obtain a target loss. Based on the target loss, the parameters of the adaptive encoder preset in the human recognition model are adjusted until the target loss converges or the number of training times reaches a preset number of training times, thereby obtaining a trained human recognition model.

[0097] In this embodiment, the human body recognition model based on triplet loss can better distinguish details, especially in image classification. When two very similar images are input, the triplet loss can learn a better representation of these two vectors with smaller differences, making the human body recognition model perform better. In order to prevent overfitting, the triplet loss is combined with the classification loss to ensure the accuracy of the model after training. Training with triplet loss is conducive to obtaining a model that pays more attention to detail distinction. The triplet loss can learn more subtle differences between images. Combined with the classification loss, it improves the model's ability to distinguish samples of different classes, thereby ensuring the accuracy of recognition.

[0098] A second input human body image and an image label of the second input human body image are obtained, and the second input human body image and the image label of the second input human body image are used to train a human body recognition model to obtain a trained human body recognition model. The trained human body recognition model is used to extract features from the human body image. The human body recognition model includes a preset adaptive encoder and an image encoder in a trained image-text model. In the present application, a human body recognition model including a preset adaptive encoder and an image encoder in a trained image-text model is trained. During the training process, the preset adaptive encoder and the image encoder in the trained image-text model are used to extract features respectively, so as to facilitate the fusion of the features extracted by the adaptive encoder into the image features, improve the discrimination of the features, and thus improve the training accuracy of the human body recognition model.

[0099] Please refer to Figure 3, which is a schematic diagram of the structure of a human body recognition model training device provided in an embodiment of the present invention. In this embodiment, the terminal includes various units configured to execute the steps described in the embodiment corresponding to Figure 2. For details, please refer to the relevant description of the embodiment corresponding to Figure 2. For ease of illustration, only the portions relevant to this embodiment are shown. Referring to Figure 3, training device 30 includes an acquisition module 31 and a training module 32.

[0100] The acquisition module 31 is configured to acquire a second input human body image and an image label of the second input human body image.

[0101] The training module 32 is used to train the human body recognition model using the second input human body image and the image label of the second input human body image to obtain a trained human body recognition model. The trained human body recognition model is used to extract features from the human body image. The human body recognition model includes a preset adaptive encoder and an image encoder in the trained image-text model.

[0102] Optionally, the training device 30 further includes:

[0103] The labeling module is used to perform pseudo-label labeling on the first input human image using a pre-trained image-text model and preset text prompt words to obtain a pseudo label of the first input human image.

[0104] The supervised training module is used to perform supervised training on the pre-trained image-text model according to the first input human image and the pseudo label, so as to update the parameters of the pre-trained image-text model.

[0105] Optionally, the annotation module includes:

[0106] The image encoding unit is used to use the image encoder in the pre-trained image-text model to perform feature encoding on the first input human body image to obtain image encoding features.

[0107] The text encoding unit is used to use the text encoder in the pre-trained image-text model to perform feature encoding on the preset text prompt words to obtain text encoding features.

[0108] The matching unit is used to match the target text prompt word to the first input human image according to the image coding feature and the text coding feature, and determine the target text prompt word as a pseudo label of the first input human image.

[0109] Optionally, the matching unit includes:

[0110] The calculation subunit is used to calculate the similarity between the image encoding feature and each text encoding feature to obtain a similarity value.

[0111] The matching subunit is used to match the target text prompt word to the first input human body image according to the similarity value.

[0112] Optionally, the training module 32 includes:

[0113] The extraction unit is used to use the human body recognition model to perform feature extraction on the second input human body image to obtain image features.

[0114] The classification unit is used to classify the image features into human body identifications to obtain human body identification classification results.

[0115] The classification loss calculation unit is used to calculate the classification loss based on the human body identification classification results and image labels.

[0116] The training unit is used to perform supervised training on the human body recognition model according to the classification loss, adjust the parameters of the adaptive encoder preset in the human body recognition model, and obtain a trained human body recognition model.

[0117] Optionally, the extraction unit includes:

[0118] The first extraction subunit is configured to perform feature extraction on the second input human image using the image encoder in the human body recognition model to obtain first image features.

[0119] The second extraction subunit is used to use a preset adaptive encoder in the human body recognition model to perform feature extraction on the second input human body image to obtain second image features.

[0120] The fusion subunit is used to perform feature fusion on the first image feature and the second image feature to obtain an image feature.

[0121] Optionally, the training module 32 further includes:

[0122] The sample acquisition unit is used to acquire positive sample images with the same label as the image and negative sample images with different labels from the image.

[0123] The feature extraction unit is used to use the human body recognition model to extract features from the positive sample image to obtain positive sample image features, and to use the human body recognition model to extract features from the negative sample image to obtain negative sample image features.

[0124] The triplet loss calculation unit is used to calculate the triplet loss based on the image features, the positive sample image features and the negative sample image features.

[0125] The adjustment unit is used to adjust the parameters of the adaptive encoder preset in the human recognition model according to the classification loss and the triplet loss to obtain a trained human recognition model.

[0126] It should be noted that the information interaction, execution process, etc. between the above-mentioned modules, units, and sub-units are based on the same concept as the embodiment of the method of the present invention. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.

[0127] Figure 4 is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. As shown in Figure 4, the computer device of this embodiment includes: at least one processor (only one is shown in Figure 4), a memory, and a computer program stored in the memory and executable on the at least one processor. When the processor executes the computer program, it implements the steps of any of the above-mentioned human body recognition model training method embodiments.

[0128] The computer device may include, but is not limited to, a processor and a memory. Those skilled in the art will appreciate that FIG4 is merely an example of a computer device and does not limit the computer device. The computer device may include more or fewer components than shown in the figure, or may combine certain components or have different components.

[0129] The processor may be a CPU, other general-purpose processors, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0130] Memory includes readable storage media, internal memory, and the like. Internal memory can be the internal memory of a computer device, providing an environment for the operation of the operating system and computer-readable instructions stored in the readable storage medium. The readable storage medium can be the computer device's hard drive. In other embodiments, it can also be an external storage device, such as a plug-in hard drive, a Smart Media Card (SMC), a Secure Digital (SD) card, or a flash memory card. Furthermore, memory can include both the computer device's internal storage unit and external storage devices. Memory is used to store the operating system, application programs, boot loaders, data, and other programs, such as the program code of computer programs. Memory can also be used to temporarily store data that has been output or is about to be output.

[0131] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other and are not used to limit the scope of protection of the present invention. The specific working process of the units and modules in the above-mentioned device can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention can implement all or part of the process steps in the above-mentioned method embodiments by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When executed by a processor, the computer program can implement the steps of the above-mentioned method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file, or some intermediate form. Computer-readable media can include at least: any entity or device capable of carrying computer program code, recording media, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signals, telecommunications signals, and software distribution media. Examples include USB flash drives, removable hard drives, magnetic disks, or optical disks. In some jurisdictions, based on legislation and patent practice, computer-readable media cannot be electric carrier signals or telecommunications signals.

[0132] The present invention may implement all or part of the processes in the above-mentioned method embodiments, and may also be completed through a computer program product. When the computer program product runs on a computer device, the computer device can implement the steps in the above-mentioned method embodiments when executing the computer program product.

[0133] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0134] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.

[0135] In the embodiments provided by the present invention, it should be understood that the disclosed apparatus / computer equipment and methods can be implemented in other ways. For example, the apparatus / computer equipment embodiments described above are merely illustrative. For example, the division of modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0136] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0137] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.

Claims

1. A training method for a human body recognition model, characterized in that The training method includes: Obtaining a second input human body image and an image label of the second input human body image; Using the second input human body image and the image label of the second input human body image to train a human body recognition model, obtaining a trained human body recognition model, where the trained human body recognition model is used to extract features from a human body image, and the human body recognition model includes a preset adaptive encoder and an image encoder in a trained image-text model.

2. The training method according to claim 1, wherein Before using the second input human body image and the image label of the second input human body image to train the human body recognition model, it further includes: Performing pseudo-label annotation on a first input human body image through a pre-trained image-text model and a preset text prompt, obtaining a pseudo-label of the first input human body image; According to the first input human body image and the pseudo-label, performing supervised training on the pre-trained image-text model to update the parameters of the pre-trained image-text model.

3. The training method according to claim 2, characterized in that, The performing pseudo-label annotation on the first input human body image through a pre-trained image-text model and a preset text prompt, obtaining a pseudo-label of the first input human body image, includes: Using the image encoder in the pre-trained image-text model to perform feature encoding on the first input human body image, obtaining an image encoding feature; Using the text encoder in the pre-trained image-text model to perform feature encoding on the preset text prompt, obtaining a text encoding feature; According to the image encoding feature and the text encoding feature, matching a target text prompt for the first input human body image, and determining the target text prompt as the pseudo-label of the first input human body image.

4. The training method according to claim 3, wherein The matching a target text prompt for the first input human body image according to the image encoding feature and the text encoding feature includes: Calculating the similarity between the image encoding feature and each text encoding feature, obtaining a similarity value; According to the similarity value, matching a target text prompt for the first input human body image.

5. The training method according to claim 1, characterized in that, The using the second input human body image and the image label of the second input human body image to train the human body recognition model, obtaining a trained human body recognition model, includes: Using the human body recognition model to extract features from the second input human body image, obtaining an image feature; Performing human body identification classification on the image feature, obtaining a human body identification classification result; Calculating a classification loss according to the human body identification classification result and the image label; According to the classification loss, performing supervised training on the human body recognition model, adjusting the parameters of the preset adaptive encoder in the human body recognition model, obtaining a trained human body recognition model.

6. The training method according to claim 5, characterized in that The using the human body recognition model to extract features from the second input human body image, obtaining an image feature, includes: Using the image encoder in the human body recognition model to extract features from the second input human body image, obtaining a first image feature; Use the preset adaptive encoder in the human body recognition model to extract features from the second input human body image, obtaining second image features; Perform feature fusion on the first image features and the second image features to obtain image features.

7. The training method according to claim 5, wherein The training of the human body recognition model using the second input human body image and the image label of the second input human body image to obtain a trained human body recognition model further includes: Obtain positive sample images identical to the image label and negative sample images different from the image label; Use the human body recognition model to extract features from the positive sample images, obtaining positive sample image features, and use the human body recognition model to extract features from the negative sample images, obtaining negative sample image features; Calculate the triplet loss according to the image features, the positive sample image features, and the negative sample image features; Adjust the parameters of the preset adaptive encoder in the human body recognition model according to the classification loss and the triplet loss to obtain a trained human body recognition model.

8. A training device for a human body recognition model, characterized in that The training device includes: An acquisition module for acquiring a second input human body image and the image label of the second input human body image; A training module for training the human body recognition model using the second input human body image and the image label of the second input human body image to obtain a trained human body recognition model, where the trained human body recognition model is used to extract features from human body images, and the human body recognition model includes a preset adaptive encoder and an image encoder in a trained image-text model.

9. A computer device, characterized in that, The computer device includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the training method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the training method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Training method and device for displaying object recognition model and electronic equipment

    CN115100472A

  • Pedestrian re-identification method based on contrast language image pre-training model CLIP

    CN115393902A

  • Multi-modal unsupervised pedestrian re-identification method, device and equipment and storage medium

    CN116524543A

  • Human body recognition model training method and device, equipment and medium

    CN117975501A

  • Method and system for human re-identification, and non-transitory computer-readable storage medium

    US20230316798A1

Cited By

  • Retrieval method and device, equipment, medium and product

    CN121743488A