A method for constructing a dataset for pedestrian retrieval tasks based on text description

By constructing a text description-based data set in pedestrian retrieval task, and using large language models and diffusion models to generate diversified image data, the problem of insufficient data set dependence and diversity in the prior art is solved, and high-quality and efficient pedestrian retrieval model training is achieved.

CN119807466BActive Publication Date: 2025-05-16SUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510294112.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-05-16
Estimated Expiration
2045-03-13

AI Technical Summary

Technical Problem

The existing technology relies heavily on the original data set when building new data sets, resulting in low image diversity and poor resolution, which in turn leads to low retrieval accuracy of the trained pedestrian retrieval model.

Method used

The pedestrian retrieval task dataset construction method based on text description is adopted. The basic template is constructed by using the pedestrian character characteristics and scene characteristics as placeholders, and fill words are generated using a large language model, and the template is randomly filled to generate a prompt template. Then, the initial image data is generated using the diffusion model, and the image data is edited through local, global, and non-rigid editing models to generate a diverse image data set.

Benefits of technology

It does not rely entirely on original data, reduces privacy risks and avoids qualification problems. The generated images have high resolution and strong diversity, and can train pedestrian retrieval models more comprehensively and improves model recognition accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119807466B_ABST
    Figure CN119807466B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of data generation technology, and discloses a method for constructing a data set for a pedestrian retrieval task based on text description, including directly using the character features of pedestrians and the scene features of the scene where the pedestrians are located as placeholders to construct a basic template, and generating corresponding prompt words after filling the basic template; using a diffusion model to generate image data based on the prompt words, which is completely independent of the original data, greatly reducing privacy risks and avoiding eligibility issues. At the same time, the present invention uses a local editing model, a global editing model, and a non-rigid editing model to selectively edit the features of corresponding attributes in the image data directly based on the generated initial image data to obtain edited image data. The obtained edited image data has high resolution, and the image generation has good generalization and high degree of freedom, which greatly improves the diversity of the generated image data, can train the pedestrian retrieval model more comprehensively, and improve the model recognition accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data generation, and in particular to a method for constructing a pedestrian retrieval task dataset based on text description. Background Art

[0002] The text-based person retrieval (TBPR) task refers to the task of finding images of specific pedestrians that match a given text description from a large number of pedestrian images captured by a camera.

[0003] Currently, there are three mainstream datasets for this task, namely CUHK-PEDES, ICFG-PEDES, and RSTPReid; these datasets are composed of different numbers of image-text pairs. In order to improve the performance of the TBPR task model, the existing technology hopes to improve the model performance by building a new dataset and providing training data. When building a dataset, if real camera sampling is chosen, it will cause people to continue to worry about personal privacy, and manual annotation of the collected images requires high labor costs; therefore, the existing technology uses generative models to build datasets.

[0004] Most of the existing data generation methods are based on the original dataset. For example, based on the images and text annotations of the original dataset, modify specific words in the text annotations, such as changing "white shirt" to "blue shirt", and then use the fine-tuned diffusion model to generate new images based on the original images for the modified words. Based on the images and text annotations of the original dataset, for the specific attributes of the characters in the text annotations, such as "hair" and "skirt", use the image segmentation model to detect in the corresponding original image and segment out the mask of the corresponding attribute; then use the available clothing and accessories images to edit the segmented attributes on the original image; finally, a new image is obtained, and the corresponding new text annotation is changed using the large language model according to the clothing name of the "clothing and accessories image" used in the image editing; for example, when the words for the clothing used to replace the ReferenceAttribute Image part are gold hair, white skirt, and sandals, then the large language model will be used to match which parts of the original image annotation these three words replace, and replace them to obtain a new image. The text annotations of the original dataset are directly used to guide the diffusion model to generate new images, and then the cross-modal BLIP model is used to generate text annotations for the generated images, and the richness of the text annotations is improved through additional attribute annotations; that is, many types of attributes are proposed, and the BLIP model is allowed to match the generated images against these attributes to see if they contain these attributes. If so, they are added to the generated text annotations to improve the richness of the annotations.

[0005] In summary, the existing technology only makes limited modifications in the adjacent space of the original data to obtain new data. Since the data set is limited, the generated images lack sufficient diversity and are limited to certain fixed features. When the image resolution of the original data set is low, the existing technology directly edits the image of the original data set. The resolution of the image obtained after editing is naturally low, resulting in poor generated data. In addition, the existing technology is heavily dependent on the data of the original data set, which may lead to privacy risks and compliance issues. If there is a lack of a large-scale data set, the existing technology will not be able to construct a new data set, which will result in the inability to obtain a TBPR task model that meets the retrieval accuracy requirements after training the TBPR task model using the constructed data set when completing the pedestrian retrieval task. The model performance is poor, resulting in low completion of the pedestrian retrieval task. Summary of the invention

[0006] To this end, the technical problem to be solved by the present invention is to overcome the problem that the prior art is extremely dependent on the original dataset when constructing a new dataset, resulting in low image diversity and poor resolution of the new dataset, which in turn leads to low retrieval accuracy of the trained pedestrian retrieval model.

[0007] In order to solve the above technical problems, the present invention provides a method for constructing a pedestrian retrieval task dataset based on text description, comprising:

[0008] Using the pedestrian's character features and the scene features of the pedestrian's scene as placeholders, a basic template is constructed, and a specific filler word set is constructed for each placeholder;

[0009] Based on the specific filler word set corresponding to each placeholder in the basic template, randomly fill each placeholder in the basic template with a corresponding specific filler word from the specific filler word set corresponding to each placeholder to obtain a filled basic template;

[0010] Use the filled basic template as the input of the large language model and output the corresponding prompt template;

[0011] Taking the prompt template as the input of the diffusion model, obtaining multiple initial image data corresponding to each prompt template;

[0012] For each piece of initial image data, based on the preset editing text, using the local editing model, the global editing model and the non-rigid editing model, the attributes of each piece of initial image data are edited respectively to obtain multiple pieces of edited image data corresponding to each piece of initial image data;

[0013] All the initial image data and all the edited image data form an image data set;

[0014] Use a multimodal language model to construct corresponding text annotations for each image data in the image dataset;

[0015] Based on each image data and its corresponding text annotation, a text description-based pedestrian retrieval task dataset is formed.

[0016] Preferably, the character characteristics of the pedestrian include the pedestrian's gender, age, appearance characteristics and clothing characteristics; the scene characteristics of the scene where the pedestrian is located include the pedestrian's movements, geographical environment, building facilities and weather conditions.

[0017] Preferably, the step of taking the prompt template as the input of the diffusion model and obtaining the multiple image data corresponding to each prompt template includes:

[0018] For each prompt template, it is used as the input of the diffusion model, and multiple image data corresponding to each prompt template are obtained by adjusting the generation guidance coefficient and random seed of the diffusion model.

[0019] Preferably, after all the initial image data and all the edited image data are formed into an image data set, the method further comprises:

[0020] Perform human posture recognition on each image data in the image data set, and identify the nose, left eye, right eye, left ear, right ear, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip, right hip, left knee, right knee, left ankle and right ankle of the pedestrian in each image data as human key points;

[0021] The key points of the human body are divided into regions, the nose, left eye, right eye, left ear and right ear are divided into the head region, the left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip and right hip are divided into the body region, and the left knee, right knee, left ankle and right ankle are divided into the leg region;

[0022] For each image data in the image data set, if the same human key points exist in the image data, it is determined that there is not only a single pedestrian in the image data, and the image data is deleted;

[0023] For each image data in the image data set, if the head region, the body region and the leg region do not exist at the same time in the image data, it is determined that there is no complete pedestrian in the image data, and the image data is deleted;

[0024] The image dataset after the image is deleted is obtained as the updated image dataset.

[0025] Preferably, a specific filler word set is constructed for each placeholder based on a large language model, including:

[0026] Inputting the character features of the pedestrian into the large language model, and issuing instructions to the large language model to make the large language model output multiple specific filler words, thereby generating a specific filler word set corresponding to the character features of the pedestrian;

[0027] The scene features of the scene where the pedestrian is located are input into the large language model, and an instruction is issued to the large language model to make the large language model output multiple specific filler words, thereby generating a specific filler word set corresponding to the scene features of the scene where the pedestrian is located.

[0028] Preferably, the filled basic template is used as the input of the Qwen2 model, and the corresponding prompt template is output, including:

[0029] Input the filled basic template into the tokenizer of the Qwen2 model and convert it into an input index;

[0030] Input the input index into the backbone network of the Qwen2 model and convert it into a text vector;

[0031] The text vector is input into the decoder layer of the Qwen2 model. After the text vector is normalized by the mean square layer and the attention mechanism, it is residually connected with the input text vector to obtain the first hidden state. The first hidden state is normalized by the mean square layer and sent to the multilayer perceptron. The output of the multilayer perceptron is residually connected with the first hidden state to output the decoding feature.

[0032] The decoded features are input into the output layer, linearly transformed, and the corresponding prompt template is obtained.

[0033] Preferably, the initial image data is locally edited using the local editing model, including:

[0034] Based on the attention mechanism, obtain the cross-attention map of the initial image data;

[0035] Input the initial image data into the local editing model for denoising and compare the current denoising times With preset constants :

[0036] like , then obtain the cross-attention map of the preset edited text as the target cross-attention map;

[0037] like , then obtain the cross-attention map corresponding to the initial image data and the cross-attention map of the preset edited text as the target cross-attention map;

[0038] Until the preset denoising times are reached, the initial image data is denoised using the target cross-attention map to obtain the local edited image data corresponding to the initial image data.

[0039] Preferably, the global editing model is used to globally edit the initial image data, including:

[0040] Based on the attention mechanism, obtain the cross-attention map of the initial image data;

[0041] The initial image data is fed into the global editing model for denoising;

[0042] Compare the words at the same position of the preset edit text and the text corresponding to the initial image data to see if they are the same:

[0043] If they are the same, the value at the corresponding position in the cross-attention map will not be changed;

[0044] If they are different, change the value at the corresponding position in the cross-attention map;

[0045] After all word units are compared, the current cross-attention map is obtained as the target cross-attention map, the initial image data is denoised, and the global edited image data corresponding to the initial image data is obtained.

[0046] Preferably, the non-rigid editing model is used to perform non-rigid editing on the initial image data, including:

[0047] Input the initial image data into the non-rigid editing model for denoising and compare the current denoising times With preset constants :

[0048] like , then calculate the attention ;

[0049] like , then calculate the attention ;

[0050] Until the preset denoising times are reached, the current attention is obtained, the initial image data is edited, and the non-rigid edited image data corresponding to the initial image data is obtained;

[0051] in, ; Indicates Second denoising, Indicates a preset constant; represents the mapping vector representation of the original image data, The mapping vector representation of the text corresponding to the initial image data, express The length of the vector; The mapping vector representation of the locally edited image data, Represents the mapped vector representation of the preset edit text.

[0052] Preferably, a multimodal language model is used to construct corresponding text annotations for each image data in the image data set, including: constructing text annotations based on fixed instructions and constructing text annotations based on variable instructions;

[0053] The method of constructing text annotations based on fixed instructions is as follows: using a multimodal language model to read image data, and filling corresponding attributes for a fixed output template based on preset annotation instructions to generate text annotations;

[0054] The method of constructing text annotations based on variable instructions is: using a multimodal language model to read image data, and directly generating text annotations based on a fixed output template.

[0055] The above technical solution of the present invention has the following beneficial effects compared with the prior art:

[0056] The method for constructing a data set for pedestrian retrieval tasks based on text descriptions described in the present invention does not need to be derived from the original data, but directly uses the character features of pedestrians and the scene features of the scenes where pedestrians are located as placeholders to construct a basic template, and after filling the basic template, generates corresponding prompt words; uses a diffusion model to generate image data based on prompt words, which is completely independent of the original data, greatly reduces privacy risks and avoids eligibility issues; and the diffusion model can generate high-quality images more stably through gradual denoising, and improves the resolution of the generated images. At the same time, the present invention uses a local editing model, a global editing model and a non-rigid editing model, directly based on the generated initial image data, selectively edits the features of the corresponding attributes in the image data, obtains the edited image data, and makes the image generation have good generalization and high degree of freedom, greatly improves the diversity of the generated image data, can train the pedestrian retrieval model more comprehensively, and improves the model recognition accuracy.

[0057] The present invention utilizes human posture recognition to identify key points of the human body in each data image, and based on the identified key points, deletes images with poor recognition effect and that do not meet the task requirements, screens out low-quality images, improves the quality of images in the image data set, and further improves the data quality in the pedestrian retrieval task data set. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] In order to make the content of the present invention more clearly understood, the present invention is further described in detail below according to specific embodiments of the present invention in conjunction with the accompanying drawings, wherein:

[0059] Figure 1 It is a flowchart of the steps of the method for constructing a pedestrian retrieval task dataset provided by the present invention;

[0060] Figure 2 Image editing diagram

[0061] Figure 3 It is a schematic diagram of fine-tuning the diffusion model using Dreambooth;

[0062] Figure 4 is a schematic diagram comparing the initial image data and the local edited image data;

[0063] Figure 5 (a) is a schematic diagram of the missing key points in the head area. Figure 5 (b) is a schematic diagram of the missing key points in the head and leg regions. Figure 5 (c) is a schematic diagram of the missing key points in the head area. Figure 5 (d) is a schematic diagram of the missing key points in the leg area;

[0064] Figure 6(a) is the first schematic diagram with multiple complete pedestrians. Figure 6 (b) is the second schematic diagram with multiple complete pedestrians;

[0065] Figure 7 It is a cutting diagram. DETAILED DESCRIPTION

[0066] The present invention is further described below in conjunction with the accompanying drawings and specific embodiments so that those skilled in the art can better understand the present invention and implement it, but the embodiments are not intended to limit the present invention.

[0067] Reference Figure 1 As shown, the flowchart of the method for constructing a pedestrian retrieval task dataset of the present invention includes the following specific steps:

[0068] S101: Using the character features of the pedestrian and the scene features of the scene where the pedestrian is located as placeholders, a basic template is constructed, and a specific filler word set is constructed for each placeholder;

[0069] Based on the semantic preset construction, according to the attributes of the placeholder, multiple specific filler words that meet the attributes are artificially constructed to generate a specific filler word set;

[0070] The large language model-based construction is to input the attributes of the placeholder into the large language model, and issue instructions to the large language model, so that the large language model outputs multiple specific filler words, and generates a specific filler word set, including: inputting the character features of the pedestrian into the large language model, and issuing instructions to the large language model, so that the large language model outputs multiple specific filler words, and generates a specific filler word set corresponding to the character features of the pedestrian; inputting the scene features of the scene where the pedestrian is located into the large language model, and issuing instructions to the large language model, so that the large language model outputs multiple specific filler words, and generates a specific filler word set corresponding to the scene features of the scene where the pedestrian is located.

[0071] S102: Based on the basic template and the specific filling word set corresponding to each placeholder, randomly fill each placeholder in the basic template with a corresponding specific filling word from the specific filling word set corresponding to each placeholder to obtain a filled basic template;

[0072] S103: using the filled basic template as the input of the large language model, and outputting the corresponding prompt template;

[0073] S104: using the prompt template as the input of the diffusion model, obtaining a plurality of initial image data corresponding to each prompt template;

[0074] S105: for each piece of initial image data, based on the preset editing text, using the local editing model, the global editing model and the non-rigid editing model, respectively perform attribute editing on each piece of initial image data, and obtain multiple pieces of edited image data corresponding to each piece of initial image data;

[0075] S106: All the initial image data and all the edited image data are combined into an image data set;

[0076] S107: construct corresponding text annotations for each image data in the image dataset using a multimodal language model;

[0077] S108: Based on each image data and its corresponding text annotation, a text description-based pedestrian retrieval task dataset is formed.

[0078] In step S101, the character features of the pedestrian include the pedestrian's gender, age, appearance features and clothing features; the scene features of the pedestrian include the pedestrian's movements, geographical environment, building facilities and weather conditions. Among them, the appearance features include hairstyle, body shape, skin color and facial contour; clothing features include whether to wear accessories, top type and color, bottom type and color.

[0079] Specifically, in step S103, the filled basic template is used as the input of the large language model, and the corresponding prompt template is output; the large language model selected in this embodiment is the Qwen2 model, and the filled basic template is used as the input of the Qwen2 model, and the corresponding prompt template is output, including:

[0080] S103-1: Input the filled basic template into the word segmenter of the Qwen2 model and convert it into an input index;

[0081] S103-2: Input the input index into the backbone network of the Qwen2 model and convert it into a text vector;

[0082] S103-3: Input the text vector into the decoder layer of the Qwen2 model, subject the text vector to mean square layer normalization and attention mechanism in turn, and perform residual connection with the input text vector to obtain a first hidden state; subject the first hidden state to mean square layer normalization and send it to a multilayer perceptron; perform residual connection between the output of the multilayer perceptron and the first hidden state, and output the decoding feature;

[0083] S103-4: Input the decoded features into the output layer, perform linear transformation, and obtain the corresponding prompt template.

[0084] Specifically, in step S104, for each prompt template, it is used as the input of the diffusion model, and the multiple image data corresponding to each prompt template is obtained by adjusting the generation guidance coefficient and the random seed of the diffusion model. Before the prompt word is used as the input of the diffusion model to obtain the multiple image data corresponding to each prompt word, this embodiment also includes: using Dreambooth technology to fine-tune the diffusion model.

[0085] Specifically, in step S106, after all the initial image data and all the edited image data are combined into an image data set, the following further includes:

[0086] Perform human posture recognition on each image data in the image data set, and identify the nose, left eye, right eye, left ear, right ear, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip, right hip, left knee, right knee, left ankle and right ankle of the pedestrian in each image data as human key points;

[0087] The key points of the human body are divided into regions, the nose, left eye, right eye, left ear and right ear are divided into the head region, the left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip and right hip are divided into the body region, and the left knee, right knee, left ankle and right ankle are divided into the leg region;

[0088] For each image data in the image data set, if the same human key points exist in the image data, it is determined that there is not only a single pedestrian in the image data, and the image data is deleted;

[0089] For each image data in the image data set, if the head region, the body region and the leg region do not exist at the same time in the image data, it is determined that there is no complete pedestrian in the image data, and the image data is deleted;

[0090] The image dataset after the image is deleted is obtained as the updated image dataset.

[0091] The present invention utilizes human posture recognition to identify key points of the human body in each data image, and based on the identified key points, deletes images with poor recognition effect and that do not meet the task requirements, screens out low-quality images, improves the quality of images in the image data set, and further improves the data quality in the pedestrian retrieval task data set.

[0092] Specifically, this embodiment provides three image editing models for constructing edited image data based on the initial image data, including a local editing model, a global editing model, and a non-rigid editing model; this embodiment can arbitrarily combine these three editing models, so as to obtain multiple edited image data for one initial image data; for example, any one of the models can be selected for editing, any two models can be selected for editing, or all three models can be combined to obtain edited image data. Figure 2 The figure shows a schematic diagram of image editing.

[0093] (1) Using the local editing model, the initial image data is locally edited, including:

[0094] Based on the attention mechanism, obtain the cross-attention map of the initial image data;

[0095] Input the initial image data into the local editing model for denoising and compare the current denoising times With preset constants :

[0096] like , then obtain the cross-attention map of the preset edited text as the target cross-attention map;

[0097] like , then obtain the cross-attention map corresponding to the initial image data and the cross-attention map of the preset edited text as the target cross-attention map;

[0098] Until the preset denoising times are reached, the initial image data is denoised using the target cross-attention map to obtain the local edited image data corresponding to the initial image data.

[0099] It can be expressed as: ;

[0100] ;

[0101] ; ;

[0102] in, The target cross-attention map representing the local edit, represents the cross attention map corresponding to the initial image data, Indicates Second denoising, Indicates a preset constant; represents the mapping vector representation of the original image data, The mapping vector representation of the text corresponding to the initial image data, express The length of the vector; The mapping vector representation of the locally edited image data, Represents the mapped vector representation of the preset edit text.

[0103] (2) Using the global editing model, the initial image data is globally edited, including:

[0104] Based on the attention mechanism, obtain the cross-attention map of the initial image data;

[0105] The initial image data is fed into the global editing model for denoising;

[0106] Compare the words at the same position of the preset edit text and the text corresponding to the initial image data to see if they are the same:

[0107] If they are the same, the value at the corresponding position in the cross-attention map will not be changed;

[0108] If they are different, change the value at the corresponding position in the cross-attention map;

[0109] After all word units are compared, the current cross-attention map is obtained as the target cross-attention map, the initial image data is denoised, and the global edited image data corresponding to the initial image data is obtained.

[0110] It can be expressed as: ;

[0111] in, Represents the value of the jth text word at the i-th image pixel in the matrix corresponding to the cross-attention map corresponding to the target image data of global editing, Represents the value of the jth text word at the i-th image pixel in the matrix corresponding to the cross-attention map corresponding to the initial image data; represents the jth word in the preset edit text, Indicates that the preset edit text has more words than the text corresponding to the initial image data.

[0112] Using the non-rigid editing model, the initial image data is non-rigidly edited, including:

[0113] Input the initial image data into the non-rigid editing model for denoising and compare the current denoising times With preset constants :

[0114] like , then calculate the attention ;

[0115] like , then calculate the attention ;

[0116] Until the preset denoising times are reached, the current attention is obtained, the initial image data is edited, and the non-rigid edited image data corresponding to the initial image data is obtained;

[0117] in, represents the mapping vector representation of the original image data, The mapping vector representation of the text corresponding to the initial image data, The mapping vector representation of the locally edited image data, Represents the mapped vector representation of the preset edit text.

[0118] It can be expressed as: ;

[0119] in, represents the attention between the initial noise image after non-rigid editing and the text data of the target image data, Represents the attention corresponding to the text data of the initial noise image and the initial image data after non-rigid editing; Indicates Second denoising, Indicates a preset constant.

[0120] Specifically, after all the initial image data and all the edited image data are combined to form an image data set, a multimodal language model is used to construct corresponding text annotations for each image data in the image data set, including: constructing text annotations based on fixed instructions and constructing text annotations based on variable instructions; constructing text annotations based on fixed instructions is: using a multimodal language model to read image data, and filling corresponding attributes for a fixed output template based on preset annotation instructions to generate text annotations; constructing text annotations based on variable instructions is: using a multimodal language model to read image data, and directly generating text annotations based on a fixed output template.

[0121] The method for constructing a data set for pedestrian retrieval tasks based on text descriptions described in the present invention does not need to be derived from the original data, but directly uses the character features of pedestrians and the scene features of the scenes where pedestrians are located as placeholders to construct a basic template, and after filling the basic template, generates corresponding prompt words; uses a diffusion model to generate image data based on prompt words, which is completely independent of the original data, greatly reduces privacy risks and avoids eligibility issues; and the diffusion model can generate high-quality images more stably through gradual denoising, and improves the resolution of the generated images. At the same time, the present invention uses a local editing model, a global editing model and a non-rigid editing model, directly based on the generated initial image data, selectively edits the features of the corresponding attributes in the image data, obtains the edited image data, and makes the image generation have good generalization and high degree of freedom, greatly improves the diversity of the generated image data, can train the pedestrian retrieval model more comprehensively, and improves the model recognition accuracy.

[0122] Based on the above embodiment, in an embodiment of the present invention, the method for constructing a pedestrian retrieval task dataset based on text description provided by the present invention is used to configure a specific model and construct a dataset, which specifically includes:

[0123] S201: construct prompt words;

[0124] ①. Build the basic template:

[0125] The basic template of prompt words is constructed from multiple perspectives, such as character characteristics (such as gender, age, appearance, clothing, etc.), scene or background information (such as the character's environment, occupation, action, etc.), and divided into different categories. Among them, placeholders need to be provided in the basic template.

[0126] The basic template constructed by example is as follows:

[0127] and A {gender} person, in the {location}.

[0128] Among them, Age, gender, hair, and upper_clothes_adjective are placeholders. For the filler words, the direct way is to construct them manually. For example, the distinction of hair can be short hair or long hair, and the filler words are directly written according to semantic habits. The indirect way is to construct with the help of a large language model. Still taking the hair placeholder as an example, provide the instruction Please give possible values ​​for hair in the format of'{age} = [young, middle-aged, old]' to the large language model. The output of the large language model is similar to {hair}=[shorthair,long hair].

[0129] ②、Build placeholder filler words:

[0130] According to the placeholder attributes of the basic template, construct the specific content of each placeholder attribute;

[0131] The filling process is actually to use Python programming language directly, based on the basic template, randomly select from the filler words corresponding to the corresponding placeholders. For example, if the basic template is A {age} {gender} person, with{hair},{upper_clothes_adjective} {upper_clothes} with {sleeve}. The filler words corresponding to the placeholders are:

[0132] {age}=[young, middle-aged, old];

[0133] {gender}=[female, male];

[0134] {hair}=[short hair, long hair];

[0135] {upper_clothes_adjective}=[stripe pattern, cartoon pattern, geometricpatterns, letters and logo patterns, photography patterns];

[0136] {upper}=[Denim jacket, Knit sweater, T-shirt, Hoodie, Blouse, Tanktop, Polo shirt, Cardigan, Button-down shirt, Tunic, Vest];

[0137] {sleeve}=[long sleeve, short sleeve]…

[0138] Write a program to set the placeholders Age, gender, hair, upper_clothes_adjective, upper, and sleeve to be able to choose any of the above fill-in words to fill in the basic template.

[0139] ③、Expand the specific content of the basic template:

[0140] With the help of the extension directive, the base template after filling the placeholders is expanded using the large language model.

[0141] For example, using the Qwen2 large language model, provide it with extended instructions:

[0142] In the style of “The [{man / woman}] is [].” This is input: “{base_template}”. You can remove some clothing or add some accessories. Expand thisdescription by including details on the person's actions, emotions, andsurroundings. Keep the final description under 77 words.

[0143] Among them, base_template is the basic template after filling in the placeholders.

[0144] For example, we can set base_template as the base template and fill in the placeholders to an example sentence: "A youngmale person, with short hair, cartoon pattern T-shirt with short sleeves, stripe pattern trousers, a pair of football shoes, sun-glasses, in the bird's-eye view.". The output of Qwen2 is "The man is lounging carefree. A youngmale, with short hair, dons a cartoon T-shirt, shorts, and football shoes. Sunglasses shield his eyes as he sprawls on a checkered picnic blanket. Surrounded by nature, he chuckles softly, feeding crumbs to nearby birds. His easy smile and relaxed posture reflect pure contentment. The day is warm, skies clear, embodying perfect leisure.".

[0145] In this way, each basic template and specific instruction are provided to the large language model, and the output of the large language model is obtained as the constructed prompt word.

[0146] The large language model of this embodiment uses the Qwen2 model. The entire model adopts a decoder-only architecture. The general process is as follows:

[0147] 1.Tokenizer processes text input;

[0148] ②.Embedding converts the index into a vector;

[0149] ③. After multiple Transformer layers (decoder layers), deep feature extraction is performed using self-attention, residual connection, and feedforward network (MLP);

[0150] ④. The output layer generates the final prediction.

[0151] For example, given the text of an instruction, the text will be processed by the Tokenizer as the text input as an index, and then the index will be converted into a vector in the Embedding layer. The vector represents the features of the text, and the specific features are extracted by multiple Transformer layers. Finally, the prediction result is obtained as the output.

[0152] S202: image generation;

[0153] ①. No reliance on additional images:

[0154] Use the diffusion model to generate according to the prompt words obtained above, and adjust the generation guidance coefficient and random seed by yourself to obtain any amount of image data;

[0155] ②. Relying on very small amounts of data:

[0156] Select real-world images, use Dreambooth to fine-tune the diffusion model, use the fine-tuned diffusion model to generate images according to the constructed prompt words, and obtain image data;

[0157] Specifically, real-world images can be selected from existing data sets, or pedestrian images can be collected by oneself. Only a small number of images are required, such as setting the selection to 8 images.

[0158] Reference Figure 3 The figure shows a schematic diagram of fine-tuning the diffusion model using Dreambooth; the text corresponding to the reconstruction loss is "A sks person", and the text corresponding to the class-specific prior retention loss is "A person". The original weights of the diffusion model are fine-tuned by setting different learning rates and iterations. [V] in the figure represents the selected real-world image.

[0159] S203: Editing the generated image data;

[0160] ①. Select the attributes to be edited: Select the appropriate editing attributes for different attributes of the character, such as posture, clothing, environment, image style, etc.;

[0161] ②、Select the method to edit specific attributes, which can be divided into three categories:

[0162] For local attributes such as clothing and environment, select the local editing model;

[0163] For global attributes such as style, select the global editing model;

[0164] For non-rigidly changing attributes of characters such as posture, select the non-rigid editing model;

[0165] ③. Use the selected editing model to edit multiple attributes of different images to obtain more image data.

[0166] In this embodiment, the optional multimodal language models include: QwenVL, MIniCPM and InternVL; the local editing model is FreeCustom; the global editing model is Prompt-to-Prompt; and the non-rigid editing model is MasaCtrl.

[0167] Specifically, this embodiment selects the environment where the characters are located in the edited image as the local feature, adopts the local editing technology, such as selecting the "NMG+Prompt to Prompt" model to perform local editing of the environment, and obtains the local editing image data. Figure 4 The figure shows a comparison diagram between the initial image data and the locally edited image data.

[0168] S204: screening low-quality images in the image data generated in step S202 and step S203;

[0169] ①、Filter:

[0170] Select a human posture recognition model, such as the Yolov8-pose model, to divide the human body into multiple key points, and then divide these key points into the head area, the body area, and the leg area;

[0171] The head area includes the nose, left eye, right eye, left ear and right ear; the body area includes the left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip and right hip; the leg area includes the left knee, right knee, left ankle and right ankle;

[0172] The screening logic used in this embodiment is: to ensure that a complete person appears, that is, the three areas must ensure that key points are detected in each area; to ensure that there is only one person, that is, multiple identical key points cannot appear in one picture;

[0173] Figure 5 Schematic diagram of multiple images that are filtered out due to incomplete pedestrian key points; Figure 5 (a) is a schematic diagram of the missing key points in the head area. Figure 5 (b) is a schematic diagram of the missing key points in the head and leg regions. Figure 5 (c) is a schematic diagram of the missing key points in the head area. Figure 5 (d) is a schematic diagram of the missing key points in the leg area. Figure 6 A schematic diagram of multiple images that are filtered out because they are not single complete pedestrians; Figure 6 (a) is the first schematic diagram with multiple complete pedestrians. Figure 6(b) is the second schematic diagram with multiple complete pedestrians.

[0174] Based on this screening logic, the images that do not meet the requirements are filtered out and image data with a single pedestrian is obtained. The complete single pedestrian image is more conducive to the training of the TBPR task model.

[0175] ②, cutting:

[0176] According to the human posture recognition model, a complete single pedestrian image area is identified and cropped out, that is, redundant background information is removed, so that the TBPR model can better learn pedestrian features.

[0177] Reference Figure 7 The figure shows a cutting diagram.

[0178] S205: text annotation generation;

[0179] Constructing instructions for guiding the visual-linguistic multimodal large language model, and guiding the multimodal large language model to perform text annotation on the image data remaining after the screening and cropping in step S204 based on the instructions;

[0180] Since the output style of a multimodal large language model is consistent for a single instruction, which will lead to a lack of diversity in the text annotations of the corresponding image, there are two types of instructions;

[0181] ①、Fixed instructions:

[0182] Design a fixed instruction that requires a multimodal large language model to read an image and generate detailed annotations for it, by requiring it to fully summarize multiple specific attributes of the person;

[0183] For example, only use the command "Don't mention the background environment of the people in the image. Please provide a detailed description of this person's age, gender, top(including color and style), bottom(including color and style), hair(including color and style), shoes(including color and style), and belongings (including color and style). Finally, combine all the details into a single sentence."

[0184] ②、Change type instruction:

[0185] Design a template for the desired output: Design a fixed output template based on the different attributes of the characters and the collocation of different English words. And hope that the final output result of the model will strictly follow this template to give an answer; "Wearing [clothing description], the [person / woman / man] also has [hair description] and is carrying [belongings description]".

[0186] Design instructions for the large language model: require it to give text annotations for pedestrian images according to the given template; for example, "Generate a description about the overall appearance of the person, including the clothing, shoes, hairstyle, generation and belongings, in a styles similar to the template: '{template}'. If some requirements in the template are not visible, you can ignore. Do not imagine any contents that are not in the image.". Among them, {template} is the output template constructed above.

[0187] Specifically, for the aforementioned Figure 7 The cropped image in the figure uses the QwenVL model to generate text annotations based on fixed instructions and variable instructions, including:

[0188] ① Using the QwenVL model, the output obtained using fixed-type instructions is as follows:

[0189] The person appears to be an adult male wearing a plain white short-sleeve t-shirt, bright blue shorts with a simple design, and a wide-brimmedstraw hat covering their hair. The individual is barefoot and wearing a blacksmartwatch on their left wrist.

[0190] ②Using the QwenVL model, the output obtained by using the variable instruction is as follows:

[0191] Wearing a plain white short-sleeve t-shirt and bright blue shorts, the man also has a wide-brimmed straw hat and is carrying a black smartwatch on his left wrist.

[0192] The present invention proposes a method for constructing prompt words for an image generation model with good generalization and high degree of freedom, which greatly improves the diversity of generated data; and does not rely on or relies on a small amount of original data, greatly reducing privacy risks and avoiding eligibility issues. At the same time, it proposes to use a human posture detection model to screen out low-quality images and improve the average quality of generated data; and it can control the generation of higher-resolution images, which has a better effect in the pre-training of the TBPR task model.

[0193] The method for constructing a data set for pedestrian retrieval tasks based on text descriptions described in the present invention does not need to be derived from the original data, but directly uses the character features of pedestrians and the scene features of the scenes where pedestrians are located as placeholders to construct a basic template, and after filling the basic template, generates corresponding prompt words; uses a diffusion model to generate image data based on prompt words, which is completely independent of the original data, greatly reduces privacy risks and avoids eligibility issues; and the diffusion model can generate high-quality images more stably through gradual denoising, and improves the resolution of the generated images. At the same time, the present invention uses a local editing model, a global editing model and a non-rigid editing model, directly based on the generated initial image data, selectively edits the features of the corresponding attributes in the image data, obtains the edited image data, and makes the image generation have good generalization and high degree of freedom, greatly improves the diversity of the generated image data, can train the pedestrian retrieval model more comprehensively, and improves the model recognition accuracy. The present invention utilizes human posture recognition to identify key points of the human body in each data image, and based on the identified key points, deletes images with poor recognition effect and that do not meet the task requirements, screens out low-quality images, improves the quality of images in the image data set, and further improves the data quality in the pedestrian retrieval task data set.

[0194] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0195] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0196] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0197] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0198] Obviously, the above embodiments are merely examples for the purpose of clear explanation and are not intended to limit the implementation methods. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the implementation methods here. The obvious changes or modifications derived therefrom are still within the scope of protection of the present invention.

Claims

1. A method for constructing a pedestrian retrieval task dataset based on text description, characterized in that: include: Using the pedestrian's character features and the scene features of the pedestrian's scene as placeholders, a basic template is constructed, and a specific filler word set is constructed for each placeholder; Based on the specific filler word set corresponding to each placeholder in the basic template, randomly fill each placeholder in the basic template with a corresponding specific filler word from the specific filler word set corresponding to each placeholder to obtain a filled basic template; Use the filled basic template as the input of the large language model and output the corresponding prompt template; Taking the prompt template as the input of the diffusion model, obtaining multiple initial image data corresponding to each prompt template; For each piece of initial image data, based on the preset editing text, using the local editing model, the global editing model and the non-rigid editing model, the attributes of each piece of initial image data are edited respectively to obtain multiple pieces of edited image data corresponding to each piece of initial image data; All the initial image data and all the edited image data form an image data set; Use a multimodal language model to construct corresponding text annotations for each image data in the image dataset; Based on each image data and its corresponding text annotation, a text description-based pedestrian retrieval task dataset is formed.

2. The method for constructing a pedestrian retrieval task dataset based on text description according to claim 1 is characterized in that: The pedestrian's character characteristics include the pedestrian's gender, age, appearance characteristics and clothing characteristics; The scene features of the scene where the pedestrian is located include the pedestrian's movements, geographical environment, building facilities and weather conditions.

3. The method for constructing a pedestrian retrieval task dataset based on text description according to claim 1 is characterized in that: The method of taking the prompt template as the input of the diffusion model and obtaining multiple image data corresponding to each prompt template includes: For each prompt template, it is used as the input of the diffusion model, and multiple image data corresponding to each prompt template are obtained by adjusting the generation guidance coefficient and random seed of the diffusion model.

4. The method for constructing a pedestrian retrieval task dataset based on text description according to claim 1 is characterized in that: After all the initial image data and all the edited image data are combined into an image data set, it also includes: Perform human posture recognition on each image data in the image data set, and identify the nose, left eye, right eye, left ear, right ear, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip, right hip, left knee, right knee, left ankle and right ankle of the pedestrian in each image data as human key points; The key points of the human body are divided into regions, the nose, left eye, right eye, left ear and right ear are divided into the head region, the left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip and right hip are divided into the body region, and the left knee, right knee, left ankle and right ankle are divided into the leg region; For each image data in the image data set, if the same human key points exist in the image data, it is determined that there is not only a single pedestrian in the image data, and the image data is deleted; For each image data in the image data set, if the head region, the body region and the leg region do not exist at the same time in the image data, it is determined that there is no complete pedestrian in the image data, and the image data is deleted; The image dataset after the image is deleted is obtained as the updated image dataset.

5. The method for constructing a pedestrian retrieval task dataset based on text description according to claim 1 is characterized in that: A specific set of filler words is constructed for each placeholder based on a large language model, including: Inputting the character features of the pedestrian into the large language model, and issuing instructions to the large language model to make the large language model output multiple specific filler words, thereby generating a specific filler word set corresponding to the character features of the pedestrian; The scene features of the scene where the pedestrian is located are input into the large language model, and an instruction is issued to the large language model to make the large language model output multiple specific filler words, thereby generating a specific filler word set corresponding to the scene features of the scene where the pedestrian is located.

6. The method for constructing a pedestrian retrieval task dataset based on text description according to claim 1 is characterized in that: The filled basic template is used as the input of the Qwen2 model, and the corresponding prompt template is output, including: Input the filled basic template into the tokenizer of the Qwen2 model and convert it into an input index; Input the input index into the backbone network of the Qwen2 model and convert it into a text vector; The text vector is input into the decoder layer of the Qwen2 model. After the text vector is normalized by the mean square layer and the attention mechanism, it is residually connected with the input text vector to obtain the first hidden state. The first hidden state is normalized by the mean square layer and sent to the multilayer perceptron. The output of the multilayer perceptron is residually connected with the first hidden state to output the decoding feature. The decoded features are input into the output layer, linearly transformed, and the corresponding prompt template is obtained.

7. The method for constructing a pedestrian retrieval task dataset based on text description according to claim 1 is characterized in that: Using the local editing model, the initial image data is locally edited, including: Based on the attention mechanism, obtain the cross-attention map of the initial image data; Input the initial image data into the local editing model for denoising and compare the current denoising times With preset constants : like , then obtain the cross-attention map of the preset edited text as the target cross-attention map; like , then obtain the cross-attention map corresponding to the initial image data and the cross-attention map of the preset edited text as the target cross-attention map; Until the preset denoising times are reached, the initial image data is denoised using the target cross-attention map to obtain the local edited image data corresponding to the initial image data.

8. The method for constructing a pedestrian retrieval task dataset based on text description according to claim 1 is characterized in that: Using the global editing model, the initial image data is globally edited, including: Based on the attention mechanism, obtain the cross-attention map of the initial image data; The initial image data is fed into the global editing model for denoising; Compare the words at the same position of the preset edit text and the text corresponding to the initial image data to see if they are the same: If they are the same, the value at the corresponding position in the cross-attention map will not be changed; If they are different, change the value at the corresponding position in the cross-attention map; After all word units are compared, the current cross-attention map is obtained as the target cross-attention map, the initial image data is denoised, and the global edited image data corresponding to the initial image data is obtained.

9. The method for constructing a pedestrian retrieval task dataset based on text description according to claim 1 is characterized in that: Using the non-rigid editing model, the initial image data is non-rigidly edited, including: Input the initial image data into the non-rigid editing model for denoising and compare the current denoising times With preset constants : like , then calculate the attention ; like , then calculate the attention ; Until the preset denoising times are reached, the current attention is obtained, the initial image data is edited, and the non-rigid edited image data corresponding to the initial image data is obtained; in, ; Indicates Second denoising, Indicates a preset constant; represents the mapping vector representation of the original image data, The mapping vector representation of the text corresponding to the initial image data, express The length of the vector; The mapping vector representation of the locally edited image data, Represents the mapped vector representation of the preset edit text.

10. The method for constructing a pedestrian retrieval task dataset based on text description according to claim 1, characterized in that: Using a multimodal language model to construct corresponding text annotations for each image data in the image dataset, including: constructing text annotations based on fixed instructions and constructing text annotations based on variable instructions; The method of constructing text annotations based on fixed instructions is as follows: using a multimodal language model to read image data, and filling corresponding attributes for a fixed output template based on preset annotation instructions to generate text annotations; The method of constructing text annotations based on variable instructions is: using a multimodal language model to read image data, and directly generating text annotations based on a fixed output template.

Citation Information

Patent Citations

  • Scene text recognition method based on robustness representation learning

    CN113343707A

  • Security data generation method and device, storage medium and electronic device

    CN119416761A