Character re-identification method and device, terminal equipment and storage medium
Through multi-dimensional feature extraction and fusion, including the fusion of character attribute features, image features and semantic correlation features, the problem of insufficient recognition of characters in traditional technology is solved, and higher accuracy and more comprehensive feature description are achieved.
Patent Information
- Application Number
- CN202510177287.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-06
AI Technical Summary
Traditional deep learning technology is difficult to describe the target character from different angles in character recognizing, and cannot effectively connect and analyze the character attributes with the image area, resulting in insufficient feature extraction and easy misjudgment, which reduces the accuracy of character recognizing.
Through the extraction and fusion of multi-dimensional features, including the fusion of character attribute features, image features and semantic correlation features, a more comprehensive target character features are generated. This method enhances the ability to capture key information in the image and understand the content of the image area through region division and semantic correlation features.
It improves the accuracy of character re-recognition, reduces misidentification caused by insufficient single features, and can describe the characters more comprehensively, and describe the characteristics of the characters to be identified from multiple angles.
Smart Images

Figure CN120107889A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of person re-identification, and in particular to a person re-identification method, device, terminal equipment and storage medium. Background Art
[0002] Person Re-identification (Re ID), also known as pedestrian re-identification, refers to identifying the same person or pedestrian at different times, locations or camera angles. With the development of deep neural networks, pedestrian features can be automatically extracted and cross-camera retrieval or database comparison can be performed to achieve recognition and identity confirmation of the same person. However, traditional deep learning technology often relies on only one feature for feature extraction when extracting features from people in images. For example, after only using the attribute features of the person to derive the features of the person, subsequent feature comparison and person identity confirmation are performed. The traditional feature extraction method cannot describe the target person from different angles, and cannot connect and analyze the attributes of the person with the various regions of the image. It not only cannot accurately reflect the characteristics of the target person, but also cannot understand the content of the image region, resulting in the inability to refine the description of the person when extracting features, which is prone to misjudgment, making the accuracy of person re-identification low. Summary of the invention
[0003] The embodiments of the present invention provide a person re-identification method, apparatus, terminal device and storage medium. By extracting and fusing multi-dimensional features, a person can be described more comprehensively, reducing misidentification caused by insufficient single features. Moreover, by generating semantically associated features, the ability to capture key information in an image and the ability to understand the content of an image region can be enhanced, thereby improving the accuracy of person re-identification.
[0004] An embodiment of the present invention provides a method for person re-identification, comprising:
[0005] Obtain a target image containing a person to be identified;
[0006] Input the target image into a preset image recognition model so that the image recognition model extracts the character attribute features corresponding to the person to be recognized in the target image and the image features corresponding to the target image, divides the target image into regions according to the character attribute features and extracts the semantic association features between the character attribute features and each region, and then fuses the semantic association features, the character attribute features and the image features to generate the target character features corresponding to the person to be recognized;
[0007] Calculate the similarity between the target person feature and each preset person feature in the database; wherein different preset person features correspond to different preset images; the preset image includes a preset person and identity information corresponding to the preset person;
[0008] The identity information corresponding to the preset image with the highest similarity is marked as the identity information corresponding to the person to be identified in the target image.
[0009] Preferably, the character attribute features include: appearance features, posture features, clothing features, and background features used to represent the environment background of the character; each area corresponds to an appearance feature, a posture feature, a clothing feature, or a background feature;
[0010] The extracting of the semantic association features between the character attribute features and each region includes:
[0011] For each region, according to the character attribute characteristics corresponding to the region, the visual features corresponding to the region are extracted; wherein the visual features include: hair color, action, clothing style or scene;
[0012] Each region and the visual features corresponding to each region are used as nodes, and then each node is connected to generate an association graph for representing the semantic relationship between the character attribute features and the regions; wherein the connection edges in the association graph represent the semantic relationship between the character attribute features and the regions;
[0013] For each node, update the node features corresponding to each node according to the features of the neighboring nodes connected to the node and the semantic relationship of the connecting edges connected to the node;
[0014] Generate, based on each updated node feature, a spatial correlation feature for representing the location of a person's attribute feature in a region, an attribute correlation feature for representing the similarity of the person's attribute features between different regions, a scene correlation feature for representing the usage scene of the person's attribute features in a region, and a usage correlation feature between different person's attribute features;
[0015] The spatial association features, attribute association features, scene association features and usage association features are output as semantic association features.
[0016] Preferably, the regions in the target image include: a background region, a head region, a limb region, and a clothing region;
[0017] The step of dividing the target image into regions according to the character attribute features includes:
[0018] According to the background features, the target image is divided into a background area and a human body area;
[0019] Extracting a head region from the human body region according to the appearance features;
[0020] Extracting a limb region from the human body region according to the posture feature;
[0021] A clothing region is extracted from the human body region according to the clothing features.
[0022] Preferably, the spatial association feature includes: a position feature of hair color in the face area; the attribute association feature includes: a similarity feature between the hair color in the head area and the clothing style in the clothing area; the scene association feature includes: a usage scene feature of the clothing style in the background area; the use association feature includes: a use feature predicted based on the clothing style and the scene;
[0023] The step of fusing the semantic association feature, the character attribute feature and the image feature to generate a target character feature corresponding to the character to be identified includes:
[0024] Generate a number of different text prompts based on the attribute characteristics of each character and a preset prompt template; wherein the text prompt is used to describe the character's expression and the character's posture, or to describe the character's clothing, the character's expression, the character's posture and the scene in which the character is located;
[0025] Splicing each text prompt with the image feature to generate a spliced image feature;
[0026] The position features of hair color in the facial area, the similarity features between hair color and clothing style between the head area and clothing area, the usage scenario features of clothing style in the background area, the usage features predicted based on clothing style and scene, and the spliced image features are fused to generate the target person features corresponding to the person to be identified.
[0027] Preferably, the appearance features include: smiling facial features; the posture features include: crossed hands action features; the clothing features include: top features for describing a top as a green jacket and bottom features for describing trousers as black sports trousers; the background features include: a sports field; the text prompts are used to describe the character's clothing, the character's expression, the character's posture and the scene where the character is located;
[0028] According to the attribute characteristics of each character and the preset prompt template, a number of different text prompts are generated, including:
[0029] A text prompt is generated according to the facial expression features of a smile, the action features of crossing hands, the features of an upper garment, the features of a lower garment, a sports field and a preset prompt template; wherein the text prompt includes: a person wearing a green jacket and black pants is standing on a sports field, smiling, with his hands crossed in front of his chest.
[0030] Preferably, the step of splicing the text prompts with the image features to generate spliced image features includes:
[0031] Based on the similarity function, the similarity score between each text prompt and the image feature is calculated respectively; wherein the similarity score is used to characterize the degree of association between the text prompt vector and the image feature;
[0032] Based on the similarity scores, a weight vector is generated between each text hint and image feature;
[0033] The spliced image features are generated according to each weight vector, each text prompt and the image features.
[0034] Preferably, the training process of the image recognition model includes:
[0035] Taking a number of image samples and the actual character features corresponding to each image sample as input, and taking the predicted character features of each image sample as output, the image recognition model to be trained is iteratively trained until the model converges, thereby generating a preset image recognition model;
[0036] In each iterative training, an image sample is input into the image recognition model so that the image recognition model generates a prediction result of a person feature corresponding to the image sample;
[0037] The predicted character feature results are compared with the actual character feature results, and the network parameters of the image recognition model are adjusted according to the comparison results.
[0038] Based on the above method embodiments, the present invention provides corresponding device embodiments.
[0039] An embodiment of the present invention provides a person re-identification device, comprising: an image acquisition module, a person feature generation module, a similarity calculation module, and an identity information determination module;
[0040] The image acquisition module is used to acquire a target image containing a person to be identified;
[0041] The character feature generation module is used to input the target image into a preset image recognition model so that the image recognition model extracts character attribute features corresponding to the character to be recognized in the target image and image features corresponding to the target image, divides the target image into regions according to the character attribute features and extracts semantic association features between the character attribute features and each region, and then fuses the semantic association features, the character attribute features and the image features to generate target character features corresponding to the character to be recognized;
[0042] The similarity calculation module is used to calculate the similarity between the target person feature and each preset person feature in the database; wherein different preset person features correspond to different preset images; and the preset image includes a preset person and identity information corresponding to the preset person;
[0043] The identity information determination module is used to mark the identity information corresponding to the preset image with the highest similarity as the identity information corresponding to the person to be identified in the target image.
[0044] Based on the above method embodiments, the present invention provides corresponding terminal device embodiments.
[0045] Another embodiment of the present invention provides a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, and when the processor executes the computer program, a person re-identification method described in the above-mentioned embodiment of the invention is implemented.
[0046] Based on the above method embodiments, the present invention provides corresponding storage medium item embodiments.
[0047] Another embodiment of the present invention provides a storage medium, wherein the computer-readable storage medium includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute a person re-identification method described in the above-mentioned embodiment of the invention.
[0048] The following beneficial effects are achieved by implementing the present invention:
[0049] The embodiment of the present invention provides a person re-identification method, device, terminal device and storage medium. After acquiring the target image, the present invention not only extracts the person attribute features of the person to be identified in the target image through the image recognition model, but also extracts the image features of the target image itself, thereby increasing the dimension of feature extraction and making the feature description richer and more comprehensive. Further, the present invention can also divide the target image into regions according to the extracted person attribute features, and calculate the semantic association features between each region and the person attribute, so as to capture the semantic association features between the person attribute and the image region through the obtained semantic association features, and realize the connection and analysis of the person attribute and each region of the image. Finally, the semantic association features, the person attribute features and the image features are fused to achieve feature complementarity, and the features of the person to be identified can be described from multiple angles through multi-dimensional feature fusion, thereby improving the accuracy of person recognition. Compared with the prior art, the present invention can more comprehensively describe the person through the extraction and fusion of multi-dimensional features, reduce the misidentification caused by the insufficiency of a single feature, and enhance the ability to capture key information in the image and the ability to understand the content of the image region through the generation of semantic association features, thereby improving the accuracy of person re-identification. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 It is a flowchart of a person re-identification method provided by an embodiment of the present invention.
[0051] Figure 2 It is a structural schematic diagram of a person re-identification device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0052] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0053] like Figure 1 FIG. 1 is a flow chart of a method for person re-identification provided by an embodiment of the present invention, wherein the method for person re-identification comprises:
[0054] Step S1: Acquire a target image containing a person to be identified;
[0055] Step S2: inputting the target image into a preset image recognition model, so that the image recognition model extracts the character attribute features corresponding to the person to be recognized in the target image and the image features corresponding to the target image, divides the target image into regions according to the character attribute features and extracts the semantic association features between the character attribute features and each region, and then fuses the semantic association features, the character attribute features and the image features to generate the target character features corresponding to the person to be recognized;
[0056] Step S3: Calculate the similarity between the target person feature and each preset person feature in the database; wherein different preset person features correspond to different preset images; and the preset image includes a preset person and identity information corresponding to the preset person;
[0057] Step S4: marking the identity information corresponding to the preset image with the highest similarity as the identity information corresponding to the person to be identified in the target image.
[0058] For step S1, in a preferred embodiment, the target image acquired by the present invention can be an image captured by a surveillance camera in a security monitoring scene, or an image captured by a network camera in a remote video conference or online education scene, so as to confirm and verify the identity information of the person in the target.
[0059] For step S2, in a preferred embodiment, the training process of the image recognition model includes:
[0060] Taking a number of image samples and the actual character features corresponding to each image sample as input, and taking the predicted character features of each image sample as output, the image recognition model to be trained is iteratively trained until the model converges, thereby generating a preset image recognition model;
[0061] In each iterative training, an image sample is input into the image recognition model so that the image recognition model generates a prediction result of a person feature corresponding to the image sample;
[0062] The predicted character feature results are compared with the actual character feature results, and the network parameters of the image recognition model are adjusted according to the comparison results.
[0063] Illustratively, the present invention enables the image recognition model to gradually learn the mapping relationship between the character features in the image samples and the actual character features through iterative training, so that when processing new images subsequently, the character features can be more accurately extracted, thereby improving the accuracy of recognition.
[0064] Illustratively, based on a trained image recognition model, the present invention can extract the person attribute features of the person in the target image based on the image recognition model, thereby generating multiple text prompts based on multiple different person attribute features, and performing refined area division of the image to more accurately perform person re-identification, which not only improves the performance of person re-identification, but also enhances the model's adaptability in complex scenes and changing conditions.
[0065] In a preferred embodiment, the character attribute features include appearance features, posture features, clothing features, and background features used to represent the environmental background of the character.
[0066] Then, multiple text prompts can be generated based on multiple different character attribute features, thereby providing diverse image information for subsequent character recognition, such as:
[0067] Generate a number of different text prompts based on the attribute characteristics of each character and a preset prompt template; wherein the text prompt is used to describe the character's expression and the character's posture, or to describe the character's clothing, the character's expression, the character's posture and the scene in which the character is located;
[0068] Each text prompt is spliced with the image feature to generate a spliced image feature.
[0069] If the appearance features include: smiling facial features; the posture features include: crossed hands action features; the clothing features include: top features used to describe the top as a green jacket and bottom features used to describe the pants as black sweatpants; the background features include: sports field; the text prompts are used to describe the character's clothing, the character's expression, the character's posture, and the scene where the character is located, then:
[0070] According to the attributes of each character and the preset prompt template, several different text prompts are generated, including:
[0071] A text prompt is generated according to the facial expression features of a smile, the action features of crossing hands, the features of an upper garment, the features of a lower garment, a sports field and a preset prompt template; wherein the text prompt includes: a person wearing a green jacket and black pants is standing on a sports field, smiling, with his hands crossed in front of his chest.
[0072] It is understandable that the present invention can extract the character attribute features in the image, including appearance features (such as smiling expression features), posture features (such as the action features of crossing hands), clothing features (such as the features of a green jacket and black sweatpants) and background features (such as a sports field). Further, the extracted character attribute features are filled into the corresponding positions in the template one by one according to the structure and grammatical requirements of the preset prompt template. For example, for a template describing the character's clothing, expression, posture and the scene in which the character is located, the smiling expression features, the action features of crossing hands, the features of the top and bottom clothes and the background features of the sports field can be filled in sequentially. After filling in all the necessary information, according to the grammatical rules and expression habits of the template, the model of the present invention can also make necessary adjustments and polishes to the generated text to ensure that it is accurate, fluent and easy to understand, and finally obtain one or more text prompts describing the character attribute features.
[0073] In complex scenes and changing conditions, the appearance, posture, clothing, background and other features of a person may change. The present invention can enhance the model's adaptability to these changes by generating text prompts containing these features, so that it can achieve good recognition results in different scenes. Finally, the text prompts can be spliced with image features to further enrich the feature representation and improve the model's ability to understand and utilize image information.
[0074] In a preferred embodiment, the step of splicing the text prompts with the image features to generate spliced image features includes:
[0075] Based on the similarity function, the similarity score between each text prompt and the image feature is calculated respectively; wherein the similarity score is used to characterize the degree of association between the text prompt vector and the image feature;
[0076] Based on the similarity scores, a weight vector is generated between each text hint and image feature;
[0077] The spliced image features are generated according to each weight vector, each text prompt and the image features.
[0078] Schematically, by generating a weight vector, the contribution of different text cues in the image feature representation can be quantified. For example, the model can pay more attention to text cues that are highly relevant to image features, thereby improving recognition accuracy and efficiency.
[0079] In complex scenes, the features of people in the image may be affected by various factors such as lighting, occlusion, angle, etc. The embodiment of the present invention introduces a weight vector, so that the model can handle these changes more flexibly, thereby enhancing the robustness and adaptability of the model.
[0080] The final generated spliced image features contain more information, which not only include the visual information of the image itself, but also incorporate the semantic information in the text prompts, so that the model can understand the characteristics of the characters in the image more comprehensively.
[0081] Schematically, after generating a variety of prompt texts, the present invention can fuse the prompt texts with image features, thereby achieving cross-modal alignment and effectively fusing image and text features. For example, the image and text data are encoded separately by a text encoder and an image encoder, and mapped to the same feature space. In this way, the features (high-dimensional vectors) of each prompt text and image can be mapped to a unified semantic space, and a weight vector between each prompt and the corresponding image feature can be generated through an attention mechanism, and the features of each prompt text are weighted, and then spliced together with the corresponding image features to obtain the features after the text and image are fused, ensuring that the semantic information between the image and the text can be effectively combined, thereby providing a richer and more accurate feature representation.
[0082] Furthermore, the present invention can also use the semantic segmentation network layer of the image recognition model to further analyze the image, that is, to accurately divide different regions according to the attributes of the person, and associate each region with the specific attributes of the person. For example, it is possible to correspond the different parts of the image (such as hair, clothes, facial expressions, etc.) with the attribute information one by one, and provide a clear spatial position for each region. Through such segmentation, the details reflected in different regions of the image can be better understood, thereby improving the fine-grained feature extraction capability of person re-identification.
[0083] Specifically, after obtaining the character attribute features including appearance features, posture features, clothing features, and background features for representing the environmental background of the character, the present invention can also divide each area of the image based on the above-mentioned character attribute features to obtain more detailed image features and image descriptions, then:
[0084] The regions in the target image include: a background region, a head region, a limb region, and a clothing region, and the region division of the target image according to the character attribute features includes:
[0085] According to the background features, the target image is divided into a background area and a human body area;
[0086] Extracting a head region from the human body region according to the appearance features;
[0087] Extracting a limb region from the human body region according to the posture feature;
[0088] A clothing region is extracted from the human body region according to the clothing features.
[0089] Wherein, each region corresponds to an appearance feature, a posture feature, a clothing feature or a background feature. Then, when extracting the semantic association feature between the character attribute feature and each region, it can be specifically:
[0090] For each region, according to the character attribute characteristics corresponding to the region, the visual features corresponding to the region are extracted; wherein the visual features include: hair color, action, clothing style or scene; schematically, the visual feature of the background region is: scene; the visual feature of the head region is: hair color; the visual feature of the limb region is: action; the visual feature of the clothing region is: clothing style;
[0091] Each region and the visual features corresponding to each region are used as nodes, and then each node is connected to generate an association graph for representing the semantic relationship between the character attribute features and the regions; wherein the connection edges in the association graph represent the semantic relationship between the character attribute features and the regions;
[0092] For each node, update the node features corresponding to each node according to the features of the neighboring nodes connected to the node and the semantic relationship of the connecting edges connected to the node;
[0093] Generate, based on each updated node feature, a spatial correlation feature for representing the location of a person's attribute feature in a region, an attribute correlation feature for representing the similarity of the person's attribute features between different regions, a scene correlation feature for representing the usage scene of the person's attribute features in a region, and a usage correlation feature between different person's attribute features;
[0094] The spatial association features, attribute association features, scene association features and usage association features are output as semantic association features.
[0095] Schematically, the present invention can more accurately capture the attribute features of people in the image and their semantic association with the image area by finely dividing the image into regions and extracting visual features. Moreover, based on the generation of the association graph and the update of the node features, the semantic association features finally generated cover multiple dimensions such as space, attributes, scenes and uses, thereby providing a comprehensive perspective for the model to understand the characteristics of people in the image.
[0096] In a preferred embodiment, for each segmented attribute region, the present invention can extract the visual features of each region, thereby reflecting the specific performance of each attribute in the image, such as the color and shape of hair, the style and fabric of clothes, etc. By extracting the above-mentioned local visual features, the details in the image can be captured more accurately, making the expression of each attribute richer.
[0097] After extracting the local features of each attribute area, the present invention can further use the graph neural network layer to construct a semantic relationship graph between the image area and the character attributes. Specifically, based on the above target image, each area in the image and its corresponding character attribute features are regarded as nodes in the graph, and the nodes are connected by edges, and the semantic relationship between them is represented.
[0098] Through the message passing mechanism, each node can exchange information with neighboring nodes. Specifically, each node will perform weighted summation based on the features of its connected neighboring nodes, and normalize the results of the weighted summation to ensure that the update of each node has a relatively consistent scale. After weighted summation and normalization, the node's features will be updated to new features, thereby capturing the semantic information of neighboring nodes and improving the understanding of nodes and their contexts. For example, the spatial relationship between clothing and color, hair and face, or the functional connection between sneakers and sportswear can be identified.
[0099] By generating the above-mentioned association graph, the various regions and attributes of the image can be effectively expressed in the same semantic space, and the potential associations between different regions and attributes in the image can be captured, thereby enhancing the understanding and description of the image content and further improving the model's ability to understand and distinguish fine-grained features.
[0100] Finally, the present invention can combine semantic association features, character attribute features and image features to generate final fusion features, that is, target character features corresponding to the character to be identified, which include visual information of the image and a variety of semantic information.
[0101] Schematically, the spatial association feature includes: the position feature of hair color in the face area; the attribute association feature includes: the similarity feature between the hair color in the head area and the clothing style in the clothing area; the scene association feature includes: the usage scene feature of the clothing style in the background area; the use association feature includes: the use feature predicted according to the clothing style and the scene;
[0102] The semantic association features, the character attribute features and the image features are then integrated to generate target character features corresponding to the character to be identified, specifically including:
[0103] Generate a number of different text prompts based on the attribute characteristics of each character and a preset prompt template; wherein the text prompt is used to describe the character's expression and the character's posture, or to describe the character's clothing, the character's expression, the character's posture and the scene in which the character is located;
[0104] Splicing each text prompt with the image feature to generate a spliced image feature;
[0105] The position features of hair color in the facial area, the similarity features between hair color and clothing style between the head area and clothing area, the usage scenario features of clothing style in the background area, the usage features predicted based on clothing style and scene, and the spliced image features are fused to generate the target person features corresponding to the person to be identified.
[0106] Specifically, when the target image shows a young woman standing in a park, wearing a red dress, smiling, and gently waving her hands, the present invention can extract and fuse her various features to generate a comprehensive feature description of the target person.
[0107] First, the character attribute features are extracted. The appearance features are: smiling expression features, the posture features are: the action features of gently swinging hands, the clothing features are: red dress features, and the background features are: park scenes.
[0108] Based on the appearance features and posture features, a text prompt can be generated: "A smiling woman with her hands gently waving", or based on the clothing features and background features, another text prompt can be generated: "A woman in a red dress standing in the park".
[0109] Then, the visual features of the image are extracted, such as color, texture, shape, etc. Further, after the image is divided into regions, the extracted semantic association features can be:
[0110] Spatial association features: The location features of the hair color (assumed to be black) in the facial area, that is, the hair color is close to the scalp and covers the entire head;
[0111] Attribute association features: If there is no obvious similarity between the black hair in the head area and the red dress in the clothing area, they can be recorded as different attribute features;
[0112] Scene-related features: The use scene features of the red dress in the park scene are described as casual or comfortable;
[0113] Purpose-related features: Based on the red dress and the park scene, the predicted purpose features are walking or leisure activities.
[0114] Finally, the generated text prompts "a smiling woman, gently waving her hands" and "a woman in a red dress standing in the park" are concatenated with the image features to form a concatenated image feature containing text and visual information;
[0115] The spatial correlation features (the position of black hair in the facial area), attribute correlation features (the different attributes of black hair and red dress), scene correlation features (the leisure scene of red dress in the park) and purpose correlation features (the predicted purpose of walking and leisure activities) are further integrated with the spliced image features to generate the target person features.
[0116] Indicatively, since the dress matches the leisure scene in the park, it means that she should be taking a walk or doing leisure activities. The image shows her smiling expression, swaying movements, the color and texture of the dress and other visual features. The target person features finally fused can be described as: "A smiling woman with her hands gently swaying, wearing a red dress and black hair, standing in the park.
[0117] Therefore, the present invention not only extracts visual information from the image, but also integrates a variety of semantic information, including the person's expression, posture, clothing, scene, and the associated features between them. The accuracy of person re-identification can be improved through comprehensive feature description, and even in complex and changing application scenarios, the person's characteristics in the image can be identified for accurate identity recognition and matching.
[0118] For step S3 and step S4, in a preferred embodiment, the present invention can select cosine similarity as a measurement method to calculate the similarity. Then, for the target person feature vector and each preset person feature vector in the database, a similarity measurement method (such as cosine similarity) is used to calculate to obtain a series of similarity scores.
[0119] Sort the calculated similarity scores from high to low. Select the preset image with the highest similarity score and mark the preset person corresponding to it as the most similar to the person in the target image. Finally, obtain the corresponding identity information (such as name, age, gender, etc.), and mark this identity information as the identity information corresponding to the person to be identified in the target image.
[0120] Since the present invention not only extracts the character attribute features of the person to be identified in the target image, but also extracts the image features of the target image itself, the dimension of feature extraction can be increased, making the feature description richer and more comprehensive, and through the obtained semantic association features, the semantic association features between character attributes and image regions can be captured. Furthermore, in the feature extraction process of the present invention, multi-prompt learning is also used to generate diversified prompt information related to character features, and the image and text features are effectively fused through a cross-modal alignment mechanism. Semantic segmentation and attribute matching are used to perform fine-grained attribute segmentation on character images, and the features of each attribute region are independently extracted and fused, thereby improving the extraction accuracy of character features, and by modeling the semantic relationship between image regions and attributes through graph neural networks, the model's ability to understand and distinguish fine-grained features can be further enhanced, thereby improving the accuracy of character re-identification.
[0121] like Figure 2 As shown, based on the above-mentioned various embodiments of the person re-identification method, the present invention provides a corresponding device embodiment;
[0122] An embodiment of the present invention provides a person re-identification device, comprising: an image acquisition module, a person feature generation module, a similarity calculation module, and an identity information determination module;
[0123] The image acquisition module is used to acquire a target image containing a person to be identified;
[0124] The character feature generation module is used to input the target image into a preset image recognition model so that the image recognition model extracts character attribute features corresponding to the character to be recognized in the target image and image features corresponding to the target image, divides the target image into regions according to the character attribute features and extracts semantic association features between the character attribute features and each region, and then fuses the semantic association features, the character attribute features and the image features to generate target character features corresponding to the character to be recognized;
[0125] The similarity calculation module is used to calculate the similarity between the target person feature and each preset person feature in the database; wherein different preset person features correspond to different preset images; and the preset image includes a preset person and identity information corresponding to the preset person;
[0126] The identity information determination module is used to mark the identity information corresponding to the preset image with the highest similarity as the identity information corresponding to the person to be identified in the target image.
[0127] It should be noted that the device embodiments described above are merely schematic, wherein the modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules, and may be located in one place or distributed on multiple network modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, in the accompanying drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines. Those of ordinary skill in the art can understand and implement it without paying creative labor.
[0128] Those skilled in the art can clearly understand that, for the sake of convenience and simplicity, the specific working process of the device described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0129] Based on the above-mentioned various embodiments of the person re-identification method, the present invention provides corresponding embodiments of terminal equipment items.
[0130] An embodiment of the present invention provides a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, and when the processor executes the computer program, a person re-identification method described in any method embodiment of the present invention is implemented.
[0131] The terminal device may be a computing terminal device such as a desktop computer, a notebook, a palm computer, a cloud server, etc. The terminal device may include, but is not limited to, a processor and a memory.
[0132] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the terminal device, and uses various interfaces and lines to connect various parts of the entire terminal device.
[0133] The memory can be used to store the computer program, and the processor realizes various functions of the terminal device by running or executing the computer program stored in the memory and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function, etc.; the data storage area can store data created according to the use of the mobile phone, etc. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Med i aCard, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device or other volatile solid-state storage device.
[0134] Based on the above-mentioned various embodiments of the person re-identification method, the present invention provides a corresponding storage medium item embodiment.
[0135] An embodiment of the present invention provides a storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute a person re-identification method described in any method embodiment of the present invention.
[0136] The storage medium is a computer-readable storage medium, and the computer program is stored in the computer-readable storage medium. When the computer program is executed by the processor, the steps of each of the above-mentioned method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.
[0137] The above is a preferred embodiment of the present invention. It should be pointed out that a person skilled in the art can make several improvements and modifications without departing from the principle of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A person re-identification method, characterized in that: include: Obtain a target image containing a person to be identified; Input the target image into a preset image recognition model so that the image recognition model extracts the character attribute features corresponding to the person to be recognized in the target image and the image features corresponding to the target image, divides the target image into regions according to the character attribute features and extracts the semantic association features between the character attribute features and each region, and then fuses the semantic association features, the character attribute features and the image features to generate the target character features corresponding to the person to be recognized; Calculate the similarity between the target person feature and each preset person feature in the database; wherein different preset person features correspond to different preset images; the preset image includes a preset person and identity information corresponding to the preset person; The identity information corresponding to the preset image with the highest similarity is marked as the identity information corresponding to the person to be identified in the target image.
2. A person re-identification method as claimed in claim 1, characterized in that: The character attribute features include: appearance features, posture features, clothing features, and background features used to represent the environment background of the character; each area corresponds to an appearance feature, a posture feature, a clothing feature, or a background feature; The extracting of the semantic association features between the character attribute features and each region includes: For each region, according to the character attribute characteristics corresponding to the region, the visual features corresponding to the region are extracted; wherein the visual features include: hair color, action, clothing style or scene; Each region and the visual features corresponding to each region are used as nodes, and then each node is connected to generate an association graph for representing the semantic relationship between the character attribute features and the regions; wherein the connection edges in the association graph represent the semantic relationship between the character attribute features and the regions; For each node, update the node features corresponding to each node according to the features of the neighboring nodes connected to the node and the semantic relationship of the connecting edges connected to the node; Generate, based on each updated node feature, a spatial correlation feature for representing the location of a person's attribute feature in a region, an attribute correlation feature for representing the similarity of the person's attribute features between different regions, a scene correlation feature for representing the usage scene of the person's attribute features in a region, and a usage correlation feature between different person's attribute features; The spatial association features, attribute association features, scene association features and usage association features are output as semantic association features.
3. A person re-identification method as claimed in claim 2, characterized in that: The regions in the target image include: a background region, a head region, a limb region, and a clothing region; The step of dividing the target image into regions according to the character attribute features includes: According to the background features, the target image is divided into a background area and a human body area; Extracting a head region from the human body region according to the appearance features; extracting a limb region from the human body region according to the posture feature; A clothing region is extracted from the human body region according to the clothing features.
4. A person re-identification method as claimed in claim 3, characterized in that: The spatial association features include: the position features of hair color in the facial area; the attribute association features include: the similarity features between the hair color in the head area and the clothing style in the clothing area; the scene association features include: the usage scene features of the clothing style in the background area; the use association features include: the use features predicted based on the clothing style and the scene; The step of fusing the semantic association feature, the character attribute feature and the image feature to generate a target character feature corresponding to the character to be identified includes: Generate a number of different text prompts based on the attribute characteristics of each character and a preset prompt template; wherein the text prompt is used to describe the character's expression and the character's posture, or to describe the character's clothing, the character's expression, the character's posture and the scene in which the character is located; Splicing each text prompt with the image feature to generate a spliced image feature; The position features of hair color in the facial area, the similarity features between hair color and clothing style between the head area and clothing area, the usage scenario features of clothing style in the background area, the usage features predicted based on clothing style and scene, and the spliced image features are fused to generate the target person features corresponding to the person to be identified.
5. A person re-identification method as claimed in claim 4, characterized in that: The appearance features include: smiling facial features; the posture features include: crossed hands action features; the clothing features include: top features used to describe a green jacket and bottom features used to describe black sweatpants; the background features include: sports field; the text prompts are used to describe the character's clothing, facial expressions, postures, and scenes; According to the attribute characteristics of each character and the preset prompt template, a number of different text prompts are generated, including: A text prompt is generated according to the facial expression features of a smile, the action features of crossing hands, the features of an upper garment, the features of a lower garment, a sports field and a preset prompt template; wherein the text prompt includes: a person wearing a green jacket and black pants is standing on a sports field, smiling, with his hands crossed in front of his chest.
6. A person re-identification method as claimed in claim 5, characterized in that: The step of splicing each text prompt with the image feature to generate the spliced image feature includes: Based on the similarity function, the similarity score between each text prompt and the image feature is calculated respectively; wherein the similarity score is used to characterize the degree of association between the text prompt vector and the image feature; Based on the similarity scores, a weight vector is generated between each text hint and image feature; The spliced image features are generated according to each weight vector, each text prompt and the image features.
7. A person re-identification method as claimed in claim 1, characterized in that: The training process of the image recognition model includes: Taking a number of image samples and the actual character features corresponding to each image sample as input, and taking the predicted character features of each image sample as output, the image recognition model to be trained is iteratively trained until the model converges, thereby generating a preset image recognition model; In each iterative training, an image sample is input into the image recognition model so that the image recognition model generates a prediction result of a person feature corresponding to the image sample; The predicted character feature results are compared with the actual character feature results, and the network parameters of the image recognition model are adjusted according to the comparison results.
8. A person re-identification device, characterized in that: include: Image acquisition module, character feature generation module, similarity calculation module and identity information determination module; The image acquisition module is used to acquire a target image containing a person to be identified; The character feature generation module is used to input the target image into a preset image recognition model so that the image recognition model extracts character attribute features corresponding to the character to be recognized in the target image and image features corresponding to the target image, divides the target image into regions according to the character attribute features and extracts semantic association features between the character attribute features and each region, and then fuses the semantic association features, the character attribute features and the image features to generate target character features corresponding to the character to be recognized; The similarity calculation module is used to calculate the similarity between the target person feature and each preset person feature in the database; wherein different preset person features correspond to different preset images; and the preset image includes a preset person and identity information corresponding to the preset person; The identity information determination module is used to mark the identity information corresponding to the preset image with the highest similarity as the identity information corresponding to the person to be identified in the target image.
9. A terminal device, characterized in that: The method comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, a person re-identification method as claimed in any one of claims 1 to 7 is implemented.
10. A storage medium, characterized in that: The storage medium includes a stored computer program, wherein when the computer program is executed, the device where the storage medium is located is controlled to execute a person re-identification method as described in any one of claims 1 to 7.