Object identification method and device, storage medium and electronic device

By fusing prompt text and image features using a continuously stacked target fusion module, semantic control weight vectors are generated, and target image features are activated, which solves the problem of pre-defined attribute sets in the prior art, and realizes efficient object attribute recognition and target re-recognition.

CN119693632BActive Publication Date: 2025-05-16ZHEJIANG DAHUA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510201639.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-05-16
Estimated Expiration
2045-02-24

AI Technical Summary

Technical Problem

In the prior art, object attribute recognition methods require predefined attribute sets, resulting in high training costs and limited recognition scope, and the inability to effectively identify new attributes, resulting in low recognition rate and low efficiency.

Method used

The continuously stacked target fusion module is used to extract the prompt text and the features of the image to be recognized, fuse and map to obtain a semantic control weight vector, which is used to activate the target image features and realize attribute recognition and target re-recognition.

Benefits of technology

Without predefined attribute vocabulary, image features are dynamically adjusted to adapt to specific attribute retrieval and target re-identification tasks, realizing open-set attribute retrieval and efficient target recognition, improving the flexibility and efficiency of recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119693632B_ABST
    Figure CN119693632B_ABST
Patent Text Reader

Abstract

The present application discloses an object recognition method and device, storage medium and electronic device. The method includes: obtaining a prompt text and an image to be recognized; using a continuously stacked target fusion module to extract the prompt text features of the prompt text and the image features to be recognized of the image to be recognized, fusing the prompt text features and the image features to be recognized to obtain a fusion feature, and performing a mapping operation on the fusion feature to determine a semantic control weight vector; using the semantic control weight vector to activate the target image feature to obtain a text activation image feature; performing attribute recognition tasks and target re-recognition tasks based on the text activation image feature to generate a target recognition result. The present application solves the technical problem in the related art that it is necessary to predefine the object attribute values ​​in order to recognize the specified object, resulting in a limited object recognition rate and too low recognition efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computers, and in particular to an object recognition method and device, a storage medium, and an electronic device. Background Art

[0002] In related technologies, object attribute recognition methods usually rely on predefined attribute sets, such as color, gender, age, etc., to identify and retrieve objects with specific attributes. This closed-set recognition method requires a large number of attributes and their variants to be stored in the training dataset, resulting in high model training costs and limited recognition scope to known attributes.

[0003] Furthermore, when encountering new attributes that do not appear in the training data, traditional object attribute recognition systems cannot effectively identify them, resulting in limited object recognition rates and low recognition efficiency. For example, if the training data set only covers common clothing colors, it cannot respond to queries for specific clothing patterns.

[0004] In summary, the related art has a technical problem that it is necessary to predefine object attribute values ​​in order to identify a specified object, resulting in a limited object recognition rate and low recognition efficiency.

[0005] To address the above-mentioned problems, no effective solution has been proposed yet. Summary of the invention

[0006] The embodiments of the present application provide an object recognition method and device, a storage medium and an electronic device to at least solve the technical problem in the related art that object attribute values ​​need to be predefined to identify a specified object, resulting in a limited object recognition rate and low recognition efficiency.

[0007] According to one aspect of an embodiment of the present application, a method for identifying an object is provided, comprising: obtaining a prompt text and an image to be identified; using a continuously stacked target fusion module to extract a prompt text feature of the prompt text and an image feature to be identified of the image to be identified, fusing the prompt text feature and the image feature to be identified to obtain a fused feature, and performing a mapping operation on the fused feature to determine a semantic control weight vector, wherein the semantic control weight vector is used to represent the weight value of each image part in the image to be identified, and the target fusion module represents a module for processing the prompt text feature and the image feature to be identified based on a state space model; using the semantic control weight vector to activate the target image feature to obtain the text Activate image features, wherein the target image features are determined by the target image, the target image is different from the image to be identified, and the text-activated image features represent the image features of the target image integrated with the prompt text features; based on the text-activated image features, perform attribute recognition tasks and target re-recognition tasks respectively to generate target recognition results, wherein the target recognition result is used to indicate whether the target image includes a target object, the image to be identified includes a target object, the prompt text is used to indicate the target object attributes of the target object, the target re-recognition task is used to identify target objects in different time and space in the target image, and the attribute recognition task is used to determine the target object whose object attributes in the target image are the target object attributes.

[0008] According to another aspect of the embodiment of the present application, there is also provided an object recognition device, comprising: an acquisition module, used to acquire a prompt text and an image to be recognized; a fusion module, used to use a continuously stacked target fusion module to extract a prompt text feature of the prompt text and an image feature to be recognized of the image to be recognized, fuse the prompt text feature and the image feature to be recognized to obtain a fusion feature, and perform a mapping operation on the fusion feature to determine a semantic control weight vector, wherein the semantic control weight vector is used to represent the weight value of each image part in the image to be recognized, and the target fusion module represents a module for processing the prompt text feature and the image feature to be recognized based on a state space model; an activation module, used to activate the target image using the semantic control weight vector The target image feature is determined by the target image, the target image is different from the image to be identified, and the text activation image feature represents the image feature of the target image integrated with the prompt text feature; a generation module is used to perform attribute recognition tasks and target re-identification tasks based on the text activation image feature to generate a target recognition result, wherein the target recognition result is used to indicate whether the target image includes a target object, the image to be identified includes a target object, the prompt text is used to indicate the target object attribute of the target object, the target re-identification task is used to identify the target objects in different time and space in the target image, and the attribute recognition task is used to determine the target object whose object attribute in the target image is the target object attribute.

[0009] Optionally, the device is used to extract prompt text features of the prompt text and extract image features to be identified of the image to be identified using continuously stacked target fusion modules in the following manner: using a text feature extraction network model to extract initial text features of the prompt text; using a visual feature extraction network model to segment the image to be identified into a group of image regions, and performing a linear mapping operation on a group of image regions to obtain initial image features; using a continuously stacked first fusion module to process the initial text features to determine the prompt text features, and using a continuously stacked second fusion module to process the initial image features to determine the image features to be identified, wherein the target fusion module includes a first fusion module and a second fusion module.

[0010] Optionally, the device is also used for at least one of the following: performing a vector forward transformation operation on the initial text feature to obtain a first text feature vector represented by a one-dimensional vector, inputting the first text feature vector into a plurality of self-attention modules with different configuration parameters to obtain a plurality of second text feature vectors, performing a vector reverse transformation operation on the plurality of second text feature vectors to obtain a plurality of third text feature vectors represented by a multi-dimensional vector, performing a vector addition operation on the plurality of third text feature vectors to obtain text feature parameters, performing an addition operation on a dot product result of the text feature parameters and the initial text feature and the initial text feature to obtain a prompt text feature, wherein the first fusion module It includes a self-attention module; performing a vector forward transformation operation on the initial image feature to obtain a first image feature vector represented by a one-dimensional vector, inputting the first image feature vector into multiple self-attention modules with different configuration parameters to obtain multiple second image feature vectors, performing a vector reverse transformation operation on the multiple second image feature vectors to obtain multiple third image feature vectors represented by multi-dimensional vectors, performing a vector addition operation on the multiple third image feature vectors to obtain image feature parameters, performing an addition operation on the dot product result of the image feature parameters and the initial image feature and the initial image feature to obtain the image feature to be identified, wherein the second fusion module includes the self-attention module.

[0011] Optionally, the device is used to activate the target image feature according to the semantic control weight vector to obtain the text activation image feature in the following manner: determine the parameter at the first position in the semantic control weight vector as the first semantic control weight parameter, and determine the parameter at the second position in the semantic control weight vector as the second semantic control weight parameter, wherein the first position is before the second position; multiply the first semantic control weight parameter and the target image feature element by element to obtain the initial activation feature; and add the second semantic control weight parameter and the initial activation feature element by element to obtain the text activation image feature.

[0012] Optionally, the device is also used to: obtain sample prompt text and sample to-be-recognized image; use continuously stacked target fusion modules to extract sample prompt text features of the sample prompt text, and extract sample to-be-recognized image features of the sample to-be-recognized image, multiply the sample prompt text features and the sample to-be-recognized image features element by element to obtain sample fusion features, and perform mapping operations on the sample fusion features to determine a sample semantic control weight vector, wherein the sample semantic control weight vector represents the weight values ​​of each sample image part in the sample to-be-recognized image; activate the sample target image features according to the sample semantic control weight vector to obtain sample text activation image features, wherein the sample target image features are determined by the sample target image, and the sample target image is different from the sample to-be-recognized image; train an initial recognition network based on the sample text activation image features to obtain a target recognition network, wherein the target recognition network is used to perform at least one of an attribute recognition task and a target re-recognition task based on the prompt text, the image to be recognized and the target image.

[0013] Optionally, the device is used to train an initial recognition network based on sample text activated image features to obtain a target recognition network in the following manner: use a cross entropy loss function to calculate text loss values ​​of sample text activated image features and sample initial text features, and use a triple loss function to calculate image loss values ​​of sample text activated image features and sample initial image features, wherein the sample initial text features are text features of sample prompt text extracted by a text feature extraction network model, and the sample initial image features are image features of a sample image to be recognized extracted by a visual feature extraction network model; when the text loss value satisfies a first convergence condition and the image loss value satisfies a second convergence condition, the initial recognition network is determined as a target recognition network.

[0014] Optionally, the device is also used to: obtain prompt text and an image to be identified; use a text feature extraction network model to extract initial text features of the prompt text; use a visual feature extraction network model to segment the image to be identified into a group of image regions, and perform a linear mapping operation on the group of image regions to obtain initial image features; use a continuously stacked first fusion module to process the initial text features to determine prompt text features, and use a continuously stacked second fusion module to process the initial image features to determine image features to be identified, wherein the target fusion module includes a first fusion module and a second fusion module; multiply the prompt text features and the image features to be identified element by element to obtain fused features, and perform mapping on the fused features Operation, determine the semantic control weight vector; determine the parameter at the first position in the semantic control weight vector as the first semantic control weight parameter, and determine the parameter at the second position in the semantic control weight vector as the second semantic control weight parameter, wherein the first position is before the second position; multiply the first semantic control weight parameter and the target image feature element by element to obtain the initial activation feature, wherein the target image feature is obtained by processing the target image by a visual feature extraction network model; add the second semantic control weight parameter and the initial activation feature element by element to obtain the text activation image feature; perform the attribute recognition task and the target re-recognition task based on the text activation image feature respectively to generate a target recognition result.

[0015] According to another aspect of the embodiments of the present application, a computer-readable storage medium is provided, in which a computer program is stored, wherein the computer program is configured to execute the above-mentioned object recognition method when running.

[0016] According to another aspect of the embodiment of the present application, a computer program product or a computer program is provided, the computer program product or the computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the above object recognition method.

[0017] According to another aspect of the embodiments of the present application, there is further provided an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the object recognition method through the computer program.

[0018] In an embodiment of the present application, a prompt text and an image to be identified are obtained; a continuously stacked target fusion module is used to extract prompt text features of the prompt text and image features to be identified of the image to be identified, the prompt text features and the image features to be identified are fused to obtain fused features, and a mapping operation is performed on the fused features to determine a semantic control weight vector. Through text-guided image feature extraction and activation, the purpose of dynamically adjusting image features to adapt to specific attribute retrieval and target re-identification tasks without the need for predefined attribute vocabulary is achieved, thereby achieving the technical effect of open set attribute retrieval and efficient target recognition, and further, solving the technical problem in the related art that object attribute values ​​need to be predefined to identify specified objects, resulting in limited object recognition rate and low recognition efficiency.

[0019] Specifically, the target fusion module in this embodiment can not only process the features of both image and text modalities and generate fusion features, but also, through the mapping and activation mechanism of the semantic control weight vector, enable the model to adaptively focus on and extract information in the image that matches the text description according to the real-time prompt text. This effectively overcomes the limitations of traditional closed-set recognition methods, allows the system to process attribute retrieval of open vocabulary, and improves the flexibility and extensiveness of recognition.

[0020] At the same time, the activation mechanism of the target image features using the semantic control weight vector in this embodiment ensures that the model can efficiently extract key features from the target image, speeds up the recognition process, and improves the recognition efficiency. This enables the embodiment of the present application to more accurately and quickly recognize the target object, solving the problem of limited object recognition rate and low recognition efficiency in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0022] Figure 1 is a schematic diagram of an application environment of an optional object recognition method according to an embodiment of the present application;

[0023] Figure 2 is a flow chart of an optional object recognition method according to an embodiment of the present application;

[0024] Figure 3 is a schematic diagram of an optional object recognition method according to an embodiment of the present application;

[0025] Figure 4 is a schematic diagram of another optional object recognition method according to an embodiment of the present application;

[0026] Figure 5 is a schematic diagram of another optional object recognition method according to an embodiment of the present application;

[0027] Figure 6 is a schematic structural diagram of an optional object recognition device according to an embodiment of the present application;

[0028] Figure 7 is a schematic structural diagram of an optional object identification product according to an embodiment of the present application;

[0029] Figure 8 It is a schematic diagram of the structure of an optional electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0030] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.

[0031] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0032] The present application is described below in conjunction with embodiments:

[0033] According to one aspect of an embodiment of the present application, a method for identifying an object is provided. Optionally, in this embodiment, the method for identifying an object can be applied to: Figure 1 In the hardware environment composed of the server 101 and the terminal device 103 shown in FIG. Figure 1As shown, the server 101 is connected to the terminal device 103 via a network, and can be used to provide services for the terminal device or an application 107 installed on the terminal device. The application can be a video application, an instant messaging application, a browser application, an educational application, a game application, etc. A database 105 may be set up on the server or independently of the server to provide data storage services for the server 101, for example, a game data storage server. The above-mentioned network may include but is not limited to: a wired network, a wireless network, wherein the wired network includes: a local area network, a metropolitan area network and a wide area network, and the wireless network includes: Bluetooth, WIFI and other networks that realize wireless communication. The terminal device 103 may be a terminal configured with an application, and may include but is not limited to at least one of the following: a mobile phone (such as an Android phone, an iOS phone, etc.), a laptop computer, a tablet computer, a PDA, a MID (Mobile Internet Devices), a PAD, a desktop computer, a smart TV, an intelligent voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, a virtual reality (Virtual Reality, VR for short) terminal, an augmented reality (Augmented Reality, AR for short) terminal, a mixed reality (Mixed Reality, MR for short) terminal and other computer devices. The above-mentioned server may be a single server, or a server cluster consisting of multiple servers, or a cloud server.

[0034] Combination Figure 1 As shown, the above-mentioned object recognition method can be executed by an electronic device, which can be a terminal device or a server. The above-mentioned object recognition method can be implemented by the terminal device or the server respectively, or by the terminal device and the server together.

[0035] The above is only an example and is not specifically limited in this embodiment.

[0036] Optionally, as an optional implementation, as Figure 2 As shown, the above-mentioned object recognition method includes:

[0037] S202, obtaining a prompt text and an image to be recognized;

[0038] Optionally, in the embodiment of the present application, the above-mentioned image to be identified refers to an image that needs to be identified by an attribute retrieval task (the above-mentioned attribute identification task) or a ReID task (the above-mentioned target re-identification task). It can be a real-time snapshot from video surveillance or a historical image stored in a database.

[0039] It should be noted that the sources and types of images to be identified are diverse, including but not limited to high-definition images, low-resolution images, images under different lighting conditions, and images taken from different viewing angles, and this application does not limit this.

[0040] It should also be noted that the structure and content of the prompt text can have many variations. In addition to direct attribute descriptions, it can also contain relationship information, scene descriptions or time information, such as "a person who appears at the same place as a known target" or "a person riding an electric bike photographed at night." At the same time, the format and quality of the images to be identified may also differ, such as different formats such as JPEG, PNG, GIF, and images under different viewing conditions such as high definition, blur, and occlusion. The embodiments of the present application are intended to handle these diverse situations to ensure that no matter how the text prompts change or the image quality is, the model can effectively perform attribute recognition and ReID tasks, thereby having a high degree of flexibility and robustness. In addition, the application scenarios and specific forms of the prompt texts and images to be identified involved in the embodiments of the present application are not fixed, and can change to adapt to different needs and environmental changes.

[0041] S204, using a continuously stacked target fusion module to extract a prompt text feature of the prompt text and an image feature to be recognized of the image to be recognized, fuse the prompt text feature and the image feature to be recognized to obtain a fused feature, and perform a mapping operation on the fused feature to determine a semantic control weight vector, wherein the semantic control weight vector is used to represent a weight value of each image part in the image to be recognized, and the target fusion module represents a module for processing the prompt text feature and the image feature to be recognized based on a state space model;

[0042] Optionally, in an embodiment of the present application, the target fusion module refers to a multimodal feature fusion component based on a deep learning structure, including but not limited to using a Mamba architecture or a similar state-space model processing mechanism for extracting and integrating prompt text features and image features to be identified. The target fusion module can perform deep processing and interactive fusion of text and image features through continuously stacked neural network layers to generate a unified fusion feature representation.

[0043] Optionally, in an embodiment of the present application, the above-mentioned semantic control weight vector refers to a vector mapped from the fusion feature, which is used to adjust and control the activation degree of different parts of the target image features, so that the model can dynamically adjust the attention to key features in the image according to the attribute description of the prompt text.

[0044] Optionally, in an embodiment of the present application, the above-mentioned state space model refers to a mathematical model that describes the dynamic changes and interactions of text features and image features in the fusion process by defining a set of state variables and state transfer functions when processing multimodal information, for example, the SSM model (Sequential Self-Attention Model), linear ordinary differential equations, etc.

[0045] It should be noted that the structure and parameter settings of the target fusion module can have multiple variations, such as using different numbers of Mamba Blocks for stacking, or adjusting the connection mode between neural network layers and the dimension of feature mapping to adapt to image data and text descriptions of different scales and complexities. In addition, the method of feature fusion can also be flexibly changed, and operations such as dot multiplication and weighted summation can be used, or attention mechanisms can be introduced for feature selection, which is not limited in this application.

[0046] For example, in the embodiment of the present application, a continuously stacked target fusion module is used to first extract the features of the prompt text and convert the text into an understandable numerical representation. Then, the visual features of the image to be identified are extracted to capture the visual information in the image. The extracted text features and image features are fused to generate fused features containing the interactive information of the text and image. Subsequently, a mapping operation is performed on the fused features to convert them into a semantic control weight vector.

[0047] S206, using the semantic control weight vector to activate the target image feature to obtain a text activated image feature, wherein the target image feature is determined by the target image, the target image is different from the image to be recognized, and the text activated image feature represents the image feature of the target image fused with the prompt text feature;

[0048] Optionally, in an embodiment of the present application, the semantic control weight vector includes but is not limited to a parameter set for adjusting the importance of each region in the target image feature, and can dynamically change the activation mode of the target image feature according to the semantic information of the prompt text.

[0049] Optionally, in an embodiment of the present application, the above-mentioned target image features refer to feature representations generated by a target image (another image different from the image to be identified) through an image feature extraction network, which contain content information of the target image, including but not limited to pixel-level features, regional features or global features of the image, and can capture details and contextual information in the target image for subsequent attribute recognition and ReID tasks.

[0050] It should be noted that the relationship between the target image and the image to be identified can be diverse. The target image is the object used by the system for analysis and comparison. It can be an image from different cameras, at different time periods, or an image that has been preprocessed and enhanced. This application does not limit this.

[0051] Exemplarily, the process of activating target image features using the semantic control weight vector includes but is not limited to:

[0052] First, the weight vector is element-wise multiplied by the target image feature, and then the text-activated image feature is obtained through operations such as weighted summation. This enables the model to selectively enhance or suppress relevant or irrelevant parts of the target image feature according to the attribute information in the text description, thereby obtaining the text-activated image feature that incorporates the prompt text feature.

[0053] In this way, even when there are differences between the target image being processed and the image to be identified, the model can accurately capture features that match the text description, improving the accuracy and robustness of attribute recognition and ReID tasks, making feature extraction more targeted and adaptable, and suitable for various image processing and object recognition scenarios.

[0054] S208, based on the text-activated image features, respectively perform attribute recognition tasks and target re-recognition tasks to generate a target recognition result, wherein the target recognition result is used to indicate whether the target image includes a target object, the image to be recognized includes a target object, the prompt text is used to indicate the target object attributes of the target object, the target re-recognition task is used to recognize target objects in different time and space in the target image, and the attribute recognition task is used to determine the target object whose object attributes in the target image are target object attributes.

[0055] Optionally, in an embodiment of the present application, the above-mentioned target recognition result refers to a comprehensive judgment result obtained after analysis based on text-activated image features, indicating whether the target image contains the target object and the specific attributes of the target object, including but not limited to attribute recognition results and target re-identification (ReID) results, which can be used to determine the attribute information and spatiotemporal continuity of the target object.

[0056] Optionally, in an embodiment of the present application, the above-mentioned attribute recognition task refers to the process of determining whether the attributes of the object in the target image match the attributes of the target object described in the prompt text based on the text-activated image features, including but not limited to recognizing the clothing color, gender, behavioral characteristics, etc. of the target object. These attributes can be open vocabulary, that is, the system can recognize attributes that are not predefined in the training phase.

[0057] Optionally, in an embodiment of the present application, the above-mentioned target re-identification task refers to the process of identifying whether the target object in the target image at different time and space is the same entity as the target object in the image to be identified, including but not limited to pedestrian re-identification, vehicle re-identification, etc. By analyzing the text-activated image features, the system can transcend the limitations of time and space and determine the identity consistency of the same target object in different scenarios.

[0058] It should be noted that the types and expressions of the target object attributes can be diverse, and can be simple attribute words or complex description statements, and this application does not limit this. At the same time, the different time and space in the target re-identification task are not limited to scenes captured by different cameras, but can also be images captured by the same camera at different time points, or even image scenes under different lighting and different viewing angles, and this application also does not limit this.

[0059] Exemplarily, based on the text-activated image features, attribute recognition and target re-identification analysis can be performed simultaneously. In the attribute recognition task, the system will compare the similarity between the text-activated image features and the preset attribute classification to determine whether the target object has the attributes described in the prompt text; in the target re-identification task, the distance between the text-activated image features and the features of other images in the database can be calculated to identify whether the object in the target image is the same entity as the target object in the image to be identified. Through multimodal feature processing and fusion, the two key problems of attribute recognition and target re-identification can be solved simultaneously, greatly improving the recognition efficiency and accuracy, and is suitable for various security monitoring, identity authentication and intelligent search scenarios.

[0060] In an exemplary embodiment, taking the application scenario of pedestrian identification and attribute retrieval as an example, Figure 3It is a schematic diagram of an optional object recognition method according to an embodiment of the present application, wherein Image Features is used to represent the target image, with a size of H'W'C, H' is the image height, W' is the image width, and C is the number of image channels; Controlweights is used to represent the semantic control weight vector, with a size of 21D, D is the number of feature channels; Activated ImageFeatures is used to represent the text activated image features, with a size of H1W1C, H1 is the image height, W1 is the image width, and C is the number of image channels; Image Features is used to represent the image to be recognized, with a size of H0W0C0, H0 is the image height, W0 is the image width, and C is the number of image channels; ViT Backbone is used to represent the visual feature extraction network model; Mamba Controller is used to represent the target fusion module; Text Features is used to represent the initial text features, with a size of TC, T is the text length, and C is the number of feature channels; CrossEntropy Loss is used to represent the cross entropy loss function; BEAT is used to represent the text feature extraction network model; Triplet loss is used to represent the triple loss function, including but not limited to:

[0061] S1, obtain prompt text and image to be recognized: First, the system receives a text description from the operator or user, such as "a woman wearing a gray top, black pants, and holding a red backpack", and a pedestrian image extracted from a real-time or historical database of video surveillance.

[0062] S2, extract features and fuse them: Figure 4 is a schematic diagram of another optional object recognition method according to an embodiment of the present application. The processing flow of the target fusion module can be as follows Figure 4 As shown in the figure, Hybrid Features is used to represent the fused features, the size is HWT, H is the image height, W is the image width, and T is the text length; Average+Linear is used to represent the "average+linear mapping operation", and the target fusion module Mamba block is used. This module consists of continuously stacked multimodal processing units, and each unit performs deep processing and fusion of the prompt text features and the image features to be identified based on the state space model. Figure 5 is a schematic diagram of another optional object recognition method according to an embodiment of the present application, the Mamba block architecture is as follows Figure 5 As shown, using Flatten i Represents a vector forward transformation operation; using SSM i Represents the self-attention module; using Unflatten i Indicates a vector reverse conversion operation; use "F Out∈Image Features” represents the image features to be identified. Figure 5 For example, the value range of i is [1, 4], the size is HWC, H is the image height, W is the image width, and C is the number of feature channels.

[0063] Combination Figure 3 For example, the text description is fed into the BERT network for encoding to extract text features; at the same time, the image to be recognized is processed through the ViT network to extract image features. After that, the two features are fused to generate a fused feature and a mapping operation is performed to determine the semantic control weight vector, which is used to adjust the importance of different parts of the image for attribute recognition and target re-recognition tasks.

[0064] S3, activate the target image features: use the semantic control weight vector to activate the features of another target image. The "target image" here refers to the image used by the system for comparison, which can come from different cameras or time points. The target image is sent to the same image feature extraction network (such as ViT) to extract its feature representation, and then the semantic control weight vector is used to adjust these features to obtain the text-activated image feature, which better reflects the image information that matches the text description.

[0065] S4, perform attribute recognition and target re-identification tasks: Based on the text-activated image features, the system performs attribute recognition tasks and target re-identification tasks respectively. In the attribute recognition task, the system compares the text-activated image features with the templates of known attribute categories to determine whether the pedestrian in the target image has the same attributes as described in the prompt text. In the target re-identification task, the system uses the image features fused with text information to identify whether the pedestrian in the target image is the same pedestrian described in the image to be recognized, even in different environments, postures or lighting conditions.

[0066] S5, generate target recognition result: Based on the results of comprehensive attribute recognition and target re-recognition, the system generates a target recognition result, indicating whether the target image contains a target object with the attributes described in the prompt text, and whether the object is the same entity as the target object in the image to be recognized.

[0067] In addition, it should be noted that according to actual business needs, only the attribute recognition task or the target re-identification task can be performed, rather than both at the same time. This means that in some scenarios, such as when you only need to retrieve objects that meet specific attribute descriptions without caring about their identity continuity, you can use only the prompt text for attribute recognition. At this time, you can focus on understanding the attribute information described in the text and adjust the output of the image feature extraction network accordingly to enhance the image features associated with the attribute description. In this case, the generation of the semantic control weight vector will rely entirely on the prompt text without referring to the image to be recognized.

[0068] Through the embodiments of the present application, a prompt text and an image to be identified are obtained; a continuously stacked target fusion module is used to extract the prompt text features of the prompt text and the image features to be identified of the image to be identified, the prompt text features and the image features to be identified are fused to obtain fused features, and a mapping operation is performed on the fused features to determine the semantic control weight vector. Through text-guided image feature extraction and activation, the purpose of dynamically adjusting image features to adapt to specific attribute retrieval and target re-identification tasks without pre-defining attribute vocabulary is achieved, thereby achieving the technical effect of open set attribute retrieval and efficient target recognition, and further, solving the technical problem in the related art that object attribute values ​​need to be pre-defined to identify specified objects, resulting in limited object recognition rate and low recognition efficiency.

[0069] Specifically, the target fusion module in this embodiment can not only process the features of both image and text modalities and generate fusion features, but also, through the mapping and activation mechanism of the semantic control weight vector, enable the model to adaptively focus on and extract information in the image that matches the text description according to the real-time prompt text. This effectively overcomes the limitations of traditional closed-set recognition methods, allows the system to process attribute retrieval of open vocabulary, and improves the flexibility and extensiveness of recognition.

[0070] At the same time, the activation mechanism of the target image features using the semantic control weight vector in this embodiment ensures that the model can efficiently extract key features from the target image, speeds up the recognition process, and improves the recognition efficiency. This enables the embodiment of the present application to more accurately and quickly recognize the target object, solving the problem of limited object recognition rate and low recognition efficiency in the prior art.

[0071] As an optional solution, a continuously stacked target fusion module is used to extract prompt text features of the prompt text and to extract image features to be identified of the image to be identified, including: using a text feature extraction network model to extract initial text features of the prompt text; using a visual feature extraction network model to segment the image to be identified into a group of image regions, and performing a linear mapping operation on a group of image regions to obtain initial image features; using a continuously stacked first fusion module to process the initial text features to determine the prompt text features, and using a continuously stacked second fusion module to process the initial image features to determine the image features to be identified, wherein the target fusion module includes a first fusion module and a second fusion module.

[0072] Optionally, in an embodiment of the present application, the above-mentioned target fusion module refers to a deep learning component including a first fusion module and a second fusion module, which is used to process and integrate multimodal information, including but not limited to a state space model processing mechanism based on the Mamba architecture, and a feature extraction module using neural networks such as ViT and BERT.

[0073] It should be noted that the specific implementation methods of the first fusion module and the second fusion module can be diversified. For example, different versions of the Transformer architecture can be used, or convolutional neural networks (CNNs) can be combined to preprocess image features. Other advanced neural network structures such as the Mamba architecture can also be used for feature fusion. This application does not limit this.

[0074] Exemplarily, a text feature extraction network model such as BERT is responsible for converting the prompt text into initial text features, which carry the semantic information of the prompt text; a visual feature extraction network model such as ViT is used to segment the image to be identified into multiple image regions and perform linear mapping operations on these regions to obtain initial image features, which reflect the visual content of the image to be identified.

[0075] Through the embodiments of the present application, a continuously stacked target fusion module is adopted to achieve deep processing and efficient fusion of prompt text features and image features to be identified, which can accurately understand the text description and combine it with the image content to generate a more semantically expressive image feature representation, thereby achieving the purpose of significantly improving recognition accuracy and flexibility in attribute recognition tasks and target re-identification tasks, especially when processing open vocabulary attributes and cross-scene re-identification, showing stronger adaptability and generalization capabilities.

[0076] As an optional solution, the method further includes at least one of the following: performing a vector forward transformation operation on the initial text feature to obtain a first text feature vector represented by a one-dimensional vector, inputting the first text feature vector into a plurality of self-attention modules with different configuration parameters to obtain a plurality of second text feature vectors, performing a vector reverse transformation operation on the plurality of second text feature vectors to obtain a plurality of third text feature vectors represented by a multi-dimensional vector, performing a vector addition operation on the plurality of third text feature vectors to obtain text feature parameters, performing an addition operation on the dot product result of the text feature parameters and the initial text feature and the initial text feature to obtain a prompt text feature, wherein the first fusion The module includes a self-attention module; performing a vector forward transformation operation on the initial image feature to obtain a first image feature vector represented by a one-dimensional vector, inputting the first image feature vector into multiple self-attention modules with different configuration parameters to obtain multiple second image feature vectors, performing a vector reverse transformation operation on the multiple second image feature vectors to obtain multiple third image feature vectors represented by multi-dimensional vectors, performing a vector addition operation on the multiple third image feature vectors to obtain image feature parameters, performing an addition operation on the dot product result of the image feature parameters and the initial image feature and the initial image feature to obtain the image feature to be identified, wherein the second fusion module includes the self-attention module.

[0077] Optionally, in the embodiment of the present application, the above-mentioned vector forward conversion operation refers to the process of converting the initial feature into a one-dimensional vector representation, that is, Figure 5 The Flatten1 (forward transformation) stage in the model is designed to facilitate further processing of features and interaction between modules, including but not limited to feature dimensionality reduction and reorganization using operations such as linear transformation, pooling or fully connected layers.

[0078] Optionally, in an embodiment of the present application, the self-attention module included in the first fusion module and the second fusion module refers to a neural network component that performs feature learning and fusion based on the attention mechanism in the Transformer architecture, for example, Figure 5 The SSM module in , through the self-attention module with different configuration parameters, processes the features multiple times iteratively, thereby enhancing the representation ability of the features and the integration effect of multimodal information.

[0079] It should be noted that the configuration parameters of the self-attention module may include but are not limited to the number of attention heads, feature dimensions, hidden layer size, etc. The adjustment of these parameters can affect the granularity and depth of feature learning. The system can flexibly configure the parameters of the self-attention module according to specific task requirements and computing resources to optimize the feature processing process. At the same time, the specific implementation methods of the vector forward conversion and reverse conversion operations can be diversified, such as using linear layers, LSTM or other recurrent neural network units for sequence feature encoding and decoding, which is not limited in this application.

[0080] For example, through the vector forward conversion and reverse conversion operations, multimodal features can be effectively processed and converted to adapt to the input requirements of the self-attention module. At the same time, by configuring multiple self-attention modules, features can be deeply learned from different angles and levels, and finally, through vector addition operations and dot product results, prompt text features and image features to be identified with more semantic relevance and representation capabilities are generated.

[0081] Through the embodiments of the present application, continuously stacked self-attention modules are used for feature processing and fusion, which achieves deep optimization of initial text features and initial image features, significantly enhances the semantic expression of features and the representation ability of image content, and achieves the purpose of improving recognition accuracy and model generalization ability in attribute recognition and target re-identification tasks.

[0082] As an optional scheme, the target image feature is activated according to the semantic control weight vector to obtain the text activation image feature, including: determining the parameter at the first position in the semantic control weight vector as the first semantic control weight parameter, and determining the parameter at the second position in the semantic control weight vector as the second semantic control weight parameter, wherein the first position is before the second position; multiplying the first semantic control weight parameter and the target image feature element by element to obtain the initial activation feature; and adding the second semantic control weight parameter and the initial activation feature element by element to obtain the text activation image feature.

[0083] Optionally, in an embodiment of the present application, the above-mentioned process of activating target image features according to the semantically controlled weight vector refers to adjusting the weights of each part of the image features by performing mathematical operations on specific parameters in the weight vector and the target image features to enhance the expression of information related to the retrieval task, including but not limited to activating and enhancing features through element-by-element multiplication and addition operations.

[0084] It should be noted that the specific values ​​and modes of action of the first semantic control weight parameter and the second semantic control weight parameter can be diversified. For example, the first parameter can be used to enhance local information in the image features, while the second parameter is used to adjust the overall features, or the two parameters can act together on different aspects of the image features. It depends on the design strategy of the semantic control weight vector and the requirements of the target task, and this application does not limit this.

[0085] Exemplarily, the embodiments of the present application can be implemented through the following two key steps:

[0086] First, the first semantic control weight parameter is used to perform element-by-element multiplication with the target image feature, the purpose of which is to strengthen the part of the image feature that matches the text description attribute and generate the initial activation feature;

[0087] Subsequently, the second semantic control weight parameter is added element-by-element to the initial activation feature to further adjust the feature to ensure that the final text activation image feature not only contains local information related to the text attributes, but also takes into account the integrity of the global features of the image, thereby better supporting attribute recognition and target re-identification tasks.

[0088] In an exemplary embodiment, the application scenarios of pedestrian attribute retrieval and identity authentication are taken as an example, including but not limited to:

[0089] Firstly, the features of pedestrian images and text descriptions related to retrieval are extracted to generate a semantic control weight vector.

[0090] Next, the first semantic control weight parameter in the semantic control weight vector is multiplied element-by-element by the pedestrian image feature to highlight the attributes described in the image (e.g., the part wearing a gray top) to obtain the initial activation feature.

[0091] Then, by adding the second semantic control weight parameter in the semantic control weight vector to the initial activation feature element by element, the image features can be adjusted as a whole to ensure that even under complex conditions such as lighting and posture changes, the image features can still accurately reflect the attributes and identity information of the target object, and finally obtain the text-activated image features for attribute retrieval and target re-identification.

[0092] Through the embodiments of the present application, a hierarchical application mechanism of semantic control weight vectors is adopted to achieve fine adjustment and enhancement of target image features, which can more accurately capture and express the attribute information of the text description while maintaining the integrity and robustness of the image features, thereby achieving the purpose of improving the accuracy and efficiency of attribute recognition and target re-identification tasks.

[0093] As an optional solution, the method also includes: obtaining sample prompt text and sample to-be-recognized image; using a continuously stacked target fusion module to extract sample prompt text features of the sample prompt text, and extract sample to-be-recognized image features of the sample to-be-recognized image, multiplying the sample prompt text features and the sample to-be-recognized image features element by element to obtain sample fusion features, and performing a mapping operation on the sample fusion features to determine a sample semantic control weight vector, wherein the sample semantic control weight vector represents the weight value of each sample image part in the sample to-be-recognized image; activating the sample target image features according to the sample semantic control weight vector to obtain sample text activation image features, wherein the sample target image features are determined by the sample target image, and the sample target image is different from the sample to-be-recognized image; training an initial recognition network based on the sample text activation image features to obtain a target recognition network, wherein the target recognition network is used to perform at least one of an attribute recognition task and a target re-recognition task based on the prompt text, the image to be recognized, and the target image.

[0094] Optionally, in an embodiment of the present application, the above process of obtaining sample prompt text and sample images to be identified refers to extracting representative samples from a data set during a training phase for learning how to associate text descriptions with image content, including but not limited to extracting sample pairs related to attribute retrieval and target re-identification tasks from a large-scale labeled image data set.

[0095] Optionally, in an embodiment of the present application, the target fusion module processes the sample prompt text features and the sample image features to be identified by continuous stacking, aiming to build a deeper feature fusion and understanding capability, including but not limited to using the Mamba structure or the Transformer-based self-attention mechanism for feature integration.

[0096] It should be noted that the extraction methods of sample prompt text features and sample image features to be identified can be diversified. For example, pre-trained BERT models and ViT models can be used for feature extraction, or other neural network architectures or technical means, such as CNN, RNN, etc., can be combined to adapt to the needs of different tasks and data sets. This application does not limit this.

[0097] For example, this training process may include but is not limited to:

[0098] S1, the system selects sample pairs from the data set, including sample prompt text and sample images to be recognized.

[0099] S2, uses the target fusion module to extract and fuse features of these samples, generates sample fusion features, and determines the sample semantic control weight vector through mapping operation, which is used to represent the relative importance of different parts in the image.

[0100] S3, the weight vector is used to activate another set of sample target image features to obtain sample text activation image features.

[0101] S4, using sample text to activate image features to adjust the parameters of the initial recognition network, optimize the network performance through the training process, and obtain the final target recognition network.

[0102] In an exemplary embodiment, taking pedestrian attribute retrieval and identity authentication in a video surveillance system as an example:

[0103] S1, the system obtains sample prompt text (such as "male wearing glasses") and sample images to be identified (pedestrian images taken by a surveillance camera) from a training set containing pedestrian attribute labels and cross-temporal and spatial identity information;

[0104] S2, uses the target fusion module to extract and fuse features of these samples to generate sample fusion features, and determines the sample semantic control weight vector through mapping, which emphasizes the importance of glasses and male features in the image;

[0105] S3, using this weight vector to activate the features of another sample target image (i.e., the image that the system will perform attribute retrieval and identity authentication on), and obtain the sample text activation image features;

[0106] S4, based on these sample texts to activate image features, train the initial recognition network to learn how to perform attribute retrieval and target re-identification tasks based on text descriptions and image content. After training and network parameter optimization, a target recognition network that can effectively perform retrieval tasks is obtained.

[0107] Through the embodiments of the present application, a sample-based deep feature fusion and control weight learning mechanism is adopted to achieve the ability to automatically learn the association between text and images from training data, and to build a highly adaptive and generalized model that can not only handle predefined attribute retrieval, but can also be extended to unseen attribute descriptions and cross-temporal and spatial target re-identification tasks, thereby achieving the purpose of greatly improving the accuracy and flexibility of the recognition network in multi-attribute retrieval and identity authentication in complex scenarios.

[0108] As an optional scheme, an initial recognition network is trained based on sample text activated image features to obtain a target recognition network, including: using a cross entropy loss function to calculate text loss values ​​of sample text activated image features and sample initial text features, and using a triple loss function to calculate image loss values ​​of sample text activated image features and sample initial image features, wherein the sample initial text features are text features of sample prompt text extracted by a text feature extraction network model, and the sample initial image features are image features of the sample to be recognized extracted by a visual feature extraction network model; when the text loss value satisfies a first convergence condition and the image loss value satisfies a second convergence condition, the initial recognition network is determined as a target recognition network.

[0109] Optionally, in an embodiment of the present application, the above-mentioned process of training the initial recognition network based on sample text activated image features refers to the process of adjusting the initial recognition network parameters to minimize the loss function using the sample text activated image features as supervision signals, including but not limited to using the cross entropy loss function and the triplet loss function Triplet loss to quantify the difference between the model prediction and the true label, and updating the network parameters.

[0110] Optionally, in an embodiment of the present application, the use of the above-mentioned cross entropy loss function and triple loss function refers to the specific quantification of the prediction deviation and similarity loss between the sample text activation image features and the sample initial text features and the sample initial image features to guide the training of the model, including but not limited to optimizing the accuracy of the attribute classification task through the cross entropy loss function and improving the accuracy of the target re-identification task through the triple loss function.

[0111] It should be noted that the specific settings of the first convergence condition and the second convergence condition are diverse. For example, the loss function value may be lower than a certain threshold, the loss value change rate may be less than a threshold within several consecutive iteration cycles, or a predetermined training round may be reached, etc. It depends on the model training objectives and performance requirements, and this application does not limit this.

[0112] For example, the system iteratively adjusts the network parameters by calculating the text loss value and the image loss value until the two types of loss values ​​meet the convergence conditions, thereby determining the final target recognition network. During the training process, the model continuously learns how to make more accurate predictions in attribute recognition and target re-identification tasks through sample text activation image features and sample initial text features and sample initial image features, ultimately enabling the model to demonstrate higher recognition performance when processing unknown data.

[0113] In an exemplary embodiment, taking pedestrian attribute retrieval and identity authentication in security monitoring as an example, the system first generates sample text activation image features through a training data set. These features contain a deep association between the text description and the image content. Then, the cross-entropy loss function is used to calculate the classification error between the activation features and the sample initial text features (i.e., the true attribute label), and the triple loss function is used to evaluate the similarity loss between the activation features and the sample initial image features (the same pedestrian image from different cameras). During the training process, when the classification error and similarity loss reach the preset convergence criteria respectively, the model training is completed, and the obtained target recognition network can achieve high accuracy and stability in attribute retrieval and target re-identification tasks.

[0114] Through the embodiments of the present application, a joint optimization strategy of cross entropy loss and triple loss is adopted to achieve collaborative training and performance improvement of the model in attribute recognition and target re-identification tasks, and it is possible to build a target recognition network that has a deep understanding of multimodal information. Even in the face of complex scenarios and open vocabularies, it can accurately perform attribute classification and identity authentication, thereby achieving the purpose of improving the practicality, adaptability and generalization ability of the model.

[0115] As an optional solution, the method also includes: obtaining a prompt text and an image to be recognized; extracting initial text features of the prompt text using a text feature extraction network model; segmenting the image to be recognized into a group of image regions using a visual feature extraction network model, and performing a linear mapping operation on the group of image regions to obtain initial image features; processing the initial text features using a continuously stacked first fusion module to determine the prompt text features, and processing the initial image features using a continuously stacked second fusion module to determine the image features to be recognized, wherein the target fusion module includes the first fusion module and the second fusion module; element-by-element multiplication of the prompt text features and the image features to be recognized to obtain fused features, and performing A mapping operation is performed to determine a semantic control weight vector; a parameter at a first position in the semantic control weight vector is determined as a first semantic control weight parameter, and a parameter at a second position in the semantic control weight vector is determined as a second semantic control weight parameter, wherein the first position is before the second position; the first semantic control weight parameter and the target image feature are multiplied element by element to obtain an initial activation feature, wherein the target image feature is obtained by processing the target image with a visual feature extraction network model; the second semantic control weight parameter and the initial activation feature are added element by element to obtain a text activation image feature; attribute recognition tasks and target re-recognition tasks are respectively performed based on the text activation image feature to generate a target recognition result.

[0116] It should be noted that attribute retrieval (the above-mentioned attribute recognition task) is usually used to search for people with a certain type of features, such as men wearing black tops, people riding electric bikes, etc., while ReID (the above-mentioned target re-identification task) is usually used to retrieve the same target object at different times and places. Conventional retrieval systems usually use two deep learning models to implement attribute retrieval and ReID functions respectively. It is impossible to unify the deployment of the two tasks into one system, and attribute retrieval generally requires pre-defined attribute names, such as top color, gender, age, etc., and only pre-defined attribute results can be output during retrieval, which belongs to a closed set retrieval. The ReID model only uses image information as input. There is a lack of a general framework that supports attribute retrieval and ReID in the related technology.

[0117] Based on this, the embodiment of the present application improves the network structure, unifies the training and deployment of attribute recognition tasks and ReID tasks, introduces a text-image fusion prompt structure, supports the processing of multimodal information of images and texts, and specifies attributes through texts without pre-defining attribute values ​​to achieve open set retrieval. Specifically, a deep neural network based on Mamba and Transformer is used as a feature extraction network, wherein the visual feature extraction network uses ViT (or Swin Transformer and CNN), and the text feature extraction network uses BERT (or other text encoders). The overall framework draws on CLIP (processing cross-modal tasks) and adds a Mamba-based image-text fusion structure, which outputs a vector for activating ViT intermediate features (the above-mentioned target image features) to support ViT in extracting text-oriented image features based on text prompts. This feature can be used for attribute retrieval and ReID tasks (person re-identification, but also widely used in other fields such as vehicle re-identification, etc.).

[0118] For example, the embodiments of the present application may include but are not limited to the following steps:

[0119] S1, obtaining the image to be identified, which can be an image of a pedestrian, a non-motor vehicle or other target. In the video surveillance scenario, the image to be identified generally comes from the output of the target detection network;

[0120] S2, image preprocessing and enhancement. Preprocess the input image to be recognized, including normalization, resizing, etc. Data enhancement is required during training, including translation, rotation, etc.;

[0121] S3, extract text features. Enter the prompt text related to attribute retrieval, such as "the color of the top is gray, the color of the pants is black, and the gender is female". The text is processed by the BERT network to obtain text features, corresponding to Figure 1 The TextFeatures (numeric vectors that capture the semantic information of words in the text and their relationships) in , the dimension is TC, where T is the length of the text, is the number of words or characters in the sentence, is the number of tokens, C is the number of feature channels, which can be 1024, and C is the length of the embedding vector of each token.

[0122] S4, obtains the semantic control weight vector through the image to be recognized and the prompt text, corresponding to Figure 3 The structure of Mamba Controller is as follows Figure 4As shown. The initial input image resolution is H0W0C0, where H0 is the image height, W0 is the image width, and C is the number of image channels. The image is divided into patches and linearly mapped to obtain the features of H`W`C (for details, please refer to the original ViT processing method. ViT divides the input image into non-overlapping areas of fixed size, each of which is called a "patch". For example, a 224-pixel image can be divided into 14 16-pixel patches, where 14 is the number of patches the image is divided into horizontally and vertically, and 16 is the size of each patch);

[0123] Next, four consecutive stacked Mamba Blocks are used to extract deep feature information. The TextFeature obtained in step S3 is also processed in the same way. The structure of Image Mamba Block and Text Mamba Block can adopt the classic Vision Mamba structure. For the specific structure, refer to Figure 5 The image features and text features (the deep feature information mentioned above) output by Mamba Block are multiplied element by element, averaged and linearly mapped to obtain the parameters of the semantic controller, corresponding to Figure 2 Control weights in , where D is the number of feature channels of Control weights, which needs to be the same as the number of channels C of image features, and can be 1024.

[0124] S5, extracts the image features of another target image (different from the above target image), the purpose of which is to compare the similarity between the target image and the above target image, complete the REID task or attribute recognition task, and use the semantic control weight (the parameters of the above semantic controller) for activation. The image feature extraction network adopts the classic ViT structure. The specific operation of using semantic control weights to activate image features is:

[0125] First, the first semantic control parameter is element-wise multiplied with the target image feature to obtain the initial activation feature.

[0126] Next, the initial activation features are added element-by-element to the second semantic control parameter, and the output is the activated image features that incorporate text information, with the same dimension as the input image features.

[0127] Finally, the text activation image features output by ViT are obtained, and the dimension is H1W1C.

[0128] S6, attribute retrieval and ReID downstream tasks. The text-activated image features extracted in step S5 can be used for both ReID tasks and attribute recognition tasks.

[0129] On the one hand, in the ReID task, Triplet loss is used (a loss function commonly used in contrastive learning or metric learning, mainly used to train the model to learn how to distinguish different entities or instances in the feature space. In the ReID (person re-identification) task, the goal is to enable the model to recognize the same person from images taken by different cameras or at different times).

[0130] On the other hand, in the attribute recognition task, the image features and text features can be used to calculate the cross entropy loss (CLIP uses images and corresponding text descriptions to train the model in the pre-training stage, so that the image and text features can be correctly matched in the same space. In the attribute recognition task, the model uses image features and attribute-related text descriptions (such as "black top") to calculate the cross entropy loss, guiding the model to learn how to map image features to the correct attribute categories).

[0131] S7, the difference between training and inference. During training, images come from the training set images, text comes from the text labels of the attribute recognition dataset, and the attribute recognition task and the ReID task are trained synchronously. During inference, the text prompts come from the text information entered by the operator.

[0132] Through the embodiments of the present application, a joint training and deployment framework for attribute retrieval and ReID tasks is implemented. The network architecture is based on Mamba and Transformer. The framework supports open vocabulary attribute recognition and can use text to enhance the ReID effect. An image-text prompt fusion module based on the Mamba structure (the above-mentioned target fusion module) is also implemented. The semantic control weights output by this module can be used to activate relevant information in image features. The obtained text-activated image features are adaptive to downstream tasks and there is no need to re-fine-tune the training for downstream tasks.

[0133] It is understandable that in the specific implementation of this application, related data such as user information is involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions.

[0134] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.

[0135] According to another aspect of the embodiments of the present application, a device for identifying an object for implementing the above-mentioned object identification method is also provided. Figure 6 As shown, the device comprises:

[0136] An acquisition module 602 is used to acquire a prompt text and an image to be recognized;

[0137] A fusion module 604 is used to extract the prompt text features of the prompt text and the image features to be recognized of the image to be recognized by using the continuously stacked target fusion modules, fuse the prompt text features and the image features to be recognized to obtain fused features, and perform a mapping operation on the fused features to determine a semantic control weight vector, wherein the semantic control weight vector is used to represent the weight value of each image part in the image to be recognized, and the target fusion module represents a module for processing the prompt text features and the image features to be recognized based on the state space model;

[0138] An activation module 606 is used to activate the target image feature using the semantic control weight vector to obtain a text activated image feature, wherein the target image feature is determined by the target image, the target image is different from the image to be recognized, and the text activated image feature represents the image feature of the target image fused with the prompt text feature;

[0139] A generation module 608 is used to perform attribute recognition tasks and target re-recognition tasks based on text-activated image features to generate target recognition results, wherein the target recognition results are used to indicate whether the target image includes a target object, the image to be recognized includes a target object, the prompt text is used to indicate the target object attributes of the target object, the target re-recognition task is used to identify target objects in different time and space in the target image, and the attribute recognition task is used to determine that the object attributes in the target image are target objects whose attributes are target object attributes.

[0140] As an optional solution, the above-mentioned device is also used to: obtain prompt text and an image to be recognized; use a text feature extraction network model to extract initial text features of the prompt text; use a visual feature extraction network model to divide the image to be recognized into a group of image regions, and perform a linear mapping operation on the group of image regions to obtain initial image features; use a continuously stacked first fusion module to process the initial text features to determine prompt text features, and use a continuously stacked second fusion module to process the initial image features to determine image features to be recognized, wherein the target fusion module includes a first fusion module and a second fusion module; multiply the prompt text features and the image features to be recognized element by element to obtain fused features, and perform a linear mapping operation on the fused features. A row mapping operation is performed to determine a semantic control weight vector; a parameter at a first position in the semantic control weight vector is determined as a first semantic control weight parameter, and a parameter at a second position in the semantic control weight vector is determined as a second semantic control weight parameter, wherein the first position is before the second position; the first semantic control weight parameter and the target image feature are multiplied element by element to obtain an initial activation feature, wherein the target image feature is obtained by processing the target image with a visual feature extraction network model; the second semantic control weight parameter and the initial activation feature are added element by element to obtain a text activation image feature; attribute recognition tasks and target re-recognition tasks are respectively performed based on the text activation image feature to generate a target recognition result.

[0141] As an optional solution, the above-mentioned device is used to extract prompt text features of the prompt text and extract image features to be identified of the image to be identified by using continuously stacked target fusion modules in the following manner: using a text feature extraction network model to extract initial text features of the prompt text; using a visual feature extraction network model to segment the image to be identified into a group of image regions, and performing a linear mapping operation on a group of image regions to obtain initial image features; using a continuously stacked first fusion module to process the initial text features to determine the prompt text features, and using a continuously stacked second fusion module to process the initial image features to determine the image features to be identified, wherein the target fusion module includes a first fusion module and a second fusion module.

[0142] As an optional solution, the above-mentioned device is also used for at least one of the following: performing a vector forward transformation operation on the initial text feature to obtain a first text feature vector represented by a one-dimensional vector, inputting the first text feature vector into a plurality of self-attention modules with different configuration parameters to obtain a plurality of second text feature vectors, performing a vector reverse transformation operation on the plurality of second text feature vectors to obtain a plurality of third text feature vectors represented by a multi-dimensional vector, performing a vector addition operation on the plurality of third text feature vectors to obtain text feature parameters, performing an addition operation on the dot product result of the text feature parameters and the initial text feature and the initial text feature to obtain a prompt text feature, wherein the first fusion The fusion module includes a self-attention module; performing a vector forward transformation operation on the initial image feature to obtain a first image feature vector represented by a one-dimensional vector, inputting the first image feature vector into multiple self-attention modules with different configuration parameters to obtain multiple second image feature vectors, performing a vector reverse transformation operation on the multiple second image feature vectors to obtain multiple third image feature vectors represented by multi-dimensional vectors, performing a vector addition operation on the multiple third image feature vectors to obtain image feature parameters, performing an addition operation on a dot product result of the image feature parameters and the initial image feature and the initial image feature to obtain an image feature to be identified, wherein the second fusion module includes a self-attention module.

[0143] As an optional scheme, the above-mentioned device is used to activate the target image feature according to the semantic control weight vector in the following manner to obtain the text activation image feature: determine the parameter at the first position in the semantic control weight vector as the first semantic control weight parameter, and determine the parameter at the second position in the semantic control weight vector as the second semantic control weight parameter, wherein the first position is before the second position; multiply the first semantic control weight parameter and the target image feature element by element to obtain the initial activation feature; and add the second semantic control weight parameter and the initial activation feature element by element to obtain the text activation image feature.

[0144] As an optional solution, the above-mentioned device is also used to: obtain sample prompt text and sample to-be-recognized image; use continuously stacked target fusion modules to extract sample prompt text features of the sample prompt text, and extract sample to-be-recognized image features of the sample to-be-recognized image, multiply the sample prompt text features and the sample to-be-recognized image features element by element to obtain sample fusion features, and perform mapping operations on the sample fusion features to determine a sample semantic control weight vector, wherein the sample semantic control weight vector represents the weight values ​​of each sample image part in the sample to-be-recognized image; activate the sample target image features according to the sample semantic control weight vector to obtain sample text activation image features, wherein the sample target image features are determined by the sample target image, and the sample target image is different from the sample to-be-recognized image; train an initial recognition network based on the sample text activation image features to obtain a target recognition network, wherein the target recognition network is used to perform at least one of an attribute recognition task and a target re-recognition task based on the prompt text, the image to be recognized and the target image.

[0145] As an optional solution, the above-mentioned device is used to train an initial recognition network based on sample text activated image features to obtain a target recognition network in the following manner: use a cross entropy loss function to calculate the text loss values ​​of the sample text activated image features and the sample initial text features, and use a triple loss function to calculate the image loss values ​​of the sample text activated image features and the sample initial image features, wherein the sample initial text features are text features of the sample prompt text extracted by a text feature extraction network model, and the sample initial image features are image features of the sample to be recognized extracted by a visual feature extraction network model; when the text loss value satisfies the first convergence condition and the image loss value satisfies the second convergence condition, the initial recognition network is determined as the target recognition network.

[0146] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0147] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0148] According to one aspect of the present application, a computer program product is provided. The computer program product includes a computer program.

[0149] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0150] Figure 7 The structure block diagram of a computer system for implementing an electronic device according to an embodiment of the present application is schematically shown.

[0151] It should be noted that Figure 7 The computer system 700 of the electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0152] like Figure 7 As shown, the computer system 700 includes a central processing unit 701 (CPU), which can perform various appropriate actions and processes according to the program stored in the read-only memory 702 (ROM) or the program loaded from the storage part 708 to the random access memory 703 (RAM). Various programs and data required for system operation are also stored in the random access memory 703. The central processing unit 701, the read-only memory 702 and the random access memory 703 are connected to each other through a bus 704. The input / output interface 705 (Input / Output interface, i.e., I / O interface) is also connected to the bus 704.

[0153] The following components are connected to the input / output interface 705: an input section 706 including a keyboard, a mouse, etc.; an output section 707 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, a modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the input / output interface 705 as needed. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 710 as needed so that a computer program read therefrom is installed into the storage section 708 as needed.

[0154] In particular, according to an embodiment of the present application, the process described in each method flow chart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer readable medium, and the computer program contains a program code for executing the method shown in the flow chart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 709, and / or installed from the removable medium 711. When the computer program is executed by the central processor 701, various functions defined in the system of the present application are executed.

[0155] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 709, and / or installed from the removable medium 711. When the computer program is executed by the central processor 701, various functions provided by the embodiment of the present application are performed.

[0156] According to another aspect of the embodiment of the present application, an electronic device for implementing the above-mentioned object recognition method is also provided. The electronic device may be Figure 1 The terminal device or server shown in the figure. This embodiment is described by taking the electronic device as a terminal device as an example. Figure 8 As shown, the electronic device includes a memory 802 and a processor 804. The memory 802 stores a computer program, and the processor 804 is configured to execute the steps in any of the above method embodiments through the computer program.

[0157] Optionally, in this embodiment, the electronic device may be located in at least one network device among a plurality of network devices of a computer network.

[0158] Optionally, in this embodiment, the above-mentioned processor can be configured to execute the methods in each embodiment of the present application through a computer program.

[0159] Alternatively, a person skilled in the art may understand that: Figure 8 The structure shown is for illustration only. Figure 8 The structure of the electronic device is not limited. Figure 8 More or fewer components (such as network interfaces, etc.) as shown in, or with Figure 8 Different configurations are shown.

[0160] Among them, the memory 802 can be used to store software programs and modules, such as the program instructions / modules corresponding to the object recognition method and device in the embodiment of the present application. The processor 804 executes various functional applications and data processing by running the software programs and modules stored in the memory 802, that is, realizing the above-mentioned object recognition method. The memory 802 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 802 may further include a memory remotely located relative to the processor 804, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. Among them, the memory 802 can be specifically, but not limited to, used to store information such as text-activated image features. As an example, such as Figure 8 As shown, the memory 802 may include but is not limited to the acquisition module 602, fusion module 604, activation module 606 and generation module 608 in the object recognition device. In addition, it may also include but is not limited to other module units in the object recognition device, which will not be repeated in this example.

[0161] Optionally, the transmission device 806 is used to receive or send data via a network. Specific examples of the network may include a wired network and a wireless network. In one example, the transmission device 806 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices and routers via a network cable so as to communicate with the Internet or a local area network. In one example, the transmission device 806 is a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0162] In addition, the electronic device further includes: a display 808 for displaying the prompt text and the image to be recognized, the target image, etc.; and a connection bus 810 for connecting the various module components in the electronic device.

[0163] In other embodiments, the terminal device or server may be a node in a distributed system, wherein the distributed system may be a blockchain system, and the blockchain system may be a distributed system formed by connecting the multiple nodes through network communication. The nodes may form a peer-to-peer network, and any form of computing device, such as a server, terminal or other electronic device, may become a node in the blockchain system by joining the peer-to-peer network.

[0164] According to one aspect of the present application, a computer-readable storage medium is provided, and a processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the object recognition method provided in various optional implementations of the above-mentioned object recognition aspects.

[0165] Optionally, in this embodiment, the above-mentioned computer-readable storage medium can be configured to store data for executing the methods in various embodiments of the present application.

[0166] Optionally, in this embodiment, a person of ordinary skill in the art may understand that all or part of the steps in the various methods of the above embodiments may be completed by instructing hardware related to the terminal device through a program, and the program may be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0167] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0168] If the integrated units in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in the above computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for one or more electronic devices to execute all or part of the steps of the methods described in each embodiment of the present application.

[0169] In the above embodiments of the present application, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0170] In the several embodiments provided in the present application, it should be understood that the disclosed application can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0171] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0172] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0173] The above is only a preferred implementation of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A method for identifying an object, characterized in that: include: Get the prompt text and the image to be recognized; Extracting the prompt text features of the prompt text and the image features to be recognized of the image to be recognized by using a continuously stacked target fusion module, fusing the prompt text features and the image features to be recognized to obtain a fused feature, and performing a mapping operation on the fused feature to determine a semantic control weight vector, wherein the semantic control weight vector is used to represent the weight value of each image part in the image to be recognized, and the target fusion module represents a module for processing the prompt text features and the image features to be recognized based on a state space model; Performing element-by-element multiplication and element-by-element addition on the semantic control weight vector and the target image feature to activate the target image feature to obtain a text activation image feature, wherein the target image feature is determined by a target image, the target image is different from the image to be recognized, and the text activation image feature represents an image feature of the target image fused with the prompt text feature; A target recognition network is used to respectively perform an attribute recognition task and a target re-recognition task based on the text-activated image features to generate a target recognition result, wherein the target recognition result is used to indicate whether the target image includes a target object, the image to be recognized includes the target object, the prompt text is used to indicate the target object attribute of the target object, the target re-recognition task is used to recognize the target object in different time and space in the target image, the attribute recognition task is used to determine the target object whose object attribute in the target image is the target object attribute, and the target recognition network is obtained by training an initial recognition network based on sample prompt text, sample image to be recognized, and sample text-activated image features.

2. The method according to claim 1, characterized in that: The extracting of the prompt text features of the prompt text and the extracting of the image features to be recognized of the image to be recognized by using the continuously stacked target fusion modules includes: Extracting initial text features of the prompt text using a text feature extraction network model; Using a visual feature extraction network model to segment the image to be identified into a group of image regions, and performing a linear mapping operation on the group of image regions to obtain initial image features; The initial text features are processed using a first fusion module stacked continuously to determine the prompt text features, and the initial image features are processed using a second fusion module stacked continuously to determine the image features to be identified, wherein the target fusion module includes the first fusion module and the second fusion module.

3. The method according to claim 2, characterized in that The method further comprises at least one of the following: Performing a vector forward transformation operation on the initial text feature to obtain a first text feature vector represented by a one-dimensional vector, inputting the first text feature vector into a plurality of self-attention modules with different configuration parameters to obtain a plurality of second text feature vectors, performing a vector reverse transformation operation on the plurality of second text feature vectors to obtain a plurality of third text feature vectors represented by a multi-dimensional vector, performing a vector addition operation on the plurality of third text feature vectors to obtain text feature parameters, performing the addition operation on a dot product result of the text feature parameters and the initial text feature and the initial text feature to obtain the prompt text feature, wherein the first fusion module includes the self-attention module; A vector forward transformation operation is performed on the initial image feature to obtain a first image feature vector represented by a one-dimensional vector, the first image feature vector is respectively input into a plurality of self-attention modules with different configuration parameters to obtain a plurality of second image feature vectors, a vector reverse transformation operation is respectively performed on the plurality of second image feature vectors to obtain a plurality of third image feature vectors represented by a multi-dimensional vector, a vector addition operation is performed on the plurality of third image feature vectors to obtain image feature parameters, the dot product result of the image feature parameters and the initial image feature is performed on the initial image feature, and the addition operation is performed to obtain the image feature to be identified, wherein the second fusion module includes the self-attention module.

4. The method according to claim 1, characterized in that The step of activating the target image feature according to the semantic control weight vector to obtain the text activated image feature comprises: Determine a parameter at a first position in the semantic control weight vector as a first semantic control weight parameter, and determine a parameter at a second position in the semantic control weight vector as a second semantic control weight parameter, wherein the first position is before the second position; Multiplying the first semantic control weight parameter and the target image feature element by element to obtain an initial activation feature; The second semantic control weight parameter and the initial activation feature are added element by element to obtain the text activation image feature.

5. The method according to claim 1, characterized in that: The method further comprises: Obtaining the sample prompt text and the sample image to be recognized; Using the target fusion modules stacked in succession to extract sample prompt text features of the sample prompt text, and extract sample image features of the sample image to be recognized, multiplying the sample prompt text features and the sample image features to be recognized element by element to obtain sample fusion features, and performing the mapping operation on the sample fusion features to determine a sample semantic control weight vector, wherein the sample semantic control weight vector represents the weight value of each sample image part in the sample image to be recognized; Activating the sample target image feature according to the sample semantic control weight vector to obtain the sample text activation image feature, wherein the sample target image feature is determined by a sample target image, and the sample target image is different from the sample to-be-recognized image; The initial recognition network is trained based on the sample text to activate the image features to obtain the target recognition network, wherein the target recognition network is used to perform at least one of the attribute recognition task and the target re-recognition task based on the prompt text, the image to be recognized and the target image.

6. The method according to claim 5, characterized in that The initial recognition network is trained based on the sample text activated image features to obtain a target recognition network, including: Using a cross entropy loss function to calculate the text loss value of the sample text activation image feature and the sample initial text feature, and using a triple loss function to calculate the image loss value of the sample text activation image feature and the sample initial image feature, wherein the sample initial text feature is the text feature of the sample prompt text extracted by a text feature extraction network model, and the sample initial image feature is the image feature of the sample to be recognized extracted by a visual feature extraction network model; When the text loss value satisfies a first convergence condition and the image loss value satisfies a second convergence condition, the initial recognition network is determined as the target recognition network.

7. The method according to claim 1, characterized in that The method further comprises: Obtaining the prompt text and the image to be recognized; Extracting initial text features of the prompt text using a text feature extraction network model; Using a visual feature extraction network model to segment the image to be identified into a group of image regions, and performing a linear mapping operation on the group of image regions to obtain initial image features; Using a first fusion module stacked continuously to process the initial text features, determine the prompt text features, and using a second fusion module stacked continuously to process the initial image features, determine the image features to be identified, wherein the target fusion module includes the first fusion module and the second fusion module; Multiplying the prompt text feature and the to-be-recognized image feature element by element to obtain the fused feature, and performing a mapping operation on the fused feature to determine the semantic control weight vector; Determine a parameter at a first position in the semantic control weight vector as a first semantic control weight parameter, and determine a parameter at a second position in the semantic control weight vector as a second semantic control weight parameter, wherein the first position is before the second position; Multiplying the first semantic control weight parameter and the target image feature element by element to obtain an initial activation feature, wherein the target image feature is obtained by processing the target image by the visual feature extraction network model; Adding the second semantic control weight parameter and the initial activation feature element by element to obtain the text activation image feature; The attribute recognition task and the target re-recognition task are respectively performed based on the text-activated image features to generate the target recognition result.

8. An object recognition device, characterized in that: include: An acquisition module is used to acquire the prompt text and the image to be recognized; A fusion module, used for extracting the prompt text features of the prompt text and the image features to be recognized of the image to be recognized by using the continuously stacked target fusion modules, fusing the prompt text features and the image features to be recognized to obtain fused features, and performing a mapping operation on the fused features to determine a semantic control weight vector, wherein the semantic control weight vector is used to represent the weight value of each image part in the image to be recognized, and the target fusion module represents a module for processing the prompt text features and the image features to be recognized based on a state space model; an activation module, configured to perform element-by-element multiplication and element-by-element addition on the semantic control weight vector and the target image feature to activate the target image feature and obtain a text activation image feature, wherein the target image feature is determined by a target image, the target image is different from the image to be recognized, and the text activation image feature represents an image feature of the target image fused with the prompt text feature; A generation module is used to use a target recognition network to respectively perform an attribute recognition task and a target re-recognition task based on the text activation image feature to generate a target recognition result, wherein the target recognition result is used to indicate whether the target image includes a target object, the image to be recognized includes the target object, the prompt text is used to indicate the target object attribute of the target object, the target re-recognition task is used to recognize the target object in the target image at different time and space, the attribute recognition task is used to determine the target object whose object attribute in the target image is the target object attribute, and the target recognition network is obtained by training an initial recognition network based on sample prompt text, sample image to be recognized, and sample text activation image features.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored computer program, wherein the computer program can be executed by an electronic device to perform the method described in any one of claims 1 to 7.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 7 are implemented.

11. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 7 through the computer program.

Citation Information

Patent Citations

  • Data processing method and device, electronic equipment and storage medium

    CN115115913A

  • Target detection method and device

    CN117392379A