Artificial intelligence-based target recognition method and model training method and device
By employing a multi-head attention mechanism for target recognition, combined with text and image feature extraction networks, efficient pet identification was achieved, solving the problem of high identification difficulty in pet management and improving the convenience and accuracy of identification.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-30
- Publication Date
- 2026-03-24
AI Technical Summary
When implementing intelligent management of pets in cities, the diverse range of pet species and their low visual differentiation make it costly to build a large database of identity features. Furthermore, it is difficult to collect facial and limb features of pets, making identification challenging.
A target recognition method based on multi-head attention mechanism is adopted. By extracting and fusing features from the target text and the initial image, a deep learning model is used to determine the target image of the pet, including a text feature extraction network, an image feature extraction network, a fusion network and a recognition network, so as to realize the pet identification.
It improves the convenience and accuracy of pet identification, reduces the difficulty of object identification in urban scenarios, and simplifies the pet identification process.
Smart Images

Figure CN115909357B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the field of artificial intelligence, in particular to the field of image recognition and video analysis, and can be applied to smart city, city governance, emergency management and the like. More specifically, the present disclosure provides a target recognition method, a training method of a target recognition model, an apparatus, an electronic device and a storage medium. BACKGROUND
[0002] With the development of artificial intelligence technology, images or videos collected by video collection devices can be recognized to determine the position and category of objects in the images and videos. SUMMARY
[0003] The present disclosure provides a target recognition method, a training method of a target recognition model, an apparatus, an electronic device and a storage medium.
[0004] According to an aspect of the present disclosure, a target recognition method is provided, which includes: in response to obtaining a target text, performing feature extraction on the target text to obtain a target text feature, wherein the target text is related to a target object; performing feature extraction on at least one initial image related to the target text to obtain at least one initial image feature; obtaining at least one query feature, at least one key feature and at least one value feature according to the target text feature and the at least one initial image feature; fusing the at least one query feature, the at least one key feature and the at least one value feature to obtain at least one target fusion feature; determining at least one recognition result corresponding to the at least one initial image according to the at least one target fusion feature; and determining a target image related to the target object from the at least one initial image according to the at least one recognition result.
[0005] According to another aspect of the present disclosure, a training method of a target recognition model is provided, the target recognition model including an image feature extraction network, a text feature extraction network, a fusion network and a recognition network, the method including: inputting a sample text into the text feature extraction network to obtain a sample text feature, wherein the sample text is related to a sample object; inputting a sample image into the image feature extraction network to obtain a sample image feature; obtaining a query feature, a key feature and a value feature according to the sample text feature and the sample image feature; inputting the query feature, the key feature and the value feature into the fusion network to obtain a sample fusion feature; inputting the sample fusion feature into the recognition network to obtain a sample recognition result corresponding to the sample image; and training the image recognition model according to a label of the sample image and the sample recognition result.
[0006] According to another aspect of the present disclosure, there is provided a target identification device, comprising: a first feature extraction model configured to perform feature extraction on a target text to obtain target text features in response to obtaining the target text, wherein the target text is related to a target object; a second feature extraction module configured to perform feature extraction on at least one initial image related to the target text to obtain at least one initial image feature; a first obtaining module configured to obtain at least one query feature, at least one key feature and at least one value feature according to the target text features and the at least one initial image feature; a fusion module configured to fuse the at least one query feature, the at least one key feature and the at least one value feature to obtain at least one target fusion feature; a first determining module configured to determine at least one identification result corresponding to the at least one initial image according to the at least one target fusion feature; and a second determining module configured to determine a target image related to the target object from the at least one initial image according to the at least one identification result.
[0007] According to another aspect of the present disclosure, there is provided a training device of a target identification model, the target identification model comprising an image feature extraction network, a text feature extraction network, a fusion network and an identification network, the device comprising: a second obtaining module configured to input a sample text into the text feature extraction network to obtain sample text features, wherein the sample text is related to a sample object; a third obtaining module configured to input a sample image into the image feature extraction network to obtain sample image features; a fourth obtaining module configured to obtain query features, key features and value features according to the sample text features and the sample image features; a fifth obtaining module configured to input the query features, the key features and the value features into the fusion network to obtain sample fusion features; a sixth obtaining module configured to input the sample fusion features into the identification network to obtain a sample identification result corresponding to the sample image; and a training module configured to train the image identification model according to a label of the sample image and the sample identification result.
[0008] According to another aspect of the present disclosure, there is provided an electronic device, comprising: at least one processor; and a memory communicatively connected with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method provided by the present disclosure.
[0009] According to another aspect of the present disclosure, there is provided a non-transitory computer readable storage medium storing computer instructions for causing a computer to perform the method provided by the present disclosure.
[0010] According to another aspect of the present disclosure, there is provided a computer program product comprising a computer program which, when executed by a processor, implements the method provided by the present disclosure.
[0011] It should be understood that the contents described in this part are not intended to identify key or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0012] The accompanying drawings are used to better understand the present scheme and do not limit the present disclosure. Among them:
[0013] Figure 1 is an exemplary system architecture schematic diagram to which the target recognition method and device according to one embodiment of the present disclosure can be applied;
[0014] Figure 2 is a flowchart of a target recognition method according to one embodiment of the present disclosure;
[0015] Figure 3 is a schematic diagram of a target recognition model according to one embodiment of the present disclosure;
[0016] Figure 4A is a schematic diagram of an initial image according to one embodiment of the present disclosure;
[0017] Figure 4B is a schematic diagram of a recognition result according to one embodiment of the present disclosure;
[0018] Figure 5 is a flowchart of a training method of a target recognition model according to another embodiment of the present disclosure;
[0019] Figure 6 is a schematic diagram of a training method of a target recognition model according to one embodiment of the present disclosure;
[0020] Figure 7 is a block diagram of a target recognition device according to one embodiment of the present disclosure;
[0021] Figure 8 is a block diagram of a training device of a target recognition model according to one embodiment of the present disclosure; and
[0022] Figure 9 is a block diagram of an electronic device to which a target recognition method according to one embodiment of the present disclosure can be applied. DETAILED DESCRIPTION
[0023] Exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings, which include various details of the embodiments of the present disclosure to help understanding, and should be considered as merely exemplary. Thus, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Also, descriptions of well-known functions and structures are omitted from the following description for clarity and conciseness.
[0024] With the continuous development of urbanization, super cities gradually highlight the characteristics of population intensification. Based on the widely deployed cameras in the city, computer vision technology can be used to determine the face and body features to facilitate city management and control. At the same time, pets have gradually become an element that cannot be ignored in the city. However, the intelligent management scheme for pets is still in its infancy. Fine intelligent management of pets helps to improve the city's appearance and improve the city's public health.
[0025] In the process of pedestrian identity recognition, computer vision technology can be used to collect a large number of face or body pictures for feature extraction to construct an identity data library. In the process of recognition, the identity of the object in the collected image is determined according to the similarity between the features of the collected image and the features in the identity data library.
[0026] In the process of pet recognition, computer vision technology can also be used to perform category detection on the collected image, and then extract features based on the detection frame and perform feature similarity calculation with the features in the pet identity data library to perform pet identity recognition.
[0027] However, the species of pets are diverse and the visual distinction is low, so it is costly to construct a large identity feature library for pet recognition. In addition, when establishing the pet identity feature library, the cooperation degree of the pet is low, and the face and body features of the pet are difficult to collect.
[0028] Figure 1 is an exemplary system architecture schematic diagram according to an embodiment of the present disclosure, which can apply the target recognition method and device. It should be noted that, Figure 1 The figure shown is only an example of a system architecture that can apply the embodiments of the present disclosure to help those skilled in the art understand the technical content of the present disclosure, but does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios.
[0029] As Figure 1 shown, the system architecture 100 according to the embodiment can include terminal devices 101, 102, 103, a network 104 and a server 105. The network 104 is used to provide a communication link medium between the terminal devices 101, 102, 103 and the server 105. The network 104 can include various connection types, such as wired and / or wireless communication links, etc.
[0030] The user can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. The terminal devices 101, 102, 103 can be various electronic devices with display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, desktop computers, and the like.
[0031] The server 105 can be a server providing various services, such as a background management server (only as an example) providing support for a website browsed by the user using the terminal devices 101, 102, 103. The background management server can analyze and process received user requests and the like, and feed back the processing results (such as a webpage, information, or data, etc. obtained or generated according to the user request) to the terminal device.
[0032] It should be noted that the target identification method provided by the embodiments of the present disclosure can generally be executed by the server 105. Accordingly, the target identification apparatus provided by the embodiments of the present disclosure can generally be arranged in the server 105. The target identification method provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103 and / or the server 105. Accordingly, the target identification apparatus provided by the embodiments of the present disclosure can also be arranged in a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103 and / or the server 105.
[0033] Figure 2 is a flowchart of a target identification method according to an embodiment of the present disclosure.
[0034] As shown in Figure 2 , the method 200 can include operations S210 to S260.
[0035] At operation S210, in response to obtaining the target text, feature extraction is performed on the target text to obtain target text features.
[0036] In the embodiments of the present disclosure, the target text is related to a target object. For example, the target object can be a pet. The pet can be an animal such as a cat or a dog. For another example, the target object can also be other objects, which are not limited in the present disclosure.
[0037] In the embodiments of the present disclosure, the target text can be used to describe semantic information of the target object. For example, the target text can be "gray teddy". The target text can also be "10-year-old Alaskan".
[0038] In the embodiments of the present disclosure, the target text can be extracted by various ways.
[0039] In operation S220, feature extraction is performed on the at least one initial image related to the target text, to obtain at least one initial image feature.
[0040] In the embodiments of the present disclosure, the image captured by the video capture device can be used as the at least one initial image. For example, the video capture device can be a camera. The at least one initial image can be determined from the images captured by the plurality of cameras.
[0041] In the embodiments of the present disclosure, the time point at which the target text is obtained can be used as the target time point. The image captured within the preset time period before the target time point can be used as the at least one initial image.
[0042] In the embodiments of the present disclosure, the at least one initial image can be J. J can be an integer greater than or equal to 1.
[0043] In the embodiments of the present disclosure, the initial image can be extracted in various ways. For example, the jth initial image can be extracted to obtain the jth initial image feature. J can be an integer greater than or equal to 1 and less than or equal to J.
[0044] In operation S230, at least one query feature, at least one key feature, and at least one value feature are obtained according to the target text feature and the at least one initial image feature.
[0045] In the embodiments of the present disclosure, the target text feature can be used as the query feature, the key feature, or the value feature. The initial image feature can also be used as the query feature, the key feature, or the value feature. For example, the target text feature can be used as the jth key feature, and the jth initial image feature can be used as the jth query feature and the jth value feature.
[0046] In operation S240, the at least one query feature, the at least one key feature, and the at least one value feature are fused to obtain at least one target fusion feature.
[0047] In the embodiments of the present disclosure, the query feature, the key feature, and the value feature can be fused based on the multi-head attention mechanism to obtain the target fusion feature. For example, the jth query feature, the jth key feature, and the jth value feature can be fused based on the multi-head attention mechanism to obtain the jth target fusion feature.
[0048] In operation S250, at least one recognition result corresponding to the at least one initial image is determined according to the at least one target fusion feature.
[0049] In this embodiment of the disclosure, the recognition result may include candidate detection boxes and category confidence scores of the initial image. For example, the j-th recognition result can be obtained based on the j-th target fusion feature. The j-th recognition result may include the candidate detection boxes of the j-th initial image and the confidence scores of multiple categories of the object in the j-th initial image. For example, the confidence scores of multiple categories may include the confidence score of the category "tabby cat" and the confidence score of the category "Alaska".
[0050] In operation S260, a target image related to the target object is determined from at least one initial image based on at least one recognition result.
[0051] For example, the target text could be "10-year-old Alaska" as mentioned above. The recognition result with the highest confidence in the category "Alaska" can be used as the target recognition result. The initial image corresponding to the target recognition result can be used as the target image.
[0052] Through the embodiments of this disclosure, a target image is determined from images captured by a video capture device based on the target text, making it more convenient to identify target images related to objects. In situations where it is difficult to obtain relevant features of the object, fusing target text features and initial image features based on a multi-head attention mechanism helps to quickly determine the target image matching the target text, significantly expanding the application scenarios of object recognition and reducing the difficulty of object (e.g., pet) recognition in urban scenarios.
[0053] As can be understood, the method flow of this disclosure has been described above. In the embodiments of this disclosure, a deep learning model can be used to implement the above method, which will be described in detail below.
[0054] Figure 3 This is a schematic diagram of a target recognition model according to an embodiment of the present disclosure.
[0055] like Figure 3 As shown, the target recognition model 300 may include a text feature extraction network 310, an image feature extraction network 320, a fusion network 330, and a recognition network 340.
[0056] In the embodiment of the present disclosure, in the operation S210, the target text 301 can be input into the text feature extraction network 310 to obtain a target text feature. For example, the text feature extraction network can be a robustly optimized bidirectional encoder representations from transformers (RoBERTa). For another example, the target text can be "gray teddy". In an example, the target text can be tokenized to obtain a token sequence of the target text. The token sequence of the target text is input into the text feature extraction network to obtain the target text feature.
[0057] As shown in FIG. 3, the image feature extraction network 320 can include a first feature extraction unit 321 and a second feature extraction unit 322. The second feature extraction unit 322 can include K feature extraction layers. The K feature extraction layers can include a feature extraction layer 3221, a feature extraction layer 3222, and a feature extraction layer 3223. Figure 3
[0058] In the embodiment of the present disclosure, in the operation S220, the K-level feature extraction is performed on the jth initial image related to the target text to obtain a K-level initial image feature of the jth initial image. For example, j can be an integer greater than or equal to 1 and less than or equal to J, and K can be an integer greater than or equal to 1. The jth initial image 302 can be input into the first feature extraction unit 321 to obtain a first initial image feature. The first initial image feature can be input into the feature extraction layer 3221 to obtain a first initial image feature of the jth initial image. The first initial image feature of the jth initial image can be input into the feature extraction layer 3222 to obtain a second initial image feature of the jth initial image. The second initial image feature of the jth initial image can be input into the feature extraction layer 3223 to obtain a third initial image feature of the jth initial image. It can be understood that, in the embodiment, K can be 3. Thus, the K initial image features of the jth initial image are obtained. Next, the target text feature and the K initial image features can be fused respectively to obtain K target fusion features. Through the embodiment of the present disclosure, different scale features of the image can be extracted, and effective information can be extracted from the image sufficiently, so as to obtain a more accurate recognition result.
[0059] In this embodiment of the disclosure, in the above-described operation S230, key features can be obtained based on target text features. Query features and value features can be obtained based on initial image features. For example, the first query feature and the first value feature of the j-th initial image can be obtained based on the first initial image features of the j-th initial image. The second query feature and the second value feature of the j-th initial image can be obtained based on the second initial image features of the j-th initial image. The third query feature and the third value feature of the j-th initial image can be obtained based on the third initial image features of the j-th initial image. Furthermore, the target text features can be respectively used as the first key feature, the second key feature, and the third key feature corresponding to the j-th initial image.
[0060] like Figure 3 As shown, the fusion network 330 may include I fusion units. The fusion units may be constructed based on the Transformer model. The I fusion units may include fusion unit 331, fusion unit 332, and fusion unit 333. I is an integer greater than 1. It can be understood that fusion unit 333 can be a level I fusion unit; in this embodiment, I can be 3. It can also be understood that fusion unit 331 and fusion unit 332 can be i-th level fusion units, where i can be an integer greater than or equal to 1 and less than 1, and the value of i can be 1 or 2.
[0061] In this embodiment of the disclosure, in the above-described operation S240, at least one level of fusion can be performed on at least one query feature, at least one key feature, and at least one value feature to obtain at least one target fusion feature. For example, I fusion units can be used to perform I-level fusion on the first query feature of the j-th initial image, the first key feature corresponding to the j-th initial image, and the first value feature of the j-th initial image to obtain the first target fusion feature corresponding to the j-th initial image. As another example, I fusion units can be used to perform I-level fusion on the second query feature of the j-th initial image, the second key feature corresponding to the j-th initial image, and the second value feature of the j-th initial image to obtain the second target fusion feature corresponding to the j-th initial image. As yet another example, I fusion units can be used to perform I-level fusion on the third query feature of the j-th initial image, the third key feature corresponding to the j-th initial image, and the third value feature of the j-th initial image to obtain the third target fusion feature corresponding to the j-th initial image. The following example, using the process of obtaining the first target fusion feature corresponding to the j-th initial image, further illustrates the fusion network 330.
[0062] In the embodiments of the present disclosure, the query feature, the key feature and the value feature can be taken as the first-level query feature, the first-level key feature and the first-level value feature respectively. Based on the multi-head attention mechanism, the first-level query feature, the first-level key feature and the first-level value feature are fused to obtain the first-level intermediate fusion feature. For example, the first query feature of the jth initial image can be taken as the first-level query feature. The first key feature corresponding to the jth initial image can be taken as the first-level key feature. The first value feature of the jth initial image can be taken as the first-level value feature. The first-level query feature, the first-level key feature and the first-level value feature are input into the fusion unit 331, and the first-level intermediate fusion feature can be obtained.
[0063] In the embodiments of the present disclosure, the ith intermediate fusion feature can be fused with the target text feature and the initial image feature respectively to obtain the (i+1)th text fusion feature and the (i+1)th image fusion feature. The (i+1)th key feature is obtained according to the (i+1)th text fusion feature. The (i+1)th query feature and the (i+1)th value feature are obtained according to the (i+1)th image fusion feature. Based on the multi-head attention mechanism, the (i+1)th query feature, the (i+1)th key feature and the (i+1)th value feature are fused to obtain the (i+1)th intermediate fusion feature. The ith intermediate fusion feature is taken as the target fusion feature. For example, the above-mentioned first-level intermediate fusion feature can be fused with the target text feature to obtain the second-level text fusion feature. The above-mentioned first-level intermediate fusion feature can be fused with the first initial image feature of the jth initial image to obtain the second-level image fusion feature. The second-level text fusion feature can be taken as the second-level key feature. The second-level image fusion feature can be taken as the second-level query feature and the second-level value feature. The second-level query feature, the second-level key feature and the second-level value feature are input into the fusion unit 332, and the second-level intermediate fusion feature can be obtained. For another example, the above-mentioned second-level intermediate fusion feature can be fused with the target text feature to obtain the third-level text fusion feature. The above-mentioned second-level intermediate fusion feature can be fused with the first initial image feature of the jth initial image to obtain the third-level image fusion feature. The third-level text fusion feature can be taken as the third-level key feature. The third-level image fusion feature can be taken as the third-level query feature and the third-level value feature. The third-level query feature, the third-level key feature and the third-level value feature are input into the fusion unit 333, and the third-level intermediate fusion feature can be obtained. The third-level intermediate fusion feature can be taken as the first target fusion feature corresponding to the jth initial image. Through the embodiments of the present disclosure, the text feature and the image feature can be fully fused based on the multi-head attention mechanism. In the target recognition scene (especially in the pet recognition scene), the information of the text and the image can be more fully obtained, which is helpful to conveniently and accurately determine the target image from the image collected by the video collection device.
[0064] As Figure 3As shown, the recognition network 340 can process target fusion features and output recognition results.
[0065] In this embodiment of the disclosure, in the above-described operation S250, the target fusion feature can be convolved at least once to obtain a recognition result. For example, the recognition network 340 can perform at least one convolution on the first target fusion feature corresponding to the j-th initial image to obtain a first recognition result 341 corresponding to the j-th initial image. As another example, the recognition network 340 can perform at least one convolution on the second target fusion feature corresponding to the j-th initial image to obtain a second recognition result 342 corresponding to the j-th initial image. As yet another example, the recognition network 340 can perform at least one convolution on the third target fusion feature corresponding to the j-th initial image to obtain a third recognition result 343 corresponding to the j-th initial image.
[0066] In this embodiment of the disclosure, J initial images are input into the target recognition model 300, and J×K recognition results can be obtained.
[0067] It is understood that the target recognition model of this disclosure has been explained above, and the following will provide further explanation based on the recognition results of this disclosure.
[0068] Figure 4A This is a schematic diagram of an initial image according to an embodiment of the present disclosure.
[0069] like Figure 4A As shown, the initial image 401' may include an object. The actual category of the object can be "Golden Retriever".
[0070] Figure 4B This is a schematic diagram of the identification result according to an embodiment of the present disclosure.
[0071] In this embodiment of the disclosure, the recognition result includes candidate detection boxes and target category confidence scores for the initial image. For example, the first recognition result of the initial image 401' can be implemented as a vector (H, W, x, y, w, h, score). H and W represent the number of regions into which the initial image is divided, respectively. Figure 4B As shown, H can be 5, and W can also be 5. (x, y) can represent the coordinates of the center point of the candidate detection box. w is the width of the candidate detection box. h is the height of the candidate detection box. socre is the confidence information of the object in the initial image. For example, the confidence information can include the confidence of multiple categories. Multiple categories can include categories such as "Golden Retriever", "Teddy", etc. Each category corresponds to one confidence level. It can be understood that for an initial image, there can be K recognition results. H and W can be different for different recognition results.
[0072] In the embodiments of the present disclosure, the target category can be determined according to the target text. For example, the target text can be “big golden dog”. Based on this, the category “golden retriever” can be determined as the target category.
[0073] In the embodiments of the present disclosure, in operation S260 described above, in response to determining that the target category confidence is greater than or equal to the preset confidence threshold, the initial image corresponding to the recognition result is determined as the target image. For example, as shown in FIG. 4B, the confidence of the target category “golden retriever” in the first recognition result of the initial image 401’ can be greater than the preset confidence threshold. The initial image 401 can be determined as a target image. Figure 4B
[0074] It can be understood that the target recognition method of the present disclosure is described above. In the embodiments of the present disclosure, the target recognition model described above can be trained, which will be described in detail below.
[0075] Figure 5 FIG. 5 is a flowchart of a method for training a target recognition model according to another embodiment of the present disclosure.
[0076] As shown in FIG. 5, the method 500 can include operation S510 to operation S560. Figure 5
[0077] In the embodiments of the present disclosure, the target recognition model can include an image feature extraction network, a text feature extraction network, a fusion network, and a recognition network.
[0078] In operation S510, the sample text is input into the text feature extraction network to obtain a sample text feature.
[0079] In the embodiments of the present disclosure, the sample text is related to a sample object. For example, the sample object can be a pet. The pet can be an animal such as a cat or a dog. For example, the sample text can be “gray teddy”. The sample text can also be “10-year-old Alaskan”. For another example, the sample object can also be other objects, which are not limited in the present disclosure.
[0080] In the embodiments of the present disclosure, the text extraction network can be various feature extraction networks. For example, the text extraction network can be the strong optimized bidirectional encoder representation model based on Transformer described above.
[0081] In operation S520, the sample image is input into the image feature extraction network to obtain a sample image feature.
[0082] In the embodiments of the present disclosure, an image including a sample object can be used as a sample image. For example, an image including a related object can be selected as a sample image. It can be understood that the sample text can be artificially made. The sample text can describe the semantic information of the object in the sample image.
[0083] In the embodiments of the present disclosure, the image feature extraction network can be various feature extraction networks.
[0084] In operation S530, query features, key features and value features are obtained according to the sample text features and the sample image features.
[0085] In the embodiments of the present disclosure, the sample text features can be used as the query features, the key features or the value features. The sample image features can also be used as the query features, the key features or the value features.
[0086] In operation S540, the query features, the key features and the value features are input into a fusion network to obtain sample fusion features.
[0087] In the embodiments of the present disclosure, the query features, the key features and the value features can be fused based on a multi-head attention mechanism to obtain the sample fusion features.
[0088] In operation S550, the sample fusion features are input into a recognition network to obtain sample recognition results corresponding to the sample image.
[0089] In the embodiments of the present disclosure, the recognition results can include candidate bounding boxes and class confidences of the initial image. For example, the sample recognition results are obtained according to the sample fusion features. The sample recognition results can include candidate bounding boxes of the sample image and confidences of multiple classes of the sample object in the sample image. For another example, the confidences of the multiple classes can include a confidence of a class “Lilac Cat” and a confidence of a class “Alaskan”.
[0090] In operation S560, the image recognition model is trained according to the labels of the sample image and the sample recognition results.
[0091] In the embodiments of the present disclosure, the labels can include an annotated bounding box of the sample object in the sample image and an annotated class of the sample object. For example, the annotated class of the sample object can include annotated confidences of multiple classes. In the annotated confidences of the multiple classes, the annotated confidence of the real class of the sample object can be 1, and the annotated confidences of other classes can be 0.
[0092] It can be understood that the training method of the target recognition model of the present disclosure is described above, and the training method of the target recognition model of the present disclosure will be further described in combination with related embodiments.
[0093] Figure 6 is a schematic diagram of the training method of the target recognition model according to an embodiment of the present disclosure.
[0094] As Figure 6As shown, the target recognition model 600 may include a text feature extraction network 610, an image feature extraction network 620, a fusion network 630, and a recognition network 640.
[0095] In this embodiment of the disclosure, during the above-described operation S510, the sample text 601 can be input into the text feature extraction network 610 to obtain sample text features. For example, the sample text could be "gray teddy bear".
[0096] like Figure 6 As shown, the image feature extraction network 620 may include a first feature extraction unit 621 and a second feature extraction unit 622. The second feature extraction unit 622 may include K feature extraction layers. The K feature extraction layers may include feature extraction layer 6221, feature extraction layer 6222, and feature extraction layer 6223.
[0097] In this embodiment, during operation S620, the sample image can be input into the first feature extraction unit to obtain the first sample image feature. The first sample image feature can be input into the first-level feature extraction layer to obtain the first sample image feature. The k-th level sample image feature can be input into the (k+1)-th level feature extraction layer to obtain the (k+1)-th sample image feature. For example, k is an integer greater than or equal to 1 and less than K. For example, the sample image 602 can be input into the first feature extraction unit 621 to obtain the first sample image feature. The first sample image feature can be input into the feature extraction layer 6221 to obtain the first sample image feature of the sample image. The first sample image feature of the sample image can be input into the feature extraction layer 6222 to obtain the second sample image feature of the sample image. The second sample image feature of the sample image can be input into the feature extraction layer 6223 to obtain the third sample image feature of the sample image. It can be understood that in this embodiment, K can be 3, and the value of k can be 1 or 2; the feature extraction layer 6223 can be the K-th level feature extraction layer. Thus, K sample image features of the sample image are obtained. Next, the sample text features and K sample image features can be fused separately to obtain K sample fused features.
[0098] In this embodiment of the disclosure, in operation S530 described above, key features are obtained based on sample text features. Query features and value features are obtained based on sample image features. For example, the first query feature and the first value feature of the sample image can be obtained based on the first sample image feature. The second query feature and the second value feature of the sample image can be obtained based on the second sample image feature. The third query feature and the third value feature of the sample image can be obtained based on the third sample image feature. As another example, sample text features can be used as the first key feature, the second key feature, and the third key feature, respectively.
[0099] like Figure 6 As shown, the fusion network 630 may include I fusion units. The fusion units may be constructed based on the Transformer model. The I fusion units may include fusion unit 631, fusion unit 632, and fusion unit 633. I is an integer greater than 1. It can be understood that fusion unit 633 can be used as the I-th level fusion unit. In this embodiment, I can be 3, and i can be an integer greater than or equal to 1 and less than I, with i taking values of 1 and 2. Fusion units 631 and fusion unit 632 can be used as the i-th level fusion units.
[0100] In this embodiment of the disclosure, in the above-described operation S540, at least one fusion unit can be used to perform at least one level of fusion on the query feature, key feature, and value feature to obtain a sample fusion feature. For example, I fusion units can be used to perform I-level fusion on the first query feature, the first key feature, and the first value feature of the sample image to obtain a first sample fusion feature. As another example, I fusion units can be used to perform I-level fusion on the second query feature, the second key feature, and the second value feature of the sample image to obtain a second sample fusion feature. As yet another example, I fusion units can be used to perform I-level fusion on the third query feature, the third key feature, and the third value feature of the sample image to obtain a third sample fusion feature. The process of obtaining the first sample fusion feature will be used as an example to further explain the fusion network 630.
[0101] In the embodiments of the present disclosure, the query feature, the key feature and the value feature can be respectively regarded as the first-level query feature, the first-level key feature and the first-level value feature. The first-level query feature, the first-level key feature and the first-level value feature can be input into the first-level fusion unit to obtain the first-level intermediate fusion feature. For example, the first query feature of the sample image can be regarded as the first-level query feature. The first key feature can be regarded as the first-level key feature. The first value feature of the sample image can be regarded as the first-level value feature. The first-level query feature, the first-level key feature and the first-level value feature are input into the fusion unit 631, and the first-level intermediate fusion feature can be obtained.
[0102] In the embodiments of the present disclosure, the i-level intermediate fusion feature can be fused with the sample text feature and the sample image feature respectively to obtain the i+1-level sample text fusion feature and the i+1-level sample image fusion feature. The i+1-level key feature can be obtained according to the i+1-level sample text fusion feature. The i+1-level query feature and the i+1-level value feature can be obtained according to the i+1-level image fusion feature. The i+1-level query feature, the i+1-level key feature and the i+1-level value feature are input into the i+1-level fusion unit to obtain the i+1-level intermediate fusion feature. The i-level intermediate fusion feature is regarded as the sample fusion feature. For example, the first-level intermediate fusion feature described above can be fused with the sample text feature to obtain the second-level sample text fusion feature. The first-level intermediate fusion feature described above can be fused with the first sample image feature of the sample image to obtain the second-level sample image fusion feature. The second-level sample text fusion feature can be regarded as the second-level key feature. The second-level sample image fusion feature can be regarded as the second-level query feature and the second-level value feature. The second-level query feature, the second-level key feature and the second-level value feature are input into the fusion unit 632, and the second-level intermediate fusion feature can be obtained. For another example, the second-level intermediate fusion feature described above can be fused with the sample text feature to obtain the third-level sample text fusion feature. The second-level intermediate fusion feature described above can be fused with the first sample image feature of the sample image to obtain the third-level sample image fusion feature. The third-level sample text fusion feature can be regarded as the third-level key feature. The third-level sample image fusion feature can be regarded as the third-level query feature and the third-level value feature. The third-level query feature, the third-level key feature and the third-level value feature are input into the fusion unit 633, and the third-level intermediate fusion feature can be obtained. The third-level intermediate fusion feature can be regarded as the first sample fusion feature corresponding to the sample image.
[0103] As shown in FIG. 6, the recognition network 640 can process the target fusion feature and output a recognition result. Figure 6
[0104] In this embodiment of the disclosure, in the above-described operation S550, the recognition network can perform at least one convolution on the sample fusion features to obtain the sample recognition result. For example, the recognition network 640 can perform at least one convolution on the first sample fusion feature of the sample image to obtain the first sample recognition result 641 of the sample image. As another example, the recognition network 640 can perform at least one convolution on the second sample fusion feature of the sample image to obtain the second sample recognition result 642 of the sample image. As yet another example, the recognition network 640 can perform at least one convolution on the third sample fusion feature of the sample image to obtain the third sample recognition result 643 of the sample image.
[0105] For example, the sample recognition results include candidate detection boxes for the initial image and sample category confidence.
[0106] Next, based on the identification results of the first sample 641, the second sample 642, the third sample 643, and the label 603, various loss functions can be used to determine the loss value. This loss value can then be used to train the target recognition model. For example, the parameters of the target recognition model can be adjusted to make the loss value converge.
[0107] Figure 7 This is a block diagram of a target recognition device according to an embodiment of the present disclosure.
[0108] like Figure 7 As shown, the device 700 may include a first feature extraction module 710, a second feature extraction module 720, a first acquisition module 730, a fusion module 740, a first determination module 750, and a second determination module 760.
[0109] The first feature extraction model 710 is used to extract features from the target text in response to obtaining the target text, thereby obtaining the target text features. For example, the target text is related to the target object;
[0110] The second feature extraction module 720 is used to extract features from at least one initial image related to the target text to obtain at least one initial image feature.
[0111] The first obtaining module 730 is used to obtain at least one query feature, at least one key feature, and at least one value feature based on the target text features and at least one initial image feature.
[0112] The fusion module 740 is used to fuse at least one query feature, at least one key feature and at least one value feature to obtain at least one target fusion feature.
[0113] The first determining module 750 is used to determine at least one recognition result corresponding to at least one initial image based on at least one target fusion feature.
[0114] The second determining module 760 is configured to determine, according to the at least one recognition result, a target image related to the target object from the at least one initial image.
[0115] In some embodiments, the first obtaining module comprises: a first obtaining sub-module configured to obtain at least one key feature according to the target text feature; and a second obtaining sub-module configured to obtain at least one query feature and at least one value feature according to the at least one initial image feature.
[0116] In some embodiments, the fusion module comprises a fusion sub-module configured to perform at least one level of fusion on the at least one query feature, the at least one key feature, and the at least one value feature to obtain at least one target fusion feature.
[0117] In some embodiments, the fusion sub-module comprises: a first obtaining unit configured to take the query feature, the key feature, and the value feature as a first-level query feature, a first-level key feature, and a first-level value feature, respectively; and a first fusion unit configured to fuse the first-level query feature, the first-level key feature, and the first-level value feature based on a multi-head attention mechanism to obtain a first-level intermediate fusion feature.
[0118] In some embodiments, the fusion sub-module further comprises: a second fusion unit configured to fuse the ith-level intermediate fusion feature with the target text feature and the initial image feature, respectively, to obtain an (i+1)th-level text fusion feature and an (i+1)th-level image fusion feature, where i is an integer greater than or equal to 1 and less than I, and I is an integer greater than 1; a second obtaining unit configured to obtain an (i+1)th-level query feature and an (i+1)th-level value feature according to the (i+1)th-level text fusion feature; a third fusion unit configured to fuse the (i+1)th-level query feature, the (i+1)th-level key feature, and the (i+1)th-level value feature based on the multi-head attention mechanism to obtain an (i+1)th-level intermediate fusion feature; and a third obtaining unit configured to take the Ith-level intermediate fusion feature as the target fusion feature.
[0119] In some embodiments, the at least one initial image is J, and J is an integer greater than or equal to 1, and the second feature extraction module comprises: a feature extraction sub-module configured to perform K-level feature extraction on the jth initial image related to the target text to obtain K initial image features of the jth initial image, where j is an integer greater than or equal to 1 and less than or equal to J, and K is an integer greater than or equal to 1.
[0120] In some embodiments, the first determining module comprises: a convolution sub-module configured to perform at least one convolution on the target fusion feature to obtain a recognition result, wherein the recognition result comprises a candidate detection frame of the initial image and a target class confidence.
[0121] In some embodiments, the second determining module comprises a determining submodule configured to determine the initial image corresponding to the recognition result as the target image in response to determining that the target category confidence is greater than or equal to the preset confidence threshold.
[0122] Figure 8 is a block diagram of a training device of a target recognition model according to another embodiment of the present disclosure.
[0123] The target recognition model comprises an image feature extraction network, a text feature extraction network, a fusion network, and a recognition network.
[0124] As shown in Figure 8 The apparatus 800 can comprise a second obtaining module 810, a third obtaining module 820, a fourth obtaining module 830, a fifth obtaining module 840, a sixth obtaining module 850, and a training module 860.
[0125] The second obtaining module 810 is configured to input the sample text into the text feature extraction network to obtain a sample text feature. For example, the sample text is related to a sample object.
[0126] The third obtaining module 820 is configured to input the sample image into the image feature extraction network to obtain a sample image feature.
[0127] The fourth obtaining module 830 is configured to obtain a query feature, a key feature, and a value feature according to the sample text feature and the sample image feature.
[0128] The fifth obtaining module 840 is configured to input the query feature, the key feature, and the value feature into the fusion network to obtain a sample fusion feature.
[0129] The sixth obtaining module 850 is configured to input the sample fusion feature into the recognition network to obtain a sample recognition result corresponding to the sample image.
[0130] The training module 860 is configured to train the image recognition model according to a label of the sample image and the sample recognition result.
[0131] In some embodiments, the fourth obtaining module comprises a third obtaining submodule configured to obtain the key feature according to the sample text feature. The fourth obtaining submodule is configured to obtain the query feature and the value feature according to the sample image feature.
[0132] In some embodiments, the fusion network comprises at least one fusion unit, and the fifth obtaining module comprises a second fusion submodule configured to perform at least one level of fusion on the query feature, the key feature, and the value feature by using the at least one fusion unit to obtain the sample fusion feature.
[0133] In some embodiments, the second fusion submodule includes: a fourth obtaining unit, configured to use query features, key features, and value features as first-level query features, first-level key features, and first-level value features, respectively; and a fifth obtaining unit, configured to input the first-level query features, first-level key features, and first-level value features into the first-level fusion unit to obtain first-level intermediate fusion features.
[0134] In some embodiments, the second fusion submodule further includes: a fourth fusion unit, configured to fuse the i-th level intermediate fusion feature with the sample text feature and the sample image feature respectively, to obtain the (i+1)-th level sample text fusion feature and the (i+1)-th level sample image fusion feature, where i is an integer greater than or equal to 1 and less than 1, and 1 is an integer greater than 1. A sixth obtaining unit, configured to obtain the (i+1)-th level key feature based on the (i+1)-th level sample text fusion feature. A seventh obtaining unit, configured to obtain the (i+1)-th level query feature and the (i+1)-th level value feature based on the (i+1)-th level sample image fusion feature. An eighth obtaining unit, configured to input the (i+1)-th level query feature, the (i+1)-th level key feature, and the (i+1)-th level value feature into the (i+1)-th level fusion unit to obtain the (i+1)-th level intermediate fusion feature. A ninth obtaining unit, configured to use the i-th level intermediate fusion feature as the sample fusion feature.
[0135] In some embodiments, the image feature extraction network includes a first feature extraction unit and a second feature extraction unit. The second feature extraction unit includes K feature extraction layers, where K is an integer greater than 1. The third obtaining module includes: a tenth obtaining module, used to input a sample image into the first feature extraction unit to obtain a first sample image feature; an eleventh obtaining module, used to input the first sample image feature into a first-level feature extraction layer to obtain a first sample image feature; and a twelfth obtaining module, used to input the k-th level sample image feature into the (k+1)-th level feature extraction layer to obtain the (k+1)-th sample image feature, where k is an integer greater than or equal to 1 and less than K.
[0136] In some embodiments, the sixth obtaining module includes: a second convolution submodule, used to perform at least one convolution on the sample fusion features using a recognition network to obtain sample recognition results, wherein the sample recognition results include candidate detection boxes and sample category confidence of the initial image.
[0137] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0138] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0139] Figure 9A schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0140] like Figure 9 As shown, device 900 includes a computing unit 901, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 902 or a computer program loaded from storage unit 908 into random access memory (RAM) 903. RAM 903 may also store various programs and data required for the operation of device 900. The computing unit 901, ROM 902, and RAM 903 are interconnected via bus 904. Input / output (I / O) interface 905 is also connected to bus 904.
[0141] Multiple components in device 900 are connected to I / O interface 905, including: input unit 906, such as keyboard, mouse, etc.; output unit 907, such as various types of monitors, speakers, etc.; storage unit 908, such as disk, optical disk, etc.; and communication unit 909, such as network card, modem, wireless transceiver, etc. Communication unit 909 allows device 900 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0142] The computing unit 901 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 901 performs various methods and processes described above, such as the object recognition method and / or the training method of the object recognition model. For example, in some embodiments, the object recognition method and / or the training method of the object recognition model can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded onto the RAM 903 and executed by the computing unit 901, one or more steps of the object recognition method and / or the training method of the object recognition model described above can be performed. Alternatively, in other embodiments, the computing unit 901 can be configured to perform the object recognition method and / or the training method of the object recognition model by any other appropriate means, such as by means of firmware.
[0143] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a complex programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0144] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces a means for implementing the functions / acts specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine and partially on a remote machine or entirely on a remote machine or server.
[0145] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of a computer program code, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0146] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) monitor or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0147] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0148] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0149] It should be understood that the various forms of flow shown above can be used to reorder, add, or delete steps. For example, the steps described in the present disclosure can be performed in parallel, in series, or in a different order, as long as the desired results of the technology disclosed in the present disclosure can be achieved, which is not limited herein.
[0150] The above detailed description does not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A target recognition method, comprising: In response to obtaining the target text, feature extraction is performed on the target text to obtain target text features, wherein the target text is related to the target object; Feature extraction is performed on at least one initial image related to the target text to obtain at least one initial image feature; Based on the target text features and at least one initial image feature, at least one query feature, at least one key feature, and at least one value feature are obtained; By fusing at least one of the query features, at least one of the key features, and at least one of the value features, at least one target fusion feature is obtained; Based on at least one of the target fusion features, determine at least one recognition result corresponding to at least one of the initial images; and Based on at least one of the recognition results, a target image related to the target object is determined from at least one of the initial images. The step of obtaining at least one query feature, at least one key feature, and at least one value feature based on the target text features and at least one initial image feature includes: Based on the target text features, at least one of the key features is obtained; and Based on at least one of the initial image features, at least one of the query features and at least one of the value features are obtained. The step of fusing at least one query feature, at least one key feature, and at least one value feature to obtain at least one target fusion feature includes: At least one level of fusion is performed on at least one of the query features, at least one of the key features, and at least one of the value features to obtain at least one target fusion feature. The step of fusing at least one level of at least one query feature, at least one key feature, and at least one value feature to obtain at least one target fusion feature includes: The query feature, the key feature, and the value feature are respectively designated as the first-level query feature, the first-level key feature, and the first-level value feature; and Based on a multi-head attention mechanism, the query features, key features, and value features described in Level 1 are fused to obtain the intermediate fused features in Level 1. The step of fusing at least one level of at least one query feature, at least one key feature, and at least one value feature to obtain at least one target fusion feature further includes: The intermediate fusion feature at level i is fused with the target text feature and the initial image feature respectively to obtain the (i+1)th level text fusion feature and the (i+1)th level image fusion feature, where i is an integer greater than or equal to 1 and less than 1, and 1 is an integer greater than 1. Based on the text fusion features described at level i+1, the key features described at level i+1 are obtained; Based on the image fusion features described at level i+1, the query features and value features described at level i+1 are obtained. Based on a multi-head attention mechanism, the query feature, key feature, and value feature at level i+1 are fused to obtain the intermediate fused feature at level i+1; and The intermediate fusion feature described in Level I is used as the target fusion feature.
2. The method according to claim 1, wherein, There must be at least one initial image of J, where J is an integer greater than or equal to 1. The step of extracting features from at least one initial image related to the target text to obtain at least one initial image feature includes: K-level feature extraction is performed on the j-th initial image related to the target text to obtain K initial image features of the j-th initial image, where j is an integer greater than or equal to 1 and less than or equal to J, and K is an integer greater than or equal to 1.
3. The method according to claim 1, wherein, The step of determining at least one recognition result corresponding to at least one of the initial images based on at least one of the target fusion features includes: The target fusion features are convolved at least once to obtain the recognition result, wherein the recognition result includes the candidate detection boxes and target category confidence of the initial image.
4. The method according to claim 3, wherein, Determining the target image related to the target object from at least one of the initial images based on at least one of the recognition results includes: In response to determining that the confidence level of the target category is greater than or equal to a preset confidence threshold, the initial image corresponding to the recognition result is determined as the target image.
5. A method for training a target recognition model, the target recognition model comprising an image feature extraction network, a text feature extraction network, a fusion network, and a recognition network, the method comprising: The sample text is input into the text feature extraction network to obtain sample text features, wherein the sample text is related to the sample object; The sample image is input into the image feature extraction network to obtain the sample image features; Based on the sample text features and the sample image features, query features, key features, and value features are obtained; The query features, the key features, and the value features are input into the fusion network to obtain sample fusion features; The sample fusion features are input into the recognition network to obtain the sample recognition result corresponding to the sample image; and Based on the labels of the sample images and the sample recognition results, the image recognition model is trained. The step of obtaining query features, key features, and value features based on the sample text features and the sample image features includes: Based on the sample text features, the key features are obtained; Based on the features of the sample image, the query features and the value features are obtained. The fusion network includes at least one fusion unit. The step of inputting the query features, the key features, and the value features into the fusion network to obtain sample fusion features includes: The query feature, the key feature, and the value feature are fused at least once using at least one of the fusion units to obtain the sample fusion feature. The step of performing at least one level of fusion on the query feature, the key feature, and the value feature using at least one of the fusion units includes: The query feature, the key feature, and the value feature are respectively designated as the first-level query feature, the first-level key feature, and the first-level value feature; The query features, key features, and value features described in Level 1 are input into the fusion unit described in Level 1 to obtain the intermediate fusion features of Level 1. The step of performing at least one level of fusion on the query feature, the key feature, and the value feature using at least one of the fusion units further includes: The intermediate fusion feature at level i is fused with the sample text feature and the sample image feature respectively to obtain the sample text fusion feature at level i+1 and the sample image fusion feature at level i+1, where i is an integer greater than or equal to 1 and less than I, and I is an integer greater than 1. Based on the sample text fusion features described at level i+1, the key features described at level i+1 are obtained; Based on the sample image fusion features of level i+1, the query features and value features of level i+1 are obtained; The query feature, key feature, and value feature of level i+1 are input into the fusion unit of level i+1 to obtain the intermediate fusion feature of level i+1. The intermediate fusion features described in Level I are used as the sample fusion features.
6. The method according to claim 5, wherein, The image feature extraction network includes a first feature extraction unit and a second feature extraction unit. The second feature extraction unit includes K feature extraction layers, where K is an integer greater than 1. The step of inputting the sample image into the image feature extraction network to obtain sample image features includes: The sample image is input into the first feature extraction unit to obtain the first sample image features; The features of the first sample image are input into the first-level feature extraction layer to obtain the features of the first sample image; The sample image features of level k are input into the feature extraction layer of level k+1 to obtain the sample image features of level k+1, where k is an integer greater than or equal to 1 and less than K.
7. The method according to claim 5, wherein, The step of inputting the sample fusion features into the recognition network to obtain the sample recognition result corresponding to the sample image includes: The sample fusion features are convolved at least once using the recognition network to obtain the sample recognition result, wherein the sample recognition result includes candidate detection boxes and sample category confidence of the initial image.
8. A target recognition device, comprising: A first feature extraction model is used to extract features from the target text in response to obtaining the target text, thereby obtaining target text features, wherein the target text is related to the target object; The second feature extraction module is used to extract features from at least one initial image related to the target text to obtain at least one initial image feature; The first obtaining module is configured to obtain at least one query feature, at least one key feature, and at least one value feature based on the target text features and at least one initial image feature; The fusion module is used to fuse at least one of the query features, at least one of the key features, and at least one of the value features to obtain at least one target fusion feature; The first determining module is configured to determine at least one recognition result corresponding to at least one of the initial images based on at least one of the target fusion features; and The second determining module is configured to determine a target image related to the target object from at least one of the initial images based on at least one of the recognition results. The step of obtaining at least one query feature, at least one key feature, and at least one value feature based on the target text features and at least one initial image feature includes: Based on the target text features, at least one of the key features is obtained; and Based on at least one of the initial image features, at least one of the query features and at least one of the value features are obtained. The step of fusing at least one query feature, at least one key feature, and at least one value feature to obtain at least one target fusion feature includes: At least one level of fusion is performed on at least one of the query features, at least one of the key features, and at least one of the value features to obtain at least one target fusion feature. The step of fusing at least one level of at least one query feature, at least one key feature, and at least one value feature to obtain at least one target fusion feature includes: The query feature, the key feature, and the value feature are respectively designated as the first-level query feature, the first-level key feature, and the first-level value feature; and Based on a multi-head attention mechanism, the query features, key features, and value features described in Level 1 are fused to obtain the intermediate fused features in Level 1. The step of fusing at least one level of at least one query feature, at least one key feature, and at least one value feature to obtain at least one target fusion feature further includes: The intermediate fusion feature at level i is fused with the target text feature and the initial image feature respectively to obtain the (i+1)th level text fusion feature and the (i+1)th level image fusion feature, where i is an integer greater than or equal to 1 and less than 1, and 1 is an integer greater than 1. Based on the text fusion features described at level i+1, the key features described at level i+1 are obtained; Based on the image fusion features described at level i+1, the query features and value features described at level i+1 are obtained. Based on a multi-head attention mechanism, the query feature, key feature, and value feature at level i+1 are fused to obtain the intermediate fused feature at level i+1; and The intermediate fusion feature described in Level I is used as the target fusion feature.
9. A training apparatus for a target recognition model, the target recognition model comprising an image feature extraction network, a text feature extraction network, a fusion network, and a recognition network, the apparatus comprising: The second acquisition module is used to input the sample text into the text feature extraction network to obtain sample text features, wherein the sample text is related to the sample object; The third acquisition module is used to input the sample image into the image feature extraction network to obtain the sample image features; The fourth obtaining module is used to obtain query features, key features, and value features based on the sample text features and the sample image features; The fifth obtaining module is used to input the query features, the key features, and the value features into the fusion network to obtain sample fusion features; The sixth obtaining module is used to input the sample fusion features into the recognition network to obtain the sample recognition result corresponding to the sample image; and The training module is used to train the image recognition model based on the labels of the sample images and the sample recognition results. The step of obtaining query features, key features, and value features based on the sample text features and the sample image features includes: Based on the sample text features, the key features are obtained; Based on the features of the sample image, the query features and the value features are obtained. The fusion network includes at least one fusion unit. The step of inputting the query features, the key features, and the value features into the fusion network to obtain sample fusion features includes: The query feature, the key feature, and the value feature are fused at least once using at least one of the fusion units to obtain the sample fusion feature. The step of performing at least one level of fusion on the query feature, the key feature, and the value feature using at least one of the fusion units includes: The query feature, the key feature, and the value feature are respectively designated as the first-level query feature, the first-level key feature, and the first-level value feature; The query features, key features, and value features described in Level 1 are input into the fusion unit described in Level 1 to obtain the intermediate fusion features of Level 1. The step of performing at least one level of fusion on the query feature, the key feature, and the value feature using at least one of the fusion units further includes: The intermediate fusion feature at level i is fused with the sample text feature and the sample image feature respectively to obtain the sample text fusion feature at level i+1 and the sample image fusion feature at level i+1, where i is an integer greater than or equal to 1 and less than I, and I is an integer greater than 1. Based on the sample text fusion features described at level i+1, the key features described at level i+1 are obtained; Based on the sample image fusion features of level i+1, the query features and value features of level i+1 are obtained; The query feature, key feature, and value feature of level i+1 are input into the fusion unit of level i+1 to obtain the intermediate fusion feature of level i+1. The intermediate fusion features described in Level I are used as the sample fusion features.
10. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 7.
11. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 7.
12. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Combined query image retrieval method based on multi-order adversarial feature learning
CN112818157A
Text-based image retrieval method and device and readable storage medium
CN114357231A
Image text processing method and device, readable medium and electronic equipment
CN115331228A