Target identification method and device and computer storage medium
By performing hierarchical segmentation and weight fusion calculation on image features, the problem of insufficient local semantic analysis in cross-modal target recognition is solved, thereby improving the accuracy of target recognition.
Patent Information
- Application Number
- CN202510765201.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-11-07
AI Technical Summary
In existing technologies, when matching full image features with text features in cross-modal target recognition tasks, local semantic analysis is ignored, resulting in low target recognition accuracy.
By dividing the global features of the image to be recognized into layers and blocks, feature sets at different levels are obtained, and the similarity between each layer of features and text features is calculated. The fusion weights are determined using Gaussian distribution, so as to achieve complementary use of local semantic information and finally obtain multi-level target recognition results.
It improves the semantic accuracy of cross-modal matching, enhances the accuracy of target recognition, and achieves improved matching accuracy from single-level to multi-level.
Smart Images

Figure CN120912845A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of image processing, and in particular to a target recognition method and device and computer storage medium. BACKGROUND
[0002] In the fields of intelligent security, multimedia retrieval, human-computer interaction, etc., it has important application value to identify the category of a target in an image, such as accurately identifying the target in a to-be-processed image as a cat, a car, a house, etc. With the development of cross-modal technology, the cross-modal target recognition task, such as cross-text and visible light image, has become a research hotspot, aiming to match the text features of the target category with the image features of the visible light image to realize the target recognition task.
[0003] In related technologies, the global image features and text features are usually directly matched. Although the global image features can summarize the overall information of the image, they ignore the layered local semantic analysis of the image, which leads to the inability to accurately capture the local discriminative features that play a key role in target recognition, and further makes the semantic correspondence in cross-modal matching not accurate enough, ultimately affecting the accuracy of target recognition. Therefore, how to improve the accuracy of target recognition accuracy has become a technical problem to be solved in related technologies. SUMMARY
[0004] To solve the above technical problems, the present application provides a target recognition method, device and computer storage medium.
[0005] To solve the above technical problems, the present application provides a target recognition method, which comprises: extracting global features of a to-be-identified image;
[0006] obtaining at least two sets of features divided based on the global features, wherein the number of feature blocks in a first set of features in the at least two sets of features is different from the number of feature blocks in a second set of features;
[0007] obtaining a similarity set of each set of features and text features;
[0008] determining a first fusion confidence of each set of features based on the similarity set;
[0009] fusing at least two first fusion confidences corresponding to the at least two sets of features to obtain a second fusion confidence of the global features and the text features;
[0010] determining a target recognition result of the to-be-identified image according to the second fusion confidence.
[0011] wherein determining a first fusion confidence of each set of features based on the similarity set comprises:
[0012] determining a target coordinate of a target position in the image to be recognized;
[0013] generating a fusion weight corresponding to each patch feature in the current feature set according to the target coordinate and a target variance corresponding to the image to be recognized, to obtain a first weight set, wherein the target variance comprises a first variance for controlling a horizontal weight decay speed and a second variance for controlling a vertical weight decay speed;
[0014] fusing the similarity set according to the first weight set to obtain the first fused confidence.
[0015] The generating of the fusion weight corresponding to each patch feature in the current feature set according to the target coordinate and the target variance corresponding to the image to be recognized to obtain the first weight set comprises:
[0016] For each patch feature in the current feature set, a position coordinate corresponding to the current patch feature in the image to be recognized is obtained;
[0017] determining the fusion weight corresponding to the current patch feature according to the position coordinate, the target coordinate and the target variance, wherein the fusion weight corresponding to the patch feature in the current feature set is in a Gaussian distribution.
[0018] The fusing of the similarity set according to the first weight set comprises:
[0019] obtaining a foreground region where the target is located in the image to be recognized;
[0020] determining a filtering weight corresponding to each patch feature in the current feature set according to a relative position between a region corresponding to the patch feature in the image to be recognized and the foreground region, to obtain a second weight set;
[0021] fusing the similarity set according to the first weight set and the second weight set to obtain the first fused confidence.
[0022] The fusing of the similarity set according to the first weight set and the second weight set to obtain the first fused confidence comprises:
[0023] for each similarity in the similarity set, determining a fusion weight corresponding to the current similarity, a filtering weight and a first product between the similarities;
[0024] adding all the first products corresponding to the similarities in the current feature set to obtain a first sum;
[0025] obtain a hierarchical parameter corresponding to the current feature set, and determine the first fusion confidence according to the hierarchical parameter and the first sum, wherein the hierarchical parameter is positively correlated with the number of block features in the feature set.
[0026] wherein the target coordinate of the target position in the to-be-identified image is determined, comprising:
[0027] a coordinate of a center pixel point of the to-be-identified image is determined as the target coordinate; or, a foreground region in the to-be-identified image is obtained; a coordinate of a center pixel point of the foreground region is determined as the target coordinate.
[0028] wherein the at least two first fusion confidences corresponding to the at least two feature sets are fused to obtain a second fusion confidence of the global feature and the text feature, comprising:
[0029] obtain a total number of all block features in the at least two feature sets;
[0030] a ratio of a sum of the at least two first fusion confidences to the total number is determined as the second fusion confidence.
[0031] wherein the target recognition method further comprises:
[0032] in the case of multiple text features, a second fusion confidence of the global feature and each text feature is determined respectively to obtain a target confidence set;
[0033] a class corresponding to a text feature corresponding to a maximum value in the target confidence set is determined as the target recognition result.
[0034] To solve the above technical problems, the present application further provides a target recognition device, which comprises a memory and a processor coupled with the memory; wherein the memory is used to store program data, and the processor is used to execute the program data to realize the target recognition method as described above.
[0035] To solve the above technical problems, the present application further provides a computer storage medium, which is used to store program data, and the program data is used to realize the target recognition method as described above when executed by a computer.
[0036] Compared with the prior art, the application has the beneficial effects that: the to-be-recognized image is divided into layers and blocks. Due to the difference in the number of blocks in each layer feature set, different levels of feature sets can capture local semantic information of different granularities, improving the capturing ability of high-discriminative local semantics in the target recognition process; the similarity between the block features in each layer feature set and the text features is calculated respectively to obtain a similarity set, improving the accuracy of the semantic in cross-modal matching and the accuracy of target recognition; the similarity set corresponding to each layer can reflect the matching degree of each block feature and the text feature, and the first fusion confidence of each layer feature set is determined according to the similarity set, realizing the complementary use of local semantic information. By fusing the first fusion confidence corresponding to the multi-layer feature set, the second fusion confidence is obtained, which can comprehensively utilize the multi-dimensional semantic information from the local to the whole in the to-be-recognized image, improve the matching precision of the to-be-recognized image features and the text features from a single level to multi-level fusion, and further improve the accuracy of cross-text and image target recognition. BRIEF DESCRIPTION OF DRAWINGS
[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0038] Among them:
[0039] Figure 1 is a framework schematic diagram of an embodiment of the target recognition model provided by the present application;
[0040] Figure 2 is a flowchart of a first embodiment of the target recognition method provided by the present application;
[0041] Figure 3 is a flowchart of determining the first fusion confidence of the current feature set provided by the present application;
[0042] Figure 4 is a schematic diagram of determining the filtering weight in an optional embodiment provided by the present application;
[0043] Figure 5 is a schematic diagram of determining the weight in an embodiment provided by the present application;
[0044] Figure 6 is a structural schematic diagram of an embodiment of the target recognition device provided by the present application;
[0045] Figure 7 is a structural schematic diagram of an embodiment of the target recognition device provided by the present application;
[0046] Figure 8 is a structural schematic diagram of an embodiment of the computer storage medium provided in the present application. DETAILED DESCRIPTION
[0047] The technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work belong to the scope of protection of the present application.
[0048] The terms "first", "second", "third", "fourth" and the like (if any) in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to the process, method, product or device.
[0049] The target recognition method of the present application is applied to a target recognition model, please refer to Figure 1 , for details. Figure 1 is a framework schematic diagram of an embodiment of the target recognition model provided in the present application.
[0050] As shown in Figure 1 , the target recognition model of the present application includes an image feature extraction network, a text feature extraction network, a block module, a similarity calculation module and a confidence fusion module.
[0051] The image feature extraction network in the target recognition model is input with the image to be recognized, and the image feature extraction network extracts the global features of the image to be recognized. The text sentence of the descriptive category is input into the text feature extraction network, and the text feature extraction network extracts the text embedding features of the text sentence to generate text features. In the case of multiple categories, features are extracted for the text sentence corresponding to each category to generate a text feature sequence, wherein the text feature sequence includes the text features corresponding to each category, and each text feature is represented by a feature vector.
[0052] The block module is configured to divide the global feature, and the block module includes at least two block layers, each block layer is configured to perform block processing on the global feature to obtain a feature set corresponding to the block layer, and the block module outputs at least two feature sets, wherein, when each block layer performs block processing on the global feature, different division strategies are adopted, so that the number of block features in different layer feature sets is different.
[0053] The similarity calculation module is configured to calculate the similarity between each block feature in each layer feature set and the text feature to obtain at least two similarity sets, and the at least two layer feature sets and the at least two similarity sets correspond to each other, that is, a single similarity set records the similarity set of each block feature in the corresponding single layer feature set.
[0054] The confidence fusion module is configured to fuse the similarities in the similarity set to obtain a confidence of the global feature and the text feature, that is, a second fusion confidence, so as to determine a target recognition result of the to-be-recognized image according to the second fusion confidence.
[0055] Based on the model framework of Figure 1 , please continue to refer to Figure 2 , Figure 2 is a flowchart of the first embodiment of the target recognition method provided by the present application.
[0056] The target recognition method of the present application is applied to a target recognition device, wherein the target recognition device of the present application can be a server, a terminal device, or a system in which the server and the terminal device cooperate with each other. Accordingly, each part of the target recognition device, such as each unit, sub-unit, module, and sub-module, can be all arranged in the server, all arranged in the terminal device, or arranged in the server and the terminal device respectively.
[0057] Further, the above-mentioned server can be hardware or software. When the server is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or as a single server. When the server is software, it can be implemented as multiple software or software modules, such as software or software modules for providing a distributed server, or as a single software or software module, which is not limited here.
[0058] It should be noted that the target recognition device of the present application can be an intelligent terminal loaded with the target recognition model shown in Figure 1 .
[0059] As shown in Figure 2 , the specific steps are as follows:
[0060] Step S11: extracting a global feature of a to-be-recognized image.
[0061] In the embodiment of the present application, the target recognition model extracts features of the to-be-recognized image through the image feature extraction network to obtain global features of the to-be-recognized image, wherein the global features can be represented as I a , and the target recognition model is used to identify the category of the target in the to-be-recognized image.
[0062] Step S12: obtaining at least two feature sets divided based on the global features, wherein the number of feature blocks in a first feature set in the at least two feature sets is different from the number of feature blocks in a second feature set.
[0063] In the embodiment of the present application, the at least two feature sets are obtained by a block division module based on the global features, the block division module is divided into at least two block layers, and the global features are divided into at least two layers, and each layer is divided into a feature set. Therefore, the block division module obtains at least two feature sets.
[0064] The first feature set and the second feature set are any two different feature sets in the at least two feature sets, and the number of feature blocks in the first feature set is different from the number of feature blocks in the second feature set, that is, the number of feature blocks corresponding to the feature sets in different layers in the at least two feature sets is different.
[0065] The number of feature blocks corresponding to the feature sets in different layers is different because the block division strategies in different layers are different. For example, the block division module adopts average feature block division, but the size of the feature blocks in each layer is different. For example, the block division module includes three block layers, namely a first block layer, a second block layer and a third block layer.
[0066] The size of the block features divided by the first block layer is the same as that of the global features, that is, the first block layer does not divide the global features, and the feature set corresponding to the first block layer is represented as:
[0067] I1={I a}
[0068] I1 is the feature set corresponding to the first block layer, I a is the global feature.
[0069] The block features divided by the second block layer are 1 / 9 of the global features, that is, the second block layer divides the global features into nine block features with the same size, and the feature set corresponding to the second block layer is represented as:
[0070] I2={I 21 , I 22 , …, I 2i , …, I 29}
[0071] I2 is the feature set corresponding to the second block layer, I2i is the i-th patch feature in the feature set corresponding to the second patch layer, i = 1, 2, …, 9.
[0072] The patch features divided by the third patch layer are 1 / 25 of the global feature, that is, the third patch layer divides the global feature into 25 patch features of the same size, and the feature set corresponding to the third patch layer is represented as:
[0073] I3= {I 31 , I 32 , …, I 3j , …, I 325}
[0074] I3is the feature set corresponding to the third patch layer, I 3j is the j-th patch feature in the feature set corresponding to the second patch layer, j = 1, 2, …, 25.
[0075] Step S13: Obtain a similarity set of each layer feature set and a text feature.
[0076] In the embodiments of the present application, for each layer feature set, the similarity calculation module in the target recognition model calculates the similarity between the feature set and the text feature, and obtains a similarity set, which includes the similarity between each patch feature in the feature set and the text feature, and the text feature is denoted as T1.
[0077] Step S14: Determine a first fusion confidence of each layer feature set based on the similarity set.
[0078] Step S15: Fuse at least two first fusion confidences corresponding to at least two layer feature sets to obtain a second fusion confidence of the global feature and the text feature.
[0079] In the embodiments of the present application, the similarity sets corresponding to each layer feature set are fused by the confidence fusion in the target recognition model, which is divided into two steps: each similarity set is fused to determine the first fusion confidence corresponding to each layer feature set, at least two first fusion confidences corresponding to at least two layer feature sets are fused to obtain the second fusion confidence.
[0080] Step S16, determining a target recognition result of the to-be-recognized image according to the second fusion confidence.
[0081] In the embodiments of the present application, the target recognition result of the to-be-recognized image is determined according to the second fusion confidence determined in step S15, wherein the target recognition result is used to represent the category of the target in the to-be-recognized image.
[0082] In the case of only one text feature, the category corresponding to the text feature is the first category, the second fusion confidence is used to represent the possibility that the target in the to-be-identified image is the first category, and whether the target in the to-be-identified image belongs to the first category can be determined according to the second fusion confidence. For example, in the case where the second fusion confidence is greater than or equal to a confidence threshold, the target recognition result is that the target in the to-be-identified image belongs to the first category, and otherwise, in the case where the second fusion confidence is less than the confidence threshold, the target recognition result is that the target in the to-be-identified image does not belong to the first category.
[0083] Optionally, in the case of multiple text features, the second fusion confidence of each text feature and the global feature is determined respectively to obtain a target confidence set; and the category corresponding to the text feature corresponding to the maximum value in the target confidence set is determined as the target recognition result.
[0084] In the embodiments of the present application, multiple categories are included in the text sentence, so that the text feature extraction module can extract multiple text features, each text feature representing a category, the second fusion confidence corresponding to each text feature and the global feature is calculated to obtain a target confidence set, wherein the category of the text feature corresponding to the maximum value in the target confidence set is the second category, and the target in the to-be-identified image has the highest probability of belonging to the second category among all categories. Therefore, the second category is the target recognition result, that is, the category corresponding to the text feature corresponding to the maximum value in the target confidence set is determined as the target recognition result.
[0085] The target recognition result can also be determined in combination with the target confidence set and a confidence threshold. For example, the maximum value in the target confidence set is determined, and in the case where the maximum value is greater than or equal to the confidence threshold, the category corresponding to the text feature corresponding to the maximum value in the target confidence set is determined as the target recognition result, and otherwise, in the case where the maximum value is less than the confidence threshold, the target recognition result is that it does not belong to any of the above multiple categories.
[0086] Through the above embodiment, the to-be-recognized image is divided into layers and blocks. Due to the difference in the number of blocks in each layer feature set, the feature sets at different levels can capture local semantic information of different granularities, improving the ability to capture high-discriminative local semantics in the target recognition process; the similarity between each block feature in each layer feature set and the text feature is calculated to obtain a similarity set, improving the accuracy of the semantic in cross-modal matching and improving the accuracy of target recognition; the similarity set corresponding to each layer can reflect the matching degree of each block feature and the text feature, and the first fusion confidence of each layer feature set is determined according to the similarity set, realizing the complementary use of local semantic information; the first fusion confidence corresponding to the multi-layer feature set is fused to obtain the second fusion confidence, which can comprehensively utilize the multi-dimensional semantic information from the local to the whole in the to-be-recognized image, so that the matching accuracy of the image features of the to-be-recognized image and the text features is improved from a single level to multi-level fusion, further improving the accuracy of cross-text and image target recognition.
[0087] Further, in the process of determining the first fusion confidence of each layer feature set based on the similarity set by the confidence fusion module, the same fusion processing steps are performed on the similarity sets of at least two layer feature sets to determine the first fusion confidence of each layer feature set. When processing the similarity set of one of the layer feature sets, the layer feature set is recorded as the current feature set, and the similarity set corresponding to the current feature set is recorded as the current similarity set.
[0088] The first fusion confidence of the current feature set is determined based on the current similarity set, but each block feature in the current feature set corresponds to a different spatial position in the to-be-recognized image, so the contribution degree is also different when fusing. Therefore, the contribution degree of each block feature, i.e., the fusion weight, needs to be determined when fusing, and all similarity sets in the current similarity set are fused according to the fusion weight corresponding to each block feature.
[0089] In this embodiment, the spatial attention distribution of each block feature is simulated, that is, the closer the block feature to the target position, the higher the contribution degree to the target recognition, that is, the larger the fusion weight, and the fusion weight gradually decays in the horizontal and vertical dimensions from the target position according to the Gaussian distribution.
[0090] For details, please continue to refer to Figure 3 , Figure 3 is a flowchart for determining the first fusion confidence of the current feature set provided by the present application, as shown in Figure 3 , and the specific steps are as follows:
[0091] Step S41: determining the target coordinates of the target position in the to-be-recognized image.
[0092] In the embodiment of the present application, first, the position of the block feature with the highest contribution degree to target recognition in the to-be-recognized image, i.e., the target position, is determined.
[0093] Optionally, a position of a block feature in which a center pixel point of the to-be-identified image is located is determined as the target position, a target coordinate corresponding to the target position is a coordinate of the center pixel point, that is, a coordinate of the center pixel point of the to-be-identified image is determined as the target coordinate, or a target coordinate corresponding to the target position is a coordinate of a center point of a region in which the block feature is located.
[0094] Preferably, since the target in the to-be-identified image is not necessarily located at the center of the image, in order to further improve the accuracy of target identification, a region in which the target is located, that is, a foreground region, is determined, a position of a block feature in which a center pixel point of the foreground region is located is determined as the target position, a target coordinate corresponding to the target position is a coordinate of the center pixel point of the foreground region, that is, the foreground region in the to-be-identified image is obtained, a coordinate of the center pixel point of the foreground region is determined as the target coordinate, or a target coordinate corresponding to the target position is a coordinate of a center point of a region in which the block feature is located.
[0095] Step S42: generating a fusion weight corresponding to each block feature in the current feature set according to the target coordinate and a target variance corresponding to the to-be-identified image, to obtain a first weight set, wherein the target variance includes a first variance for controlling a horizontal weight decay speed and a second variance for controlling a vertical weight decay speed.
[0096] In the embodiment of the present application, after the target coordinate is determined, the target variance is obtained, the target variance controls a speed at which the weight decays with the distance between the block feature and the target coordinate, the target variance includes a first variance and a second variance, the first variance σ1 controls the horizontal weight decay speed, the greater σ1 is, the slower the horizontal decay speed is, and the greater σ1 is, the faster the horizontal decay speed is; the second variance σ2 controls the vertical weight decay speed, the greater σ2 is, the slower the vertical decay speed is, and the greater σ2 is, the faster the vertical decay speed is, the target variance is a preset value, and a fusion weight corresponding to each block feature is determined based on the target variance and the target coordinate, to obtain the first weight set.
[0097] Optionally, generating a fusion weight corresponding to each block feature in the current feature set according to the target coordinate and a target variance corresponding to the to-be-identified image, to obtain a first weight set, includes: for each block feature in the current feature set, obtaining a position coordinate corresponding to the current block feature in the to-be-identified image; and determining a fusion weight corresponding to the current block feature according to the position coordinate, the target coordinate and the target variance, wherein the fusion weight corresponding to the block feature in the current feature set is in a Gaussian distribution.
[0098] In the embodiments of the present application, for each patch feature in the current feature set, a corresponding fusion weight is determined, and in determining the fusion weight of any patch feature, the patch feature is recorded as a current patch feature.
[0099] The position coordinates corresponding to the current patch feature in the image to be recognized are obtained, which can be specifically the coordinates of the center pixel point of the region where the patch feature is located; the position coordinates, the target coordinates and the target variance can be used to determine the fusion weight corresponding to the current patch feature, and specifically, the position coordinates, the target coordinates and the target variance are brought into the Gaussian attenuation formula to solve the corresponding fusion weight.
[0100] Optionally, the Gaussian attenuation formula is represented as:
[0101]
[0102] wherein, GW i is the fusion weight of the current patch feature, i is the serial number of the current patch feature, σ1 is the first variance, σ2 is the second variance, (x, y) is the position coordinates corresponding to the current patch feature, and (x0, y0) is the target coordinates.
[0103] In the embodiments, through the above-mentioned manner of assigning fusion weights, the patch features closer to the target position are assigned with higher weights, thereby improving the accuracy of the final cross-modal matching.
[0104] In the training phase of the model, the above-mentioned calculation manner of fusion weight is used to calculate the weighted loss function.
[0105] Step S43: fusing the similarity set according to the first weight set to obtain the first fused confidence.
[0106] After the fusion weight corresponding to each patch feature in the current feature set is calculated, the similarity set corresponding to the current feature set is fused to obtain the first fused confidence.
[0107] Optionally, each similarity in the similarity set corresponding to the current feature set is multiplied by the fusion weight corresponding to the similarity, and the products obtained are all added to obtain the first fused confidence.
[0108] Preferably, a foreground region where the target is located in the image to be recognized is obtained; a filtering weight corresponding to each patch feature in the current feature set is determined according to the relative position of the region corresponding to the patch feature in the image to be recognized and the foreground region, to obtain a second weight set; and the similarity set is fused according to the first weight set and the second weight set to obtain the first fused confidence.
[0109] In the embodiments of the present application, the text features usually only contain semantic information of the corresponding category and do not contain background information other than the category, while the global features of the image are completely different. In the image to be recognized, the target usually does not occupy the entire image, so the global features often contain a large amount of background information, which will reduce the final matching accuracy, especially when the foreground target accounts for a small proportion in the entire image to be recognized, which may cause target recognition failure.
[0110] In the embodiments, since the global features are processed by block, the region corresponding to a single block feature may be all background region, and the block feature does not contain any target information. If the block feature is matched with the text feature, it will affect the final target recognition result. In order to further improve the accuracy of target recognition, the background in the block feature is filtered out. For each block feature, the filtering weight corresponding to the block feature is determined according to the relative position of the region corresponding to the block feature and the foreground region where the target is located. For example, when the region corresponding to the block feature and the foreground region are completely overlapped, the filtering weight of the block feature is determined as zero, and when the region corresponding to the block feature and the foreground region at least partially overlap, the filtering weight of the block feature is determined as 1.
[0111] The image to be recognized is input into the segmentation network, the target in the image to be recognized is segmented from the background, and the compact real boundary of the foreground target is obtained. The boundary is input into each block layer. If all pixels in the region corresponding to a certain block feature in the block layer do not fall into the boundary, the region corresponding to the block feature and the foreground region are completely overlapped, and the filtering weight of the block feature is determined as zero. If at least one pixel in the region corresponding to a certain block feature falls into the boundary, the region corresponding to the block feature and the foreground region at least partially overlap, and the filtering weight of the block feature is determined as 1.
[0112] Taking the target in the image to be recognized as a truck as an example, please refer to Figure 4 , Figure 4 The schematic diagram for determining the filtering weight in an optional embodiment of the present application is shown in FIG. 2. Figure 4As shown, the image to be recognized is input into the segmentation network, the compact real boundary of the truck in the image to be recognized is recognized, and the real boundary is sent to each sub-block layer. There is only one sub-block feature in the feature set 1 corresponding to the first sub-block layer, and part of the pixels in the sub-block feature fall into the boundary. Therefore, the filtering weight of the sub-block feature is 1. There are 9 sub-block features in the feature set 2 corresponding to the second sub-block layer, part of the sub-block features do not have pixels falling into the boundary, and the filtering weights corresponding to the sub-block features are determined as 0. The remaining sub-block features have at least part of the pixels falling into the boundary, and the filtering weights corresponding to the sub-block features are determined as 1. There are 25 sub-block features in the feature set 3 corresponding to the third sub-block layer, part of the sub-block features do not have pixels falling into the boundary, and the filtering weights corresponding to the sub-block features are determined as 0. The remaining sub-block features have at least part of the pixels falling into the boundary, and the filtering weights corresponding to the sub-block features are determined as 1.
[0113] Preferably, the filtering weight corresponding to each sub-block feature in the current feature set is determined according to the relative position of the region corresponding to the sub-block feature in the image to be recognized and the foreground region, to obtain a second weight set, including: for each sub-block feature, determining the proportion of pixels falling into the foreground region in the region corresponding to the current sub-block feature; determining the filtering weight corresponding to the current sub-block feature according to the proportion.
[0114] In the embodiments of the present application, the filtering weight corresponding to each sub-block feature is determined, and when determining the filtering weight of the current sub-block feature, the number of common pixels between the region of the current sub-block feature in the image to be recognized and the foreground region is determined. The ratio of the number of common pixels to the total number of pixels in the region corresponding to the current sub-block feature is determined as the above-mentioned proportion, and the filtering weight corresponding to the current sub-block feature is determined according to the proportion, for example, the proportion value is directly determined as the corresponding filtering weight.
[0115] In an optional embodiment, the similarity set is fused according to the first weight set and the second weight set to obtain the first fusion confidence, including: for each similarity in the similarity set, determining the first product of the fusion weight, the filtering weight corresponding to the current similarity and the similarity; all the first products corresponding to all similarities in the current feature set are added to obtain a first sum; the hierarchical parameter corresponding to the current feature set is obtained, and the first fusion confidence is determined according to the hierarchical parameter and the first sum, wherein the hierarchical parameter is positively correlated with the number of sub-block features in the feature set.
[0116] After the first weight set and the second weight set are determined, the similarity set corresponding to the current feature set is fused based on the first weight set and the second weight set, specifically including multiplying each similarity in the similarity set by the fusion weight and the filtering weight corresponding to the similarity, and adding all the products obtained to obtain a first sum.
[0117] The hierarchical parameter corresponding to the current feature set is obtained, and the hierarchical parameter is positively correlated with the number of block features, that is, the more the number of block features, the greater the hierarchical parameter corresponding to the current feature set. Optionally, the hierarchical parameters of the hierarchical blocks can be sequentially increased from small to large according to the number of block features, for example, the hierarchical parameter of the first block layer can be 1, the hierarchical parameter of the second block layer can be 2, and the hierarchical parameter of the third block layer can be 3; or on the basis of average block division in the horizontal direction and the vertical direction, the hierarchical parameter is determined as the number of block features in the horizontal direction or the vertical direction.
[0118] The smaller the number of block features means that each block has more target features, and the larger the number of block features means that each block feature contains fewer target semantic features. The hierarchical parameter is set to give a higher weight to a feature set with fewer blocks, further improving the accuracy of target recognition.
[0119] The ratio of the first sum to the hierarchical parameter is determined as the first fusion confidence, and specifically, the first fusion confidence corresponding to the current feature set can be represented by the following formula:
[0120]
[0121] Wherein, S is the first fusion confidence corresponding to the current feature set, K is the hierarchical number corresponding to the current feature set, n represents the number of block features in the current feature set, i represents the serial number of each block feature in the current feature set, WS i is the filtering weight of the ith block feature, GW i is the fusion weight of the ith block feature, A i is the similarity between the ith block feature and the text feature.
[0122] In an optional embodiment, the at least two first fusion confidences corresponding to the at least two layers of feature sets are fused to obtain a second fusion confidence of the global feature and the text feature, including: obtaining the total number of all block features in the at least two layers of feature sets; and determining the ratio of the sum of the at least two first fusion confidences to the total number as the second fusion confidence.
[0123] In the embodiment of the application, the first fusion confidences corresponding to all feature sets, that is, the at least two layers of feature sets, are added, and the sum of all first fusion confidences is divided by the total number of all block features to fuse and normalize all first fusion confidences to obtain a second fusion confidence.
[0124] Please refer to Figure 5 , Figure 5 is a schematic diagram for determining the weight in an embodiment provided by the application, asFigure 5 As shown, the image to be identified is input into the segmentation network, each block layer in different block layers is determined based on the feature filtering strategy of the segmentation result, the corresponding filtering weight, and then the final fusion weight is generated based on the filtering weight combined with the Gaussian attenuation, the fusion weight corresponding to each block feature whose filtering weight is not 0 is determined according to the Gaussian attenuation, the fusion weight of the block feature whose filtering weight is 0 is directly set to 0, and the target weight set corresponding to each layer feature set is obtained. The first fusion confidence corresponding to at least two layer feature sets is fused to obtain the second fusion confidence.
[0125] Those skilled in the art can understand that in the above method of the specific embodiment, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.
[0126] To achieve the above target recognition method, the present application also provides a target recognition device, please refer to Figure 6 , Figure 6 is a structural schematic diagram of an embodiment of the target recognition device provided by the present application.
[0127] The target recognition device 500 of the embodiment comprises:
[0128] The extraction module 51 is configured to extract a global feature of an image to be identified.
[0129] The first acquisition module 52 is configured to acquire at least two layer feature sets divided based on the global feature, wherein the number of feature blocks in a first layer feature set is different from the number of feature blocks in a second layer feature set.
[0130] The second acquisition module 53 is configured to acquire a similarity set of each layer feature set and a text feature.
[0131] The first determination module 54 is configured to determine a first fusion confidence of each layer feature set based on the similarity set.
[0132] The fusion module 55 is configured to fuse at least two first fusion confidences corresponding to at least two layer feature sets to obtain a second fusion confidence of the global feature and the text feature.
[0133] The determination module 56 is configured to determine a target recognition result of the image to be identified according to the second fusion confidence.
[0134] To achieve the above target recognition method, the present application also provides a target recognition device, please refer to Figure 7 , Figure 7is a structural schematic diagram of an embodiment of the target identification device provided in the present application.
[0135] The target identification device 400 in the embodiment comprises a processor 41, a memory 42, an input / output device 43 and a bus 44.
[0136] The processor 41, the memory 42 and the input / output device 43 are connected to the bus 44 respectively, the memory 42 stores program data, and the processor 41 is used to execute the program data to realize the target identification method described in the above embodiments.
[0137] In the embodiment of the present application, the processor 41 can also be referred to as a CPU (Central Processing Unit). The processor 41 can be an integrated circuit chip with signal processing capability. The processor 41 can also be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor 41 can also be any conventional processor.
[0138] The present application also provides a computer storage medium, please continue to refer to Figure 8 , Figure 8 is a structural schematic diagram of an embodiment of the computer storage medium provided in the present application. The computer storage medium 600 stores a computer program 61. When the computer program 61 is executed by a processor, the target identification method described in the above embodiments is realized.
[0139] The embodiments of the present application are realized in the form of software function units and sold or used as independent products, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application or the whole or part of the technical solutions that essentially contribute to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the method described in the various embodiments of the present application. The aforementioned storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0140] The above description is only the embodiments of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation using the content of the specification and drawings, or direct or indirect application in other related technical fields, is also included in the patent protection scope of the present application.
Claims
1. A target recognition method characterized by, The target recognition method comprises: extracting global features of an image to be recognized; obtaining at least two sets of features divided based on the global features, wherein the number of feature blocks in a first set of features is different from the number of feature blocks in a second set of features; obtaining a set of similarities between each set of features and a text feature; determining a first fusion confidence of each set of features based on the set of similarities; fusing at least two first fusion confidences corresponding to the at least two sets of features to obtain a second fusion confidence of the global features and the text feature; determining a target recognition result of the image to be recognized according to the second fusion confidence.
2. The target recognition method of claim 1, wherein determining a first fusion confidence of each set of features based on the set of similarities comprises: determining target coordinates of a target position in the image to be recognized; generating a fusion weight corresponding to each block feature in a current set of features according to the target coordinates and a target variance corresponding to the image to be recognized to obtain a first weight set, wherein the target variance comprises a first variance for controlling a horizontal weight decay rate and a second variance for controlling a vertical weight decay rate; fusing the set of similarities according to the first weight set to obtain the first fusion confidence.
3. The object recognition method of claim 2, wherein, generating a fusion weight corresponding to each block feature in a current set of features according to the target coordinates and a target variance corresponding to the image to be recognized to obtain a first weight set comprises: for each block feature in the current set of features, obtaining a position coordinate corresponding to the current block feature in the image to be recognized; determining a fusion weight corresponding to the current block feature according to the position coordinate, the target coordinates and the target variance, wherein the fusion weight corresponding to each block feature in the current set of features is in a Gaussian distribution.
4. The object recognition method of claim 2, wherein, fusing the set of similarities according to the first weight set comprises: obtaining a foreground region where the target is located in the image to be recognized; determining a filtering weight corresponding to each block feature in the current set of features according to a relative position between a region corresponding to each block feature in the image to be recognized and the foreground region to obtain a second weight set; fusing the set of similarities according to the first weight set and the second weight set to obtain the first fusion confidence.
5. The target recognition method of claim 4, wherein fusing the set of similarities according to the first weight set and the second weight set to obtain the first fusion confidence comprises: for each similarity in the set of similarities, determining a fusion weight, a filtering weight corresponding to the current similarity and a first product between similarities; adding all the first products corresponding to all similarities in the current set of features to obtain a first sum; obtaining a hierarchical parameter corresponding to the current set of features and determining the first fusion confidence according to the hierarchical parameter and the first sum, wherein the hierarchical parameter is positively correlated with the number of block features in the set of features.
6. The target recognition method of claim 2, wherein The target coordinate of a target position in the image to be recognized is determined, including: The coordinate of a center pixel point of the image to be recognized is determined as the target coordinate; or, a foreground region in the image to be recognized is acquired; and the coordinate of a center pixel point of the foreground region is determined as the target coordinate.
7. The object recognition method of claim 1, wherein, The at least two first fusion confidences corresponding to the at least two layers of feature sets are fused to obtain a second fusion confidence of the global feature and the text feature, including: The total number of all block features in the at least two layers of feature sets is acquired; The ratio of the sum of the at least two first fusion confidences to the total number is determined as the second fusion confidence.
8. The object recognition method of claim 1, wherein, The target recognition method further includes: In the case of multiple text features, the second fusion confidence of the global feature and each text feature is respectively determined to obtain a target confidence set; The class corresponding to the text feature corresponding to the maximum value in the target confidence set is determined as the target recognition result.
9. A target identification device, characterized by The target recognition device includes a memory and a processor coupled with the memory; The memory is configured to store program data, and the processor is configured to execute the program data to implement the target recognition method according to any one of claims 1 to 8.
10. A computer storage medium, characterized in that, The computer storage medium is configured to store program data, and the program data, when executed by a computer, is configured to implement the target recognition method according to any one of claims 1 to 8.