Cross-modal object detection method, training method and related devices

By separating the matching branches and positioning branches in cross-modal object detection, and performing semantic information matching and positioning processing respectively, the problem of cross-modal information in the prior art combined affecting detection accuracy is solved, and a higher target detection accuracy is achieved.

CN119559643BActive Publication Date: 2025-05-30ZHEJIANG DAHUA TECH CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510132777.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-05-30
Estimated Expiration
2045-02-06

AI Technical Summary

Technical Problem

When existing cross-modal object detection technology combines multiple different mode information, it can easily affect the accuracy of the task, resulting in low accuracy of object detection.

Method used

A cross-modal object detection method is proposed. By obtaining the image to be detected and the description text, feature extraction of the image and text, using matching branches to be matched for semantic information, and using positioning branches to perform positioning processing, explicitly separating the matching branches and positioning branches to improve detection accuracy.

Benefits of technology

By separating the matching branches and positioning branches, the semantic information matching and positioning accuracy are enhanced respectively, and the accuracy of target detection is generally improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119559643B_ABST
    Figure CN119559643B_ABST
Patent Text Reader

Abstract

The present application discloses a cross-modal object detection method, a training method and related devices. The method includes: obtaining an image to be detected and a descriptive text; respectively performing feature extraction on the image to be detected and the descriptive text to obtain image features and text features; using a matching branch to perform matching processing on the image features and the text features to obtain the category of the object; and using a localization branch to perform localization processing on the image features to obtain a localization map of the object. The above solution can improve the accuracy of object detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of object detection, and in particular, to a cross-modal object detection method, a training method, and related devices. Background Art

[0002] As one of the very important tasks in the field of computer vision, object detection technology has achieved rapid development with the rapid development of deep learning methods.

[0003] Currently, cross-modal object detection technology can be used in object detection tasks. The cross-modal object detection task mainly combines multiple different modal information to perform the object detection task. However, combining multiple different modal information may affect the accuracy of the task, resulting in low accuracy of object detection. Summary of the Invention

[0004] The main technical problem to be solved by this application is to provide a cross-modal object detection method, a training method, and related devices, which can improve the accuracy of object detection.

[0005] In the first aspect of this application, a cross-modal object detection method is provided. The method includes: obtaining an image to be detected and a description text; respectively performing feature extraction on the image to be detected and the description text to obtain image features and text features; using a matching branch to perform matching processing on the image features and the text features to obtain the category of the target; and using a localization branch to perform localization processing on the image features to obtain a localization map of the target.

[0006] In the second aspect of this application, a training method for an object detection model is provided. The object detection model includes at least a matching branch and a localization branch. The method includes: obtaining an image sample to be detected and a description text sample; respectively performing feature extraction on the image sample to be detected and the description text sample to obtain image features and text features; using the matching branch to perform matching processing on the image features and the text features to obtain the category of the target, and based on the category of the target, obtaining an alignment loss value; and using the localization branch to perform localization processing on the image features to obtain a localization map of the target, and based on the localization map of the target, obtaining a localization loss value; comprehensively combining the alignment loss value and the localization loss value to train the object detection model to obtain a trained object detection model.

[0007] In the third aspect of this application, a computer device is provided. The computer device includes a memory and a processor coupled to each other. The memory stores program data, and the processor is configured to execute the program data to implement any step of the above cross-modal object detection method and the training method for the object detection model.

[0008] A fourth aspect of the present application provides a computer-readable storage medium storing program data that can be run by a processor, and the program data is used to implement any step of the above cross-modal object detection method and the training method of the object detection model.

[0009] In the above solution, by obtaining the image to be detected and the description text, feature extraction is respectively performed on the image to be detected and the description text to obtain image features and text features. Then, the matching branch and the localization branch are explicitly separated. For the matching branch that requires strong semantic features, the matching branch is used to perform matching processing on the image features and the text features to obtain the category of the target. The matching branch can fully perceive the semantic information of the text and the image, and a large amount of semantic information is added to the features of the matching branch, which is beneficial to improving the accuracy of the matching task; while the localization branch does not perceive the text semantic features, and the localization branch is used to perform localization processing on the image features to obtain the localization map of the target, which can improve the accuracy of the localization task. Overall, the accuracy of object detection can be improved.

[0010] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the technical solutions in the present application, the following will briefly introduce the drawings required in the description of the embodiments. Obviously, the following described drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. Among them:

[0012] Figure 1 is a flowchart of an embodiment of the cross-modal object detection method of the present application;

[0013] Figure 2 is a structural diagram of an embodiment of the object detection model of the present application;

[0014] Figure 3 is the present application Figure 1 is a flowchart of an embodiment of step S14 in the present application;

[0015] Figure 4 is a flowchart of the first embodiment of the training method of the object detection model of the present application;

[0016] Figure 5 is a structural diagram of another embodiment of the object detection model of the present application;

[0017] Figure 6 is the present application Figure 4 is a flowchart of an embodiment of step S24 in the present application;

[0018] Figure 7 It is a schematic flowchart of the second embodiment of the training method of the object detection model of the present application;

[0019] Figure 8 It is the present application Figure 7 A schematic flowchart of an embodiment of step S32 in it;

[0020] Figure 9 It is a schematic structural diagram of an embodiment of the cross-modal object detection device of the present application;

[0021] Figure 10 It is a schematic structural diagram of an embodiment of the training device of the object detection model of the present application;

[0022] Figure 11 It is a schematic structural diagram of an embodiment of the computer device of the present application;

[0023] Figure 12 It is a schematic structural diagram of an embodiment of the computer-readable storage medium of the present application. Detailed implementation manners

[0024] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0025] The terms "first" and "second" in the present application are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present application, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes unlisted steps or units, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.

[0026] References to "embodiments" in this application mean that the specific features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of this application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0027] As used herein, the term "and / or" is merely a description of the relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Additionally, the character " / " in this document generally indicates that the associated objects before and after are in an "or" relationship. Furthermore, "plurality" in this document means two or more than two. Additionally, the term "at least one" in this document means any one of a plurality or any combination of at least two of a plurality. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set composed of A, B, and C.

[0028] This application provides the following embodiments, and the following is a specific description of each embodiment.

[0029] It can be understood that the cross-modal object detection method and the training method of the object detection model in this application can be executed by a computer device, which can be any device with processing capabilities, such as a mobile device, a computer, a server, etc. This application places no restrictions on this. In some possible implementation manners, the cross-modal object detection method and the training method of the object detection model can also be implemented by a processor calling program data stored in a memory. This application places no restrictions on the execution device.

[0030] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of an embodiment of the cross-modal object detection method of this application. The method may include the following steps:

[0031] S11: Obtain the image to be detected and the descriptive text.

[0032] The cross-modal object detection of this application can be implemented using an object detection model, where cross-modality can refer to the process of information conversion or association between different modalities. For example, visual modality, text modality, and / or auditory modality, etc. This application places no restrictions on cross-modality. Optionally, this application takes the visual modality and the text modality as examples for illustration.

[0033] The image to be detected for target detection and the descriptive text can be obtained, where the descriptive text is the text content related to the image to be detected. For example, the descriptive text is the description text of the scene of the image to be detected (such as the target, background, color, features, category, etc.). Optionally, the descriptive text can contain at least one character or word. This application places no restrictions on the descriptive text.

[0034] Optionally, the image to be detected can be a visible light image. The scene of the image to be detected can include a target. For example, the target includes but is not limited to: objects, living bodies, moving objects, scenes, texts, etc. This application places no restrictions on the image to be detected.

[0035] Exemplarily, if the descriptive text is "A dog holds a Frisbee". The scene of the image to be detected contains a dog and a Frisbee, etc.

[0036] S12: Feature extraction is performed on the image to be detected and the descriptive text respectively to obtain image features and text features.

[0037] The target detection model can include a feature extraction network. The feature extraction network can be used to perform feature extraction on the image to be detected and the descriptive text respectively to obtain image features and text features.

[0038] Optionally, an image feature extraction network can be used to perform feature extraction on the image to be detected to obtain image features. Among them, the image feature extraction network can include but is not limited to: image encoders (such as Visual Encoder), convolutional neural networks (such as LeNet, ResNet, VGGNet, etc.), graph neural networks, etc. This application takes an image encoder as an example for illustration, and this application places no restrictions on the image feature extraction network.

[0039] Optionally, a text feature extraction network can be used to perform feature extraction on the descriptive text to obtain text features. Among them, the text feature extraction network can include but is not limited to: text encoders (such as Text Encoder), recurrent neural networks, convolutional neural networks, Transformers, etc. In the following, this application takes a text encoder as an example for illustration, and this application places no restrictions on the text feature extraction network.

[0040] S13: The matching branch is used to perform matching processing on the image features and the text features to obtain the category of the target.

[0041] Please refer to Figure 2 , the target detection model can at least include a matching branch and a localization branch. Among them, the matching branch is used to obtain the category of the target. The localization branch is used to obtain the localization map of the target, and the localization map can represent the localization box / detection box / localization coordinates, etc. of the target in the image to be detected.

[0042] Input the image features and text features into the matching branch. Through the matching branch, perform matching processing on the image features and text features to obtain the category of the target. Optionally, a deep fusion module (such as Deep Fusion) can be used to deeply fuse the image features and text features to perform cross-modal fusion on the image features and text features, obtaining cross-modal fusion text fusion features and image fusion features. The text fusion features represent text features fused with image features, and the image fusion features represent image features fused with text features. Then, use the matching module to perform feature matching on the text fusion features and image fusion features to perform category recognition or matching, obtaining the category of the target.

[0043] In some embodiments, the image features and text features can be deeply fused. Specifically, the above-obtained image features and text features After that, the image features and text features can be subjected to multi-modal fusion (such as Fusion) to obtain fusion features, such as text features fused with image features and image features fused with text features , so as to be input into the branch of text features and the branch of image features respectively.

[0044] Then, use the text features (such as ) and fusion features ( ) to perform text encoding (such as text encoding by the text encoder TextEncoder), generating the fused text fusion features . Use the image features (such as ) and fusion features (such as ) to perform image encoding (such as using the image encoder DyHead Module for image encoding), generating the fused image fusion features .

[0045] Refer to the above method for multiple deep fusions. For the text features and fusion features at the i-th (i is an integer greater than 1) time, perform text encoding to obtain the final text fusion features . For the image features and fusion features at the i-th time, perform image encoding to obtain the final image fusion features . Among them, M and N are integers greater than 1. This application does not limit the number of deep fusions.

[0046] Through the above method, text fusion features ( , ,..., ), and image fusion features ( , ,..., ). The text fusion features can represent the features of N texts (such as word segmentation, etc.), and the image fusion features can represent the features of M regions. By using the alignment network to perform feature matching on the text fusion features and the image fusion features respectively, each text and each region can be matched to obtain the matching similarity of each matching pair, so as to obtain the alignment degree of each matching pair (such as Word-Region Alignment Score). Among them, each matching pair can represent the category of a target. Matching pairs with an alignment degree greater than the preset alignment degree can be obtained to obtain the category of the target. Among them, the Word-Region Alignment Score is an index used to measure the alignment degree between words in the text and regions in the image. Specifically, it evaluates the alignment effect by calculating the similarity or matching degree between the features of text phrases and different regions in the image.

[0047] S14: Use the localization branch to perform localization processing on the image features to obtain the localization map of the target.

[0048] Only input the image features into the localization branch, perform localization processing on the image features, and detect the localization map of the target. The localization map is also the detection box / localization box, etc. of the target in the image to be detected, which can represent the position information of the target.

[0049] In some embodiments, please refer to Figure 3 , the step S14 of the above embodiment can be further expanded. Use the localization branch to perform localization processing on the image features to obtain the localization map of the target. This embodiment may include the following steps:

[0050] S141: Perform transition processing on the image features to obtain transition features.

[0051] Continue to refer to Figure 2 , the localization branch may include a localization transition layer, a first localization network, a dimension transformation network, a feature separation network, a second localization network, etc.

[0052] Use the localization transition layer to perform transition processing on the image features to obtain transition features . Among them, the transition features , where W, H, and C are the width, height, and number of channels of the features respectively. The localization transition layer can be a neural network or a convolutional network, etc., and the present application does not limit this.

[0053] Optionally, the image features may not be subjected to transition processing, that is, the image features are used to represent the transition features of this step, and the following steps can be directly executed. The present application does not limit this.

[0054] S142: Perform the first positioning on the transition feature to obtain a first positioning map.

[0055] Use the first positioning network to perform the first positioning on the transition feature to perform a rough bounding box positioning once, and obtain a first positioning map , where N is the number of output channels of the first positioning. In some embodiments, the first positioning network performs the first positioning on the transition feature to obtain a first positioning map The process can be expressed as:

[0056] .

[0057] Wherein, represents the first positioning network.

[0058] Since in principle only the positioning-related information is retained in the first positioning map. In other words, for the first positioning map, the features that are not related to positioning will be gradually filtered out during the training process. Therefore, the first positioning map can be used as a positioning feature filtering factor (such as a relatively pure positioning feature filtering factor).

[0059] S143: Use the first positioning map to perform feature separation on the transition feature to obtain separated features.

[0060] The first positioning map can be used to perform feature separation on the transition feature to separate the features that are not related to positioning in the transition feature to obtain separated features . This process can filter and separate to a certain extent the features that are not related to positioning in the transition feature of the positioning transition layer, and then the remaining features are mostly positioning-aware feature information, making the subsequent positioning more accurate.

[0061] In some embodiments, the first positioning map can be dimensionally transformed through a dimensional transformation network to make the dimension of the first positioning map factor into a feature with the same dimension as the transition feature, and obtain a feature separation factor . Wherein, the dimension of the feature separation factor is the same as the dimension of the transition feature .

[0062] Optionally, the dimensional transformation network can include: a network of an inverse prediction layer + an activation function (such as a sigmoid activation function). The present application does not limit this.

[0063] Exemplarily, perform dimensional transformation on the first positioning map to obtain a feature separation factor The process can be expressed as follows:

[0064] .

[0065] Among them, represents the anti-prediction layer, which can be composed of 2 convolutional layers. After convolution processing by the anti-prediction layer, and then processed through the sigmoid activation function, the feature separation factor can be obtained, which is also a positioning feature filtering and separation factor. Since the positioning-unrelated features in the first positioning map are gradually filtered out, the feature filtering factor generated by the anti-prediction layer generally only contains activation information related to positioning in principle.

[0066] Then, the feature separation network is adopted to use the feature separation factor to perform feature separation (such as multiplication) on the transitional feature , and the separated feature can be obtained.

[0067] Optionally, the feature separation factor and the learnable bias term bias can be used to perform feature separation on the transitional feature , and the separated feature can be obtained.

[0068] Exemplarily, the separated feature can be expressed by the following formula:

[0069] .

[0070] Among them, represents element-wise multiplication, and bias is a learnable bias term. By separating the transitional feature in the above manner, the positioning-unrelated features in the transitional feature of the positioning transitional layer are filtered and separated to a certain extent, that is, what remains is mostly positioning-aware feature information.

[0071] S144: Perform a second positioning on the separated feature to obtain a second positioning map, where the second positioning map serves as the positioning map of the target in the image to be detected.

[0072] After separating the features unrelated to positioning, a second positioning network (such as a fine coordinate regressor) is then used to perform a second positioning on the separated feature, that is, coordinate regression can be performed to obtain a more accurate second positioning map . Among them, the second positioning map can be used as the final positioning map, that is, the second positioning map can be used as the positioning map of the target in the image to be detected.

[0073] Thus, by comprehensively matching the category of the target in the matching branch and the localization map of the target in the localization branch, a cross-modal object detection result is obtained.

[0074] In the above solution, by obtaining the image to be detected and the descriptive text, feature extraction is respectively performed on the image to be detected and the descriptive text to obtain image features and text features. Then, the matching branch and the localization branch are explicitly separated. For the matching branch that requires strong semantic features, the matching branch is used to match the image features and the text features to obtain the category of the target. The matching branch can fully perceive the semantic information of the text and the image, and a large amount of semantic information is added to the features of the matching branch, which is beneficial to improving the accuracy of the matching task; while the localization branch does not perceive the text semantic features and requires more detailed edge features. The localization branch is used to perform localization processing on the image features to obtain the localization map of the target, which can improve the accuracy of the localization task. Overall, the accuracy of object detection can be improved.

[0075] In some embodiments, the above object detection model can be trained so that the trained object detection model performs cross-modal object detection. The training process of the object detection model can refer to the following embodiments.

[0076] Please refer to Figure 4 , Figure 4 which is a schematic flowchart of the first embodiment of the training method of the object detection model of the present application. The method may include the following steps:

[0077] S21: Obtain an image sample to be detected and a descriptive text sample.

[0078] A number of image samples to be detected and descriptive text samples can be obtained. The image samples to be detected and the descriptive text samples can refer to the descriptions in the above embodiments and will not be elaborated here.

[0079] S22: Respectively perform feature extraction on the image sample to be detected and the descriptive text sample to obtain image features and text features.

[0080] Please refer to Figure 5 The object detection model may include a feature extraction network (such as a text encoder Text Encoder and an image encoder Visual Encoder), a matching branch, and a localization branch.

[0081] The text encoder Text Encoder can be used to perform feature extraction on the descriptive text sample to obtain text features. And the image encoder Visual Encoder can be used to perform feature extraction on the image sample to be detected to obtain image features.

[0082] S23: Use the matching branch to match the image features and text features to obtain the category of the target, and based on the category of the target, obtain the alignment loss value.

[0083] Input the image features and text features into the matching branch to use the matching branch to match the image features and text features and obtain the category of the target. Then, calculate the loss based on the category of the target, and the alignment loss value of the matching branch can be obtained.

[0084] In some embodiments, use the matching branch to deeply fuse the image features and text features to obtain cross-modal fusion text fusion features and image fusion features. Then, perform feature matching on the text fusion features and image fusion features to obtain the alignment degree of each category of the target. Then, the alignment loss value can be obtained according to the alignment degree.

[0085] S24: Use the localization branch to perform localization processing on the image features to obtain the localization map of the target, and based on the localization map of the target, obtain the localization loss value.

[0086] Only input the image features into the localization branch, use the localization branch to perform localization processing on the image features to obtain the localization map of the target, and then calculate the loss based on the localization map of the target, and the localization loss value of the localization branch can be obtained.

[0087] In some embodiments, please refer to Figure 6 , the steps of S24 in the above embodiments can be further expanded. Use the localization branch to perform localization processing on the image features to obtain the localization map of the target, and based on the localization map of the target, obtain the localization loss value. This embodiment may include the following steps:

[0088] S241: Perform transition processing on the image features to obtain transition features.

[0089] Continue to refer to Figure 5 , the localization branch may include a localization transition layer, a first localization network, a dimension transformation network, a feature separation network, a second localization network, etc.

[0090] In this step, use the localization transition layer to perform transition processing on the image features to obtain transition features .

[0091] S242: Perform first localization on the transition features to obtain a first localization map; use the first localization map to obtain a first localization loss value.

[0092] Use the first localization network to perform first localization on the transition features to obtain a first localization map . Use the first localization map The loss calculation is performed to obtain the first localization loss value. The first localization loss value can be used to supervise the learning process. Since the first localization map is supervised and learned by the labels corresponding to the localization bounding boxes, in principle, only the information related to localization is retained in the first localization map. For the first localization map, the features unrelated to localization will be gradually filtered out during the training process.

[0093] S243: Feature separation of the transition features is performed using the first localization map to obtain separated features.

[0094] Optionally, the feature separation network performs feature separation on the transition features using the first localization map to obtain separated features.

[0095] Optionally, a dimension transformation network is used to perform dimension transformation on the first localization map to obtain a feature separation factor . Among them, the dimension of the feature separation factor is the same as the dimension of the transition features . Then, the feature separation network uses the feature separation factor to perform feature separation on the transition features to obtain separated features .

[0096] S244: Second localization is performed on the separated features to obtain a second localization map; the second localization loss value is obtained using the second localization map.

[0097] The second localization network is used to perform second localization on the separated features to obtain a second localization map .

[0098] The loss is calculated using the second localization map to obtain the second localization loss value. This enables the use of the second localization loss value to supervise the learning process.

[0099] S245: The self-distillation loss value is obtained using the transition features and the separated features.

[0100] In some embodiments, the self-distillation loss value can also be obtained. Specifically, the self-distillation loss value can be obtained using the transition features and the separated features . In this way, while the separation of the separated features is achieved, it also promotes the separation of the transition features of the localization transition layer as much as possible. Thus, for the features fed into the second localization network, the features unrelated to localization are better separated and filtered.

[0101] In some embodiments, the transition features and the separated features The pixel similarity between each pixel point. In response to the pixel similarity being less than the preset similarity, a pixel loss value is determined based on the pixel similarity of the pixel point; otherwise, the pixel loss value of the pixel point is determined to be zero. Then, the pixel loss values ​​of each pixel point can be combined to obtain the transition feature and separation features The corresponding self-distillation loss value.

[0102] Optionally, the cosine similarity or cosine distance between the pixels may be used to represent the pixel similarity, wherein the smaller the cosine distance, the more similar the two pixels are in direction; the larger the cosine distance, the less similar the two pixels are in direction.

[0103] Exemplary, transitional features and separation features Pixel loss value between each pixel (i, j) It can be expressed as:

[0104] .

[0105] Among them, (i, j) represents a pixel point at the coordinate position (i, j) on the feature map. Represents separation features A pixel on Represents transition characteristics A pixel on top. Represents a preset parameter, which is also a control parameter. When the cosine distance between two features is less than or equal to the preset parameter When (that is, the pixel similarity is not less than the preset similarity), it is considered that the two features are very close at this pixel point and there is no need to shorten the distance between them.

[0106] For example, the pixel loss value of each pixel point (i, j) is integrated , the final self-distillation loss value can be obtained ,as follows:

[0107] .

[0108] Among them, W and H represent the width and height of the feature respectively.

[0109] S246: Obtain a positioning loss value by integrating the first positioning loss value, the second positioning loss value and / or the self-distillation loss value.

[0110] Optionally, the first positioning loss value and the second positioning loss value may be combined to obtain the positioning loss value of the positioning branch.

[0111] Optionally, the localization loss value of the localization branch can be obtained by integrating the first localization loss value, the second localization loss value, and the self-distillation loss value.

[0112] In some embodiments, during the training process, to avoid the separated features from being adversely affected by the transitional features and resulting in the separated features not being significantly separated. That is, the generated filtering factor may be too trivial (the values at each pixel point and each channel are very close to 1), such that the separated features degenerate into being very similar to the transitional features before the location-irrelevant features are fully filtered, thus losing the function of feature separation. To address this, the present application also employs the following two training strategies.

[0113] In some embodiments, for the self-distillation loss value , during the training process, gradient isolation processing is performed on the separated features . That is, when the self-distillation loss value is backpropagated, it is only backpropagated to the transitional features . Therefore, the self-distillation loss value will not affect the transitional features , thereby avoiding the separated features from being adversely affected by the transitional features .

[0114] In some embodiments, two localization loss values can be used for the matching branch and the localization branch, or respectively for the first localization map and the second localization map in the localization branch, and different positive samples can be used to obtain the localization loss values to supervise the training process of the two regressions. In this way, it is avoided that the first localization map becomes a trivial (invalid) predicted localization map, thereby avoiding the generated filtering factor from being too trivial.

[0115] Continuing to refer to Figure 4 , after the above step S24, the following steps may further be included:

[0116] S25: Integrate the alignment loss value and the localization loss value to train the object detection model, and obtain the trained object detection model.

[0117] Integrate the alignment loss value of the matching branch and the localization loss value of the localization branch to train the object detection model, and adjust the network parameters of each branch of the object detection model to obtain the trained object detection model.

[0118] In the above solution, the present application introduces the idea of feature separation into cross-modal object detection, explicitly separating the matching branch and the localization branch, and generating different features for the two branches respectively. For the matching branch that requires strong semantic features, through a deep fusion module, the connection between the two modalities is increased, generating text-image fusion features rich in semantic information and inputting them into the matching module for matching. For the localization task, which does not require semantic information but only edge texture information, there is also a localization branch with feature self-distillation. By using the method of feature self-distillation, features irrelevant to localization are gradually separated and filtered, generating localization features rich in detailed edge information and enhancing the accuracy of localization. Through the idea of feature separation, different branches can perform their tasks respectively, generating task-aware features respectively, realizing the separation of tasks, reducing the mutual interference between irrelevant features, and improving the network effect.

[0119] For step S24 of the above embodiment, for different branches, or for the first localization map and the second localization map in the localization branch respectively, the corresponding training positive samples can be used to obtain the loss value.

[0120] Please refer to Figure 7 , Figure 7 which is a schematic flowchart of the second embodiment of the training method of the object detection model of the present application. The method may include the following steps:

[0121] S31: Obtain candidate positive samples.

[0122] Based on the image sample to be detected, candidate positive samples can be obtained.

[0123] Optionally, anchor boxes can be set on each pixel of the image sample to be detected, candidate boxes can be obtained, and all candidate boxes can be used as candidate positive samples.

[0124] Optionally, partitions can be randomly selected from the image sample to be detected, each partition can be used as a candidate box, and all candidate boxes can be used as candidate positive samples. The present application does not limit the method for obtaining candidate positive samples.

[0125] S32: Perform task-separated sampling on the candidate positive samples to determine the aligned positive samples selected by the matching branch and the localization positive samples selected by the localization branch; among them, the aligned positive samples and the category of the target are used to obtain the alignment loss value, and the localization positive samples and the localization map of the target are used to obtain the localization loss value.

[0126] Continue to refer to Figure 5, during the training process of the target detection model, a task separation sampling module is also included. The task separation sampling module is used to perform task separation sampling on candidate positive samples to respectively determine the aligned positive samples selected by the matching branch and the localization positive samples selected by the localization branch. Among them, the category of the aligned positive sample and the target is used to obtain the alignment loss value, and the localization positive sample and the localization map of the target are used to obtain the localization loss value.

[0127] In some embodiments, referring to Figure 8 , step S32 may include the following steps:

[0128] S321: Obtain the matching degree between the candidate positive sample and the ground truth bounding box to obtain a reference matching value.

[0129] The matching degree between each candidate positive sample and the ground truth bounding box can be obtained respectively to obtain the reference matching degree of each candidate positive sample. Among them, the matching degree can be represented by IOU (Intersection over Union), and the selection strategy of IOU is used to select the training positive samples of each branch.

[0130] Exemplarily, obtain the matching degree (IOU) between each candidate positive sample and the ground truth bounding box to obtain the reference matching value IoU corresponding to each candidate positive sample d .

[0131] S322: Determine the aligned positive samples selected by the matching branch from the candidate positive samples by using the reference matching value and the category similarity of the category.

[0132] For the matching branch, the aligned positive samples selected by the matching branch can be determined from the candidate positive samples by using the reference matching value and the category similarity of the category. Among them, the category similarity of the category is obtained based on the matching similarity of the category corresponding to the ground truth bounding box.

[0133] In some embodiments, the matching similarity can be obtained for the image fusion feature and the text fusion feature at all pixel points to obtain a corresponding similarity matrix. Among them, each matching similarity can correspond to a category. Then, according to the ground truth bounding box corresponding to each matching similarity, the matching similarity of the category corresponding to the ground truth bounding box is selected to obtain the category similarity of each category. The category similarity S is represented by the following formula:

[0134] .

[0135] Among them, (i is less than or equal to N) represents the text fusion feature, (j is less than or equal to M) represents the image fusion feature.

[0136] In some embodiments, the reference matching value IoU can be utilized d and the category similarity S of the category to obtain the matching sample evaluation value of the matching branch. Specifically, the first trade-off parameter can be used to control the contribution degree between the reference matching value IoU d and the category similarity S. Therefore, the matching sample evaluation value of the matching branch can be expressed as:

[0137] .

[0138] Then, candidate positive samples whose matching sample evaluation value meets the matching selection condition are selected to obtain aligned positive samples. Among them, the matching selection condition includes: the matching sample evaluation value is greater than the first threshold and the first quantity K (K is an integer) before sorting. That is, among all candidate samples whose matching sample evaluation value is greater than the first threshold , the first K candidate positive samples are selected to obtain the aligned positive samples of the matching branch, which can be expressed as follows:

[0139] .

[0140] Among them, represents the matching selection condition, that is, among all candidate samples whose matching sample evaluation value is greater than the first threshold , the first K candidate positive samples are selected as the aligned positive samples.

[0141] Optionally, the matching branch can use the aligned positive samples and the corresponding categories to obtain the alignment loss value.

[0142] S323: Determine the localization positive samples selected by the localization branch from the candidate positive samples by using the reference matching value and the localization similarity of the localization map.

[0143] For the localization branch, the reference matching value IoU d and the localization similarity of the localization map can be used to determine the localization positive samples selected by the localization branch from the candidate positive samples. Among them, the localization similarity of the localization map is obtained based on the localization map and the ground truth bounding box.

[0144] Optionally, the positioning map includes a first positioning map and / or a second positioning map. Correspondingly, based on the similarity between the first positioning map and the corresponding true label box, a first positioning similarity can be obtained. Based on the similarity between the second positioning map and the corresponding true label box, a second positioning similarity can be obtained. In this way, positioning positive samples can be selected according to different positioning maps or any one of them. As the positioning features are gradually separated, the positive samples for the first positioning and the second positioning are also different.

[0145] In some embodiments, taking the selection of positive samples for the first positioning map and / or the second positioning map as an example, the positioning positive samples include first positioning positive samples and / or second positioning positive samples, where the first positioning positive samples and the first positioning map are used to obtain a first positioning loss value, and the second positioning positive samples and the second positioning map are used to obtain a second positioning loss value.

[0146] In some embodiments, for the first positioning map, the reference matching value IoU d and the first positioning similarity of the first positioning map are used to obtain a first positioning evaluation value for the positioning branch. Specifically, the second weighing parameter is used to control the contribution degree between the reference matching value IoU d and the first positioning similarity. Therefore, the first positioning evaluation value of the positioning branch can be expressed as:

[0147] .

[0148] Among them, represents the second weighing parameter, which controls the contribution degree between the reference matching value IoU d and the first positioning similarity , and the first positioning similarity represents the similarity (such as IOU) between the first positioning map and the true label box.

[0149] Then, candidate positive samples whose first positioning evaluation value meets the first positioning selection condition are selected to obtain the first positioning positive samples. Among them, the first positioning selection condition includes: the first positioning evaluation value is greater than the second threshold and the second quantity before sorting.

[0150] Exemplarily, the first positioning positive samples can be expressed as:

[0151] .

[0152] Among them, represents the first positioning selection condition, that is, it means that the first positioning evaluation value is greater than the second threshold (such as Among all the candidate samples of (), select the first second quantity (such as K) of candidate positive samples as the first localization positive samples.

[0153] In some embodiments, for the second localization map, the reference matching value IoU d , the first localization similarity of the first localization map, and the second localization similarity of the second localization map can be used to obtain the second localization evaluation value of the localization branch. Specifically, the contribution of the reference matching value IoU d can be controlled by the first weight parameter, the contribution of the first localization similarity can be controlled by the second weight parameter, and the contribution of the second localization similarity and the reference matching value IoU d and the first localization similarity can be controlled by the first weight parameter and the second weight parameter to obtain the second localization evaluation value of the localization branch.

[0154] Exemplarily, the second localization evaluation value of the localization branch can be expressed as follows:

[0155] .

[0156] Wherein, and respectively represent the first weight parameter and the second weight parameter, both of which are weight parameters less than 1. The first localization similarity represents the similarity (such as IOU) between the first localization map and the true label box. The second localization similarity represents the similarity (such as IOU) between the second localization map and the true label box.

[0157] Then, select the candidate positive samples whose second localization evaluation value meets the second localization selection condition to obtain the second localization positive samples. Among them, the second localization selection condition includes: the second localization evaluation value is greater than the third threshold and the third quantity before sorting.

[0158] Exemplarily, the second localization positive samples can be expressed as:

[0159] .

[0160] Wherein, represents the second localization selection condition, that is, it means that among all the candidate samples where the second localization evaluation value is greater than the third threshold (such as ), select the first third quantity (such as K) of candidate positive samples as the second localization positive samples.

[0161] Optionally, the first threshold, the second threshold, and the third threshold referred to above may be different thresholds or the same threshold, and the first quantity, the second quantity, and the third quantity may be different quantities or the same quantity. This application takes the same threshold and the same quantity as an example for illustration, and this application does not limit this.

[0162] In the above solution, in order to separately select more suitable training positive samples for the matching task and the localization task during the training optimization stage, through the above task separation sampling strategy, different training positive samples can be separately selected for the matching branch, as well as the first localization and the second localization of the localization branch to obtain corresponding loss values. Thus, during the parameter learning process, the two branches are differentiated to achieve the separation of the matching features and the localization features. As a result, during the inference process of the network, the mutual influence between the two types of features in the cross-modal detection task is reduced, and it can accurately respond at the corresponding pixel points, improving the performance of the network.

[0163] In addition, the task separation sampling module can separately select more suitable training positive samples for the matching task and the localization task during the training optimization stage. In the matching branch, matching-aware positive samples are selected according to the characteristics of the matching task to enhance the learning of the semantic features of the matching branch. In the localization branch, according to the progress of the localization task, suitable positive samples are allocated to the localization task in two steps to supervise the learning of the detailed and edge features. Through the task separation sampling module, the separation of the features of the two branches is supervised from the perspective of network parameter learning, so that task-aware features can be learned in the two branches respectively.

[0164] It can be understood that in the above method of the specific implementation manner, the writing order of each step does not mean a strict execution order and does not impose any limitation on the implementation process. The specific execution order of each step should be determined according to its function and possible internal logic.

[0165] In some embodiments, this application further provides a cross-modal object detection device for implementing the cross-modal object detection method of any of the above embodiments.

[0166] Please refer to Figure 9 , Figure 9 FIG. is a schematic structural diagram of an embodiment of the cross-modal object detection device of this application. The cross-modal object detection device 40 includes an acquisition module 41, a feature extraction module 42, a matching branch module 43, and a localization branch module 44. Among them, each module is interconnected.

[0167] The acquisition module 41 is used to acquire the image to be detected and the description text.

[0168] The feature extraction module 42 is used to respectively extract features from the image to be detected and the description text to obtain image features and text features.

[0169] The matching branch module 43 is used to perform matching processing on the image features and text features using the matching branch to obtain the category of the target.

[0170] The localization branch module 44 is used to perform localization processing on the image features using the localization branch to obtain the localization map of the target.

[0171] In some embodiments, the present application further provides a training device for a target detection model, which is used to implement the training method of the target detection model in any of the above embodiments.

[0172] Please refer to Figure 10 , Figure 10 FIG. is a schematic structural diagram of an embodiment of the training device for the target detection model of the present application. The training device 50 for the target detection model includes a sample acquisition module 51, a feature extraction module 52, an alignment loss module 53, a localization loss module 54, and a training module 55. Among them, each module is interconnected. The target detection model includes at least a matching branch and a localization branch.

[0173] The sample acquisition module 51 is used to acquire the image sample to be detected and the description text sample;

[0174] The feature extraction module 52 is used to perform feature extraction on the image sample to be detected and the description text sample respectively to obtain image features and text features.

[0175] The alignment loss module 53 is used to perform matching processing on the image features and text features using the matching branch to obtain the category of the target, and based on the category of the target, obtain the alignment loss value.

[0176] The localization loss module 54 is used to perform localization processing on the image features using the localization branch to obtain the localization map of the target, and based on the localization map of the target, obtain the localization loss value.

[0177] The training module 55 is used to comprehensively combine the alignment loss value and the localization loss value to train the target detection model to obtain the trained target detection model.

[0178] It should be noted that the device provided in the above embodiment and the method provided in the above embodiment belong to the same concept. The specific manners in which each module and unit perform operations have been described in detail in the method embodiment, and will not be repeated here. In practical applications, the device provided in the above embodiment can, according to needs, allocate the above functions to different functional modules, that is, divide the internal structure of the device into different functional modules to complete all or part of the functions described above. The present application places no restrictions on this.

[0179] For the above embodiment, the present application provides a computer device. Please refer to Figure 11 , Figure 11It is a schematic structural diagram of an embodiment of a computer device of the present application. The computer device 60 includes a memory 61 and a processor 62. Among them, the memory 61 and the processor 62 are coupled to each other. Program data is stored in the memory 61, and the processor 62 is configured to execute the program data to implement the steps of any one of the above-mentioned cross-modal object detection methods and object detection model training methods.

[0180] In this embodiment, the processor 62 can also be referred to as a CPU (Central Processing Unit). The processor 62 may be an integrated circuit chip with signal processing capabilities. The processor 62 may also be a general-purpose processor, a digital signal processor (DSP, Digital Signal Processing), an application-specific integrated circuit (ASIC, Application Specific Integrated Circuit), a field-programmable gate array (FPGA, Field Programmable Gate Array), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor 62 may also be any conventional processor, etc.

[0181] For the method of the above embodiment, it can be implemented in the form of a computer program. Therefore, the present application proposes a computer-readable storage medium. Please refer to Figure 12 , Figure 12 It is a schematic structural diagram of an embodiment of a computer-readable storage medium of the present application. Program data 71 that can be run by a processor is stored in the computer-readable storage medium 70. The program data 71 can be executed by the processor to implement the steps of any one of the above-mentioned cross-modal object detection methods and object detection model training methods.

[0182] The computer-readable storage medium 70 of this embodiment may be a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc, etc., which can store the program data 71. Or it may also be a server storing the program data 71. The server can send the stored program data 71 to other devices for running, or it can also run the stored program data 71 by itself.

[0183] In some embodiments, the functions or modules included in the device provided in the above embodiments of the present application can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, the present application will not repeat it here.

[0184] The descriptions of the various embodiments above tend to emphasize the differences between the various embodiments. For their similarities, reference can be made to each other. For the sake of brevity, they will not be elaborated herein in this application.

[0185] In several embodiments provided in this application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.

[0186] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0187] In addition, in each embodiment of this application, the functional units can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0188] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium, which is a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in each embodiment of this application.

[0189] Obviously, those skilled in the art should understand that the various modules or steps of the present application described above can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device. Thus, they can be stored in a computer-readable storage medium and executed by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. In this way, the present application is not limited to any specific combination of hardware and software.

[0190] The above are only embodiments of the present application, and thus do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall equally be included in the patent protection scope of the present application.

Claims

1. A cross-modal target detection method, characterized in that: include: Get the image to be detected and the description text; Extracting features of the image to be detected and the description text respectively to obtain image features and text features; Using a matching branch to match the image features with the text features to obtain a target category; as well as The positioning branch is used to perform positioning processing on the image features to obtain a positioning map of the target, including: Performing transition processing on the image features to obtain transition features, including: performing transition processing on the image features using a positioning transition layer to obtain the transition features, wherein the positioning transition layer is a neural network or a convolutional network; Performing a first positioning on the transition feature to obtain a first positioning map; Performing a dimension transformation on the first positioning map to obtain a feature separation factor, wherein the dimension of the feature separation factor is the same as the dimension of the transition feature; performing feature separation on the transition feature using the feature separation factor to obtain a separation feature; Perform a second positioning on the separation feature to obtain a second positioning map; wherein the second positioning map is used as a positioning map of the target in the image to be detected.

2. The method according to claim 1, characterized in that The using a matching branch to match the image feature and the text feature to obtain a target category includes: Deeply fusing the image features and the text features to obtain cross-modal fused text fusion features and image fusion features; Feature matching is performed on the text fusion feature and the image fusion feature to obtain the category of the target.

3. A method for training a target detection model, characterized in that: The target detection model includes at least a matching branch and a positioning branch, and the training method includes: Obtain image samples to be detected and description text samples; Extracting features of the image sample to be detected and the description text sample respectively to obtain image features and text features; Using a matching branch to match the image feature and the text feature to obtain a category of the target, and based on the category of the target, obtaining an alignment loss value; and Using the positioning branch to perform positioning processing on the image features to obtain a positioning map of the target, and based on the positioning map of the target, obtaining a positioning loss value; The target detection model is trained by combining the alignment loss value and the positioning loss value to obtain a trained target detection model; The method of using the positioning branch to perform positioning processing on the image features to obtain a positioning map of the target includes: Performing transition processing on the image features to obtain transition features, including: performing transition processing on the image features using a positioning transition layer to obtain the transition features, wherein the positioning transition layer is a neural network or a convolutional network; Performing a first positioning on the transition feature to obtain a first positioning map; performing a dimension transformation on the first positioning map to obtain a feature separation factor; wherein the dimension of the feature separation factor is the same as the dimension of the transition feature; Using the feature separation factor, the transition feature is subjected to feature separation to obtain a separation feature; The separation feature is secondly positioned to obtain a second positioning map.

4. The method according to claim 3, characterized in that The obtaining of the positioning loss value based on the positioning map of the target includes: Using the first positioning map, obtaining a first positioning loss value; Using the second positioning map, obtaining a second positioning loss value; and / or Using the transition characteristic and the separation characteristic, a self-distillation loss value is obtained; The positioning loss value is obtained by integrating the first positioning loss value, the second positioning loss value and / or the self-distillation loss value.

5. The method according to claim 3, characterized in that: Also includes: Obtain candidate positive samples; Performing task separation sampling on the candidate positive samples to determine the alignment positive samples selected by the matching branch and the positioning positive samples selected by the positioning branch; The alignment positive sample and the category of the target are used to obtain an alignment loss value, and the positioning positive sample and the positioning map of the target are used to obtain a positioning loss value.

6. The method according to claim 5, characterized in that The performing task separation sampling on the candidate positive samples to determine the alignment positive samples selected by the matching branch and the positioning positive samples selected by the positioning branch includes: Obtaining the matching degree between the candidate positive sample and the real label frame to obtain a reference matching value; Determining the aligned positive sample selected by the matching branch from the candidate positive samples by using the reference matching value and the category similarity of the category; Determine the positioning positive sample selected by the positioning branch from the candidate positive samples by using the positioning similarity of the reference matching value and the positioning map; The category similarity of the category is obtained based on the matching similarity of the category corresponding to the real label frame; and the positioning similarity of the positioning map is obtained based on the positioning map and the real label frame.

7. A computer device, characterized in that: The invention comprises a memory and a processor coupled to each other, wherein the memory stores program data, and the processor is used to execute the program data to implement the steps of the method according to any one of claims 1 to 6.

8. A computer-readable storage medium, characterized in that: Program data that can be executed by a processor are stored, and the program data is used to implement the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Target detection method and device

    CN117392379A