Object detection method, related device, equipment and storage medium

Through the hierarchical fusion method of visual features and text features, the problem of insufficient fusion of visual features and text features in existing object detection is solved, and the accuracy of object detection and the operation efficiency of the model are improved.

CN120298840APending Publication Date: 2025-07-11TENCENT TECH (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510377467.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The integration of visual features and text features in the existing object detection methods is insufficient, resulting in low accuracy in object detection.

Method used

The hierarchical visual features and text features fusion method is adopted, and through wavelet transformation and multimodal feature fusion modules, the visual features and text features are fully integrated locally and globally, and a graphic and text fusion feature map is generated to improve the accuracy of object detection.

Benefits of technology

By fusion of visual and text features in fine-grained size, the accuracy of object detection is improved, the calculation amount and training cost of the model are reduced, and the multimodal alignment and generalization capabilities of the model are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298840A_ABST
    Figure CN120298840A_ABST
Patent Text Reader

Abstract

The invention discloses an object detection method, a related device, equipment and a storage medium. The method comprises the steps of obtaining to-be-detected visual data and a corresponding prompt text; performing feature extraction on a to-be-detected image frame in the to-be-detected visual data to obtain corresponding K feature maps; performing wavelet transform on the K feature maps to obtain at least one group of visual feature map set corresponding to each feature map; performing feature extraction on the prompt text to obtain original text features; fusing the original text features, the K feature maps and the at least one group of visual feature map set corresponding to each feature map to obtain an image-text fusion feature map corresponding to each feature map; and according to the K image-text fusion feature maps, generating an object detection result of the to-be-detected image frame. In the feature fusion process, hierarchical visual features and text features are fused, fine-grained fusion between the visual features and the text features is achieved, and therefore the accuracy of object detection is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer vision technology, and in particular, to a method for object detection, related devices, equipment, and storage media. Background Art

[0002] In recent years, with the rapid development of computer vision technology, its applications in various fields have become increasingly extensive and in-depth. As an interdisciplinary subject, computer vision integrates knowledge in multiple fields such as mathematics, computer science, and image processing, aiming to enable computers to understand and interpret the content in images and videos.

[0003] In order to improve the accuracy of multi-object continuous positioning, text instructions are often introduced on the basis of multi-object continuous positioning to guide object detection in video frames. In related technologies, object detection methods usually use a cross-attention mechanism to fuse visual features and text features, and perform object detection through the fused visual-text features.

[0004] In the process of researching and practicing the current technology, the inventors found that there are at least the following problems in the current solution. In the process of fusing visual features and text features, only simple cross-attention operations are performed, and the fusion of visual features and text features is insufficient, resulting in low accuracy of object detection. Summary of the Invention

[0005] Embodiments of the present application provide a method for object detection, related devices, equipment, and storage media. In the feature fusion process of the present application, hierarchical visual features and text features are fused to achieve fine-grained fusion between visual features and text features, thereby improving the accuracy of object detection.

[0006] In view of this, on the one hand, the present application provides a method for object detection, including:

[0007] Obtain visual data to be detected and corresponding prompt text, where the visual data to be detected includes an image frame sequence composed of at least one image frame to be detected;

[0008] Extract features from the image frames to be detected in the visual data to be detected to obtain K feature maps corresponding to the image frames to be detected, where K is an integer greater than or equal to 1;

[0009] Perform wavelet transform on the K feature maps to obtain at least one set of visual feature map sets corresponding to each feature map, where each set of visual feature map sets includes at least one of a global visual feature map, a vertical visual feature map, a horizontal visual feature map, and a diagonal visual feature map;

[0010] Extract features from the prompt text to obtain the original text features;

[0011] Fuse the original text features, K feature maps, and at least one set of visual feature map sets corresponding to each feature map to obtain a text-image fusion feature map corresponding to each feature map;

[0012] Generate an object detection result for the image frame to be detected based on the K text-image fusion feature maps corresponding to the K feature maps.

[0013] On the other hand, the present application provides an object detection device, including:

[0014] An acquisition module for acquiring visual data to be detected and corresponding prompt text, where the visual data to be detected includes an image frame sequence composed of at least one image frame to be detected;

[0015] An extraction module for extracting features from the image frames to be detected in the visual data to be detected to obtain K feature maps corresponding to the image frames to be detected, where K is an integer greater than or equal to 1;

[0016] A transformation module for performing wavelet transformation on the K feature maps to obtain at least one set of visual feature map sets corresponding to each feature map, where each set of visual feature map sets includes at least one of a global visual feature map, a vertical visual feature map, a horizontal visual feature map, and a diagonal visual feature map;

[0017] The extraction module is further configured to extract features from the prompt text to obtain original text features;

[0018] A fusion module for fusing the original text features, the K feature maps, and at least one set of visual feature map sets corresponding to each feature map to obtain a text-image fusion feature map corresponding to each feature map;

[0019] A detection module for generating an object detection result for the image frame to be detected based on the K text-image fusion feature maps corresponding to the K feature maps.

[0020] In a possible design, in another implementation manner of the other aspect of the embodiments of the present application,

[0021] The extraction module is specifically configured to perform convolution processing on the image frame to be detected through an initial convolution layer included in the image feature extraction network to obtain an initial feature map corresponding to the image frame to be detected;

[0022] Perform pooling processing on the initial convolution features through a feature processing layer included in the image feature extraction network to obtain a pooled feature map, where the feature processing layer includes a batch normalization layer, an activation function layer, and a max pooling layer;

[0023] Through the R residual block layers included in the image feature extraction network, feature extraction is performed on the pooled feature map to obtain R feature maps, where each residual block layer is stacked by multiple bottleneck layers, and each bottleneck layer includes a convolutional layer connected in sequence, and R is an integer greater than or equal to K;

[0024] Obtain K feature maps with different scales from the R feature maps.

[0025] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,

[0026] The extraction module is specifically configured to perform word segmentation on the prompt text to obtain at least one token;

[0027] Encode each token in the at least one token to obtain a word vector corresponding to each token;

[0028] Generate a position vector for each token according to the position of each token in the prompt text;

[0029] Generate a combined vector for each token according to the word vector and position vector corresponding to each token;

[0030] Based on the combined vector of each token, obtain the original text features through the text feature extraction network.

[0031] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,

[0032] The transformation module is specifically configured to perform convolution processing on the K feature maps respectively using a global convolution kernel, a vertical convolution kernel, a horizontal convolution kernel, and a diagonal convolution kernel, and sequentially obtain a global visual feature map, a vertical visual feature map, a horizontal visual feature map, and a diagonal visual feature map corresponding to each feature map;

[0033] Among them, the global visual feature map, the vertical visual feature map, the horizontal visual feature map, and the diagonal visual feature map constitute a set of visual feature map collections.

[0034] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,

[0035] The fusion module is specifically configured to fuse the original text features with the global visual feature map, the vertical visual feature map, the horizontal visual feature map, and the diagonal visual feature map corresponding to the feature map respectively, and sequentially obtain a global fusion feature map, a vertical fusion feature map, a horizontal fusion feature map, and a diagonal fusion feature map corresponding to the feature map;

[0036] Perform inverse wavelet transform on the merged global fusion feature map, vertical fusion feature map, horizontal fusion feature map, and diagonal fusion feature map corresponding to the feature map to obtain the target feature map corresponding to the feature map;

[0037] Fuse the original text features and the feature maps to obtain the fused feature maps corresponding to the feature maps;

[0038] Merge the fused feature maps corresponding to the feature maps and the target feature maps corresponding to the feature maps to obtain the text-image fused feature maps corresponding to the feature maps.

[0039] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,

[0040] The transformation module is specifically configured to perform convolution processing on K feature maps respectively by using a global convolution kernel, a vertical convolution kernel, a horizontal convolution kernel, and a diagonal convolution kernel, and sequentially obtain a first global visual feature map, a first vertical visual feature map, a first horizontal visual feature map, and a first diagonal visual feature map corresponding to each feature map;

[0041] Perform convolution processing on the first global visual feature map corresponding to each feature map respectively by using a global convolution kernel, a vertical convolution kernel, a horizontal convolution kernel, and a diagonal convolution kernel, and sequentially obtain a second global visual feature map, a second vertical visual feature map, a second horizontal visual feature map, and a second diagonal visual feature map corresponding to each feature map;

[0042] Wherein, the first global visual feature map, the first vertical visual feature map, the first horizontal visual feature map, and the first diagonal visual feature map constitute a set of visual feature map sets;

[0043] And, the second global visual feature map, the second vertical visual feature map, the second horizontal visual feature map, and the second diagonal visual feature map constitute another set of visual feature map sets.

[0044] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,

[0045] The fusion module is specifically configured to fuse the original text features with the first global visual feature map, the first vertical visual feature map, the first horizontal visual feature map, and the first diagonal visual feature map corresponding to the feature maps respectively, and sequentially obtain a first global fused feature map, a first vertical fused feature map, a first horizontal fused feature map, and a first diagonal fused feature map corresponding to the feature maps;

[0046] Fuse the original text features with the second global visual feature map, the second vertical visual feature map, the second horizontal visual feature map, and the second diagonal visual feature map corresponding to the feature maps respectively, and sequentially obtain a second global fused feature map, a second vertical fused feature map, a second horizontal fused feature map, and a second diagonal fused feature map corresponding to the feature maps;

[0047] Inverse wavelet transform is performed on the merged second global fusion feature map, second vertical fusion feature map, second horizontal fusion feature map, and second diagonal fusion feature map corresponding to the feature map to obtain the first target feature map corresponding to the feature map;

[0048] The first target feature map corresponding to the feature map is merged with the first global fusion feature map, first vertical fusion feature map, first horizontal fusion feature map, and first diagonal fusion feature map corresponding to the feature map respectively to obtain the first fusion feature map, second fusion feature map, third fusion feature map, and fourth fusion feature map corresponding to the feature map in sequence;

[0049] Inverse wavelet transform is performed on the merged first fusion feature map, second fusion feature map, third fusion feature map, and fourth fusion feature map corresponding to the feature map to obtain the second target feature map corresponding to the feature map;

[0050] The original text feature and the feature map are fused to obtain the fusion feature map corresponding to the feature map;

[0051] The fusion feature map corresponding to the feature map and the second target feature map corresponding to the feature map are merged to obtain the text-image fusion feature map corresponding to the feature map.

[0052] In a possible design, in another implementation manner of another aspect of the embodiment of the present application,

[0053] The transformation module is specifically configured to perform convolution processing on K feature maps by using a global convolution kernel to obtain the first global visual feature map corresponding to each feature map;

[0054] The first global visual feature map corresponding to each feature map is respectively subjected to convolution processing by using a global convolution kernel, a vertical convolution kernel, a horizontal convolution kernel, and a diagonal convolution kernel to obtain the second global visual feature map, second vertical visual feature map, second horizontal visual feature map, and second diagonal visual feature map corresponding to each feature map in sequence;

[0055] Among them, the first global visual feature maps form a set of visual feature map collections;

[0056] Moreover, the second global visual feature map, second vertical visual feature map, second horizontal visual feature map, and second diagonal visual feature map form another set of visual feature map collections.

[0057] In a possible design, in another implementation manner of another aspect of the embodiment of the present application,

[0058] The fusion module is specifically configured to fuse the original text feature and the first global visual feature map corresponding to the feature map to obtain the first global fusion feature map corresponding to the feature map;

[0059] The original text features are respectively fused with the second global visual feature map, the second vertical visual feature map, the second horizontal visual feature map, and the second diagonal visual feature map corresponding to the feature map, and the second global fusion feature map, the second vertical fusion feature map, the second horizontal fusion feature map, and the second diagonal fusion feature map corresponding to the feature map are obtained in sequence;

[0060] After merging the second global fusion feature map, the second vertical fusion feature map, the second horizontal fusion feature map, and the second diagonal fusion feature map corresponding to the feature map, inverse wavelet transform is performed to obtain the first target feature map corresponding to the feature map;

[0061] After merging the first target feature map corresponding to the feature map and the first global fusion feature map corresponding to the feature map, inverse wavelet transform is performed to obtain the second target feature map corresponding to the feature map;

[0062] The original text features and the feature map are fused to obtain the fusion feature map corresponding to the feature map;

[0063] The fusion feature map corresponding to the feature map and the second target feature map corresponding to the feature map are merged to obtain the text-image fusion feature map corresponding to the feature map.

[0064] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,

[0065] The fusion module is specifically configured to perform feature mapping on the original text features and the feature map respectively to obtain the first text feature corresponding to the original text features and the first feature map corresponding to the feature map;

[0066] Self-attention calculation is respectively performed on the first text feature and the first feature map to obtain the second text feature corresponding to the first text feature and the second feature map corresponding to the first feature map;

[0067] After cross-attention calculation is performed on the second text feature and the second feature map, skip connection is performed with the first feature map to obtain the fusion feature map corresponding to the feature map.

[0068] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,

[0069] The detection module is specifically configured to obtain N initial query vectors, where N is an integer greater than or equal to 1;

[0070] According to the text-image fusion feature map corresponding to each feature map in K feature maps, a visual text feature vector is generated;

[0071] According to the N initial query vectors and the visual text feature vector, the object detection result of the image frame to be detected is determined.

[0072] In a possible design, in another implementation of another aspect of the embodiments of the present application,

[0073] The detection module is specifically configured to process N initial query vectors according to the original text features to obtain N target query vectors, where N is an integer greater than or equal to 1;

[0074] Generate visual text feature vectors according to the text-image fusion feature maps corresponding to each of the K feature maps;

[0075] Determine the object detection result of the image frame to be detected according to the N target query vectors and the visual text feature vectors.

[0076] In a possible design, in another implementation of another aspect of the embodiments of the present application,

[0077] The detection module is specifically configured to perform self-attention calculation on the N initial query vectors to obtain N first query vectors;

[0078] Obtain M word-level features and sentence-level features according to the original text features, where M is an integer greater than or equal to 1;

[0079] Perform cross-attention calculation on the M word-level features and the N first query vectors to obtain N second query vectors;

[0080] Interact the sentence-level features with the N second query vectors to obtain N third query vectors;

[0081] Merge the N third query vectors and the N initial query vectors to obtain N target query vectors.

[0082] In a possible design, in another implementation of another aspect of the embodiments of the present application,

[0083] The acquisition module is specifically configured to acquire the video to be detected, where the video to be detected includes an image frame sequence composed of at least two image frames;

[0084] Perform equally-spaced frame extraction processing on the video to be detected to obtain the visual data to be detected;

[0085] The detection module is further configured to detect each image frame that is not frame-extracted in the video to be detected according to the object detection result of each image frame to be detected in the visual data to be detected, and obtain the object detection result of each image frame.

[0086] In a possible design, in another implementation of another aspect of the embodiments of the present application,

[0087] An acquisition module, specifically used to acquire a video to be detected, where the video to be detected includes an image frame sequence composed of at least two image frames;

[0088] Perform key frame recognition on the video to be detected to obtain visual data to be detected, where the image frame to be detected in the visual data to be detected is the key frame of the video to be detected;

[0089] The detection module is further configured to perform detection on each image frame without frame extraction in the video to be detected according to the object detection result of each image frame to be detected in the visual data to be detected, and obtain the object detection result of each image frame.

[0090] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,

[0091] The detection module is specifically configured to construct at least one pair of image frames to be matched, where each pair of image frames to be matched includes a template image frame and a search area image frame. The template image frame is an image frame to be detected for which an object detection result has been obtained, and the search area image frame is an image frame in the video to be detected associated with the template image frame;

[0092] For each pair of image frames to be matched, perform detection on the search area image frame according to the template image frame to obtain the object detection result of the search area image frame;

[0093] Use the object detection result of each search area image frame in each pair of image frames to be matched as the object detection result of each image frame without frame extraction in the video to be detected.

[0094] Another aspect of the present application provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the methods of the above aspects are implemented.

[0095] Another aspect of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the methods of the above aspects are implemented.

[0096] Another aspect of the present application provides a computer program product, including a computer program. When the computer program is executed by a processor, the methods of the above aspects are implemented.

[0097] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:

[0098] In an embodiment of the present application, a method for object detection is provided. First, visual data to be detected and a prompt text are respectively obtained. Then, feature extraction is performed on the image frames to be detected in the visual data to be detected, and a plurality of feature maps corresponding to the image frames to be detected are obtained. Based on this, wavelet transform is performed on the feature maps to obtain hierarchical visual features. At the same time, feature extraction also needs to be performed on the prompt text to obtain the original text features. Then, the original text features, the feature maps, and the hierarchical visual features are fused to obtain a text-image fusion feature map. Finally, the image frames to be detected are detected by combining the original text features and the text-image fusion feature map. In the above manner, wavelet transform is performed on the visual features of the image frames to be detected to obtain hierarchical visual features (i.e., at least one of a global visual feature map, a vertical visual feature map, a horizontal visual feature map, and a diagonal visual feature map). Thus, in the feature fusion process, the hierarchical visual features are fused with the text features to achieve fine-grained fusion between the visual features and the text features, thereby improving the accuracy of object detection. Description of the Drawings

[0099] Figure 1 It is a schematic diagram applied to a harmful item detection scenario in an embodiment of the present application;

[0100] Figure 2 It is a schematic diagram applied to a video editing scenario in an embodiment of the present application;

[0101] Figure 3 It is a schematic diagram applied to a content rewriting scenario in an embodiment of the present application;

[0102] Figure 4 It is a schematic diagram of an implementation environment for implementing an object detection method based on a network scenario in an embodiment of the present application;

[0103] Figure 5 It is a schematic diagram of an implementation environment for implementing an object detection method based on a local scenario in an embodiment of the present application;

[0104] Figure 6 It is a schematic diagram of a process for an object detection method in an embodiment of the present application;

[0105] Figure 7 It is a schematic diagram of a structure of an image feature extraction network in an embodiment of the present application;

[0106] Figure 8 It is a schematic diagram for generating a set of visual feature map sets in an embodiment of the present application;

[0107] Figure 9 It is a schematic diagram for feature fusion based on wavelet transform in an embodiment of the present application;

[0108] Figure 10A schematic diagram for generating two sets of visual feature maps in an embodiment of the present application;

[0109] Figure 11 Another schematic diagram for feature fusion based on wavelet transform in an embodiment of the present application;

[0110] Figure 12 Another schematic diagram for generating two sets of visual feature maps in an embodiment of the present application;

[0111] Figure 13 Another schematic diagram for feature fusion based on wavelet transform in an embodiment of the present application;

[0112] Figure 14 A schematic diagram for implementing feature fusion based on a convolutional fusion module in an embodiment of the present application;

[0113] Figure 15 A schematic diagram for generating a target query vector in an embodiment of the present application;

[0114] Figure 16 A schematic diagram for object detection based on equally spaced frame extraction in an embodiment of the present application;

[0115] Figure 17 A schematic diagram for key frame extraction in an embodiment of the present application;

[0116] Figure 18 An overall architecture diagram for an object detection method in an embodiment of the present application;

[0117] Figure 19 A schematic diagram for an object detection device in an embodiment of the present application;

[0118] Figure 20 A schematic structural diagram of a computer device in an embodiment of the present application. Detailed implementation manners

[0119] The embodiments of the present application provide a method, related device, equipment, and storage medium for object detection. In the process of feature fusion in the present application, hierarchical visual features and text features are fused to achieve fine-grained fusion between visual features and text features, thereby improving the accuracy of object detection.

[0120] The terms "first", "second", "third", "fourth", etc. (if any) in the description, claims, and the above-mentioned drawings of this application are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "include" and "correspond to" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0121] Multi-object tracking (MOT) is an important task in computer vision, aiming to continuously locate multiple objects simultaneously in a video sequence. First, all objects need to be detected in each frame, and then the detected objects are associated between different frames to determine which detections belong to the same object and assign a unique identifier. The output result is the bounding box, category, and identifier of the continuously located object in the corresponding frame. Adding text instructions to the MOT task to guide the model to continuously locate the objects that match the description in the text instructions is called the referring multi-object tracking (RMOT) task, and the core technology of this task is the text-visual fusion technology. RMOT is a new language understanding task that can continuously locate multiple targets that match the text description in a video sequence according to text instructions.

[0122] Currently, in the transformer referring multi-object tracking (TransRMOT) scheme, on the one hand, the visual features and text features are not fully fused, resulting in poor model performance. On the other hand, the query vector is randomly initialized, resulting in slow model convergence, requiring higher training costs, and limiting the performance ceiling of the model. In addition, the overall framework of the TransRMOT scheme has a large computational amount, slow model inference speed, and high requirements for computing power costs.

[0123] Based on this, in the embodiments of the present application, a method for object detection is provided. First, based on the local and global multimodal feature fusion module of wavelet transform, the text features and visual features are fully fused in local features and global features respectively, so as to improve the multimodal alignment ability and generalization ability of the model. Secondly, the initial query vector is respectively subjected to attention operations with the word-level features of the prompt text, and fine-grained interactions at the feature level are carried out with the sentence-level features to obtain the target query vector, thereby strengthening the model's ability to learn relevant text information features. Finally, some images are detected using this solution by frame extraction, while the remaining images are subjected to template matching based on the detection results to obtain the corresponding detection results, thereby greatly reducing the computational amount in the inference stage and improving the operation efficiency of the model.

[0124] Before introducing the specific method of the present application, an exemplary description of the application scenario of the present application is given first. It should be noted that the following application scenarios are only for illustration, and are not limited thereto in practice.

[0125] (1) Harmful item detection scenario;

[0126] In the current situation where digital content spreads rapidly, it is of great significance to detect harmful items in videos or images. Its core purpose is to accurately identify whether there are relevant harmful items in videos and images by means of technology, such as vulgar content, illegal items, violent scenes, etc. The object detection method provided by the present application can provide technical support for harmful item identification.

[0127] Exemplarily, please refer to Figure 1 , Figure 1 is a schematic diagram applied to the harmful item detection scenario in the embodiments of the present application. As shown in the figure, 101 is used to indicate the text input area. The user can input the prompt text in the text input area indicated by 101, detect the video or image based on the prompt text, and obtain the corresponding detection result. For the videos or images in which harmful items are detected, they can be directly blocked or marked for manual review.

[0128] (2) Video editing scenario;

[0129] In the field of film and television production, face replacement technology can flexibly replace the actor's image to complete the creation of special characters or make up for shooting defects. Based on the object detection method provided by the present application, it is possible to continuously locate the bad actors or vulgar content that appear in the video, locate the area where they are located, and then edit the content of the area, thereby improving the video editing efficiency.

[0130] Exemplarily, please refer to Figure 2 , Figure 2This is a schematic diagram applied to the video editing scenario in the embodiments of the present application. As shown in the figure, 201 is used to indicate the text input area. The user can select an image as the content to be replaced and input prompt text in the text input area indicated by 201, and detect the video based on the prompt text. Thus, the detected object is replaced with the replacement content provided by the user.

[0131] (3) Content rewriting scenario;

[0132] In video production, specific areas in the video are covered by blurring, mosaics or other effects to make them unable to be clearly recognized, so as to protect personal privacy. Based on the object detection method provided by the present application, objects appearing in the video (such as faces, license plates, personal information, etc.) can be coded, and the temporal consistency of the target to be rewritten between video frames can be ensured, and the located target area can be rewritten, and the consistency of cross-frame content rewriting can be achieved according to the temporality of the target area.

[0133] Exemplarily, please refer to Figure 3 , Figure 3 This is a schematic diagram applied to the content rewriting scenario in the embodiments of the present application. As shown in Figure 3 Figure (A) therein, 301 is used to indicate the text input area. The user can input prompt text in the text input area indicated by 301 and detect the video based on the prompt text. As shown in Figure 3 Figure (B) therein, assuming that the user needs to code a license plate, then after detecting the area where the license plate is located, the area is coded. Among them, 302 is used to indicate the effect after coding.

[0134] It should be noted that the above application scenarios are only examples, and the object detection method provided in this embodiment can also be applied to other scenarios, which are not limited here.

[0135] The method provided by the present application can be applied to Figure 4 or Figure 5 the implementation environment shown in Figure 4 The one shown is the method for object detection in the local scenario, Figure 5 The one shown is the method for object detection in the network scenario.

[0136] I. Local communication architecture;

[0137] Please refer to Figure 4, the implementation environment includes a terminal 401. The terminal 401 involved in this application includes but is not limited to mobile phones, tablet computers, laptop computers, desktop computers, intelligent voice interaction devices, virtual reality devices, smart home appliances, vehicle-mounted terminals, aircraft, etc. Among them, the client is deployed on the terminal 401. The client can run on the terminal 401 in the form of a browser, or can also run on the terminal 110 in the form of an independent application (APP) or a small program, etc.

[0138] Combined with the above implementation environment, in step A1, the user inputs the visual data to be detected and the prompt text. In step A2, the terminal 401 extracts features from the image frame to be detected, and obtains K feature maps of the image frame to be detected. In step A3, the terminal 401 performs wavelet transform on each feature map to obtain at least one set of visual feature map sets corresponding to each feature map. At the same time, in step A4, the terminal 401 extracts features from the prompt text to obtain the original text features. In step A5, the terminal 401 fuses the original text features, K feature maps and at least one set of visual feature map sets corresponding to each feature map to obtain the text-image fusion feature maps corresponding to each feature map (that is, K text-image fusion feature maps). In step A6, the terminal 401 generates the object detection result of the image frame to be detected according to the K text-image fusion feature maps.

[0139] II. Network communication architecture;

[0140] Please refer to Figure 5 , the implementation environment includes a terminal 501 and a server 502, and the terminal 501 and the server 502 can communicate with each other through a network 503. Among them, the network 503 uses standard communication technologies and / or protocols, usually the Internet, but can also be any network, including but not limited to any combination of Bluetooth, local area network (LAN), metropolitan area network (MAN), wide area network (WAN), mobile, private network or virtual private network). In some embodiments, custom or dedicated data communication technologies can be used to replace or supplement the above data communication technologies.

[0141] The server 502 involved in this application can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence (AI) platforms.

[0142] Combined with the above implementation environment, in step B1, the user inputs the visual data to be detected and the prompt text. In step B2, the terminal 501 sends the visual data to be detected and the prompt text to the server 502 through the network 503.

[0143] In step B3, the server 502 extracts features from the image frame to be detected, obtaining K feature maps of the image frame to be detected. In step B4, the server 502 performs wavelet transform on each feature map, obtaining at least one set of visual feature map sets corresponding to each feature map. At the same time, in step B5, the server 502 extracts features from the prompt text, obtaining the original text features. In step B6, the server 502 fuses the original text features, K feature maps, and at least one set of visual feature map sets corresponding to each feature map, obtaining the text-image fusion feature maps corresponding to each feature map (i.e., K text-image fusion feature maps). In step B7, the server 502 generates the object detection result of the image frame to be detected based on the K text-image fusion feature maps. In step B8, the server 502 sends the object detection result to the terminal 501 through the network 503. In step B9, the terminal 501 displays the object detection result.

[0144] Combined with the above introduction, the method for object detection in this application will be introduced below. Please refer to Figure 6 , in the embodiment of this application, the method for object detection can be completed independently by the server, independently by the terminal, or jointly by the terminal and the server. The method provided in this application includes:

[0145] S601. Obtain the visual data to be detected and the corresponding prompt text, where the visual data to be detected includes an image frame sequence composed of at least one image frame to be detected;

[0146] In one or more embodiments, the visual data to be detected and the prompt text are obtained, where the visual data to be detected includes at least one frame of image to be detected. If the visual data to be detected includes only one frame of image to be detected, the visual data to be detected belongs to image data. If the visual data to be detected includes at least one frame of image to be detected, the visual data to be detected belongs to video data. The prompt text is used to guide the model to locate and identify the object to be detected in the object detection task through text instructions, that is, the prompt text can indicate the object to be detected or the target to be continuously located. For example, the prompt text is "Please find a black penguin", that is, it indicates to detect or continuously locate the image area containing the black penguin in the visual data to be detected, where the object to be detected is "a black penguin".

[0147] S602. Extract features from the image frames to be detected in the visual data to be detected, and obtain K feature maps corresponding to the image frames to be detected, where K is an integer greater than or equal to 1;

[0148] In one or more embodiments, taking the processing of one image frame to be detected in the visual data to be detected as an example, the image frame to be detected is used as the input of the image feature extraction network, and K feature maps are output through the image feature extraction network. When K is greater than 1, the K feature maps have different scales respectively.

[0149] It can be understood that in practical applications, it is necessary to extract features from each image frame to be detected in the visual data to be detected to obtain K feature maps corresponding to each image frame to be detected, which will not be elaborated here.

[0150] S603. Perform wavelet transform on the K feature maps to obtain at least one set of visual feature map sets corresponding to each feature map, where each set of visual feature map sets includes at least one of a global visual feature map, a vertical visual feature map, a horizontal visual feature map, and a diagonal visual feature map;

[0151] In one or more embodiments, taking one of the K feature maps corresponding to the image frame to be detected as an example, wavelet transform is performed on this feature map to extract the corresponding at least one set of visual feature map sets. Among them, wavelet transform is a signal processing technology used to analyze and represent the local features of signals. In image processing, it can be used for feature extraction and multi-resolution analysis, and can capture the edge and texture features in images, and is suitable for image recognition and classification. In addition, wavelet transform provides multiple resolution representations, enabling the features of images at different scales to be effectively analyzed and processed.

[0152] Specifically, each set of visual feature maps includes at least one of a global visual feature map (which can also be referred to as a "low-frequency signal feature map"), a vertical visual feature map, a horizontal visual feature map, and a diagonal visual feature map. Among them, the global visual feature map retains the main contour and low-frequency information of the image frame to be detected. The vertical visual feature map captures the edges and textures in the vertical direction of the image frame to be detected. The horizontal visual feature map captures the edges and textures in the horizontal direction of the image frame to be detected. The diagonal visual feature map captures the edges and textures in the diagonal direction of the image frame to be detected.

[0153] It can be understood that in practical applications, it is necessary to perform wavelet transform on each feature map corresponding to the image frame to be detected respectively to obtain at least one set of visual feature map sets corresponding to each feature map, which will not be elaborated here.

[0154] S604. Extract features from the prompt text to obtain the original text features;

[0155] In one or more embodiments, the prompt text is used as the input of the text feature extraction network, and the original text features are output through the text feature extraction network. The original text features are represented as a matrix, where each row in the original text features corresponds to a token, and each column corresponds to a dimension of the representation vector. Exemplarily, the token "goose" corresponds to a 768-dimensional representation vector, for example, [0.4, 0.1, -0.5,..., 0.3].

[0156] S605. Fuse the original text features, K feature maps, and at least one set of visual feature map sets corresponding to each feature map to obtain the text-image fusion feature map corresponding to each feature map;

[0157] In one or more embodiments, taking one of the K feature maps as an example, after obtaining at least one set of visual feature map sets corresponding to the feature map, perform fusion processing with the original text features and the feature map to obtain the text-image fusion feature map corresponding to the feature map.

[0158] It can be understood that in practical applications, it is necessary to perform fusion processing on each feature map corresponding to the image frame to be detected respectively to obtain the text-image fusion feature map corresponding to each feature map, which will not be elaborated here. The text-image fusion feature map corresponding to the feature map is a new feature map obtained after feature interaction between the feature map and the original text features.

[0159] S606. Generate the object detection result of the image frame to be detected according to the K text-image fusion feature maps corresponding to the K feature maps.

[0160] In one or more embodiments, since there is a text-image fusion feature map corresponding to each feature map, therefore, K feature maps correspond to K text-image fusion feature maps. Based on this, the size of each text-image fusion feature map is adjusted, that is, each text-image fusion feature map is first flattened through a flatten layer, and then the K flattened features are concatenated to obtain a visual text feature vector. Finally, based on the visual text feature vector, the object detection result of the image frame to be detected is generated.

[0161] Specifically, assuming that the maximum number of detections is N (that is, at most N different types of objects can be detected), then the object detection result includes N query results. That is, each query result includes the detection probability of a class of objects, the bounding box information, and the instruction matching degree. Among them, the detection probability is used to measure whether the class of objects is the object to be detected, the bounding box information is the predicted box of the class of objects, and the instruction matching degree is used to measure whether the class of objects conforms to the description of the prompt text.

[0162] It can be understood that in actual applications, each image frame to be detected in the visual data to be detected needs to be detected to obtain the object detection result of each image frame to be detected, which will not be elaborated here.

[0163] In the embodiments of the present application, a method for object detection is provided. Through the above method, in the feature fusion process, hierarchical visual features and text features are fused to achieve fine-grained fusion between visual features and text features, thereby improving the accuracy of object detection.

[0164] Optionally, based on the above Figure 6 In another optional embodiment provided by the embodiments of the present application on the basis of the corresponding one or more embodiments, feature extraction is performed on the image frame to be detected in the visual data to be detected to obtain K feature maps corresponding to the image frame to be detected, specifically including:

[0165] Through the initial convolutional layer included in the image feature extraction network, convolutional processing is performed on the image frame to be detected to obtain the initial feature map corresponding to the image frame to be detected;

[0166] Through the feature processing layer included in the image feature extraction network, pooling processing is performed on the initial convolutional feature to obtain the pooled feature map, where the feature processing layer includes a batch normalization layer, an activation function layer, and a max pooling layer;

[0167] Through the R residual block layers included in the image feature extraction network, feature extraction is performed on the pooled feature map to obtain R feature maps, where each residual block layer is stacked by multiple bottleneck layers, and each bottleneck layer includes convolutional layers connected in sequence, and R is an integer greater than or equal to K;

[0168] Obtain K feature maps with different scales from R feature maps.

[0169] In one or more embodiments, a method for extracting feature maps based on an image feature extraction network is introduced. As can be seen from the foregoing embodiments, when K is greater than 1, K feature maps with different scales can be obtained. The feature scale can be understood as the output corresponding to different network layers in the image feature extraction network. The deeper the feature scale, the richer the semantic information contained in the corresponding visual features, and the smaller the corresponding feature size.

[0170] It should be noted that the image feature extraction network can adopt a residual network 50 (residual network 50, ResNet50), ResNet18, ResNet101, or a feature pyramid network (feature pyramid network, FPN), etc., and wavelet transform processing can also be added. Hereinafter, the case where the image feature extraction network adopts ResNet50 will be taken as an example for introduction.

[0171] Specifically, please refer to Figure 7 , Figure 7 which is a schematic structural diagram of the image feature extraction network in the embodiment of the present application. As shown in the figure, the image frame to be detected is used as the input of the initial convolutional layer, and the initial feature map is obtained through the initial convolutional layer. Next, batch normalization processing, activation processing, and max pooling processing are sequentially performed on the initial feature map to obtain a pooled feature map. Taking the image feature extraction network including 4 residual block layers (that is, R is equal to 4) as an example, the pooled feature map is input into the "residual block layer 1" to obtain the feature map F2. The feature map F2 is input into the "residual block layer 2" to obtain the feature map F3. The feature map F3 is input into the "residual block layer 3" to obtain the feature map F4. The feature map F4 is input into the "residual block layer 4" to obtain the feature map F5.

[0172] For the convenience of description, please refer to Table 1, which is a composition schematic of each residual block layer.

[0173] Table 1

[0174] Hierarchy Number of Bottleneck Layers Number of Convolution Layers Residual Block Layer 1 3 9 Residual Block Layer 2 4 12 Residual Block Layer 3 6 18 Residual Block Layer 4 3 9

[0175] Among them, the bottleneck layer is a special layer structure mainly used for feature dimension reduction and key information extraction. Each bottleneck layer includes three convolutional layers, namely a 1×1 convolutional layer, a 3×3 convolutional layer, and a 1×1 convolutional layer. Based on this, the last three-layer feature maps can be extracted from the four feature maps as the feature maps, that is, the K feature maps include the feature map F3, the feature map F4, and the feature map F5. At this time, K is equal to 3.

[0176] Secondly, in the embodiments of the present application, a method for extracting a feature map based on an image feature extraction network is provided. Through the above method, feature maps of different scales can be extracted. The small-scale feature map can focus on the detailed information in the image frame to be detected, while the large-scale feature map can capture the overall structure and semantic information in the image frame to be detected. Thus, a richer and more comprehensive feature representation can be obtained, so as to better understand the image content.

[0177] Optionally, based on one or more corresponding embodiments above, in another optional embodiment provided by the embodiments of the present application, feature extraction is performed on the prompt text to obtain the original text features, specifically including: Figure 6 Performing word segmentation on the prompt text to obtain at least one token;

[0178] Encoding each token in the at least one token to obtain a word vector corresponding to each token;

[0179] Generating a position vector for each token according to the position of each token in the prompt text;

[0180] Generating a combined vector for each token according to the word vector and the position vector corresponding to each token;

[0181] Based on the combined vector of each token, obtaining the original text features through a text feature extraction network.

[0182] In one or more embodiments, a method for extracting the original text features based on a text feature extraction network is introduced. As can be seen from the foregoing embodiments, feature extraction can be performed on the prompt text to obtain the original text features in a matrix structure.

[0183] It should be noted that the text feature extraction network can adopt bidirectional encoder representations from transformers (BERT), robustly optimized BERT approach (RoBERTa), or a lite BERT (ALBERT), etc. Hereinafter, the case where the text feature extraction network adopts RoBERTa will be taken as an example for introduction.

[0184]

[0185] ​Specifically, taking the prompt text "a black penguin" as an example, first, the prompt text is segmented into a series of tokens, namely, ["[CLS]","一","只","黑","色","的","彭","鹅","[SEP]"]. Among them, [CLS] represents a special marker at the beginning of a sentence, which is used to capture the semantic information of the entire sentence, and [SEP] represents a special marker at the end or separation of a sentence.

[0186] Next, RoBERTa converts each token (including special tags) into a fixed-size word vector that the model learned during pre-training. At the same time, a corresponding position vector is generated based on the position of each token in the prompt text. The word vector of each token is added to the corresponding position vector to obtain the combined vector of each token.

[0187] The combined vector of each token is used as the input of the text feature extraction network, and the representation vector of each token is output through the text feature extraction network. These representation vectors constitute the original text features. Among them, the text feature extraction network includes a multi-layer transformer (Transformer), which consists of multiple self-attention mechanisms and feedforward neural network layers. They can capture the long-distance dependencies between tokens and generate representation vectors containing rich contextual information.

[0188] Secondly, in the embodiment of the present application, a method for extracting original text features based on a text feature extraction network is provided. Through the above method, the text feature extraction network can capture complex semantic information, automatically learn the lexical, syntactic and semantic information in the text, capture the long-term dependency relationship and complex semantic structure between words in the text, and thus more accurately represent the semantic content of the text.

[0189] Optionally, in the above Figure 6 On the basis of one or more corresponding embodiments, in another optional embodiment provided by the embodiment of the present application, wavelet transform is performed on K feature maps to obtain at least one set of visual feature map sets corresponding to each feature map, specifically including:

[0190] The global convolution kernel, vertical convolution kernel, horizontal convolution kernel and diagonal convolution kernel are used to perform convolution processing on the K feature maps respectively, and the global visual feature map, vertical visual feature map, horizontal visual feature map and diagonal visual feature map corresponding to each feature map are obtained in turn;

[0191] Among them, the global visual feature map, the vertical visual feature map, the horizontal visual feature map and the diagonal visual feature map constitute a set of visual feature maps.

[0192] In one or more embodiments, a method for extracting a set of visual feature maps based on wavelet transform is introduced. As can be seen from the foregoing embodiments, taking the wavelet transform of a feature map corresponding to the image frame to be detected as an example, a global convolution kernel, a vertical convolution kernel, a horizontal convolution kernel, and a diagonal convolution kernel are respectively used to extract the corresponding visual feature maps. Hereinafter, the feature map "Feature Map F3" will be taken as an example for introduction.

[0193] Specifically, please refer to Figure 8 , Figure 8 , which is a schematic diagram of generating a set of visual feature maps in an embodiment of the present application. As shown in the figure, assuming that the set step is 2, exemplarily, the global convolution kernel can be expressed as

[0194]

[0195] Among them, the global convolution kernel is used to retain the main contour and low-frequency information of the image. The global convolution kernel is used to perform convolution processing on the feature map F3 to obtain the global visual feature map F3' LL .

[0196] Exemplarily, the vertical convolution kernel can be expressed as

[0197]

[0198] Among them, the vertical convolution kernel is used to capture the edges and textures of the image in the vertical direction. The vertical convolution kernel is used to perform convolution processing on the feature map F3 to obtain the vertical visual feature map F3' LH .

[0199] Exemplarily, the horizontal convolution kernel can be expressed as

[0200]

[0201] Among them, the horizontal convolution kernel is used to capture the edges and textures of the image in the horizontal direction. The horizontal convolution kernel is used to perform convolution processing on the feature map F3 to obtain the horizontal visual feature map F3' HL .

[0202] Exemplarily, the diagonal convolution kernel can be expressed as

[0203]

[0204] Among them, the diagonal convolution kernel is used to capture the edges and textures of the image in the diagonal direction. The diagonal convolution kernel is used to perform convolution processing on the feature map F3 to obtain the diagonal visual feature map F3' HH .

[0205] Based on this, a set of visual feature maps corresponding to the feature map F3 includes the global visual feature map F3' LL , the vertical visual feature map F3'LH , the horizontal visual feature map F3' HL and the diagonal visual feature map F3' HH .

[0206] It should be noted that for any one feature Figure X , a set of visual feature maps extracted by using the global convolution kernel, the vertical convolution kernel, the horizontal convolution kernel, and the diagonal convolution kernel can be expressed as:

[0207] [X LL , X LH , X HL , X HH = [WC LL (X), WC LH (X), WC HL (X), WC HH (X)]; Equation (5)

[0208] wherein, X LL represents the global visual feature map of the feature Figure X . X LH represents the vertical visual feature map of the feature Figure X . X HL represents the horizontal visual feature map of the feature Figure X . X HH represents the diagonal visual feature map of the feature Figure X .

[0209] Secondly, in the embodiments of the present application, a method for extracting a set of visual feature maps based on wavelet transform is provided. Through the above method, the image frame to be detected is decomposed into different scales and frequencies based on wavelet transform for analysis. Thus, the features of the image frame to be detected at different scales can be captured simultaneously, which is convenient for a more comprehensive understanding and processing of the image frame to be detected.

[0210] Optionally, based on one or more corresponding embodiments described above, in another optional embodiment provided by the embodiments of the present application, the original text feature, K feature maps, and at least one set of visual feature map sets corresponding to each feature map are fused to obtain a text-image fusion feature map corresponding to each feature map, which specifically includes: Figure 6 The original text feature is respectively fused with the global visual feature map, the vertical visual feature map, the horizontal visual feature map, and the diagonal visual feature map corresponding to the feature map to sequentially obtain a global fusion feature map, a vertical fusion feature map, a horizontal fusion feature map, and a diagonal fusion feature map corresponding to the feature map;

[0211]

[0212] ​Inverse wavelet transform is performed after merging the global fusion feature map, vertical fusion feature map, horizontal fusion feature map, and diagonal fusion feature map corresponding to the feature map to obtain the target feature map corresponding to the feature map;

[0213] The original text feature and the feature map are fused to obtain the fusion feature map corresponding to the feature map;

[0214] The fusion feature map corresponding to the feature map and the target feature map corresponding to the feature map are merged to obtain the text-image fusion feature map corresponding to the feature map.

[0215] In one or more embodiments, a method for fusing multi-modal features is introduced. As can be seen from the foregoing embodiments, based on the global fusion feature map, vertical fusion feature map, horizontal fusion feature map, and diagonal fusion feature map, the fusion of local features to global features is realized. For the sake of convenience of description, the following takes the generation of the text-image fusion feature map corresponding to a feature map as an example for introduction.

[0216] Specifically, assume that the feature map is "Feature Map F3" and the prompt text is "A black penguin". Please refer to Figure 9 , Figure 9 which is a schematic diagram of feature fusion based on wavelet transform in the embodiments of the present application. As shown in the figure, the feature map F3 and the original text feature T emd are input into the Convolutional Fusion (Conv-Fuse) module for preliminary fusion of visual features and text features to obtain the fusion feature map H3. In addition, after performing wavelet transform on the feature map F3, the global visual feature map F3' LL , vertical visual feature map F3' LH , horizontal visual feature map F3' HL , and diagonal visual feature map F3' HH are obtained.

[0217] Then, the global visual feature map F3' LL , vertical visual feature map F3' LH , horizontal visual feature map F3' HL , and diagonal visual feature map F3' HH obtained by wavelet transform, and the original text feature T emd are input into the Conv-Fuse module for in-depth fusion of visual features and text features at different levels, so as to obtain the global fusion feature map G3' LL , vertical fusion feature map G3' LH , horizontal fusion feature map G3' HL , and diagonal fusion feature map G3' HH .

[0218] Next, for the global fusion feature map G3' LL, the vertically fused feature map G3' LH , the horizontally fused feature map G3' HL and the diagonally fused feature map G3' HH are merged, that is, stacked in the channel dimension. Thus, the merged feature map is then grouped by channels, and then the global convolution kernel, vertical convolution kernel, horizontal convolution kernel, and diagonal convolution kernel are respectively used to perform inverse wavelet transform on the corresponding grouped feature maps, thereby obtaining the target feature map G3.

[0219] Finally, the fused feature map H3 and the target feature map G3 are merged, that is, element-wise addition is performed. Thus, the text-image fused feature map of the feature map F3 is obtained.

[0220] It can be understood that in practical applications, it is necessary to generate corresponding text-image fused feature maps for each feature map corresponding to the image frame to be detected, which will not be elaborated here.

[0221] Again, in the embodiments of the present application, a method for fusing multi-modal features is provided. Through the above method, the original text features can be fully fused with the global fused feature map, vertically fused feature map, horizontally fused feature map, and diagonally fused feature map respectively. Thus, the multi-modal feature fusion interaction at different levels is realized, and the expression ability of the model is improved.

[0222] Optionally, on the basis of one or more of the above Figure 6 corresponding embodiments, in another optional embodiment provided by the embodiments of the present application, wavelet transform is performed on K feature maps to obtain at least one set of visual feature map sets corresponding to each feature map, specifically including:

[0223] The global convolution kernel, vertical convolution kernel, horizontal convolution kernel, and diagonal convolution kernel are respectively used to perform convolution processing on the K feature maps, and the first global visual feature map, first vertical visual feature map, first horizontal visual feature map, and first diagonal visual feature map corresponding to each feature map are obtained in sequence;

[0224] The global convolution kernel, vertical convolution kernel, horizontal convolution kernel, and diagonal convolution kernel are respectively used to perform convolution processing on the first global visual feature map corresponding to each feature map, and the second global visual feature map, second vertical visual feature map, second horizontal visual feature map, and second diagonal visual feature map corresponding to each feature map are obtained in sequence;

[0225] Among them, the first global visual feature map, first vertical visual feature map, first horizontal visual feature map, and first diagonal visual feature map constitute a set of visual feature map sets;

[0226] And the second global visual feature map, second vertical visual feature map, second horizontal visual feature map, and second diagonal visual feature map constitute another set of visual feature map sets.

[0227] In one or more embodiments, another method for extracting a set of visual feature maps based on wavelet transform is introduced. As can be seen from the foregoing embodiments, taking the wavelet transform of a feature map corresponding to the image frame to be detected as an example, a global convolution kernel, a vertical convolution kernel, a horizontal convolution kernel, and a diagonal convolution kernel are respectively used to extract the corresponding visual feature maps. Hereinafter, the feature map "Feature Map F3" will be used as an example for introduction.

[0228] Specifically, please refer to Figure 10 , Figure 10 which is a schematic diagram for generating two sets of visual feature map sets in the embodiments of the present application. As shown in the figure, the feature map F3 is subjected to convolution processing using the global convolution kernel indicated by Equation (1) to obtain the first global visual feature map F3' LL . The feature map F3 is subjected to convolution processing using the vertical convolution kernel indicated by Equation (2) to obtain the first vertical visual feature map F3' LH . The feature map F3 is subjected to convolution processing using the horizontal convolution kernel indicated by Equation (3) to obtain the first horizontal visual feature map F3' HL . The feature map F3 is subjected to convolution processing using the diagonal convolution kernel indicated by Equation (4) to obtain the first diagonal visual feature map F3' HH .

[0229] Based on this, a set of visual feature maps corresponding to the feature map F3 includes the first global visual feature map F3' LL , the first vertical visual feature map F3' LH , the first horizontal visual feature map F3' HL , and the first diagonal visual feature map F3' HH .

[0230] The first global visual feature map F3' LL is subjected to convolution processing using the global convolution kernel indicated by Equation (1) to obtain the second global visual feature map F3'' LL . The first global visual feature map F3' LL is subjected to convolution processing using the vertical convolution kernel indicated by Equation (2) to obtain the second vertical visual feature map F3'' LH . The first global visual feature map F3' LL is subjected to convolution processing using the horizontal convolution kernel indicated by Equation (3) to obtain the second horizontal visual feature map F3'' HL . The first global visual feature map F3' LL is subjected to convolution processing using the diagonal convolution kernel indicated by Equation (4) to obtain the second diagonal visual feature map F3'' HH .

[0231] Based on this, another set of visual feature map sets corresponding to the feature map F3 includes the second global visual feature map F3” LL , the second vertical visual feature map F3” LH , the second horizontal visual feature map F3” HL and the second diagonal visual feature map F3” HH .

[0232] Secondly, in the embodiments of the present application, another method for extracting a set of visual feature maps based on wavelet transform is provided. Through the above method, the image frame to be detected is decomposed into finer scale and frequency sub-bands based on two wavelet transforms. Multiple wavelet transforms can further refine the analysis of the signal, revealing more subtle features and changes in the signal, so as to be able to describe the characteristics of the image frame to be detected more comprehensively and accurately.

[0233] Optionally, based on one or more corresponding embodiments above Figure 6 , in another optional embodiment provided by the embodiments of the present application, the original text features, K feature maps, and at least one set of visual feature map sets corresponding to each feature map are fused to obtain a text-image fusion feature map corresponding to each feature map, specifically including:

[0234] Fuse the original text features with the first global visual feature map, the first vertical visual feature map, the first horizontal visual feature map, and the first diagonal visual feature map corresponding to the feature map respectively, and sequentially obtain the first global fusion feature map, the first vertical fusion feature map, the first horizontal fusion feature map, and the first diagonal fusion feature map corresponding to the feature map;

[0235] Fuse the original text features with the second global visual feature map, the second vertical visual feature map, the second horizontal visual feature map, and the second diagonal visual feature map corresponding to the feature map respectively, and sequentially obtain the second global fusion feature map, the second vertical fusion feature map, the second horizontal fusion feature map, and the second diagonal fusion feature map corresponding to the feature map;

[0236] Perform inverse wavelet transform on the merged second global fusion feature map, the second vertical fusion feature map, the second horizontal fusion feature map, and the second diagonal fusion feature map corresponding to the feature map to obtain the first target feature map corresponding to the feature map;

[0237] Merge the first target feature map corresponding to the feature map with the first global fusion feature map, the first vertical fusion feature map, the first horizontal fusion feature map, and the first diagonal fusion feature map corresponding to the feature map respectively, and sequentially obtain the first fusion feature map, the second fusion feature map, the third fusion feature map, and the fourth fusion feature map corresponding to the feature map;

[0238] Inverse wavelet transform is performed on the first fused feature map, the second fused feature map, the third fused feature map, and the fourth fused feature map corresponding to the feature map to obtain the second target feature map corresponding to the feature map;

[0239] The original text features and the feature map are fused to obtain the fused feature map corresponding to the feature map;

[0240] The fused feature map corresponding to the feature map and the second target feature map corresponding to the feature map are merged to obtain the text-image fused feature map corresponding to the feature map.

[0241] In one or more embodiments, another way of fusing multi-modal features is introduced. As can be seen from the foregoing embodiments, based on the global fused feature map, the vertical fused feature map, the horizontal fused feature map, and the diagonal fused feature map, the fusion of local features to global features is realized. For the sake of illustration, the following takes the generation of a text-image fused feature map corresponding to a feature map as an example for introduction.

[0242] Specifically, assume that the feature map is "Feature Map F3" and the prompt text is "A black penguin". Please refer to Figure 11 , Figure 11 which is another schematic diagram of feature fusion based on wavelet transform in the embodiments of the present application. As shown in the figure, the feature map F3 and the original text feature T emd are input into the Conv-Fuse module for preliminary fusion of visual features and text features to obtain the fused feature map H3. In addition, after performing wavelet transform on the feature map F3 in cascade twice, the first global visual feature map F3' LL , the first vertical visual feature map F3' LH , the first horizontal visual feature map F3' HL , and the first diagonal visual feature map F3' HH are obtained, and also, the second global visual feature map F3” LL , the second vertical visual feature map F3” LH , the second horizontal visual feature map F3” HL , and the second diagonal visual feature map F3” HH .

[0243] Then, the second global visual feature map F3” LL , the second vertical visual feature map F3” LH , the second horizontal visual feature map F3” HL , and the second diagonal visual feature map F3” HH , together with the original text feature T emd are input into the Conv-Fuse module for in-depth fusion of visual features and text features at different levels, so as to obtain the second global fused feature map H3”. LL, the second vertical fusion feature map H3" LH , the second horizontal fusion feature map H3" HL and the second diagonal fusion feature map H3" HH .

[0244] Next, the second global fusion feature map H3", LL , the second vertical fusion feature map H3", LH , the second horizontal fusion feature map H3", HL and the second diagonal fusion feature map H3" HH are merged, that is, channel dimension stacking is performed. Then, channel grouping is performed on the merged feature map, and then inverse wavelet transform is performed on the corresponding grouped feature maps using a global convolution kernel, a vertical convolution kernel, a horizontal convolution kernel, and a diagonal convolution kernel respectively, so as to obtain the first target feature map H3".

[0245] Thus, the first global visual feature map F3', LL , the first vertical visual feature map F3', LH , the first horizontal visual feature map F3', HL and the first diagonal visual feature map F3' HH , and the original text feature T emd are input into the Conv-Fuse module to perform deep-level fusion of visual features and text features at different levels, so as to obtain the first global fusion feature map, the first vertical fusion feature map, the first horizontal fusion feature map, and the first diagonal fusion feature map. The first target feature map H3" is merged with the first global fusion feature map, the first vertical fusion feature map, the first horizontal fusion feature map, and the first diagonal fusion feature map respectively, that is, element addition is performed. Thus, the first fusion feature map H3', LL , the second fusion feature map H3', LH , the third fusion feature map H3', HL and the fourth fusion feature map H3' HH are obtained.

[0246] Then, the first fusion feature map H3', LL , the second fusion feature map H3', LH , the third fusion feature map H3', HL and the fourth fusion feature map H3' HH are merged, that is, channel dimension stacking is performed. Then, channel grouping is performed on the merged feature map, and then inverse wavelet transform is performed on the corresponding grouped feature maps using a global convolution kernel, a vertical convolution kernel, a horizontal convolution kernel, and a diagonal convolution kernel respectively, so as to obtain the second target feature map H3'.

[0247] Finally, the fused feature map H3 and the second target feature map H3' are merged, that is, element-wise addition is performed. Thus, the text-image fused feature map of the feature map F3 is obtained.

[0248] It can be understood that in practical applications, corresponding text-image fused feature maps need to be generated for each feature map corresponding to the image frame to be detected, which will not be elaborated here.

[0249] Again, in the embodiments of the present application, another way to fuse multi-modal features is provided. Through the above method, the original text features can be fully fused with the global fused feature map, the vertical fused feature map, the horizontal fused feature map, and the diagonal fused feature map respectively. Thus, cross-modal feature fusion and interaction at different levels are realized, improving the expression ability of the model.

[0250] Optionally, based on one or more of the above Figure 6 corresponding embodiments, in another optional embodiment provided by the embodiments of the present application, wavelet transform is performed on K feature maps to obtain at least one set of visual feature map sets corresponding to each feature map, specifically including:

[0251] Use a global convolutional kernel to perform convolutional processing on K feature maps to obtain a first global visual feature map corresponding to each feature map;

[0252] Respectively use a global convolutional kernel, a vertical convolutional kernel, a horizontal convolutional kernel, and a diagonal convolutional kernel to perform convolutional processing on the first global visual feature map corresponding to each feature map, and sequentially obtain a second global visual feature map, a second vertical visual feature map, a second horizontal visual feature map, and a second diagonal visual feature map corresponding to each feature map;

[0253] Among them, the first global visual feature map constitutes a set of visual feature map sets;

[0254] Moreover, the second global visual feature map, the second vertical visual feature map, the second horizontal visual feature map, and the second diagonal visual feature map constitute another set of visual feature map sets.

[0255] In one or more embodiments, another way to extract a set of visual feature maps based on wavelet transform is introduced. As can be seen from the foregoing embodiments, taking the wavelet transform of a feature map corresponding to an image frame to be detected as an example, a global convolutional kernel, a vertical convolutional kernel, a horizontal convolutional kernel, and a diagonal convolutional kernel are respectively used to extract the corresponding visual feature maps. Hereinafter, the feature map "feature map F3" will be used as an example for introduction.

[0256] Specifically, please refer to Figure 12 , Figure 12Another schematic diagram for generating two sets of visual feature map sets in the embodiments of the present application is shown in the figure. The feature map F3 is convolved using the global convolution kernel indicated by Equation (1) to obtain the first global visual feature map F3'. LL .

[0257] Based on this, a set of visual feature map sets corresponding to the feature map F3 includes the first global visual feature map F3'. LL .

[0258] The first global visual feature map F3' LL is convolved using the global convolution kernel indicated by Equation (1) to obtain the second global visual feature map F3". LL . The first global visual feature map F3' LL is convolved using the vertical convolution kernel indicated by Equation (2) to obtain the second vertical visual feature map F3". LH . The first global visual feature map F3' LL is convolved using the horizontal convolution kernel indicated by Equation (3) to obtain the second horizontal visual feature map F3". HL . The first global visual feature map F3' LL is convolved using the diagonal convolution kernel indicated by Equation (4) to obtain the second diagonal visual feature map F3". HH .

[0259] Based on this, another set of visual feature map sets corresponding to the feature map F3 includes the second global visual feature map F3", LL the second vertical visual feature map F3", LH the second horizontal visual feature map F3", HL and the second diagonal visual feature map F3". HH .

[0260] Secondly, in the embodiments of the present application, another method for extracting a set of visual feature maps based on wavelet transform is provided. Through the above method, the image frame to be detected is decomposed into finer scale and frequency sub-bands based on two wavelet transforms. On the one hand, through multiple wavelet transforms, the analysis of the signal can be further refined, revealing more subtle features and changes in the signal, so as to more comprehensively and accurately describe the characteristics of the image frame to be detected. On the other hand, only the low-frequency signal is extracted in the first wavelet transform, which can save the amount of data processing and improve the calculation efficiency.

[0261] Optionally, based on one or more of the above Figure 6 corresponding embodiments, in another optional embodiment provided by the embodiments of the present application, the original text features, K feature maps, and at least one set of visual feature map sets corresponding to each feature map are fused to obtain a text-image fusion feature map corresponding to each feature map, specifically including:

[0262] Fuse the original text features with the first global visual feature map corresponding to the feature map to obtain the first global fusion feature map corresponding to the feature map;

[0263] Fuse the original text features with the second global visual feature map, the second vertical visual feature map, the second horizontal visual feature map, and the second diagonal visual feature map corresponding to the feature map respectively to obtain the second global fusion feature map, the second vertical fusion feature map, the second horizontal fusion feature map, and the second diagonal fusion feature map corresponding to the feature map in sequence;

[0264] Perform inverse wavelet transform on the merged second global fusion feature map, second vertical fusion feature map, second horizontal fusion feature map, and second diagonal fusion feature map corresponding to the feature map to obtain the first target feature map corresponding to the feature map;

[0265] Perform inverse wavelet transform on the merged first target feature map corresponding to the feature map and the first global fusion feature map corresponding to the feature map to obtain the second target feature map corresponding to the feature map;

[0266] Fuse the original text features with the feature map to obtain the fusion feature map corresponding to the feature map;

[0267] Merge the fusion feature map corresponding to the feature map and the second target feature map corresponding to the feature map to obtain the text-image fusion feature map corresponding to the feature map.

[0268] In one or more embodiments, another way of fusing multi-modal features is introduced. As can be seen from the foregoing embodiments, based on the global fusion feature map, vertical fusion feature map, horizontal fusion feature map, and diagonal fusion feature map, the fusion of local features to global features is realized. For the sake of illustration, the following takes the generation of a text-image fusion feature map corresponding to a feature map as an example for introduction.

[0269] Specifically, assume that the feature map is "Feature Map F3" and the prompt text is "A black penguin". Please refer to Figure 13 , Figure 13 which is another schematic diagram of feature fusion based on wavelet transform in the embodiments of the present application. As shown in the figure, input the feature map F3 and the original text feature T emd into the Conv-Fuse module for preliminary fusion of visual features and text features to obtain the fusion feature map H3. In addition, after performing wavelet transform on the feature map F3 twice in cascade, the first global visual feature map F3' LL , and the second global visual feature map F3” LL , the second vertical visual feature map F3” LH , the second horizontal visual feature map F3” HL and the second diagonal visual feature map F3”HH ,

[0270] Then, the second global visual feature map F3” LL , the second vertical visual feature map F3” LH , the second horizontal visual feature map F3” HL and the second diagonal visual feature map F3” HH are input into the Conv-Fuse module together with the original text feature T emd to perform deep-level fusion of visual features and text features at different levels, thereby obtaining the second global fusion feature map H3” LL , the second vertical fusion feature map H3” LH , the second horizontal fusion feature map H3” HL and the second diagonal fusion feature map H3” HH .

[0271] Next, the second global fusion feature map H3” LL , the second vertical fusion feature map H3” LH , the second horizontal fusion feature map H3” HL and the second diagonal fusion feature map H3” HH are merged, that is, channel dimension stacking is performed. Then, the merged feature map is grouped by channels, and then inverse wavelet transforms are respectively performed on the corresponding grouped feature maps using a global convolution kernel, a vertical convolution kernel, a horizontal convolution kernel, and a diagonal convolution kernel, thereby obtaining the first target feature map H3”.

[0272] Thus, the first global visual feature map F3' LL and the original text feature T emd are input into the Conv-Fuse module to perform deep-level fusion of visual features and text features at different levels, thereby obtaining the first global fusion feature map. The first target feature map H3” and the first global fusion feature map are merged, that is, element-wise addition is performed. Thus, the first fusion feature map H3' LL is obtained. Then, an inverse wavelet transform is performed on the first fusion feature map H3' LL using a global convolution kernel, thereby obtaining the second target feature map H3'.

[0273] Finally, the fusion feature map H3 and the second target feature map H3' are merged, that is, element-wise addition is performed. Thus, the graphic-text fusion feature map of the feature map F3 is obtained.

[0274] It can be understood that in practical applications, corresponding graphic-text fusion feature maps need to be generated for each feature map corresponding to the image frame to be detected, which will not be elaborated here.

[0275] Again, in the embodiments of the present application, another way of fusing multi-modal features is provided. Through the above method, the original text features can be fully fused with the global fusion feature map, the vertical fusion feature map, the horizontal fusion feature map, and the diagonal fusion feature map respectively. Thereby, the multi-modal feature fusion interaction at different levels is realized, and the expression ability of the model is improved. In addition, only the low-frequency signals are extracted in the first wavelet transform, which can save the data processing amount and improve the calculation efficiency.

[0276] Optionally, based on one or more corresponding embodiments above, in another optional embodiment provided by the embodiments of the present application, the original text features and the feature map are fused to obtain a fusion feature map corresponding to the feature map, which specifically includes: Figure 6 On the basis of one or more corresponding embodiments above, in another optional embodiment provided by the embodiments of the present application, the original text features and the feature map are fused to obtain a fusion feature map corresponding to the feature map, which specifically includes:

[0277] Feature mapping is respectively performed on the original text features and the feature map to obtain a first text feature corresponding to the original text features and a first feature map corresponding to the feature map;

[0278] Self-attention calculation is respectively performed on the first text feature and the first feature map to obtain a second text feature corresponding to the first text feature and a second feature map corresponding to the first feature map;

[0279] After cross-attention calculation is performed on the second text feature and the second feature map, a skip connection is made with the first feature map to obtain a fusion feature map corresponding to the feature map.

[0280] In one or more embodiments, a method of feature fusion based on the Conv-Fuse module is introduced. As can be seen from the foregoing embodiments, the Conv-Fuse module is used to fuse the original text features and visual features, where the visual features can be a feature map, a first global visual feature map, a first vertical visual feature map, a first horizontal visual feature map, a first diagonal visual feature map, a second global visual feature map, a second vertical visual feature map, a second horizontal visual feature map, or a second diagonal visual feature map. Hereinafter, taking the fusion of a feature map and the original text features to obtain a fusion feature map as an example for introduction.

[0281] Specifically, please refer to Figure 14 , Figure 14 which is a schematic diagram of implementing feature fusion based on the convolutional fusion module in the embodiments of the present application. As shown in the figure, first, feature mapping is performed on the feature map through a mapping layer (Project Layer) to obtain a first feature map, and feature mapping is performed on the original text features through the Project Layer to obtain a first text feature. Among them, the first feature map and the first text feature have the same feature dimension.

[0282] Next, the following method can be used for attention calculation:

[0283]

[0284] Among them, Q represents the query vector, K represents the key vector, and V represents the value vector. d k represents the dimensions of the query vector and the key vector.

[0285] Based on this, self-attention calculation is performed on the first feature map, that is, after multiplying the first feature map and the corresponding weight matrices respectively to obtain the query vector, the key vector, and the value vector, self-attention calculation is performed using formula (1) to obtain the second feature map. Correspondingly, self-attention calculation is performed on the first text feature, that is, after multiplying the first text feature and the corresponding weight matrices respectively to obtain the query vector, the key vector, and the value vector, self-attention calculation is performed using formula (1) to obtain the second text feature.

[0286] Then, multiply the second feature map by the corresponding weight sentence to obtain the query vector, multiply the second text feature and the corresponding weight matrices respectively to obtain the key vector and the value vector, and then perform text-to-image cross-attention calculation using formula (1) to obtain the visual text feature containing text information. Finally, perform a skip connection (for example, element addition) between the first feature map passing through the Project Layer and the visual text feature containing text information. Thus, the fused feature map is obtained.

[0287] Furthermore, in the embodiments of the present application, a method for feature fusion based on the Conv-Fuse module is provided. Through the above method, wavelet transform and attention mechanism are used to achieve multi-level, deep-level, and fine-grained fusion of visual features and text features. Thus, the feature representation is enhanced, and long-range dependencies in different modality features are captured, so that the semantic information in the text features can be fused into the visual feature map.

[0288] Optionally, based on one or more corresponding embodiments described above, in another optional embodiment provided by the embodiments of the present application, according to the K text-image fusion feature maps corresponding to the K feature maps, an object detection result of the image frame to be detected is generated, which specifically includes: Figure 6 On the basis of one or more corresponding embodiments described above, in another optional embodiment provided by the embodiments of the present application, according to the K text-image fusion feature maps corresponding to the K feature maps, an object detection result of the image frame to be detected is generated, which specifically includes:

[0289] Obtain N initial query vectors, where N is an integer greater than or equal to 1;

[0290] Generate visual text feature vectors according to each text-image fusion feature map corresponding to each of the K feature maps;

[0291] Determine the object detection result of the image frame to be detected according to N initial query vectors and the visual text feature vector.

[0292] In one or more embodiments, a method for generating an object detection result is introduced. As can be seen from the foregoing embodiments, assuming that the model can detect at most N different types of objects, N initial query vectors need to be set. Among them, the initial query vector is a query vector obtained through random initialization.

[0293] Specifically, taking the extraction of feature maps of three different scales based on the image frame to be detected as an example, they are the feature map F3 with a scale of 4×4, the feature map F4 with a scale of 3×3, and the feature map F5 with a scale of 2×2.

[0294] The feature map F3 of 4×4 can be expressed as:

[0295] [[A1,B1,C1,D1],[A2,B2,C2,D2],[A3,B3,C3,D3],[A4,B4,C4,D4]]

[0296] The feature map F4 of 3×3 can be expressed as:

[0297] [[E1,F1,G1],[E2,F2,G2],[E3,F3,G3]]

[0298] The feature map F5 of 2×2 can be expressed as:

[0299] [[H1,I1],[H2,I2]

[0300] First, flatten the elements of the three feature maps and then splice them. Thus, a visual text feature vector is obtained, that is:

[0301] [A1,B1,C1,D1,A2,B2,C2,D2,A3,B3,C3,D3,A4,B4,C4,D4,E1,F1,G1,E2,F2,G2,E3,F3,G3,H1.I1,H2,I2]

[0302] Based on this, encode the visual text feature vector, and then decode the encoding result and the N initial query vectors to obtain the object detection result of the image frame to be detected. The encoding method and decoding method adopted here are similar to those adopted by deformable detection transformers (Deformable DETR).

[0303] Secondly, in the embodiment of the present application, a method for generating object detection results is provided. Through the above method, based on the visual text feature vector corresponding to the image frame to be detected, subsequent calculations can be directly performed with each randomly initialized initial query vector. Therefore, by introducing the randomly initialized initial query vector, the probability of finding a better solution can be increased, and the model can converge to a better parameter state, thereby improving performance.

[0304] Optionally, in the above Figure 6 On the basis of one or more corresponding embodiments, in another optional embodiment provided by the embodiment of the present application, generating an object detection result of the image frame to be detected according to K image-text fusion feature maps corresponding to the K feature maps specifically includes:

[0305] According to the original text features, N initial query vectors are processed to obtain N target query vectors, where N is an integer greater than or equal to 1;

[0306] Generate a visual text feature vector according to the image-text fusion feature map corresponding to each feature map in the K feature maps;

[0307] According to the N target query vectors and the visual text feature vector, an object detection result of the image frame to be detected is determined.

[0308] In one or more embodiments, another method of generating object detection results is introduced. As can be seen from the above embodiments, assuming that the model can detect at most N different types of objects, it is necessary to randomly initialize to obtain N initial query vectors. In order to further utilize the information described by the prompt text, semantic enhancement is performed on the N initial query vectors to obtain N target query vectors.

[0309] Specifically, the original text features extracted based on the prompt text are divided into word-level features and sentence-level features. Then, the N initial query vectors are respectively subjected to attention operations with the word-level features and feature-level fine-grained interactions with the sentence-level features to obtain N target query vectors.

[0310] Based on this, the elements of each feature map are first dimensionally flattened and then spliced ​​as described in the above embodiment, thereby obtaining a visual text feature vector. Then, the visual text feature vector is encoded, and the encoding result and N initial query vectors are decoded to obtain the object detection result of the image frame to be detected. The encoding and decoding methods used here are similar to those used by Deformable DETR.

[0311] Secondly, in the embodiments of the present application, another way to generate object detection results is provided. Through the above method, based on the visual text feature vector corresponding to the image frame to be detected, subsequent calculations can be directly performed with each target query vector endowed with text prior information. Thereby, strengthening the model to learn the features of relevant text content, which helps to improve the accuracy of object detection.

[0312] Optionally, based on one or more corresponding embodiments above, in another optional embodiment provided by the embodiments of the present application, N initial query vectors are processed according to the original text features to obtain N target query vectors, specifically including: Figure 6 Perform self-attention calculation on the N initial query vectors to obtain N first query vectors;

[0313] Obtain M word-level features and sentence-level features according to the original text features, where M is an integer greater than or equal to 1;

[0314] Perform cross-attention calculation on the M word-level features and the N first query vectors to obtain N second query vectors;

[0315] Interact the sentence-level feature and the N second query vectors to obtain N third query vectors;

[0316] Merge the N third query vectors and the N initial query vectors to obtain N target query vectors.

[0317] In one or more embodiments, a way to generate target query vectors is introduced. As can be seen from the foregoing embodiments, first, the N initial query vectors are randomly initialized, and at the same time, the input original text features are divided into word-level features (word feature) and sentence-level features (sentence feature), and then the semantic information in the original text features is incorporated into the N initial query vectors through operations such as the attention mechanism.

[0318] Specifically, please refer to

[0319] Figure 15 Figure 15 , is a schematic diagram for generating target query vectors in the embodiments of the present application. As shown in the figure, the original text features are feature-mapped through the Project Layer so that the mapped target text features and the initial query vectors are in the same feature space, that is, they have the same feature dimension. Then, the tokens corresponding to the prompt text are divided into M word-level features and sentence-level features. Among them, each token corresponds to a word-level feature, and [CLS] corresponds to the sentence-level feature.

[0320] Perform self-attention calculation on N initial query vectors to obtain the first query vector corresponding to each initial query vector. Among them, the N first query vectors are denoted as q, and the N first query vectors are denoted as q s Use M word-level features as key vectors and value vectors, use the N first query vectors as query vectors, then perform cross-attention calculation, endow the text feature semantics at the word level to the N first query vectors, and obtain N second query vectors. Among them, the N second query vectors are denoted as q c .

[0321] Based on this, interact the N second query vectors containing word-level semantic information with the sentence-level features to obtain N third query vectors containing text word and sentence semantic information. Among them, the N third query vectors are denoted as q t . Add the N third query vectors containing text feature semantics as excitation values to the N initial query vectors, that is, merge the N third query vectors and the N initial query vectors to obtain N target query vectors.

[0322] Again, in the embodiments of the present application, a method for generating target query vectors is provided. Through the above method, the semantic information of the original text features is added to each randomly generated initial query vector to obtain the corresponding target query vectors. Thereby, the ability to parse relevant content including text descriptions from the image-text fusion feature map during decoding is strengthened, and prior knowledge for detecting relevant content is given to the target query vectors, thereby further improving the model performance.

[0323] Optionally, based on one or more corresponding embodiments above Figure 6 In another optional embodiment provided by the embodiments of the present application, obtaining the visual data to be detected specifically includes:

[0324] Obtain the video to be detected, where the video to be detected includes an image frame sequence composed of at least two image frames;

[0325] Perform equally spaced frame extraction processing on the video to be detected to obtain the visual data to be detected;

[0326] It may also include:

[0327] According to the object detection results of each image frame to be detected in the visual data to be detected, detect each image frame in the video to be detected that has not been frame-extracted, and obtain the object detection results of each image frame.

[0328] In one or more embodiments, a method for object detection based on equidistant frame extraction is introduced. As can be seen from the foregoing embodiments, in order to improve the inference efficiency of the model and reduce the computational cost, in the inference stage, the video to be detected can be frame-extracted at equal intervals. Hereinafter, an example in which the video to be detected includes 100 image frames will be used for introduction.

[0329] Specifically, please refer to Figure 16 , Figure 16 which is a schematic diagram of object detection based on equidistant frame extraction in the embodiments of the present application. As shown in the figure, the video to be detected is represented as F = {f1, f2,..., f n}, assuming that every other frame is extracted. Thus, an image frame sequence F 奇 composed of odd-numbered image frames = {f1, f3, f5,...} and an image frame sequence F 偶 composed of even-numbered image frames = {f2, f4, f6,...} are obtained. It should be noted that in actual applications, more frames can also be skipped for image frame extraction, which is not limited herein.

[0330] The image frame sequence composed of odd frames can be used as the visual data to be detected, and combined with the prompt text to generate the object detection results of each image frame to be detected, that is, R 奇 = {r1, r3, r5,...}. Based on this, the odd image frames are used as the template image frames, and the even image frames are used as the search area image frames to construct each pair of image frames to be matched. For example, the first image frame in the video to be detected is used as the template image frame, and the second image frame in the video to be detected is used as the search area image frame to construct a pair of image frames to be matched. Thus, at least one pair of image frames to be matched is obtained, that is, P = {(r1, f2), (r3, f4), (r5, f6),...}.

[0331] The template matching algorithm is used to detect each pair of image frames to be matched respectively, and R 偶 = {r2, r4, r6,...} is obtained. Finally, the object detection results of the odd-frame images and the even-frame images are merged to obtain the continuous positioning results of the entire video to be detected.

[0332] Secondly, in the embodiments of the present application, a method for object detection based on equidistant frame extraction is provided. By the above method, the image frames in the video to be detected are extracted at fixed intervals, which can not only maintain the spatial consistency and temporal consistency of the video to be detected, but also save computational resources, thereby improving the running efficiency of the algorithm.

[0333] Optionally, in the above Figure 6Based on one or more corresponding embodiments, in another alternative embodiment provided by the embodiments of the present application, obtaining the visual data to be detected specifically includes:

[0334] Obtaining the video to be detected, where the video to be detected includes an image frame sequence composed of at least two image frames;

[0335] Performing key frame recognition on the video to be detected to obtain the visual data to be detected, where the image frames to be detected in the visual data to be detected are the key frames of the video to be detected;

[0336] It may further include:

[0337] According to the object detection results of each image frame to be detected in the visual data to be detected, detecting each image frame that has not been frame-extracted in the video to be detected to obtain the object detection results of each image frame.

[0338] In one or more embodiments, a method for object detection based on key frame extraction is introduced. As can be seen from the foregoing embodiments, in order to improve the inference efficiency of the model, during the inference stage, key frame extraction can be performed on the video to be detected. Among them, the key frame is the key image frame of a shot in the video, which can reflect the main content of a shot. Hereinafter, an example in which the video to be detected includes 100 frame images will be used for introduction.

[0339] Specifically, please refer to Figure 17 , Figure 17 which is a schematic diagram of key frame extraction in the embodiments of the present application. As shown in the figure, the video to be detected is represented as F = {f1, f2,..., f n}, and by performing key frame extraction on the video to be detected, an image frame sequence composed of key frames and an image frame sequence composed of non-key frames are obtained.

[0340] The image frame sequence composed of key frames can be used as the visual data to be detected, and the object detection results of each image frame to be detected are generated in combination with the prompt text. Based on this, the image of the key frame is used as the template image frame, and the image of the non-key frame is used as the search area image frame to construct each pair of images to be matched. For example, the first image frame in the video to be detected is a key frame and can be used as the template image frame, and the second image frame in the video to be detected is a non-key frame and can be used as the search area image frame, and a pair of images to be matched is constructed with the first image frame. The third image frame in the video to be detected is also a non-key frame and can be used as the search area image frame, and another pair of images to be matched is constructed with the first image frame.

[0341] Using the template matching algorithm to detect each pair of images to be matched respectively. Finally, the object detection results of the key frame images and the object detection results of the non-key frame images are merged to obtain the continuous positioning results of the entire video to be detected.

[0342] Secondly, in the embodiments of the present application, a method for object detection based on key-frame extraction is provided. Through the above method, since the key frames concentrate the main information and significant changes in the video to be detected, using key frames for object detection can greatly reduce the amount of data to be processed and the computational cost, thereby improving the running efficiency of the algorithm.

[0343] Optionally, based on one or more corresponding embodiments above, in another optional embodiment provided by the embodiments of the present application, according to the object detection results of each image frame to be detected in the visual data to be detected, each image frame in the video to be detected that has not been frame-extracted is detected to obtain the object detection results of each image frame, which specifically includes: Figure 6 Construct at least one pair of image frames to be matched, where each pair of image frames to be matched includes a template image frame and a search area image frame. The template image frame is an image frame to be detected for which the object detection result has been obtained, and the search area image frame is an image frame in the video to be detected associated with the template image frame;

[0344] For each pair of image frames to be matched, the search area image frame is detected according to the template image frame to obtain the object detection result of the search area image frame;

[0345] The object detection result of each search area image frame in each pair of image frames to be matched is used as the object detection result of each image frame in the video to be detected that has not been frame-extracted.

[0346] In one or more embodiments, a method for object detection based on template matching is introduced. As can be seen from the foregoing embodiments, the image to be matched is used as the template image frame, and the image for matching with the template image frame as the benchmark is the search area image frame. Thus, a template image frame and the search area image frame associated with the template image frame are used as a pair of image frames to be matched. Based on this, for each pair of image frames to be matched, the object is continuously located in the search area image frame through the template matching algorithm. Thus, the object detection result of each search area image frame is obtained.

[0347]

[0348] ​It should be noted that the template matching algorithm used in this application can be a robust object modeling framework for visual tracking (ROMTrack), or a siamese fully convolutional network (SiamFC), or a siamese region proposal network (SiamRPN), etc., which is not limited here.

[0349] Again, in the embodiments of this application, a method for object detection based on template matching is provided. Through the above method, after object detection is performed on some image frames in the video to be detected, the object detection results of other image frames can be quickly predicted based on the template matching algorithm. Thereby, the inference speed of the model is improved, and the object detection efficiency is improved.

[0350] The method provided in this solution and the existing solutions are respectively applied to the downstream task of continuously locating harmful items related to video content. Thereby, the experimental results shown in Table 2 are obtained.

[0351] Table 2

[0352] Solution HOTA DetA AssA FPS TransRMOT Solution 40.16 28.28 57.82 7.7 iKUN Solution 38.36 26.27 56.48 13.4 This Solution 42.46 31.48 58.26 12.6

[0353] It can be seen that compared with the TransRMOT and insertable knowledge unification network (iKUN) solutions, this solution is higher than the existing methods in terms of the higher order tracking association (HOTA) index, detection accuracy (DetA) index, and association accuracy (AssA) index. That is, this solution has better performance in detection and continuous location association, which proves that the multi-modal feature fusion based on wavelet transform and the query vector update method based on text guidance proposed in this solution can fully fuse visual and text features and endow the query vector with the text prior knowledge of the continuous location target. In addition, in terms of inference speed, the frames per second (FPS) index also has a certain advantage, which is similar to the inference speed of iKUN and 1.6 times that of TransRMOT.

[0354] In summary, the overall process of the object detection method will be introduced below in conjunction with the illustrations. Please refer to Figure 18 , Figure 18This is an overall architecture diagram of the object detection method in the embodiments of this application. As shown in the figure, the overall architecture includes a backbone network, a wavelet fusion module, a text-guided initialization module, an encoder, a decoder, and a prediction network. Among them, the Wavelet-Fuse module integrates wavelet transform and attention mechanism to achieve fine-grained fusion of image features and text features at both local and global levels. The Text-Guide Initialization uses the original text features to guide the initialization of the query vector, endowing the initial query vector with prior knowledge of detecting objects that conform to the text description, and enhancing the model's ability to learn objects that conform to the text description. The Encoder and Decoder can directly use the Encoder and Decoder in Deformable DETR.

[0355] In the process of performing object detection on the visual data to be detected in this solution, taking the K feature maps extracted as the feature map F3, feature map F4, and feature Figure 5 as an example, the specific object detection process can be as follows:

[0356] (1) Obtain the visual data to be detected and the prompt text.

[0357] (2) Use the image feature extraction network in the Backbone (for example, Resnet50) to process each image frame to be detected in the visual data to be detected, and obtain three layers of feature maps with different scales corresponding to each image frame to be detected, that is, denoted as feature map F3, feature map F4, and feature map F5. Use the text feature extraction network in the Backbone (for example, RoBETRa) to process the prompt text, and obtain the embedded feature representation of the prompt text, that is, the original text feature T emd .

[0358] (3) Input the feature map F3, feature map F4, feature map F5, and the original text feature T emd into the Wavelet-Fuse module respectively for wavelet transform and hierarchical visual and text feature fusion, and output the text-image fusion feature maps obtained after fusion for each feature map.

[0359] (4) Input the text-image fusion feature maps corresponding to each feature map into the Encoder for further fusion, enhance the expression ability of the features, improve the model's learning ability for the objects described in the text, and output the fused text-image features.

[0360] (5) Randomly initialize N query vectors, that is, obtain N initial query vectors. Then, input the N initial query vectors and the original text feature T emd into the Text-Guide Initialization module, and use the original text feature T emd to perform attention calculation with the N initial query vectors, thereby attaching the prior information of the text description target to each initial query vector.

[0361] (6) Input the fused image-text features output by the Encoder and the N target query vectors containing the prior information of the target text into the Decoder for decoding, and output the object query embedding that implicitly represents the continuous localization result.

[0362] (7) Input the query embedding output by the Decoder into the Predict Head for processing to obtain the final continuous localization result.

[0363] The object detection device in this application will be described in detail below. Please refer to Figure 19 , Figure 19 which is a schematic diagram of an embodiment of the object detection device in an embodiment of this application. The object detection device 190 includes:

[0364] An acquisition module 1901, configured to acquire the visual data to be detected and the corresponding prompt text, where the visual data to be detected includes an image frame sequence composed of at least one image frame to be detected;

[0365] An extraction module 1902, configured to extract features from the image frames to be detected in the visual data to be detected, and obtain K feature maps corresponding to the image frames to be detected, where K is an integer greater than or equal to 1;

[0366] A transformation module 1903, configured to perform wavelet transformation on the K feature maps to obtain at least one set of visual feature map sets corresponding to each feature map, where each set of visual feature map sets includes at least one of a global visual feature map, a vertical visual feature map, a horizontal visual feature map, and a diagonal visual feature map;

[0367] The extraction module 1902 is further configured to extract features from the prompt text to obtain the original text feature;

[0368] A fusion module 1904, configured to fuse the original text feature, the K feature maps, and at least one set of visual feature map sets corresponding to each feature map to obtain a text-image fusion feature map corresponding to each feature map;

[0369] The detection module 1905 is used to generate an object detection result of the image frame to be detected according to the K image-text fusion feature maps corresponding to the K feature maps.

[0370] Optionally, in the above Figure 19 Based on the corresponding embodiment, in another embodiment of the object detection device 190 provided in the embodiment of the present application,

[0371] The extraction module 1902 is specifically used to obtain the image frame to be detected through the initial convolution layer included in the image feature extraction network, perform convolution processing, and obtain an initial feature map corresponding to the image frame to be detected;

[0372] The initial convolutional features are pooled by a feature processing layer included in the image feature extraction network to obtain a pooled feature map, wherein the feature processing layer includes a batch normalization layer, an activation function layer, and a maximum pooling layer;

[0373] Perform feature extraction on the pooled feature map through R residual block layers included in the image feature extraction network to obtain R feature maps, wherein each residual block layer is stacked by multiple bottleneck layers, each bottleneck layer includes sequentially connected convolutional layers, and R is an integer greater than or equal to K;

[0374] Obtain K feature maps with different scales from the R feature maps.

[0375] Optionally, in the above Figure 19 Based on the corresponding embodiment, in another embodiment of the object detection device 190 provided in the embodiment of the present application,

[0376] The extraction module 1902 is specifically used to perform word segmentation processing on the prompt text to obtain at least one word element;

[0377] Encode each word unit in at least one word unit to obtain a word vector corresponding to each word unit;

[0378] Generate a position vector for each word unit according to its position in the prompt text;

[0379] Generate a combined vector for each word unit based on the word vector and position vector corresponding to each word unit;

[0380] Based on the combined vector of each word, the original text features are obtained through the text feature extraction network.

[0381] Optionally, in the above Figure 19 Based on the corresponding embodiment, in another embodiment of the object detection device 190 provided in the embodiment of the present application,

[0382] The transformation module 1903 is specifically configured to perform convolution processing on K feature maps by using a global convolution kernel, a vertical convolution kernel, a horizontal convolution kernel, and a diagonal convolution kernel respectively, and sequentially obtain a global visual feature map, a vertical visual feature map, a horizontal visual feature map, and a diagonal visual feature map corresponding to each feature map;

[0383] Among them, the global visual feature map, the vertical visual feature map, the horizontal visual feature map, and the diagonal visual feature map constitute a set of visual feature map collections.

[0384] Optionally, based on the above Figure 19 corresponding embodiment, in another embodiment of the object detection device 190 provided in the embodiments of the present application,

[0385] The fusion module 1904 is specifically configured to fuse the original text features with the global visual feature map, the vertical visual feature map, the horizontal visual feature map, and the diagonal visual feature map corresponding to the feature map respectively, and sequentially obtain a global fusion feature map, a vertical fusion feature map, a horizontal fusion feature map, and a diagonal fusion feature map corresponding to the feature map;

[0386] Perform inverse wavelet transform on the merged global fusion feature map, vertical fusion feature map, horizontal fusion feature map, and diagonal fusion feature map corresponding to the feature map to obtain a target feature map corresponding to the feature map;

[0387] Fuse the original text features and the feature map to obtain a fusion feature map corresponding to the feature map;

[0388] Merge the fusion feature map corresponding to the feature map and the target feature map corresponding to the feature map to obtain a text-image fusion feature map corresponding to the feature map.

[0389] Optionally, based on the above Figure 19 corresponding embodiment, in another embodiment of the object detection device 190 provided in the embodiments of the present application,

[0390] The transformation module 1903 is specifically configured to perform convolution processing on K feature maps by using a global convolution kernel, a vertical convolution kernel, a horizontal convolution kernel, and a diagonal convolution kernel respectively, and sequentially obtain a first global visual feature map, a first vertical visual feature map, a first horizontal visual feature map, and a first diagonal visual feature map corresponding to each feature map;

[0391] Perform convolution processing on the first global visual feature map corresponding to each feature map by using a global convolution kernel, a vertical convolution kernel, a horizontal convolution kernel, and a diagonal convolution kernel respectively, and sequentially obtain a second global visual feature map, a second vertical visual feature map, a second horizontal visual feature map, and a second diagonal visual feature map corresponding to each feature map;

[0392] Among them, the first global visual feature map, the first vertical visual feature map, the first horizontal visual feature map, and the first diagonal visual feature map form a set of visual feature map sets;

[0393] Moreover, the second global visual feature map, the second vertical visual feature map, the second horizontal visual feature map, and the second diagonal visual feature map form another set of visual feature map sets.

[0394] Optionally, based on the above Figure 19 In another embodiment of the object detection device 190 provided by the embodiment of the present application, on the basis of the corresponding embodiment,

[0395] The fusion module 1904 is specifically configured to fuse the original text features with the first global visual feature map, the first vertical visual feature map, the first horizontal visual feature map, and the first diagonal visual feature map corresponding to the feature map respectively, and sequentially obtain the first global fusion feature map, the first vertical fusion feature map, the first horizontal fusion feature map, and the first diagonal fusion feature map corresponding to the feature map;

[0396] Fuse the original text features with the second global visual feature map, the second vertical visual feature map, the second horizontal visual feature map, and the second diagonal visual feature map corresponding to the feature map respectively, and sequentially obtain the second global fusion feature map, the second vertical fusion feature map, the second horizontal fusion feature map, and the second diagonal fusion feature map corresponding to the feature map;

[0397] Perform inverse wavelet transform on the merged second global fusion feature map, second vertical fusion feature map, second horizontal fusion feature map, and second diagonal fusion feature map corresponding to the feature map to obtain the first target feature map corresponding to the feature map;

[0398] Merge the first target feature map corresponding to the feature map with the first global fusion feature map, the first vertical fusion feature map, the first horizontal fusion feature map, and the first diagonal fusion feature map corresponding to the feature map respectively, and sequentially obtain the first fusion feature map, the second fusion feature map, the third fusion feature map, and the fourth fusion feature map corresponding to the feature map;

[0399] Perform inverse wavelet transform on the merged first fusion feature map, second fusion feature map, third fusion feature map, and fourth fusion feature map corresponding to the feature map to obtain the second target feature map corresponding to the feature map;

[0400] Fuse the original text features and the feature map to obtain the fusion feature map corresponding to the feature map;

[0401] Merge the fusion feature map corresponding to the feature map and the second target feature map corresponding to the feature map to obtain the text-image fusion feature map corresponding to the feature map.

[0402] Optionally, in the above Figure 19Based on the corresponding embodiment, in another embodiment of the object detection device 190 provided by the embodiments of the present application,

[0403] The transformation module 1903 is specifically configured to perform convolution processing on K feature maps by using a global convolution kernel to obtain a first global visual feature map corresponding to each feature map;

[0404] Perform convolution processing on the first global visual feature map corresponding to each feature map by using a global convolution kernel, a vertical convolution kernel, a horizontal convolution kernel, and a diagonal convolution kernel respectively, and sequentially obtain a second global visual feature map, a second vertical visual feature map, a second horizontal visual feature map, and a second diagonal visual feature map corresponding to each feature map;

[0405] Wherein, the first global visual feature maps constitute a set of visual feature map sets;

[0406] Moreover, the second global visual feature map, the second vertical visual feature map, the second horizontal visual feature map, and the second diagonal visual feature map constitute another set of visual feature map sets.

[0407] Optionally, based on the above Figure 19 Based on the corresponding embodiment, in another embodiment of the object detection device 190 provided by the embodiments of the present application,

[0408] The fusion module 1904 is specifically configured to fuse the original text feature and the first global visual feature map corresponding to the feature map to obtain a first global fusion feature map corresponding to the feature map;

[0409] Fuse the original text feature with the second global visual feature map, the second vertical visual feature map, the second horizontal visual feature map, and the second diagonal visual feature map corresponding to the feature map respectively, and sequentially obtain a second global fusion feature map, a second vertical fusion feature map, a second horizontal fusion feature map, and a second diagonal fusion feature map corresponding to the feature map;

[0410] Perform inverse wavelet transform on the merged second global fusion feature map, second vertical fusion feature map, second horizontal fusion feature map, and second diagonal fusion feature map corresponding to the feature map to obtain a first target feature map corresponding to the feature map;

[0411] Perform inverse wavelet transform on the merged first target feature map corresponding to the feature map and the first global fusion feature map corresponding to the feature map to obtain a second target feature map corresponding to the feature map;

[0412] Fuse the original text feature and the feature map to obtain a fusion feature map corresponding to the feature map;

[0413] Merge the fusion feature map corresponding to the feature map and the second target feature map corresponding to the feature map to obtain a text-image fusion feature map corresponding to the feature map.

[0414] Optionally, based on the above Figure 19 corresponding embodiment, in another embodiment of the object detection device 190 provided by the embodiments of the present application,

[0415] The fusion module 1904 is specifically configured to perform feature mapping on the original text feature and the feature map respectively to obtain a first text feature corresponding to the original text feature and a first feature map corresponding to the feature map;

[0416] Perform self-attention calculation on the first text feature and the first feature map respectively to obtain a second text feature corresponding to the first text feature and a second feature map corresponding to the first feature map;

[0417] After performing cross-attention calculation on the second text feature and the second feature map, perform a skip connection with the first feature map to obtain a fused feature map corresponding to the feature map.

[0418] Optionally, based on the above Figure 19 corresponding embodiment, in another embodiment of the object detection device 190 provided by the embodiments of the present application,

[0419] The detection module 1905 is specifically configured to obtain N initial query vectors, where N is an integer greater than or equal to 1;

[0420] Generate visual text feature vectors according to the text-image fusion feature maps corresponding to each of the K feature maps;

[0421] Determine the object detection result of the image frame to be detected according to the N initial query vectors and the visual text feature vectors.

[0422] Optionally, based on the above Figure 19 corresponding embodiment, in another embodiment of the object detection device 190 provided by the embodiments of the present application,

[0423] The detection module 1905 is specifically configured to process the N initial query vectors according to the original text feature to obtain N target query vectors, where N is an integer greater than or equal to 1;

[0424] Generate visual text feature vectors according to the text-image fusion feature maps corresponding to each of the K feature maps;

[0425] Determine the object detection result of the image frame to be detected according to the N target query vectors and the visual text feature vectors.

[0426] Optionally, based on the above Figure 19 corresponding embodiment, in another embodiment of the object detection device 190 provided by the embodiments of the present application,

[0427] The detection module 1905 is specifically configured to perform self-attention calculation on N initial query vectors to obtain N first query vectors;

[0428] Obtain M word-level features and sentence-level features according to the original text features, where M is an integer greater than or equal to 1;

[0429] Perform cross-attention calculation on the M word-level features and the N first query vectors to obtain N second query vectors;

[0430] Interact the sentence-level feature and the N second query vectors to obtain N third query vectors;

[0431] Merge the N third query vectors and the N initial query vectors to obtain N target query vectors.

[0432] Optionally, based on the above Figure 19 In another embodiment of the object detection device 190 provided in the embodiments of the present application, on the basis of the corresponding embodiment,

[0433] The acquisition module 1901 is specifically configured to acquire a video to be detected, where the video to be detected includes a sequence of image frames composed of at least two image frames;

[0434] Perform equally spaced frame extraction on the video to be detected to obtain visual data to be detected;

[0435] The detection module 1905 is further configured to detect each image frame that is not frame-extracted in the video to be detected according to the object detection result of each image frame to be detected in the visual data to be detected, and obtain the object detection result of each image frame.

[0436] Optionally, based on the above Figure 19 In another embodiment of the object detection device 190 provided in the embodiments of the present application, on the basis of the corresponding embodiment,

[0437] The acquisition module 1901 is specifically configured to acquire a video to be detected, where the video to be detected includes a sequence of image frames composed of at least two image frames;

[0438] Perform key frame recognition on the video to be detected to obtain visual data to be detected, where the image frames to be detected in the visual data to be detected are the key frames of the video to be detected;

[0439] The detection module 1905 is further configured to detect each image frame that is not frame-extracted in the video to be detected according to the object detection result of each image frame to be detected in the visual data to be detected, and obtain the object detection result of each image frame.

[0440] Optionally, based on the above Figure 19Based on the corresponding embodiments, in another embodiment of the object detection device 190 provided by the embodiments of the present application,

[0441] The detection module 1905 is specifically configured to construct at least one pair of image frames to be matched, where each pair of image frames to be matched includes a template image frame and a search area image frame. The template image frame is a to-be-detected image frame for which an object detection result has been obtained, and the search area image frame is an image frame in the to-be-detected video associated with the template image frame;

[0442] For each pair of image frames to be matched, the search area image frame is detected according to the template image frame to obtain the object detection result of the search area image frame;

[0443] The object detection result of each search area image frame in each pair of image frames to be matched is used as the object detection result of each image frame in the to-be-detected video that has not been framed.

[0444] Figure 20 FIG. is a schematic structural diagram of a computer device provided by the embodiments of the present application. The computer device 2000 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 2022 (for example, one or more processors) and a memory 2032, and one or more storage media 2030 (for example, one or more mass storage devices) storing application programs 2042 or data 2044. Among them, the memory 2032 and the storage media 2030 may be transient storage or persistent storage. The program stored in the storage media 2030 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the computer device. Further, the central processor 2022 may be configured to communicate with the storage media 2030 and execute a series of instruction operations in the storage media 2030 on the computer device 2000.

[0445] The computer device 2000 may further include one or more power supplies 2026, one or more wired or wireless network interfaces 2050, one or more input / output interfaces 2058, and / or one or more operating systems 2041, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM and so on.

[0446] The steps performed by the computer device in the above embodiments may be based on the Figure 20 shown computer device structure.

[0447] In an embodiment of the present application, a computer-readable storage medium is further provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the methods described in the foregoing embodiments are implemented.

[0448] In an embodiment of the present application, a computer program product is further provided, including a computer program. When the computer program is executed by a processor, the steps of the methods described in the foregoing embodiments are implemented.

[0449] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0450] In an embodiment of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the function of the module or unit.

[0451] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be indirect couplings or communication connections through some interfaces, devices, or units, and can be in electrical, mechanical, or other forms.

[0452] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0453] In addition, the functional units in each embodiment of the present application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0454] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a server, a terminal device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media that can store computer programs, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0455] As described above, the above embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of various embodiments of this application.

Claims

1. A method for object detection, characterized in that, Including: Obtain the visual data to be detected and the corresponding prompt text, where the visual data to be detected includes an image frame sequence composed of at least one image frame to be detected; Extract features from the image frames to be detected in the visual data to be detected, and obtain K feature maps corresponding to the image frames to be detected, where K is an integer greater than or equal to 1; Perform wavelet transform on the K feature maps to obtain at least one set of visual feature map sets corresponding to each feature map, where each set of visual feature map sets includes at least one of a global visual feature map, a vertical visual feature map, a horizontal visual feature map, and a diagonal visual feature map; Extract features from the prompt text to obtain the original text features; Fuse the original text features, the K feature maps, and at least one set of visual feature map sets corresponding to each feature map to obtain the text-image fusion feature map corresponding to each feature map; Generate the object detection result of the image frame to be detected according to the K text-image fusion feature maps corresponding to the K feature maps.

2. The method according to claim 1, characterized in that The extracting features from the image frames to be detected in the visual data to be detected and obtaining K feature maps corresponding to the image frames to be detected includes: Through the initial convolutional layer included in the image feature extraction network, perform convolutional processing on the image frame to be detected to obtain the initial feature map corresponding to the image frame to be detected; Through the feature processing layer included in the image feature extraction network, perform pooling processing on the initial convolutional feature to obtain the pooled feature map, where the feature processing layer includes a batch normalization layer, an activation function layer, and a max pooling layer; Through the R residual block layers included in the image feature extraction network, extract features from the pooled feature map to obtain R feature maps, where each residual block layer is stacked by multiple bottleneck layers, and each bottleneck layer includes a convolutional layer connected in sequence, and R is an integer greater than or equal to K; Obtain the K feature maps with different scales from the R feature maps.

3. The method according to claim 1, wherein The extracting features from the prompt text to obtain the original text features includes: Perform word segmentation processing on the prompt text to obtain at least one token; Encode each token in the at least one token to obtain the word vector corresponding to each token; Generate the position vector of each token according to the position of each token in the prompt text; Generate the combined vector of each token according to the word vector and the position vector corresponding to each token; Based on the combined vector of each token, obtain the original text features through the text feature extraction network.

4. The method according to any one of claims 1 to 3, characterized in that, The performing wavelet transform on the K feature maps to obtain at least one set of visual feature map sets corresponding to each feature map includes: Respectively use a global convolutional kernel, a vertical convolutional kernel, a horizontal convolutional kernel, and a diagonal convolutional kernel to perform convolutional processing on the K feature maps, and sequentially obtain the global visual feature map, the vertical visual feature map, the horizontal visual feature map, and the diagonal visual feature map corresponding to each feature map; Among them, the global visual feature map, the vertical visual feature map, the horizontal visual feature map, and the diagonal visual feature map constitute a set of visual feature map sets.

5. The method according to claim 4, wherein Fusing the original text features, the K feature maps, and at least one set of visual feature map sets corresponding to each feature map to obtain the text-image fusion feature map corresponding to each feature map includes: Fusing the original text features with the global visual feature map, vertical visual feature map, horizontal visual feature map, and diagonal visual feature map corresponding to the feature map respectively to sequentially obtain the global fusion feature map, vertical fusion feature map, horizontal fusion feature map, and diagonal fusion feature map corresponding to the feature map; Performing inverse wavelet transform on the merged global fusion feature map, vertical fusion feature map, horizontal fusion feature map, and diagonal fusion feature map corresponding to the feature map to obtain the target feature map corresponding to the feature map; Fusing the original text features and the feature map to obtain the fusion feature map corresponding to the feature map; Merging the fusion feature map corresponding to the feature map and the target feature map corresponding to the feature map to obtain the text-image fusion feature map corresponding to the feature map.

6. The method according to any one of claims 1 to 3, characterized in that, Performing wavelet transform on the K feature maps to obtain at least one set of visual feature map sets corresponding to each feature map includes: Performing convolution processing on the K feature maps using a global convolution kernel, a vertical convolution kernel, a horizontal convolution kernel, and a diagonal convolution kernel respectively to sequentially obtain the first global visual feature map, first vertical visual feature map, first horizontal visual feature map, and first diagonal visual feature map corresponding to each feature map; Performing convolution processing on the first global visual feature map corresponding to each feature map using the global convolution kernel, the vertical convolution kernel, the horizontal convolution kernel, and the diagonal convolution kernel respectively to sequentially obtain the second global visual feature map, second vertical visual feature map, second horizontal visual feature map, and second diagonal visual feature map corresponding to each feature map; Among them, the first global visual feature map, the first vertical visual feature map, the first horizontal visual feature map, and the first diagonal visual feature map constitute a set of visual feature map sets; Moreover, the second global visual feature map, the second vertical visual feature map, the second horizontal visual feature map, and the second diagonal visual feature map constitute another set of visual feature map sets.

7. The method according to claim 6, wherein Fusing the original text features, the K feature maps, and at least one set of visual feature map sets corresponding to each feature map to obtain the text-image fusion feature map corresponding to each feature map includes: Fusing the original text features with the first global visual feature map, first vertical visual feature map, first horizontal visual feature map, and first diagonal visual feature map corresponding to the feature map respectively to sequentially obtain the first global fusion feature map, first vertical fusion feature map, first horizontal fusion feature map, and first diagonal fusion feature map corresponding to the feature map; Fuse the original text features with the corresponding second global visual feature map, second vertical visual feature map, second horizontal visual feature map, and second diagonal visual feature map of the feature map respectively to obtain the corresponding second global fusion feature map, second vertical fusion feature map, second horizontal fusion feature map, and second diagonal fusion feature map of the feature map in sequence; Perform inverse wavelet transform on the merged second global fusion feature map, second vertical fusion feature map, second horizontal fusion feature map, and second diagonal fusion feature map corresponding to the feature map to obtain the first target feature map corresponding to the feature map; Merge the first target feature map corresponding to the feature map with the corresponding first global fusion feature map, first vertical fusion feature map, first horizontal fusion feature map, and first diagonal fusion feature map of the feature map respectively to obtain the corresponding first fusion feature map, second fusion feature map, third fusion feature map, and fourth fusion feature map of the feature map in sequence; Perform inverse wavelet transform on the merged first fusion feature map, second fusion feature map, third fusion feature map, and fourth fusion feature map corresponding to the feature map to obtain the second target feature map corresponding to the feature map; Fuse the original text features and the feature map to obtain the fusion feature map corresponding to the feature map; Merge the fusion feature map corresponding to the feature map and the second target feature map corresponding to the feature map to obtain the text-image fusion feature map corresponding to the feature map.

8. The method according to any one of claims 1 to 3, characterized in that The wavelet transform of the K feature maps to obtain at least one set of visual feature map sets corresponding to each feature map includes: Perform convolution processing on the K feature maps using a global convolutional kernel to obtain the first global visual feature map corresponding to each feature map; Perform convolution processing on the first global visual feature map corresponding to each feature map using the global convolutional kernel, vertical convolutional kernel, horizontal convolutional kernel, and diagonal convolutional kernel respectively to obtain the corresponding second global visual feature map, second vertical visual feature map, second horizontal visual feature map, and second diagonal visual feature map of each feature map in sequence; Among them, the first global visual feature map constitutes a set of visual feature map sets; Moreover, the second global visual feature map, the second vertical visual feature map, the second horizontal visual feature map, and the second diagonal visual feature map constitute another set of visual feature map sets.

9. The method according to claim 8, wherein The fusion of the original text features, the K feature maps, and at least one set of visual feature map sets corresponding to each feature map to obtain the text-image fusion feature map corresponding to each feature map includes: Fuse the original text features and the first global visual feature map corresponding to the feature map to obtain the first global fusion feature map corresponding to the feature map; Fuse the original text features with the corresponding second global visual feature map, second vertical visual feature map, second horizontal visual feature map, and second diagonal visual feature map of the feature map respectively to obtain the corresponding second global fusion feature map, second vertical fusion feature map, second horizontal fusion feature map, and second diagonal fusion feature map of the feature map in sequence; Inverse wavelet transform is performed after merging the second global fusion feature map, the second vertical fusion feature map, the second horizontal fusion feature map, and the second diagonal fusion feature map corresponding to the feature map to obtain the first target feature map corresponding to the feature map; Inverse wavelet transform is performed after merging the first target feature map corresponding to the feature map and the first global fusion feature map corresponding to the feature map to obtain the second target feature map corresponding to the feature map; The original text feature and the feature map are fused to obtain the fusion feature map corresponding to the feature map; The fusion feature map corresponding to the feature map and the second target feature map corresponding to the feature map are merged to obtain the text-image fusion feature map corresponding to the feature map.

10. The method according to claim 5, 7 or 9, characterized in that The fusing the original text feature and the feature map to obtain the fusion feature map corresponding to the feature map includes: Feature mapping is respectively performed on the original text feature and the feature map to obtain the first text feature corresponding to the original text feature and the first feature map corresponding to the feature map; Self-attention calculation is respectively performed on the first text feature and the first feature map to obtain the second text feature corresponding to the first text feature and the second feature map corresponding to the first feature map; After cross-attention calculation is performed on the second text feature and the second feature map, skip connection is performed with the first feature map to obtain the fusion feature map corresponding to the feature map.

11. The method according to any one of claims 1 to 10, characterized in that, The generating the object detection result of the to-be-detected image frame according to the K text-image fusion feature maps corresponding to the K feature maps includes: Obtain N initial query vectors, where N is an integer greater than or equal to 1; Generate visual text feature vectors according to the text-image fusion feature map corresponding to each of the K feature maps; Determine the object detection result of the to-be-detected image frame according to the N initial query vectors and the visual text feature vectors.

12. The method according to any one of claims 1 to 10, characterized in that The generating the object detection result of the to-be-detected image frame according to the K text-image fusion feature maps corresponding to the K feature maps includes: Process the N initial query vectors according to the original text feature to obtain N target query vectors, where N is an integer greater than or equal to 1; Generate visual text feature vectors according to the text-image fusion feature map corresponding to each of the K feature maps; Determine the object detection result of the to-be-detected image frame according to the N target query vectors and the visual text feature vectors.

13. The method according to claim 12, characterized in that, The processing the N initial query vectors according to the original text feature to obtain N target query vectors includes: Perform self-attention calculation on the N initial query vectors to obtain N first query vectors; Obtain M word-level features and sentence-level features according to the original text feature, where M is an integer greater than or equal to 1; Perform cross-attention calculation on the M word-level features and the N first query vectors to obtain N second query vectors; Interact the sentence-level feature and the N second query vectors to obtain N third query vectors; Combine the N third query vectors and the N initial query vectors to obtain the N target query vectors.

14. The method according to any one of claims 1 to 13, characterized in that, The obtaining of the visual data to be detected includes: Obtain a video to be detected, where the video to be detected includes an image frame sequence composed of at least two image frames; Perform equally-spaced frame extraction on the video to be detected to obtain the visual data to be detected; The method further includes: According to the object detection results of each image frame to be detected in the visual data to be detected, detect each image frame in the video to be detected that has not been frame-extracted, and obtain the object detection results of each image frame.

15. The method according to any one of claims 1 to 13, characterized in that, The obtaining of the visual data to be detected includes: Obtain a video to be detected, where the video to be detected includes an image frame sequence composed of at least two image frames; Perform key frame recognition on the video to be detected to obtain the visual data to be detected, where the image frames to be detected in the visual data to be detected are the key frames of the video to be detected; The method further includes: According to the object detection results of each image frame to be detected in the visual data to be detected, detect each image frame in the video to be detected that has not been frame-extracted, and obtain the object detection results of each image frame.

16. The method according to claim 14 or 15, characterized in that The step of, according to the object detection results of each image frame to be detected in the visual data to be detected, detecting each image frame in the video to be detected that has not been frame-extracted, and obtaining the object detection results of each image frame, includes: Construct at least one pair of image frames to be matched, where each pair of image frames to be matched includes a template image frame and a search area image frame. The template image frame is an image frame to be detected for which an object detection result has been obtained, and the search area image frame is an image frame in the video to be detected associated with the template image frame; For each pair of image frames to be matched, detect the search area image frame according to the template image frame to obtain the object detection result of the search area image frame; Use the object detection results of each search area image frame in each pair of image frames to be matched as the object detection results of each image frame in the video to be detected that has not been frame-extracted.

17. An object detection device, characterized in that, It includes: An acquisition module, configured to acquire visual data to be detected and corresponding prompt text, where the visual data to be detected includes an image frame sequence composed of at least one image frame to be detected; An extraction module, configured to extract features from the image frames to be detected in the visual data to be detected to obtain K feature maps corresponding to the image frames to be detected, where K is an integer greater than or equal to 1; A transformation module, configured to perform wavelet transformation on the K feature maps to obtain at least one set of visual feature map sets corresponding to each feature map, where each set of visual feature map sets includes at least one of a global visual feature map, a vertical visual feature map, a horizontal visual feature map, and a diagonal visual feature map; The extraction module is further configured to extract features from the prompt text to obtain original text features; A fusion module, configured to fuse the original text features, the K feature maps, and at least one set of visual feature map sets corresponding to each feature map, to obtain a text-image fusion feature map corresponding to each feature map; A detection module, configured to generate an object detection result of the to-be-detected image frame according to the K text-image fusion feature maps corresponding to the K feature maps.

18. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the object detection method according to any one of claims 1 to 16 are implemented.

19. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the steps of the object detection method according to any one of claims 1 to 16 are implemented.

20. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the object detection method according to any one of claims 1 to 16 are implemented.