Method and device for detecting image based on text and electronic equipment

By using the pre-trained target model, combining the target text and image features, object detection of multiple application scenarios is realized, and the problems of high detection cost and low efficiency in the prior art are solved, and the effect of reducing costs and improving efficiency is achieved.

CN120125916AActive Publication Date: 2025-06-10CHINA TOWER CO LTD +1

Patent Information

Application Number
CN202510595768.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-06-10
Estimated Expiration
2045-05-09

AI Technical Summary

Technical Problem

In the prior art, when detecting objects in different application scenarios, it is necessary to model individually for each application scenario, resulting in high detection cost and low detection efficiency.

Method used

By acquiring the target text and the target image, a pre-trained target model is used. The model includes an encoder, a query module and a decoder. The encoder is used for feature extraction, the query module is used to determine the similarity between text and image features, and the decoder interacts with text and image features through a cross attention mechanism, and then determines the detection result.

Benefits of technology

It realizes object detection for multiple application scenarios without the need for separate modeling, reducing detection costs and improving detection efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125916A_ABST
    Figure CN120125916A_ABST
Patent Text Reader

Abstract

The invention discloses a method and device for detecting an image based on a text and electronic equipment. The method comprises the following steps: acquiring a target text and a target image; the target text and the target image are input to a target model, the target model comprises an encoder, a query module and a decoder, the encoder is at least used for carrying out feature extraction on the target image and the target text, the query module is at least used for determining the similarity between the extracted text features and the extracted image features, and the decoder is used for decoding the target text. The decoder is at least used for interacting the text features and the image features based on a cross attention mechanism; and determining a detection result through the target model. According to the object detection method and device, the technical problems of high detection cost and low detection efficiency caused by independent modeling for each application scene when object detection is carried out on different application scenes in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image detection, and in particular, to a method, an apparatus, and an electronic device for detecting an image based on text. Background Art

[0002] With the progress of technology, object detection technology plays an increasingly important role in many fields such as land monitoring, water conservancy supervision, and environmental protection monitoring. For the detection requirements of specific application scenarios, the traditional approach is to separately develop and deploy dedicated detection models for each application scenario. For example, in the water conservancy supervision application scenario, a set of models dedicated to identifying specific objects such as floods and dam cracks is required; in the environmental protection monitoring application scenario, another set of models is needed to monitor forest fires, water pollution, etc.

[0003] However, this customized model development strategy for a single scenario, although it can provide a certain detection accuracy for specific objects, its limitations are becoming increasingly prominent. First, it is difficult to share the dedicated models developed for different application scenarios, and the generalization ability of the established dedicated models is insufficient. When new tasks are added, the existing models cannot be directly reused, and model design and training must start from scratch, resulting in a waste of repeated investment of resources and R & D time. Second, the development of dedicated models for each application scenario separately not only requires a large amount of time and computing resources in the model training stage, but also faces high cost pressure in the model maintenance, update, and deployment stages. Especially when facing the rapid change of application scenarios and the continuous update of detection requirements, it will further cause technical problems such as high detection cost and low detection efficiency.

[0004] For the above problems, no effective solution has been proposed yet. Summary of the Invention

[0005] The present application provides a method, an apparatus, and an electronic device for detecting an image based on text, so as to at least solve the technical problems of high detection cost and low detection efficiency caused by separately modeling for each application scenario in the prior art when performing object detection on different application scenarios.

[0006] According to one aspect of the present application, a method for detecting an image based on text is provided, including: obtaining a target text and a target image, where the target text is used to describe an object to be detected in the target image through a preset language, and the target image is a scene image corresponding to any one of E application scenarios, and E is a positive integer; inputting the target text and the target image into a target model, where the target model includes an encoder, a query module, and a decoder, the encoder is at least used to extract features from the target image and the target text, the query module is at least used to determine the similarity between the extracted text features and image features, and the decoder is at least used to interact the text features and image features based on a cross-attention mechanism; determining a detection result through the target model, where the detection result is used to represent the position of the object described by the target text in the target image.

[0007] Optionally, obtaining the target text and the target image includes: periodically collecting scene images corresponding to E application scenarios through E image acquisition devices to obtain an image dataset; using any image in the image dataset as the target image; performing data cleaning on the original text corresponding to the target image to obtain the target text, where the data cleaning is used to remove stop words in the original text, and the original text is used to describe objects that are likely to appear in the application scenario corresponding to the target image through a preset language.

[0008] Optionally, determining the detection result through the target model includes: determining semantic information and context information included in the target text through the encoder in the target model, and using the semantic information and context information included in the target text as the first text features corresponding to the target text; extracting features from the target image through the encoder in the target model based on a multi-scale feature fusion strategy to obtain M first image features corresponding to the target image, where M is a positive integer; determining the detection result based on the first text features and the M first image features.

[0009] Optionally, determining the detection result based on the first text features and the M first image features includes: inputting the first text features and the M first image features into the query module in the target model; enhancing the features of the first text features through the query module to obtain second text features, and enhancing the features of the M first image features to obtain M second image features; obtaining a score matrix corresponding to the M second image features, where the scores in the score matrix are used to represent the similarity between each second image feature and the second text feature; sorting the M second image features according to the scores in the score matrix to obtain an image feature sequence; using the top P second image features in the image feature sequence as P indexes, where the P indexes are used to determine the position of the object described by the target text in the target image, and P is a positive integer less than or equal to M; determining the detection result based on the second text features and the P indexes.

[0010] Optionally, determining a detection result based on the second text feature and P indexes, including: inputting the second text feature and the P indexes into a decoder in a target model; updating the second text feature through a text cross-attention layer in the decoder to obtain a third text feature, where the text cross-attention layer is used to enhance the features of the second text feature; updating the P indexes through an image cross-attention layer in the decoder to obtain P third image features, where the image cross-attention layer is used to enhance the features of the P indexes; determining the detection result according to the similarity between each of the P third image features and the third text feature.

[0011] Optionally, the target model is trained through the following steps: preprocessing X open-source images obtained from a network to obtain Y open-source training images, where both X and Y are positive integers, and Y is greater than X, and the preprocessing is at least used to perform random flipping, random cropping, random scaling, and normalization operations on the open-source images; using the Y open-source training images, Z bounding boxes, and W object types as a first data set, where both Z and W are positive integers, and the W object types are described by texts corresponding to a preset language; iteratively training an initial model based on the first data set through a linear learning rate warm-up strategy and a learning rate jump strategy to obtain a first model; determining the target model through the first model.

[0012] Optionally, determining the target model through the first model, including: preprocessing E historical scenario images corresponding to E application scenarios to obtain U historical training images, where U is a positive integer and U is greater than E; using the U historical training images and the object types that appear in the historical training images as a second data set; iteratively training the first model based on the second data set through a progressive learning rate decay strategy to obtain the target model.

[0013] According to another aspect of the present application, there is also provided a device for detecting an image based on text, including: an acquisition unit, configured to acquire a target text and a target image, where the target text is used to describe an object to be detected in the target image through a preset language, and the target image is a scenario image corresponding to any one of E application scenarios, and E is a positive integer; an input unit, configured to input the target text and the target image into a target model, where the target model includes an encoder, a query module, and a decoder, the encoder is at least configured to extract features from the target image and the target text, the query module is at least configured to determine the similarity between the extracted text features and image features, and the decoder is at least configured to interact the text features and the image features based on a cross-attention mechanism; a determination unit, configured to determine a detection result through the target model, where the detection result is used to represent the position of the object described by the target text in the target image.

[0014] According to another aspect of the present application, there is also provided a computer program product, in which a computer program is stored. When the computer program runs, it controls the computer program product to execute the method for detecting an image based on text as described in any one of the above.

[0015] According to another aspect of the present application, there is also provided an electronic device. The electronic device includes one or more processors and a memory. The memory is used to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the method for detecting an image based on text as described in any one of the above.

[0016] In the present application, first, a target text and a target image are obtained. The target text is used to describe the object to be detected in the target image through a preset language. The target image is a scene image corresponding to any one of E application scenarios, where E is a positive integer. Then, the present application inputs the target text and the target image into a target model. The target model includes an encoder, a query module, and a decoder. The encoder is at least used to extract features from the target image and the target text. The query module is at least used to determine the similarity between the extracted text features and image features. The decoder is at least used to interact the text features and the image features based on a cross-attention mechanism. Finally, the present application determines a detection result through the target model. The detection result is used to represent the position of the object described by the target text in the target image.

[0017] As can be seen from the above, the present application pre-trains a target model. The target model includes an encoder, a query module, and a decoder. The decoder can interact the text features and the image features based on a cross-attention mechanism, achieving the purpose of aligning the text features and the image features, thereby overcoming the influence of the image complexity from different scenes. Therefore, the target model trained by the present application can perform object detection on the scene image corresponding to any one of E application scenarios, so that it is not necessary to perform separate modeling for different application scenarios, and thus the purpose of reducing the object detection cost and improving the object detection efficiency is achieved.

[0018] Thus, the present application realizes the purpose of avoiding separate modeling for different application scenarios from which the target image comes by the way of performing object detection on the target image based on the target text by the target model, thereby achieving the technical effects of reducing the object detection cost and improving the object detection efficiency, and further solving the technical problems of high detection cost and low detection efficiency caused by the need for separate modeling for each application scenario when performing object detection on different application scenarios in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The accompanying drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0020] Figure 1 It is a schematic diagram of a dedicated model corresponding to an optional single application scenario;

[0021] Figure 2 It is a flowchart of an optional open-set object detection method;

[0022] Figure 3 It is a flowchart of an optional method for detecting an image based on text according to an embodiment of the present application;

[0023] Figure 4 It is a structural diagram of an optional detection system according to an embodiment of the present application;

[0024] Figure 5 It is a flowchart of another optional method for detecting an image based on text according to an embodiment of the present application;

[0025] Figure 6 It is a working flowchart of an optional target model according to an embodiment of the present application;

[0026] Figure 7 It is a schematic diagram of an optional device for detecting an image based on text according to an embodiment of the present application;

[0027] Figure 8 It is a schematic diagram of an optional electronic device according to an embodiment of the present application. Detailed implementation manners

[0028] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0029] It should be noted that the terms "first", "second", etc. in the description, claims and the above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products or devices.

[0030] It should also be noted that the relevant information (including but not limited to the information for display and analysis) and data (including but not limited to image data, text data, and scene data) involved in this application are all information and data authorized by the user or fully authorized by all parties. For example, an interface is set between this system and relevant users or institutions. Before obtaining relevant information, it is necessary to send a request for acquisition to the aforementioned users or institutions through the interface, and after receiving the consent information feedback from the aforementioned users or institutions, obtain the relevant information.

[0031] In addition, the processes of collecting, storing, using, processing, transmitting, providing, disclosing, and applying the relevant information and relevant data involved in this application all comply with the relevant laws, regulations and standards of the relevant regions, and necessary confidentiality measures are taken, and they do not violate public order and good customs. In addition, this application provides a corresponding operation entry for users to choose to agree to authorize or refuse to authorize. If the user chooses to refuse to authorize, it will enter the corresponding expert decision-making process.

[0032] In an alternative embodiment, a dedicated model for detecting a single application scenario in engineering construction, smoke recognition, or personnel swimming is provided. Figure 1 It is a schematic diagram of a dedicated model corresponding to an alternative single application scenario. The training method of the dedicated model includes:

[0033] (1) Data collection and collation: Collect data in a specific actual scenario and perform preprocessing.

[0034] (2) Model training: According to the specific characteristics of the specific scenario data, design a training process, modify the model structure and parameters, and complete model training.

[0035] (3) Model deployment: Deploy the models trained for different application scenarios separately.

[0036] In an alternative embodiment, an open-set object detection method is provided.Figure 2 is a flowchart of an optional open-set object detection method, as Figure 2 shown, the method includes:

[0037] (1) Data preprocessing: Unify the pre-training of object detection and phrase localization tasks so that the data of the two tasks can complement each other, thereby jointly improving the zero-shot object detection ability of the model. In the data preprocessing stage, unify the datasets of these two tasks into the same form of text input and image input. Among them, the form of text input is the name of the object to be recognized separated by full stops, and the image input is the image for object detection or phrase localization.

[0038] (2) Encoding different modality data: Adopt a two-tower structure, and use a pre-trained text encoder and image encoder to encode the input data respectively to obtain encoded text features and visual features.

[0039] (3) Cross-modal feature fusion and enhancement: Interact the text features and visual features through a cross-attention mechanism to achieve deep fusion of multiple modalities.

[0040] (4) Feature similarity calculation: Calculate the similarity between the image region features and the text prompt features, match the regions with high similarity and the corresponding text to obtain the detection results.

[0041] However, the deficiencies of the technical solutions corresponding to the above two embodiments are as follows:

[0042] (1) The existing research and development solutions for dedicated models for multi-scenario detection tasks require separate research and development of models for each task, which results in high research and development costs, low efficiency, and difficulty in meeting the requirements of complex multi-scenario tasks.

[0043] (2) Most of the existing open-set object detection models only support English text input. When applied in a Chinese environment, additional translation steps are required, which not only increases the time overhead but also affects the detection accuracy due to translation errors.

[0044] According to the embodiments of the present application, an embodiment of a method for detecting an image based on text is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0045] The present application provides a system for detecting an image based on text (abbreviated as detection system) for executing the method for detecting an image based on text in the present application. Figure 3 is a flowchart of an optional method for detecting an image based on text according to the embodiments of the present application, asFigure 3 As shown in Figure 3 , the method includes the following steps:

[0046] Step S301: Obtain a target text and a target image.

[0047] In step S301, the target text is used to describe, in a preset language, the object to be detected in the target image, and the target image is a scene image corresponding to any one of E application scenarios, where E is a positive integer.

[0048] Optionally, the above-mentioned E application scenarios at least include: land monitoring, water conservancy supervision, and environmental protection monitoring scenarios.

[0049] Optionally, the above-mentioned preset language is Chinese.

[0050] Optionally, the above-mentioned target image can be collected by a camera installed at a medium or high position in the E application scenarios, so as to ensure that the collected target image includes the scene data corresponding to a relatively large monitoring area in the application scenario.

[0051] Step S302: Input the target text and the target image into a target model.

[0052] In step S302, the target model includes an encoder, a query module, and a decoder. The encoder is at least used to extract features from the target image and the target text. The query module is at least used to determine the similarity between the extracted text features and image features. The decoder is at least used to interact the text features and the image features based on a cross-attention mechanism.

[0053] In this embodiment, the target text can be described in a preset language, and the preset language can be set to Chinese. That is, the target model in this application can support Chinese recognition. In other words, the encoder in this application supports feature extraction of the target text described in the preset language, omitting the process of translating Chinese into English required by the models created in the prior art, thereby avoiding the errors that occur in the translation process, and further improving the accuracy of object detection for the target image.

[0054] In addition, the decoder in this embodiment uses a cross-attention mechanism to enable in-depth interaction between the text features and the image features. The cross-attention mechanism helps the target model to comprehensively consider the visual cues included in the target image and the semantic information corresponding to the target text during the detection process, thereby improving the detection accuracy of the target model.

[0055] Step S303: Determine a detection result through the target model.

[0056] In step S303, the detection result is used to represent the position of the object described by the target text in the target image.

[0057] Optionally, the detection system marks the position of the object described in the target text in the target image in the form of an anchor box, and returns the marking result as the detection result to the user or relevant technical personnel, so that the user or relevant technical personnel can intuitively see the actual position of the object corresponding to the target text in the form of an image, thereby improving the user experience.

[0058] Furthermore, the present application can also view the detection result in the form of browser access. When the user needs to perform object detection on the application scenario, first log in to the browser corresponding to the detection system, and input the target text, target image, and confidence level. Among them, the setting of the confidence level will directly affect the accuracy of the detection result. A lower confidence threshold can detect more targets, but it can increase the probability of false detection. After that, the present application can also determine the category of the object corresponding to the target text through the target model, and output the category of the object as a label, and compare the label with the label set included in the application scenario corresponding to the target image, so as to judge whether the application scenario division of the target image is correct according to the comparison result.

[0059] As can be seen from the above, the present application has pre-trained a target model, which includes an encoder, a query module, and a decoder. Among them, the decoder can interact with the text feature and the image feature based on the cross-attention mechanism, achieving the purpose of aligning the text feature and the image feature, thereby overcoming the influence of the image complexity from different scenarios. Therefore, the target model trained by the present application can perform object detection on the scene image corresponding to any one of the E application scenarios, so that it is not necessary to perform separate modeling for different application scenarios, and thus the purpose of reducing the object detection cost and improving the object detection efficiency is achieved.

[0060] Thus, the present application realizes the purpose of avoiding separate modeling for different application scenarios from which the target image is sourced by the method of performing object detection on the target image based on the target text by the target model, thereby achieving the technical effects of reducing the object detection cost and improving the object detection efficiency, and further solving the technical problems of high detection cost and low detection efficiency caused by the need for separate modeling for each application scenario when performing object detection on different application scenarios in the prior art.

[0061] In an optional embodiment, the detection system first regularly collects the scene images corresponding to the E application scenarios through E image acquisition devices to obtain an image data set; then, the detection system takes any image in the image data set as the target image; then, the detection system performs data cleaning on the original text corresponding to the target image to obtain the target text, where the data cleaning is used to remove the stop words in the original text, and the original text is used to describe the objects that are likely to appear in the application scenario corresponding to the target image through a preset language.

[0062] Optionally, Figure 4 is a structural diagram of an optional detection system according to an embodiment of the present application. As Figure 4 shown, the detection system deploys E mid-high point image acquisition devices in E application scenarios respectively, such as fire recognition, engineering construction, and personnel swimming. After that, the image data sets collected by the E mid-high point image acquisition devices are centrally stored in the data storage server through the data management server. When object detection is required, the data management server retrieves according to user requirements. After that, the detected image data is processed and then transmitted to the high-performance computing server (i.e., the target model) as the target image. Among them, the high-performance computing server is a mid-high point multi-scenario open-set object detection model that supports Chinese input. Scene detection is performed through the target model. After an anomaly is detected, the detection result is input to the alarm management system. Finally, the alarm information is input to the monitoring center display screen that can interact with users or technicians through the alarm management system.

[0063] Optionally, the E mid-high point image acquisition devices can also collect images of their respective monitoring areas at preset time intervals, such as once every 10 minutes, to ensure that image data fully reflecting scene changes can be collected.

[0064] Optionally, when the target scene is an engineering construction application scenario, the objects described in the target text are construction equipment such as excavators, trucks, or tower cranes.

[0065] In the above embodiment, the present application performs data cleaning on the original text corresponding to the target image, removes the stop words and other noise information without semantics included in the original text, making the target text input to the target model more refined. Therefore, the target text can directly point to the target object to be concerned in the image, thereby reducing the redundant calculation amount of the target model when processing the target text, accelerating the speed of feature extraction and matching of the target model, and thus improving the overall efficiency of the target model for object detection.

[0066] In addition, the target text after data cleaning focuses on key semantic information, avoids causing incorrect training directions due to the noise information in the original text, enables the target model to more accurately understand the text description, and thus can correctly locate and identify objects in the target image, further improving the accuracy of the detection results output by the target model.

[0067] In an alternative embodiment, the detection system determines the semantic information and context information included in the target text through the encoder in the target model, and uses the semantic information and context information included in the target text as the first text feature corresponding to the target text. Then, the detection system extracts features from the target image based on a multi-scale feature fusion strategy through the encoder in the target model, and obtains M first image features corresponding to the target image, where M is a positive integer. Then, the detection system determines the detection result based on the first text feature and the M first image features.

[0068] Optionally, the encoder in the target model includes a text encoder and a visual encoder, where the text encoder is an encoder that supports preset language text prompts.

[0069] For example, when the preset language is Chinese, the target model uses BERT (Bidirectional Encoder Representations from Transformers) trained based on Chinese corpus to process the target text and generate the first text feature. The text encoder includes 12 layers of Transformer (a neural network for processing sequence data) structures, the dimension of the hidden layer is 768, and the number of attention heads is 12. After the target text is input, it first undergoes a tokenization operation by the tokenizer in the text encoder, then is converted into a vector representation through the word embedding layer in the text encoder, and finally, feature extraction is performed through the multi-layer Transformer structure in the text encoder, so as to achieve the purpose of capturing the semantic information and context relationship corresponding to the Chinese text.

[0070] For example, the detection system uses Swin-Transformer (window multi-head self-attention transformer) as the visual encoder. The structure of this visual encoder includes 4 stages, each stage contains 2, 2, 6, and 2 Transformer structures respectively. The feature dimension captured by the visual encoder starts from 96 and doubles stage by stage, and finally reaches 768. The number of heads in the multi-head attention is [3, 6, 12, 24]. Moreover, the visual encoder adopts a multi-scale feature fusion strategy, so as to achieve the purpose of improving the accuracy of detecting objects of different sizes included in the target image.

[0071] In an alternative embodiment, the detection system first inputs the first text feature and M first image features into the query module in the target model. Then, the detection system enhances the features of the first text feature through the query module to obtain a second text feature, and enhances the features of the M first image features to obtain M second image features. Next, the detection system obtains the score matrix corresponding to the M second image features, where the scores in the score matrix are used to represent the similarity between each second image feature and the second text feature.

[0072] In addition, the detection system sorts the M second image features according to the scores in the score matrix to obtain an image feature sequence. Then, the detection system uses the top P second image features in the image feature sequence as P indices, where the P indices are used to determine the position of the object described by the target text in the target image, and P is a positive integer less than or equal to M. Finally, the detection system determines the detection result based on the second text feature and the P indices.

[0073] Optionally, the calculation formula of the score matrix is shown in the following formula (1).

[0074] (1)

[0075] In the above formula (1), L is the score matrix, is the second image feature, is the first text feature, b is the batch size, is the number of tokens in the image, is the number of tokens in the text, d is the feature dimension, i is the i-th sample in the batch, j is the j-th token in the image feature, k is the k-th token in the text feature, and l is the l-th dimension in the feature dimension.

[0076] Optionally, the above token is a fixed-length feature vector converted from each small block obtained by segmenting the image or text.

[0077] Next, the maximum score corresponding to each image token is shown in the following formula (2):

[0078] (2)

[0079] Finally, as shown in the following formula (3), the P second image features with the largest scores are selected as the P indices.

[0080] (3)

[0081] In the above embodiments, the query module can enhance text features through a general self-attention mechanism and enhance image features through a deformable self-attention mechanism, obtaining second text features and M second image features, thereby achieving the purpose of strengthening the robustness and discriminability of text feature and image feature representations based on the self-attention mechanism, enabling text features to better generalize and describe objects, and image features to more accurately capture object details, thus improving the detection accuracy of the target model.

[0082] In the above embodiments, each score in the score matrix represents the matching score between a second image feature and the second text feature. The higher the value, the stronger the correlation between the two. According to the obtained score matrix, the detection system sorts the M second image features from high to low according to the scores. Then, the target model can perform a preliminary query based on the sorting result, that is, select the top P second image features as P indices, and the P indices are used to indicate the positions where the object described by the target text appears in the target image. Finally, the detection system performs an accurate detection, that is, by comprehensively analyzing the second text feature and the P indices, further determining the accurate position of the object described by the target text. In the above recognition process, the present application utilizes the guidance of the second text feature and the pointing of the P indices to achieve the purpose of narrowing the detection range of the accurate detection, thereby improving the detection accuracy of the target model.

[0083] In an alternative embodiment, the detection system inputs the second text feature and the P indices into the decoder in the target model. Then, the detection system updates the second text feature through the text cross-attention layer in the decoder to obtain a third text feature, where the text cross-attention layer is used to enhance the features of the second text feature. Then, the detection system updates the P indices through the image cross-attention layer in the decoder to obtain P third image features, where the image cross-attention layer is used to enhance the features of the P indices. Finally, the detection system determines the detection result based on the similarity between each third image feature in the P third image features and the third text feature.

[0084] Optionally, the text cross-attention layer is as shown in the following formula (4):

[0085] (4)

[0086] Optionally, the image cross-attention layer is as shown in the following formula (5):

[0087] (5)

[0088] In the above formulas (4) and (5), Q represents the query vector, K represents the key vector, and V represents the feature vector that needs to be weighted and summed by the attention mechanism.

[0089] In the above embodiments, the text cross-attention layer allows the target model to introduce the semantic information of the text when processing the image. This not only deepens the target model's understanding of the semantic meaning of the target text, but also enables the image features to be adjusted according to the guidance of the target text, so that the target model can more accurately capture the image regions that match the target text description. The image cross-attention layer enables the first text features to absorb the corresponding image visual cues, thereby enhancing its ability to represent specific scenes and objects. This two-way feature enhancement process promotes the deep interaction between the target text and the target image, thus improving the accuracy of the target model detection.

[0090] In addition, objects in different application scenarios are allowed to have similar visual appearances but completely different semantics. By introducing the text cross-attention layer and the image cross-attention layer, the detection system enables the target model to consider both the visual similarity and semantic difference perspectives during detection, enabling the target model to quickly adapt in various application scenarios. Even when faced with unseen detection objects or application scenarios, it can perform detections based on existing knowledge and text descriptions, thereby improving the generalization performance of the target model.

[0091] In an alternative embodiment, the training steps of the target model include: First, the detection system preprocesses X open-source images obtained from the network to obtain Y open-source training images, where both X and Y are positive integers and Y is greater than X. The preprocessing is at least used to perform random flipping, random cropping, random scaling, and normalization operations on the open-source images. Then, the detection system uses the Y open-source training images, Z bounding boxes, and W object types as the first dataset, where both Z and W are positive integers, and the W object types are described by texts in a preset language. Then, the detection system iteratively trains the initial model based on the first dataset through a linear learning rate warm-up strategy and a learning rate jump strategy to obtain the first model. Finally, the detection system determines the target model through the first model.

[0092] In the above embodiments, accelerating model convergence: The linear learning rate warm-up strategy can gradually increase the learning rate at the initial stage of the model, which helps the model to explore the solution space faster, avoiding problems such as slow training, gradient disappearance, or gradient explosion caused by improper initial learning rate settings. Moreover, the warm-up strategy can smoothly start the training process, enabling the model in training to quickly adapt to the data distribution in the first dataset and shortening the training time.

[0093] In addition, the learning rate jump strategy, i.e., reducing the learning rate after a specific number of training rounds, helps to prevent the model from overfitting the training data in the later stage of training. The detection system can adjust the weights more precisely with a lower learning rate, thereby reducing the overreaction to noisy data. While maintaining the complexity of the trained model, it can also improve its generalization ability, enabling the target model to maintain good detection performance on unseen Chinese data and image data, and thus improving the detection accuracy of the trained model.

[0094] Moreover, the first dataset in this application is a large-scale sample set. By initially training the detection system on the large-scale sample set, the trained model can learn rich expressions of the Chinese language. Combining the linear warm-up and learning rate jump strategies, the features learned by the model in multiple stages will be more abundant, thereby improving the robustness of the trained model.

[0095] In an optional embodiment, the training steps of the target model further include: First, the detection system preprocesses E historical scenario images corresponding to E application scenarios to obtain U historical training images, where U is a positive integer and U is greater than E. Then, the detection system uses the U historical training images and the object types appearing in the historical training images as the second dataset. Finally, the detection system iteratively trains the first model based on the second dataset through the progressive learning rate decay strategy to obtain the target model.

[0096] In the above embodiment, the detection system iteratively trains the first model based on the second dataset through the progressive learning rate decay strategy, and a relatively small learning rate is adopted during the training process. The functions are as follows:

[0097] (1) The progressive learning rate decay strategy ensures that the model can gradually adjust its weights during training, from quickly learning basic features at the beginning to fine-tuning at the end, making it more suitable for the data distribution in the small-sample second dataset. This enables the target model to be more sensitive to object detection in specific scenarios, even when the data volume is relatively limited, thereby improving the detection accuracy of the target model.

[0098] (2) Using a relatively small learning rate for fine-tuning can limit the update amplitude of the weights during model training, thereby reducing the over-reliance on specific data features in the small-sample second dataset, and thus improving the generalization ability of the target model.

[0099] (3) By fine-tuning and training based on the small-sample second dataset through progressive learning rate decay, and using a relatively small learning rate in the later stage of training, unnecessary iterations can be reduced, the training cost of the target model can be lowered, and the overall training efficiency can be improved.

[0100] As can be seen from the above, a target model is pre-trained in this application. The target model includes an encoder, a query module, and a decoder. Among them, the decoder can interact with text features and image features based on the cross-attention mechanism, achieving the purpose of aligning text features and image features, thereby overcoming the influence of the complexity of images from different scenarios. Therefore, the target model trained in this application can perform object detection on the scene image corresponding to any one of the E application scenarios, eliminating the need for separate modeling for different application scenarios, and thus achieving the purpose of reducing the cost of object detection and improving the efficiency of object detection.

[0101] Thus, by using the target model to perform object detection on the target image based on the target text, this application achieves the purpose of avoiding separate modeling for different application scenarios where the target image comes from, thereby achieving the technical effects of reducing the cost of object detection and improving the efficiency of object detection, and further solving the technical problems of high detection cost and low detection efficiency caused by the need for separate modeling for each application scenario when performing object detection on different application scenarios in the prior art.

[0102] In an alternative embodiment, Figure 5 is a flowchart of another alternative method for detecting an image based on text according to an embodiment of the present application. As Figure 5 shown, the method includes: inputting a Chinese text and an image to be detected into the target model; performing data preprocessing on the Chinese text and the image to be detected by the target model. After that, encoding the text and the image respectively by the target model, then, performing query selection guided by the Chinese text, and performing cross-modal decoding based on the selection result. Finally, outputting the detection result.

[0103] In an alternative embodiment, Figure 6 is a flowchart of the operation of an alternative target model according to an embodiment of the present application. As Figure 6 shown, the target model structure includes: a text encoder, a visual encoder, a selection module (i.e., the above-mentioned query module), and a decoder.

[0104] Optionally, after extracting text features and image features through the text encoder and the visual encoder, the present application inputs the text features and the image features into a cross-modal feature fusion module for feature enhancement. This module includes multiple feature enhancer layers. The deformable self-attention is used to enhance the image features through the cross-modal feature fusion module, and the ordinary self-attention is used to enhance the text features. In addition, the target model also adds an image-to-text cross-attention module and a text-to-image cross-attention module for fusing the text features and the image features, so as to align the text features and the image features. Then, the image features more relevant to the input text are selected by the selection module as the queries of the decoder, and 900 indices are output. The queries are initialized based on the selected 900 indices.

[0105] Optionally, the decoder is a cross-modal decoder. Among them, the cross-modal decoder adopts a 6-layer Transformer structure, with a hidden layer dimension of 256 and 8 attention heads. During the processing, each query first passes through the self-attention layer, and then interacts with the image features and the text features through the text cross-attention layer and the image cross-attention layer respectively. Then, the interaction results are sent to the feed-forward neural network layer (with an intermediate layer dimension of 1024) for further processing of the queries, so as to improve the representation ability of the features. Finally, the position of the object described by the text is determined by calculating the similarity between the text features and the image features.

[0106] In the above embodiment, this design enables the cross-modal decoder to establish closer connections between different modalities, thereby improving the performance of the target model in visual and language understanding tasks.

[0107] Optionally, the target model in the present application includes two training stages: pre-training and fine-tuning. Among them, the pre-training stage is a key step for the model to learn general features. A large-scale object detection dataset is used in the pre-training process. This dataset contains 365 common categories, 2 million images, and 30 million bounding boxes. The pre-training stage enables the model to learn the object detection ability in diverse scenarios, so as to provide a basis for general recognition ability for subsequent fine-tuning. The specific data processing flow of the pre-training is as follows:

[0108] (1) Data preprocessing: It includes techniques such as random flipping, random cropping, random selection of scaling, and normalization processing. Among them, the random flipping technique flips the image horizontally with a probability of 0.5, increasing the diversity of the training data. The random cropping technique is to crop a part of the image area to simulate the situation of partial occlusion of the target. The random scaling technique performs object detection at different scales, thereby enhancing the performance of the model at different resolutions. The normalization processing technique normalizes the pixel values of the image to meet the input requirements of the pre-trained model.

[0109] (2)Optimization strategies in the pre-training stage include: using the AdamW (Adaptive Moment Estimation with Decoupled Weight Decay, an adaptive moment estimation algorithm with decoupled weight decay) optimizer with an initial learning rate of 0.0004, adopting a linear learning rate warm-up strategy (i.e., gradually increasing to the set value within the first 1000 steps) and a learning rate jump strategy (i.e., reducing the learning rate by 10 times in the 19th and 26th rounds), using weight decay to prevent overfitting, and clipping the gradient to limit the maximum value of the gradient norm to 0.1.

[0110] (3)In each batch of training in the pre-training stage, 144 images are used, and distributed training is carried out on multiple graphics processors through 144 images, so as to accelerate the training process and improve the training efficiency of the model.

[0111] (4)The loss functions in the pre-training stage include: the first loss function and the second loss function. Among them, the first loss function is used for classification loss, which can handle the problem of sample class imbalance and ensure the detection effect of small-class targets; the second loss function is used for bounding box regression loss, so as to ensure that the bounding box of the target can match the real position and improve the detection accuracy.

[0112] (5)During the pre-training stage, the English labels included in the large-scale object detection dataset were converted into Chinese. Due to the semantic differences between English and Chinese, the labels also need to be double-checked during the pre-training stage, so as to ensure that the objects marked in the dataset accurately correspond to the Chinese labels semantically, and thus avoid reducing the detection effect of the target model in the Chinese scenario due to translation errors.

[0113] Optionally, most settings in the fine-tuning stage are similar to those in the pre-training stage. In the fine-tuning stage, a fine-tuning dataset is established using historical images sampled from different application scenarios and the annotations corresponding to the historical images. A smaller learning rate is adopted in the fine-tuning stage to avoid destroying the general features learned in the pre-training stage. At the same time, for professional terms in specific fields, the labels corresponding to the professional terms are appropriately converted and adjusted in the fine-tuning stage to ensure that the target model can accurately understand and detect industry-specific targets. The specific data processing flow for fine-tuning is as follows:

[0114] (1)Fine-tuning dataset construction: In the fine-tuning stage, images sampled from different scenarios (such as illegal mining of minerals, firework recognition, dust monitoring, etc.) are used for annotation to construct a fine-tuning dataset for a specific field. These datasets cover typical applications of object detection in various industries, ensuring that the model can handle industry-specific detection requirements.

[0115] (2)Sample preprocessing: Similar to the pre-training stage, it includes random flipping, random cropping, and random selection of scaling techniques.

[0116] (3)Learning rate adjustment: In the fine-tuning stage, a relatively small learning rate is adopted to avoid destroying the general features learned during pre-training. Through a progressive learning rate decay strategy, the learning rate is gradually reduced, enabling the model to stably optimize on the fine-tuning dataset and better adapt to the targets specific to the application scenario.

[0117] (4)Label conversion and adjustment: For the professional terms and expressions in a specific field, the labels are appropriately converted and adjusted. To ensure that the model can accurately understand and detect the targets unique to the industry, domain-specific terms are particularly considered during the label conversion process, and detailed inspections and adjustments are carried out. This process can ensure that the model can accurately understand the industry-specific markings in object detection.

[0118] In the above embodiments, the functions of multi-scene object detection by the target model are as follows:

[0119] (1)The target model has general object detection capabilities and supports multiple downstream tasks, such as engineering construction, firework recognition, dust monitoring, etc. There is no need to develop a dedicated model for each task, improving the generality of the target model.

[0120] (2)The target model structure includes a text encoder that supports Chinese prompts and provides a complete Chinese training process, solving the problem that existing technologies can only process English inputs, avoiding losses and time overhead caused by translation, and thus improving the detection efficiency in Chinese scenarios.

[0121] (3)The target model adopts a training process of pre-training and fine-tuning, which not only ensures the generality of the model but also optimizes it for specific domain application scenarios during the fine-tuning process, thereby improving the detection accuracy and application effect.

[0122] According to another aspect of the embodiments of the present application, there is also provided a device for detecting an image based on text. Figure 7 It is a schematic diagram of an optional device for detecting an image based on text according to the embodiments of the present application. As Figure 7 shown, the device for detecting an image based on text includes: an acquisition unit 701, an input unit 702, and a determination unit 703.

[0123] Optionally, an acquisition unit is configured to acquire a target text and a target image, where the target text is used to describe an object to be detected in the target image in a preset language, and the target image is a scene image corresponding to any one of E application scenarios, where E is a positive integer; an input unit is configured to input the target text and the target image into a target model, where the target model includes an encoder, a query module, and a decoder, the encoder is at least configured to extract features from the target image and the target text, the query module is at least configured to determine the similarity between the extracted text features and image features, and the decoder is at least configured to interact the text features and the image features based on a cross-attention mechanism; a determination unit is configured to determine a detection result through the target model, where the detection result is used to represent the position of the object described by the target text in the target image.

[0124] In an alternative embodiment, the acquisition unit includes: a timing acquisition subunit, a first determination subunit, and a data cleaning subunit.

[0125] Optionally, the timing acquisition subunit is configured to periodically acquire scene images corresponding to E application scenarios through E image acquisition devices to obtain an image data set; the first determination subunit is configured to use any image in the image data set as the target image; the data cleaning subunit is configured to perform data cleaning on the original text corresponding to the target image to obtain the target text, where the data cleaning is used to remove stop words in the original text, and the original text is used to describe objects that are likely to appear in the application scenario corresponding to the target image in a preset language.

[0126] In an alternative embodiment, the determination unit includes: a second determination subunit, a feature extraction subunit, and a third determination subunit.

[0127] Optionally, the second determination subunit is configured to determine the semantic information and context information included in the target text through the encoder in the target model, and use the semantic information and context information included in the target text as the first text features corresponding to the target text; the feature extraction subunit is configured to extract features from the target image based on a multi-scale feature fusion strategy through the encoder in the target model to obtain M first image features corresponding to the target image, where M is a positive integer; the third determination subunit is configured to determine the detection result based on the first text features and the M first image features.

[0128] In an alternative embodiment, the third determination subunit includes: an input module, a feature enhancement module, an acquisition module, a sorting module, a first determination module, and a second determination module.

[0129] Optionally, an input module is configured to input the first text feature and M first image features into a query module in a target model; a feature enhancement module is configured to enhance the first text feature through the query module to obtain a second text feature, and enhance the M first image features to obtain M second image features; an acquisition module is configured to acquire a score matrix corresponding to the M second image features, where the scores in the score matrix are used to represent the similarity between each second image feature and the second text feature; a sorting module is configured to sort the M second image features according to the scores in the score matrix to obtain an image feature sequence; a first determination module is configured to use the top P second image features in the image feature sequence as P indexes, where the P indexes are used to determine the position of the object described by the target text in the target image, and P is a positive integer less than or equal to M; a second determination module is configured to determine a detection result according to the second text feature and the P indexes.

[0130] In an alternative embodiment, the second determination module includes: an input sub-module, a first update sub-module, a second update sub-module, and a determination sub-module.

[0131] Optionally, the input sub-module is configured to input the second text feature and the P indexes into a decoder in the target model; the first update sub-module is configured to update the second text feature through a text cross-attention layer in the decoder to obtain a third text feature, where the text cross-attention layer is configured to enhance the second text feature; the second update sub-module is configured to update the P indexes through an image cross-attention layer in the decoder to obtain P third image features, where the image cross-attention layer is configured to enhance the P indexes; the determination sub-module is configured to determine a detection result according to the similarity between each third image feature in the P third image features and the third text feature.

[0132] In an alternative embodiment, the apparatus for detecting an image based on text further includes: a first preprocessing unit, a first data set determination unit, a first iterative training unit, and a model determination unit.

[0133] Optionally, a first preprocessing unit is configured to preprocess X open-source images obtained from a network to obtain Y open-source training images, where both X and Y are positive integers, and Y is greater than X. The preprocessing is at least used to perform random flipping, random cropping, random scaling, and normalization operations on the open-source images; a first dataset determination unit is configured to use the Y open-source training images, Z bounding boxes, and W object types as a first dataset, where both Z and W are positive integers, and the W object types are described by texts corresponding to a preset language; a first iterative training unit is configured to iteratively train an initial model based on the first dataset through a linear learning rate warm-up strategy and a learning rate jump strategy to obtain a first model; a model determination unit is configured to determine a target model through the first model.

[0134] In an alternative embodiment, the apparatus for detecting an object in an image based on text further includes: a second preprocessing unit, a second dataset determination unit, and a second iterative training unit.

[0135] Optionally, a second preprocessing unit is configured to preprocess E historical scenario images corresponding to E application scenarios to obtain U historical training images, where U is a positive integer and U is greater than E; a second dataset determination unit is configured to use the U historical training images and the object types appearing in the historical training images as a second dataset; a second iterative training unit is configured to iteratively train the first model based on the second dataset through a progressive learning rate decay strategy to obtain a target model.

[0136] As can be seen from the above, a target model is pre-trained in this application. The target model includes an encoder, a query module, and a decoder. Among them, the decoder can interact the text features and image features based on the cross-attention mechanism, achieving the purpose of aligning the text features and image features, thereby overcoming the influence of the image complexity from different scenarios. Therefore, the target model trained in this application can detect objects in the scenario image corresponding to any one of the E application scenarios, so that there is no need to perform separate modeling for different application scenarios, and thus the purpose of reducing the object detection cost and improving the object detection efficiency is achieved.

[0137] In addition, the encoder in this application supports feature extraction of the target text described in a preset language, omitting the process in the existing technology where the created model needs to translate the preset language into English, thereby avoiding the errors that occur during the translation process, and further improving the accuracy of object detection in images.

[0138] It can be seen that through the method of object detection on the target image by the target model based on the target text, this application achieves the purpose of avoiding separate modeling for different application scenarios where the target image comes from, thereby achieving the technical effects of reducing the cost of object detection and improving the efficiency of object detection, and further solving the technical problems of high detection cost and low detection efficiency caused by the need for separate modeling for each application scenario in the prior art when performing object detection on different application scenarios.

[0139] According to another aspect of the embodiments of the present application, there is also provided a computer program product. The computer program product includes a stored computer program. Among them, when the computer program runs, it controls the computer program product to execute the method of detecting an image based on text in any one of the above.

[0140] According to another aspect of the embodiments of the present application, there is also provided an electronic device, including: a processor; and a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the method of detecting an image based on text in any one of the above via executing the executable instructions.

[0141] Optionally, Figure 8 is a schematic diagram of an optional electronic device according to an embodiment of the present application. As Figure 8 shown, the embodiments of the present application provide an electronic device. The electronic device includes a processor, a memory, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements the method of detecting an image based on text in any one of the above.

[0142] The above-described embodiments or examples disclosed in the present application are not exhaustive. They are only illustrations of some embodiments or examples and do not serve as specific limitations on the protection scope disclosed in the present application. Without contradiction, each step in a certain embodiment or example in the present application can be implemented as an independent embodiment, and the steps can be combined arbitrarily. For example, the solution after removing some steps in a certain embodiment or example can also be implemented as an independent embodiment, and the order of the steps in a certain embodiment or example can be exchanged arbitrarily. In addition, the optional ways or optional examples in a certain embodiment or example can be combined arbitrarily; furthermore, the various embodiments or individual embodiments can be combined arbitrarily. For example, some or all of the steps of different embodiments or examples can be combined arbitrarily, and a certain embodiment or example can be combined arbitrarily with the optional ways or optional examples of other embodiments or examples.

[0143] In the above embodiments of the present application, the descriptions of the various embodiments have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0144] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate a means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or a means for implementing the functions specified in one or more of the blocks.

[0145] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction means that implements the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or a means for implementing the functions specified in one or more of the blocks.

[0146] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or a means for implementing the functions specified in one or more of the blocks.

[0147] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory. The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.

[0148] A computer-readable medium includes permanent and non-permanent, removable and non-removable media that can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassette tapes, magnetic tape disk storage, or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media, such as modulated data signals and carrier waves.

[0149] It should also be noted that the term "comprising," "including," or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but also other elements not expressly listed, or elements that are inherent to such process, method, article, or apparatus. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0150] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, system, or computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0151] The above are only embodiments of the present application and are not used to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A method for detecting an image based on text, characterized in that: include: Obtain a target text and a target image, wherein the target text is used to describe the object to be detected in the target image in a preset language, and the target image is a scene image corresponding to any one of E application scenarios, where E is a positive integer; Inputting the target text and the target image into a target model, wherein the target model includes an encoder, a query module, and a decoder, wherein the encoder is at least used to extract features from the target image and the target text, the query module is at least used to determine the similarity between the extracted text features and the image features, and the decoder is at least used to interact the text features with the image features based on a cross-attention mechanism; A detection result is determined by the target model, wherein the detection result is used to characterize a position of an object described by the target text in the target image.

2. The method for detecting images based on text according to claim 1, characterized in that: Get the target text and target image, including: By using E image acquisition devices, scene images corresponding to the E application scenes are acquired at regular intervals to obtain an image data set; Taking any image in the image dataset as the target image; The target text is obtained by performing data cleaning on the original text corresponding to the target image, wherein the data cleaning is used to remove stop words in the original text, and the original text is used to describe objects that appear with probability in the application scenario corresponding to the target image in the preset language.

3. The method for detecting images based on text according to claim 1, characterized in that: Determining a detection result by using the target model includes: Determining semantic information and context information included in the target text through an encoder in the target model, and using the semantic information and context information included in the target text as a first text feature corresponding to the target text; Extracting features of the target image based on a multi-scale feature fusion strategy through an encoder in the target model to obtain M first image features corresponding to the target image, where M is a positive integer; The detection result is determined according to the first text feature and the M first image features.

4. The method for detecting images based on text according to claim 3, characterized in that: Determining the detection result according to the first text feature and the M first image features includes: Inputting the first text feature and the M first image features into a query module in the target model; Performing feature enhancement on the first text feature through the query module to obtain a second text feature, and performing feature enhancement on the M first image features to obtain M second image features; Obtaining a score matrix corresponding to the M second image features, wherein the scores in the score matrix are used to represent the similarity between each second image feature and the second text feature; Sorting the M second image features according to the scores in the score matrix to obtain an image feature sequence; Taking the top P second image features in the image feature sequence as P indexes, wherein the P indexes are used to determine the position of the object described by the target text in the target image, and P is a positive integer less than or equal to M; The detection result is determined according to the second text feature and the P indexes.

5. The method for detecting images based on text according to claim 4, characterized in that: Determining the detection result according to the second text feature and the P indexes includes: Inputting the second text feature and the P indexes into a decoder in the target model; The second text feature is updated by a text cross attention layer in the decoder to obtain a third text feature, wherein the text cross attention layer is used to perform feature enhancement on the second text feature; The P indexes are updated by an image cross attention layer in the decoder to obtain P third image features, wherein the image cross attention layer is used to perform feature enhancement on the P indexes; The detection result is determined according to a similarity between each third image feature of the P third image features and the third text feature.

6. The method for detecting images based on text according to claim 1, characterized in that: The target model is trained by the following steps: Preprocessing X open source images obtained from the Internet to obtain Y open source training images, wherein X and Y are both positive integers, Y is greater than X, and the preprocessing is at least used to perform random flipping, random cropping, random scaling and normalization operations on the open source images; The Y open source training images, Z bounding boxes, and W object types are used as a first data set, where Z and W are both positive integers, and the W object types are described by text corresponding to the preset language; Based on the first data set, the initial model is iteratively trained by a linear learning rate warm-up strategy and a learning rate jump strategy to obtain a first model; The target model is determined by using the first model.

7. The method for detecting images based on text according to claim 6, characterized in that: Determining the target model by using the first model includes: Preprocessing the E historical scene images corresponding to the E application scenarios to obtain U historical training images, where U is a positive integer and U is greater than E; Taking the U historical training images and the object types appearing in the historical training images as a second data set; The first model is iteratively trained based on the second data set through a progressive learning rate decay strategy to obtain the target model.

8. A device for detecting an image based on text, characterized in that: include: an acquisition unit, configured to acquire a target text and a target image, wherein the target text is used to describe an object to be detected in the target image in a preset language, and the target image is a scene image corresponding to any one of E application scenarios, where E is a positive integer; An input unit, used to input the target text and the target image into a target model, wherein the target model includes an encoder, a query module and a decoder, the encoder is at least used to extract features from the target image and the target text, the query module is at least used to determine the similarity between the extracted text features and the image features, and the decoder is at least used to interact the text features with the image features based on a cross-attention mechanism; A determination unit is used to determine a detection result through the target model, wherein the detection result is used to characterize the position of the object described by the target text in the target image.

9. A computer program product, characterized in that The computer program product comprises a computer program, wherein when the computer program is run, the computer program product is controlled to execute the method for detecting images based on text according to any one of claims 1 to 7.

10. An electronic device, characterized in that: It includes one or more processors and a memory, wherein the memory is used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method for detecting images based on text as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Text-guided image detection method, system, device, medium and program product

    CN118314148A

  • Hybrid expert target detection system and method

    CN118675030A

  • Infrared small target detection method based on scene text information guidance

    CN118762364A

  • Model training and application method and device for target detection and storage medium

    CN118799608A

  • Open vocabulary target detection method and system for electric power construction scene picture

    CN118898709A

Cited By

  • Sensor arrangement method and system and construction site field management method and system

    CN121637319A