Object Detection Method, System, Storage Medium and Electronic Device

Through the combination of feature extraction and large language model, the problem that traditional object detection methods cannot detect non-fixed categories is solved, and the effectiveness of open-category object detection is achieved.

CN118196775BActive Publication Date: 2025-06-13SHANGHAI MIDU DIGITAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410320989.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-20
Publication Date
2025-06-13
Estimated Expiration
2044-03-20

AI Technical Summary

Technical Problem

Traditional object detection methods are difficult to effectively detect non-fixed categories of targets, and they cannot conduct open-category object detection.

Method used

By using the feature extraction model, CLIP model and large language model based on the input, the candidate box, category feature and name feature are obtained, and the matching score is calculated to achieve object detection.

Benefits of technology

It realizes open-category target detection without limiting the detection target category, and can effectively detect targets specified by users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118196775B_ABST
    Figure CN118196775B_ABST
Patent Text Reader

Abstract

The present application provides a target detection method, system, storage medium and electronic device. The method includes the following steps: obtaining a first mapping feature map and a second mapping feature map based on an input image to be recognized; obtaining a target name text feature based on an input target name text; obtaining a candidate box, a category feature and a name feature based on the first mapping feature map, the second mapping feature map and the target name text feature; and obtaining a matching score between a detection target and the candidate box based on the category feature and the name feature to achieve target detection. By inputting a specific target name text, the present application can perform corresponding target detection without limiting the category of the detection target, realizing open-category target detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of image processing, and particularly relates to an object detection method, system, storage medium and electronic device. Background Art

[0002] Object detection is one of the core problems in the field of computer vision. Its task is to detect all objects of interest in an image and determine their categories and locations. Currently, traditional object detection methods limit the categories of detected objects, that is, under traditional methods, all categories to be detected are available during the training phase. However, for non-fixed categories, effective detection cannot be carried out, and traditional methods are difficult to perform effective detection. Summary of the Invention

[0003] The purpose of this application is to provide an object detection method, system, storage medium and electronic device that can perform open-category object detection.

[0004] In a first aspect, this application provides an object detection method, and the method includes the following steps: obtaining a first mapping feature map and a second mapping feature map based on the input image to be recognized; obtaining the target name text feature based on the input target name text; obtaining a candidate box, a category feature, and a name feature based on the first mapping feature map, the second mapping feature map, and the target name text feature; obtaining the matching score between the detected object and the candidate box based on the category feature and the name feature to achieve object detection.

[0005] In an implementation manner of the first aspect, obtaining a first mapping feature map and a second mapping feature map based on the input image to be recognized includes:

[0006] Inputting the image to be recognized into a feature extraction model to obtain a feature map;

[0007] Obtaining a first mapping feature map and a second mapping feature map based on the feature map; the first mapping feature map is used to obtain the candidate box; the second mapping feature map is used to obtain the category feature.

[0008] In an implementation manner of the first aspect, obtaining the target name text feature based on the input target name text includes:

[0009] Inputting the target name text into the CLIP model to obtain the target name text feature; the target name text feature is used to obtain the name feature.

[0010] In an implementation manner of the first aspect, obtaining a candidate box, a category feature, and a name feature based on the first mapping feature map, the second mapping feature map, and the target name text feature includes:

[0011] Input the first mapped feature map, the second mapped feature map, and the target name text feature into a large language model to obtain a first output result, a second output result, and a third output result;

[0012] Based on the first output result, the second output result, and the third output result, obtain the candidate bounding box, the class feature, and the name feature respectively.

[0013] In one implementation of the first aspect, obtaining the candidate bounding box, the class feature, and the name feature based on the first output result, the second output result, and the third output result respectively includes:

[0014] Map the dimensions of the first output result, the second output result, and the third output result respectively based on an MLP layer to correspondingly obtain the candidate bounding box, the class feature, and the name feature.

[0015] In one implementation of the first aspect, obtaining the matching score between the detection target and the candidate bounding box based on the class feature and the name feature to achieve object detection includes:

[0016] Calculate the cosine similarity between the class feature and the name feature to obtain the matching score;

[0017] Judge the candidate bounding box based on the matching score to achieve object detection.

[0018] In one implementation of the first aspect, judging the candidate bounding box based on the matching score to achieve object detection includes:

[0019] Judge the relationship between the matching score and any preset threshold. If it is higher than the preset threshold, output the corresponding candidate bounding box as the object detection bounding box to achieve object detection.

[0020] In a second aspect, the present application provides an object detection system, and the system includes:

[0021] An image module, configured to obtain a first mapped feature map and a second mapped feature map based on an input image to be recognized;

[0022] A text module, configured to obtain a target name text feature based on an input target name text;

[0023] A processing module, configured to obtain a candidate bounding box, a class feature, and a name feature based on the first mapped feature map, the second mapped feature map, and the target name text feature;

[0024] A detection module, configured to obtain the matching score between the detection target and the candidate bounding box based on the class feature and the name feature to achieve object detection.

[0025] In a third aspect, the present application provides an electronic device, which includes: a processor and a memory;

[0026] The memory is used to store a computer program;

[0027] The processor is used to execute the computer program stored in the memory, so that the electronic device executes the above-mentioned target detection method.

[0028] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by an electronic device, the above-mentioned target detection method is implemented.

[0029] As described above, the target detection method, system, storage medium and electronic device of the present application have the following beneficial effects: By inputting specific target name text, corresponding target detection can be performed, and there is no need to limit the category of the detection target, realizing open-category target detection. Description of the Drawings

[0030] Figure 1 It shows a scene schematic diagram of the electronic device of the present application in an embodiment.

[0031] Figure 2 It shows a flowchart of the target detection method described in an embodiment of the present application in an embodiment.

[0032] Figure 3 It shows a flowchart of the target detection method described in an embodiment of the present application in an embodiment.

[0033] Figure 4 It shows a flowchart of the target detection method described in an embodiment of the present application in an embodiment.

[0034] Figure 5 It shows a flowchart of the target detection method described in an embodiment of the present application in an embodiment.

[0035] Figure 6 It shows a structural schematic diagram of the target detection system described in an embodiment of the present application in an embodiment.

[0036] Figure 7 It shows a structural schematic diagram of the electronic device of the present application in an embodiment. Detailed Embodiments

[0037] The following describes the implementation manners of the present application through specific examples. Those skilled in the art can easily understand the other advantages and effects of the present application from the content disclosed in this specification. The present application can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0038] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present application in a schematic manner. Therefore, only the components related to the present application are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and proportion of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.

[0039] In addition, in the present application, descriptions such as "first" and "second" are only for descriptive purposes, and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present application.

[0040] The following embodiments of the present application provide a target detection method, which can be applied to an electronic device as Figure 1 shown. The electronic device described in the present application may include a mobile phone 11 with a wireless charging function, a tablet computer 12, a notebook computer 13, a wearable device, a vehicle-mounted device, an augmented reality (AR) / virtual reality (VR) device, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc. The specific type of the electronic device is not limited in the embodiments of the present application.

[0041] For example, the electronic device may be a station (STAION, ST) in a WLAN with wireless charging function, a cellular phone with wireless charging function, a cordless phone, a Session Initiation Protocol (SIP) phone, a Wireless Local Loop (WLL) station, a Personal Digital Assistant (PDA) device, a handheld device with wireless charging function, a computing device or other processing device, a computer, a laptop computer, a handheld communication device, a handheld computing device, and / or other devices for communicating on a wireless system, as well as next-generation communication systems, such as mobile terminals in a 5G network, mobile terminals in a future evolved Public Land Mobile Network (PLMN), or mobile terminals in a future evolved Non-terrestrial Network (NTN), etc.

[0042] For example, the electronic device may communicate with a network and other devices through wireless communication. The above wireless communication may use any communication standard or protocol, including but not limited to Global System of Mobile communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), BT, GNSS, WLAN, NFC, FM, and / or IR technology, etc. The GNSS may include Global Positioning System (GPS), Global Navigation Satellite System (GLONASS), BeiDou navigation Satellite System (BDS), Quasi-Zenith Satellite System (QZSS), and / or Satellite Based Augmentation Systems (SBAS).

[0043] The technical solutions in the embodiments of the present application will be described in detail below with reference to the accompanying drawings in the embodiments of the present application.

[0044] As Figure 2 shown, in one embodiment, the object detection method of the present application includes steps S1 - S4.

[0045] S1: Obtain a first mapped feature map and a second mapped feature map based on the input image to be recognized.

[0046] Specifically, as Figure 3 shown, step S1 includes steps S11 - S12.

[0047] S11: Input the image to be recognized into a feature extraction model to obtain a feature map.

[0048] In some embodiments, the feature extraction model can select the YOLO model as the basic model. The YOLO model is a model for object detection using a convolutional neural network. It divides the image into regions and predicts bounding boxes and probabilities for each region, realizing the application of a single neural network to the entire image. The YOLO model uses convolutional layers to downsample the feature map, which helps prevent the loss of low-level features often attributed to pooling. Step S11 extracts features from the image to be recognized through the YOLO model and outputs feature maps of multiple scales.

[0049] In some embodiments, step S11 outputs feature maps of 3 scales through the YOLO model. Assuming that the number of objects in the image to be recognized does not exceed 200, a feature map with a shape of (14 * 14, 255) is selected for subsequent operations, and the other 2 scale feature maps are not used.

[0050] S12: Obtain a first mapped feature map and a second mapped feature map based on the feature map; the first mapped feature map is used to obtain the candidate box; the second mapped feature map is used to obtain the class feature.

[0051] In some embodiments, the feature maps are respectively mapped in dimension through two MLP layers to obtain the first mapped feature map and the second mapped feature map. MLP (Multilayer Perceptron), that is, a multi-layer perceptron, is a basic artificial neural network model, and its structure is a multi-layer structure composed of multiple neurons. MLP is a feedforward neural network, usually used to solve classification and regression problems. Its basic structure includes an input layer, an output layer, and at least one or more hidden layers. Among them, each layer is composed of multiple neurons, and each neuron generates an output by performing a weighted sum on the input values and passing through an activation function. During the training process, MLP updates the weights and biases between neurons through the backpropagation algorithm to minimize the error between the predicted output and the true output.

[0052] Specifically, the first mapped feature map is used to obtain candidate boxes for object detection; the second mapped feature map is used to obtain class features.

[0053] S2: Obtain the target name text feature based on the input target name text.

[0054] In some embodiments, after the user inputs the target name text to be detected, the text encoder in the CLIP model is used to extract the target name text feature. For example, after the user inputs dog, man, or tree, the CLIP model is used to obtain dog feature, man feature, and tree feature. The CLIP model (Contrastive Language-Image Pre-training) is a multi-modal model based on contrastive text-image pairs. The CLIP model can obtain the matching relationship of text-image pairs through contrastive learning. The CLIP model includes a text encoder and an image encoder, where the text encoder is used to extract the features of the input text, and a text transformer model commonly used in NLP can be adopted.

[0055] Specifically, the target name text feature is used to obtain name features.

[0056] It should be noted that there is no sequence between step S1 and step S2, and the present application is not limited thereto.

[0057] S3: Obtain candidate boxes, class features, and name features based on the first mapped feature map, the second mapped feature map, and the target name text feature.

[0058] Specifically, as Figure 4 shown, step S3 includes steps S31 - S32.

[0059] S31: Input the first mapped feature map, the second mapped feature map, and the target name text feature into a large language model to obtain a first output result, a second output result, and a third output result.

[0060] Specifically, large language models (LLMs) are artificial intelligence models designed to understand and generate human language. By training on vast amounts of text data, they can perform a wide range of tasks, including text summarization, translation, sentiment analysis, and more. These models are typically based on deep learning architectures such as transformers, which enable them to exhibit impressive capabilities in various natural language processing tasks. For example, GPT-4 is a multimodal pre-trained large model that can receive image and text inputs, and users can specify any visual or language task for GPT-4 to generate text outputs, including natural language, code, etc. In some embodiments, the GPT-4 model can be selected to process the first mapped feature map, the second mapped feature map, and the target name text feature to obtain the first output result, the second output result, and the third output result.

[0061] Specifically, the first output result is obtained from the first mapped feature map, the second output result is obtained from the second mapped feature map, and the third output result is obtained from the target name text feature.

[0062] In some embodiments, the large language model can be trained through prompt engineering to optimize the output results.

[0063] S32: Obtain the candidate bounding boxes, the class features, and the name features respectively based on the first output result, the second output result, and the third output result.

[0064] Specifically, map the dimensions of the first output result, the second output result, and the third output result respectively based on the MLP layer to correspondingly obtain the candidate bounding boxes, the class features, and the name features.

[0065] In some embodiments, the shape of the obtained candidate bounding boxes is (14 * 14, 4), indicating that 14 * 14 = 196 candidate bounding boxes are obtained.

[0066] In some embodiments, the shape of the obtained class features is (14 * 14, 256).

[0067] In some embodiments, since the target name text input by the user is dog, man, and tree, the shape of the obtained name features is (3, 256).

[0068] S4: Obtain the matching scores between the detected objects and the candidate bounding boxes based on the class features and the name features to achieve object detection.

[0069] Specifically, as Figure 5 shown, step S4 includes steps S41 - S42.

[0070] S41: Calculate the cosine similarity between the category feature and the name feature to obtain the matching score.

[0071] Specifically, by calculating the cosine similarity between the category feature and the name feature, a similarity matrix score with a shape of (3, 14 * 14) can be obtained, which is equivalent to the matching scores of the target to be detected and 196 candidate boxes.

[0072] S42: Judge the candidate boxes based on the matching score to achieve object detection.

[0073] Specifically, judge the relationship between the matching score and any preset threshold. If it is higher than the preset threshold, output the corresponding candidate box as the object detection box to achieve object detection; if it is lower than the preset threshold, discard the corresponding candidate box.

[0074] In some embodiments, the preset threshold is set to 0.5, and the candidate boxes corresponding to the indices with a matching score > 0.5 are selected as the final output positions (object detection boxes) to achieve object detection.

[0075] In some embodiments, when the target name text input by the user is dog, man, and tree, the final obtained object detection boxes should be no less than 3. Because there may be multiple identical targets in an image, such as multiple dogs.

[0076] The protection scope of the object detection method described in the embodiments of the present application is not limited to the execution order of the steps listed in this embodiment. Any solution achieved by adding or subtracting steps of the prior art and replacing steps according to the principle of the present application is included in the protection scope of the present application.

[0077] The embodiments of the present application also provide an object detection system. The object detection system can implement the object detection method described in the present application. However, the implementation devices of the object detection system described in the present application include but are not limited to the structures of the object detection system listed in this embodiment. Any structural deformation and replacement of the prior art made according to the principle of the present application are included in the protection scope of the present application.

[0078] As Figure 6 shown, in one embodiment, the object detection system of the present application includes an image module 41, a text module 42, a processing module 43, and a detection module 44.

[0079] The image module 41 is used to obtain a first mapping feature map and a second mapping feature map based on the input image to be recognized;

[0080] A text module 42 for obtaining target name text features based on the input target name text;

[0081] A processing module 43 for obtaining candidate boxes, class features, and name features based on the first mapping feature map, the second mapping feature map, and the target name text features;

[0082] A detection module 44 for obtaining the matching scores between the detection targets and the candidate boxes based on the class features and the name features to implement target detection.

[0083] Among them, the structures and principles of the image module 41, the text module 42, the processing module 43, and the detection module 44 correspond one by one to the steps in the above target detection method, so they will not be elaborated here.

[0084] In several embodiments provided in the present application, it should be understood that the disclosed system, apparatus, or method can be implemented in other ways. For example, the apparatus embodiments described above are only illustrative. For example, the division of modules / units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or units can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces, and the indirect coupling or communication connection of devices or modules or units can be in electrical, mechanical, or other forms.

[0085] The modules / units described as separate components may or may not be physically separated, and the components displayed as modules / units may or may not be physical modules, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the modules / units can be selected according to actual needs to achieve the purpose of the embodiments of the present application. For example, in each embodiment of the present application, the various functional modules / units can be integrated in a processing module, or each module / unit can exist physically alone, or two or more modules / units can be integrated in one module / unit.

[0086] Those of ordinary skill in the art should also further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0087] The embodiments of the present application also provide a computer-readable storage medium. Those of ordinary skill in the art can understand that all or part of the steps in the methods of the above embodiments can be completed by instructing a processor through a program. The program can be stored in a computer-readable storage medium. The storage medium is a non-transitory medium, such as random access memory, read-only memory, flash memory, hard disk, solid state drive, magnetic tape, floppy disk, optical disc, and any combination thereof. The above storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center integrating one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., digital video disc (DVD)), or a semiconductor medium (e.g., solid state disk (SSD)), etc.

[0088] The embodiments of the present application also provide an electronic device. The electronic device includes a processor and a memory.

[0089] The memory is used to store a computer program.

[0090] The memory includes various media that can store program codes, such as ROM, RAM, magnetic disk, USB flash drive, memory card, or optical disc.

[0091] The processor is connected to the memory and is used to execute the computer program stored in the memory, so that the electronic device executes the above object detection method.

[0092] Preferably, the processor can be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it can also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0093] Such as Figure 7As shown, the electronic device of the present application is presented in the form of a general-purpose computing device. The components of the electronic device may include, but are not limited to: one or more processors or processing units 51, a memory 52, and a bus 53 that connects different system components (including the memory 52 and the processing unit 51).

[0094] The bus 53 represents one or more of several types of bus architectures, including a memory bus or a memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus architectures. By way of example, these architectures include, but are not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.

[0095] The electronic device typically includes a variety of computer system-readable media. These media can be any available media that can be accessed by the electronic device, including volatile and non-volatile media, removable and non-removable media.

[0096] The memory 52 may include computer system-readable media in the form of volatile memory, such as random access memory (RAM) 521 and / or cache memory 522. The electronic device may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, a storage system 523 may be used for reading and writing on non-removable, non-volatile magnetic media ( Figure 7 not shown, commonly referred to as a "hard disk drive"). Although Figure 7 not shown in the figure, a disk drive for reading and writing on a removable non-volatile disk (such as a "floppy disk") and an optical disk drive for reading and writing on a removable non-volatile optical disk (such as a CD-ROM, DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to the bus 53 through one or more data media interfaces. The memory 52 may include at least one program product having a set (e.g., at least one) of program modules that are configured to perform the functions of the embodiments of the present application.

[0097] A program / utility 524 having a set (at least one) of program modules 5241 may be stored, for example, in the memory 52. Such program modules 5241 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment. The program modules 5241 generally perform the functions and / or methods described in the embodiments of the present application.

[0098] The electronic device can also communicate with one or more external devices (such as a keyboard, a pointing device, a display, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device, and / or communicate with any device that enables the electronic device to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 54. Moreover, the electronic device can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 55. As Figure 7 shown, the network adapter 55 communicates with other modules of the electronic device through the bus 53. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0099] An embodiment of the present application can also provide a computer program product, which includes one or more computer instructions. When the computer instructions are loaded and executed on a computing device, the processes or functions described in the embodiments of the present application are fully or partially generated. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, a computer, or a data center to another website, a computer, or a data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.).

[0100] When the computer program product is executed by a computer, the computer executes the method described in the foregoing method embodiments. The computer program product can be a software installation package. In the case where the foregoing method needs to be used, the computer program product can be downloaded and executed on the computer.

[0101] The present application provides a target detection method. By inputting a specific target name text, corresponding target detection can be performed, and there is no need to limit the category of the detection target, realizing open-category target detection.

[0102] The descriptions of the processes or structures corresponding to the above-mentioned various drawings have their own emphases. For parts not detailed in a certain process or structure, reference can be made to the relevant descriptions of other processes or structures.

[0103] The above embodiments are only illustrative of the principles and effects of the present application and are not intended to limit the present application. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present application. Therefore, all equivalent modifications or changes made by those with ordinary knowledge in the technical field without departing from the spirit and technical idea disclosed by the present application should still be covered by the claims of the present application.

Claims

1. A target detection method, characterized in that: include: Acquire a first mapping feature map and a second mapping feature map based on the input image to be recognized; The first mapping feature map is used to obtain a candidate box; the second mapping feature map is used to obtain a category feature; Obtaining target name text features based on the input name text of the target to be detected; the target name text features are used to obtain name features; Acquire a candidate box, a category feature, and a name feature based on the first mapping feature map, the second mapping feature map, and the target name text feature; Obtaining a matching score between a detection target and the candidate box based on the category feature and the name feature to achieve target detection; Wherein, obtaining a candidate box, a category feature, and a name feature based on the first mapping feature map, the second mapping feature map, and the target name text feature includes: Inputting the first mapping feature graph, the second mapping feature graph and the target name text feature into a large language model to obtain a first output result, a second output result and a third output result; The candidate box, the category feature, and the name feature are respectively acquired based on the first output result, the second output result, and the third output result.

2. The target detection method according to claim 1, characterized in that: include: Acquiring a first mapping feature map and a second mapping feature map based on an input image to be recognized includes: Inputting the image to be identified into a feature extraction model to obtain a feature map; A first mapping feature map and a second mapping feature map are obtained based on the feature map.

3. The target detection method according to claim 1, characterized in that: include: The target name text features obtained based on the input name text of the target to be detected include: The target name text is input into the CLIP model to obtain the target name text features.

4. The target detection method according to claim 1, characterized in that: include: Acquiring the candidate box, the category feature, and the name feature based on the first output result, the second output result, and the third output result respectively includes: Based on the MLP layer, the first output result, the second output result and the third output result are respectively mapped in dimensions to correspondingly obtain the candidate box, the category feature and the name feature.

5. The target detection method according to claim 1, characterized in that: include: Obtaining a matching score between the detection target and the candidate box based on the category feature and the name feature to achieve target detection includes: Performing cosine similarity calculation on the category feature and the name feature to obtain the matching score; The candidate box is judged based on the matching score to achieve target detection.

6. The target detection method according to claim 5, characterized in that: Judging the candidate box based on the matching score to achieve target detection includes: The relationship between the matching score and any preset threshold is determined. If the matching score is higher than the preset threshold, the corresponding candidate box is output as the target detection box to achieve target detection.

7. A target detection system, characterized in that: include: An image module, used for acquiring a first mapping feature map and a second mapping feature map based on an input image to be recognized; The first mapping feature map is used to obtain a candidate box; the second mapping feature map is used to obtain a category feature; A text module, used for obtaining target name text features based on the input name text of the target to be detected; The target name text feature is used to obtain the name feature; A processing module, configured to obtain a candidate box, a category feature, and a name feature based on the first mapping feature map, the second mapping feature map, and the target name text feature; A detection module, configured to obtain a matching score between a detection target and the candidate box based on the category feature and the name feature to achieve target detection; wherein obtaining the candidate box, the category feature and the name feature based on the first mapping feature map, the second mapping feature map and the target name text feature comprises: Inputting the first mapping feature graph, the second mapping feature graph and the target name text feature into a large language model to obtain a first output result, a second output result and a third output result; The candidate box, the category feature, and the name feature are respectively acquired based on the first output result, the second output result, and the third output result.

8. An electronic device, characterized in that: The electronic device comprises: a processor and a memory; The memory is used to store computer programs; The processor is used to execute the computer program stored in the memory so that the electronic device performs the target detection method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by an electronic device, the target detection method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Non-rigid multiscale object detection method based on convolutional neural network

    CN107818302A

  • Image classification method and device, computer equipment and storage medium

    CN117011577A