A target detection method, device, electronic device and storage medium

By using the similarity calculation of text features and regional features in the object detection method, the problem of inconsistent classification and positioning results in the prior art is solved, and the accuracy and detection effect of object detection are improved.

CN117974971BActive Publication Date: 2025-06-27BEIJING PACTERA JINXIN TECH LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311790486.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-22
Publication Date
2025-06-27
Estimated Expiration
2043-12-22

AI Technical Summary

Technical Problem

Due to the inconsistent classification and positioning results of the existing target detection methods, the target detection accuracy is low and the detection effect is poor.

Method used

By obtaining the target image and target category information, determine the text features and feature maps, input the target image into the pre-trained target detection model, determine the target prediction border, and determine the regional features based on the border and feature map, calculate the similarity between the text features and the regional features, and select the border corresponding to the area features with the highest similarity as the target detection result.

Benefits of technology

It improves the consistency of category prediction and border prediction during the target detection process, thereby improving the accuracy and detection effect of target detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117974971B_ABST
    Figure CN117974971B_ABST
Patent Text Reader

Abstract

The present disclosure provides an object detection method, apparatus, electronic device, and storage medium. By obtaining a target image and target category information; determining the text features corresponding to the target category information and the feature map corresponding to the target image; inputting the target image into a pre-trained target detection model to determine a target prediction bounding box; determining the region features corresponding to the target prediction bounding box according to the target prediction bounding box and the feature map; determining the similarity between the text features and the region features, and using the target prediction bounding box corresponding to the region feature with the highest similarity as the object detection result. It can improve the consistency of category prediction and bounding box prediction in the object detection process, thereby improving the accuracy and detection effect of object detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technologies, and in particular, to an object detection method, apparatus, electronic device, and storage medium. Background Art

[0002] Object detection is a basic visual task. The object detection task is to find the objects of interest in an image or video and simultaneously detect their positions and sizes. Different from the image classification task, object detection not only has to solve the classification problem but also the positioning problem. Existing object detection methods are generally divided into single-stage object detection methods and two-stage object detection methods. The two-stage object detection method has better results but slower speed; the single-stage object detection method is significantly faster than the two-stage object detection method, but the detection effect is worse than that of the two-stage object detection method.

[0003] Currently, both types of methods follow the paradigm of backbone feature extraction, neck feature fusion, and head class and bounding box prediction. In single-stage object detection methods, such as yolox, the decoupled head idea is proposed and achieves impressive results on open-source datasets. Subsequently proposed yolov6, yolov7, and yolov8 models all follow the decoupled head idea. Although the decoupled head can improve the classification and positioning effects, there will be a problem of inconsistent classification and positioning results, where the bounding box effect of the correctly predicted class is poor, and the class prediction effect of the accurately predicted bounding box is poor. That is, due to the problem of inconsistent classification and positioning results in existing object detection methods, there are still problems of low object detection accuracy and poor detection effect. Summary of the Invention

[0004] The embodiments of the present disclosure at least provide an object detection method, apparatus, electronic device, and storage medium, which can improve the consistency between class prediction and bounding box prediction in the object detection process, and thereby improve the accuracy and detection effect of object detection.

[0005] The embodiments of the present disclosure provide an object detection method, including:

[0006] Obtain a target image and target category information;

[0007] Determine the text feature corresponding to the target category information and the feature map corresponding to the target image;

[0008] Input the target image into a pre-trained object detection model to determine a target prediction bounding box;

[0009] Determine the region feature corresponding to the target prediction bounding box according to the target prediction bounding box and the feature map;

[0010] Determine the similarity between the text feature and the region feature, and use the target prediction bounding box corresponding to the region feature with the highest similarity as the target detection result.

[0011] In an alternative implementation, the target detection model is trained based on the following steps:

[0012] Obtain a sample training image and the sample class label corresponding to the sample training image;

[0013] Input the sample training image and the preset sample target class information into a preset target detection model, and divide the positive samples and negative samples according to the distribution result of the predicted labels;

[0014] Determine the positive sample class text feature and the positive sample region feature corresponding to the positive sample;

[0015] Determine an auxiliary loss value according to the sample similarity between the positive sample region feature and the positive sample class text feature;

[0016] Construct a target loss function according to the auxiliary loss value, and supervise the training of the target detection model according to the target loss function.

[0017] In an alternative implementation, the positive sample region feature is determined based on the following steps:

[0018] Input the sample training image into a preset vision-language model to determine the sample feature map corresponding to the sample training image;

[0019] Perform a cropping operation on the sample feature map according to the bounding box prediction result corresponding to the positive sample to determine the bounding box region feature corresponding to the prediction bounding box in each positive sample;

[0020] Align the dimensions of each bounding box region feature to determine the positive sample region feature.

[0021] In an alternative implementation, the positive sample class text feature is determined based on the following steps:

[0022] Input the target class information into a preset vision-language model to extract the class name text feature corresponding to each target class;

[0023] Construct a text feature vector dictionary with all the class name text features;

[0024] Query the class name text feature corresponding to the class prediction result in the text feature vector dictionary according to the class prediction result corresponding to the positive sample, and use it as the positive sample class text feature.

[0025] In an alternative embodiment, the target loss function is constructed based on the following formula:

[0026] L = β1L bbox + β2L cls + β3L aux

[0027] where L represents the target loss function; L bbox represents the CIOU loss function; L cls represents the FocalLoss loss function; L aux represents the auxiliary loss value; β1 represents the loss weight corresponding to the CIOU loss function; β2 represents the loss weight corresponding to the Focal Loss loss function; β3 represents the loss weight corresponding to the auxiliary loss value.

[0028] In an alternative embodiment, the text features are extracted by the preset vision-language model;

[0029] where the preset vision-language model used to extract the text features of the class name is the same as the preset vision-language model used to extract the text features.

[0030] The embodiments of the present disclosure further provide an object detection device, including:

[0031] An acquisition module, configured to acquire a target image and target category information;

[0032] A text feature and feature map determination module, configured to determine the text features corresponding to the target category information and the feature map corresponding to the target image;

[0033] A target prediction module, configured to input the target image into a pre-trained object detection model to determine a target prediction bounding box;

[0034] A region feature determination module, configured to determine the region features corresponding to the target prediction bounding box according to the target prediction bounding box and the feature map;

[0035] A detection result determination module, configured to determine the similarity between the text features and the region features, and use the target prediction bounding box corresponding to the region feature with the highest similarity as the object detection result.

[0036] The embodiments of the present disclosure further provide an electronic device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the above object detection method, or the steps in any possible implementation manner of the above object detection method, are executed.

[0037] An embodiment of the present disclosure also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the above-mentioned object detection method or the steps in any possible implementation manner of the above-mentioned object detection method.

[0038] An embodiment of the present disclosure also provides a computer program product, including a computer program / instructions. When the computer program and instructions are executed by a processor, they implement the above-mentioned object detection method or the steps in any possible implementation manner of the above-mentioned object detection method.

[0039] An object detection method, device, electronic device and storage medium provided by an embodiment of the present disclosure include: obtaining an object image and object category information; determining text features corresponding to the object category information and a feature map corresponding to the object image; inputting the object image into a pre-trained object detection model to determine an object prediction bounding box; determining region features corresponding to the object prediction bounding box according to the object prediction bounding box and the feature map; determining the similarity between the text features and the region features, and using the object prediction bounding box corresponding to the region feature with the highest similarity as the object detection result. This can improve the consistency of category prediction and bounding box prediction in the object detection process, and further improve the accuracy and detection effect of object detection.

[0040] To make the above objects, features, and advantages of the present disclosure more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, provides detailed descriptions as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following briefly introduces the accompanying drawings required for the embodiments. The accompanying drawings are incorporated into the specification and constitute a part of this specification. These drawings show embodiments consistent with the present disclosure and, together with the specification, are used to illustrate the technical solutions of the present disclosure. It should be understood that the following drawings only show some embodiments of the present disclosure and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0042] Figure 1 Shows a flowchart of an object detection method provided by an embodiment of the present disclosure;

[0043] Figure 2 Shows a flowchart of a training method for an object detection model provided by an embodiment of the present disclosure;

[0044] Figure 3The figure shows a schematic diagram of an object detection device provided by an embodiment of the present disclosure;

[0045] Figure 4 The figure shows a schematic diagram of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners

[0046] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only some of the embodiments of the present disclosure, rather than all the embodiments. Usually, the components of the embodiments of the present disclosure described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present disclosure provided in the accompanying drawings is not intended to limit the scope of the present disclosure to be protected, but merely represents the selected embodiments of the present disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.

[0047] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0048] The term "and / or" in this article merely describes an association relationship and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the term "at least one" in this article means any one of a plurality or any combination of at least two of a plurality. For example, including at least one of A, B, and C may represent any one or more elements selected from the set composed of A, B, and C.

[0049] Through research, it is found that the existing object detection methods are roughly divided into single-stage object detection methods and two-stage object detection methods. Both types of methods follow the paradigm of backbone feature extraction, neck feature fusion, and head class and bounding box prediction. In single-stage object detection methods, such as yolox, the decoupled head idea is proposed and impressive results are achieved on open-source datasets. Subsequently, the proposed yolov6, yolov7, and yolov8 models all follow the decoupled head idea. Although the decoupled head can improve the classification and localization effects, there will be problems with inconsistent classification and localization results. The bounding box effect of the correctly predicted class is poor, while the classification prediction effect of the accurately predicted bounding box is poor. That is, due to the problem of inconsistent classification and localization results in the existing object detection methods, there are still problems of low object detection accuracy and poor detection effects.

[0050] Based on the above research, the present disclosure provides an object detection method, apparatus, electronic device, and storage medium. By obtaining a target image and target category information; determining the text features corresponding to the target category information and the feature map corresponding to the target image; inputting the target image into a pre-trained object detection model to determine a target prediction bounding box; determining the region features corresponding to the target prediction bounding box according to the target prediction bounding box and the feature map; determining the similarity between the text features and the region features, and using the target prediction bounding box corresponding to the region feature with the highest similarity as the object detection result. It is possible to improve the consistency of category prediction and bounding box prediction in the object detection process, thereby improving the accuracy and detection effect of object detection.

[0051] For ease of understanding of this embodiment, first, a detailed introduction to an object detection method disclosed in the embodiments of the present disclosure is provided. The execution subject of the object detection method provided in the embodiments of the present disclosure is generally a computer device with certain computing capabilities. Such a computer device includes, for example: a terminal device, a server, or other processing devices. The terminal device may be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. In some possible implementation manners, the object detection method may be implemented by a processor invoking computer-readable instructions stored in a memory.

[0052] See Figure 1 As shown, it is a flowchart of an object detection provided by an embodiment of the present disclosure. The method includes steps S101 to S105, where:

[0053] S101. Obtain a target image and target category information.

[0054] In a specific implementation, obtain a target image that needs to be processed for object detection, and target category information corresponding to the image content of interest to the user that needs to be detected in the target image.

[0055] Here, the target category information may be the category name of the content category of different regions obtained by segmenting the image content of the target image.

[0056] Exemplarily, the target image can be a street view image, a landscape image, a surveillance video image, an industrial product image, etc. For target images of the street view image and surveillance video image types, the target category information that the user is interested in can be image contents such as people, vehicles, buildings, etc.; for target images of the landscape image type, the target category information that the user is interested in can be image contents such as mountains, rivers, trees, etc.; for target images of the industrial product image type, the target category information that the user is interested in can be image contents such as various structural components and appearance components of industrial products.

[0057] S102. Determine the text features corresponding to the target category information and the feature map corresponding to the target image.

[0058] In a specific implementation, a vision-language model is used to extract the text features of the category name corresponding to the target category information and the feature map corresponding to the target image.

[0059] Here, the vision-language model can adopt the clip model. The target image is input into the clip model to extract the feature map corresponding to the target image; the target category information is input into the clip model to extract the text features of the category name corresponding to the target category information.

[0060] Among them, the text feature is the text feature vector of the category name corresponding to the target category information. Exemplarily, for target images of the street view image and surveillance video image types, the category names corresponding to the target category information can be names such as pedestrians, cars, bicycles, buses, etc.; for target images of the landscape image type, the category names corresponding to the target category information can be names such as mountains, rivers, trees, etc.; for target images of the industrial product image type, the category names corresponding to the target category information can be names such as bolts, nuts, crossbeams, shells, etc.

[0061] S103. Input the target image into a pre-trained target detection model to determine the target prediction bounding box.

[0062] In a specific implementation, the target image is input into a pre-trained target detection model. The target detection model detects the image content of the target image according to the target category information to predict the image content of the corresponding target category information that appears in the target image, and marks the target prediction bounding box for identifying the position and size of the image content corresponding to the target category information.

[0063] Here, the target detection model is used to implement the target detection task. According to the input target image and target category information, it identifies the image content of the target image, classifies the image content into image regions of different content categories, and uses the prediction bounding box to mark the image region that matches the target category information.

[0064] Specifically, for the training method of the target detection model, reference can be made to Figure 2 As shown in Figure 2 , it is a flowchart of a training method for a target detection model provided by an embodiment of the present disclosure. The method includes steps S201 to S205, where:

[0065] S201. Obtain sample training images and sample class labels corresponding to the sample training images.

[0066] S202. Input the sample training images and preset sample target category information into a preset target detection model, and divide positive samples and negative samples according to the assignment result of the prediction labels.

[0067] S203. Determine the category text features and sample prediction region features corresponding to the positive samples.

[0068] S204. Determine the auxiliary loss value according to the sample similarity between the positive sample region features and the positive sample category text features.

[0069] S205. Construct a target loss function according to the auxiliary loss value, and supervise the training of the target detection model according to the target loss function.

[0070] In a specific implementation, first input the sample training images into a preset auxiliary target detection model. The preset auxiliary target detection model assigns labels to the image content corresponding to the preset sample target category information and the prediction bounding boxes indicating the positions in the sample training images, and divides positive samples and negative samples according to the label assignment results.

[0071] Here, positive and negative samples are for the prediction results of target detection. Since the target detection model will give multiple prediction results for a single image, that is, the number of prediction results is often more than the number of targets in the image, it is necessary to determine which prediction results are reasonable, that is, have a high degree of coincidence with the true bounding boxes and categories. The reasonable prediction results are positive samples, and the unreasonable prediction results are negative samples.

[0072] Exemplarily, the preset auxiliary target detection model can assign the label "1" to the sample training images containing the image content corresponding to the preset sample target category information, and assign the label "0" to the sample training images without the image content corresponding to the preset sample target category information. Then, the sample training images with the label "1" are positive samples, and the sample training images with the label "0" are negative samples.

[0073] Among them, the preset auxiliary target detection model can be a target detection model based on the decoupled head idea, such as the yolox model.

[0074] Further, input the sample training image into a preset vision-language model to determine the sample feature map corresponding to the sample training image; perform feature clipping operations on the positive samples and the sample feature map, and determine the feature of the border region corresponding to the predicted border in each positive sample according to the border prediction result corresponding to the positive sample; align the dimensions of each border region feature to determine the sample prediction region feature.

[0075] Here, the process of feature clipping includes clipping operations and the processing of the SPPF module. The clipping operation clips the region feature corresponding to each predicted border in the entire sample training image, and the SPPF module aligns the dimensions of each region feature to obtain the final sample prediction region feature.

[0076] Further, input the target category information into a preset vision-language model to extract the text feature of the category name corresponding to each target category; construct a text feature vector dictionary from all the text features of the category names; according to the category prediction result corresponding to the positive sample, query the text feature of the category name corresponding to the category prediction result in the text feature vector dictionary as the sample category text feature.

[0077] Here, the true category name corresponding to the label of the training set, that is, the sample category label, is input into the preset vision-language model, and the text feature corresponding to the true category name is extracted as the text feature of the category name. A text feature vector dictionary is constructed from all the text features of the category names. The positive sample and the text feature vector dictionary are input into the query module. Based on the category prediction result obtained after the positive sample is processed in the preset auxiliary object detection model, the query module queries the corresponding sample category text feature in the text feature vector dictionary.

[0078] It should be noted that the preset vision-language model used to extract the text feature corresponding to the target category information and the feature map corresponding to the target image in step S102 needs to be the same as the preset vision-language model used to extract the sample category text feature and the sample prediction region feature.

[0079] Further, select a similarity calculation function to calculate the sample similarity between the sample prediction region feature and the sample category text feature, and use this sample similarity as the loss metric standard for the positive sample. Combine the loss metric standards corresponding to the border prediction task and the category classification task to construct an objective loss function for supervising the training of the object detection model.

[0080] It should be noted that the similarity calculation function can be set as needed and is not specifically limited here. Exemplarily, it can be the cosine similarity criterion, the l1-norm-like similarity criterion, etc.

[0081] Among them, for the cosine similarity criterion, the 1-sample similarity is used as the auxiliary loss value for each positive sample; for the l1-norm similarity criterion, the sample similarity is used as the auxiliary loss value for each positive sample.

[0082] Here, for the bounding box prediction task, the CIOU loss function is used as the loss metric; for the class classification task, the Focal Loss function is used as the loss metric; the corresponding loss weights are configured for the CIOU loss function, the Focal Loss function, and the auxiliary loss value respectively; according to the CIOU loss function, the Focal Loss function, the auxiliary loss value, and the corresponding loss weights, the target loss function is determined.

[0083] In a specific implementation, the Focal Loss function is selected to solve the potential imbalance problem between class detection and bounding box detection, and the target loss function is constructed based on the following formula:

[0084] L = β1L bbox + β2L cls + β3L aux

[0085] Among them, L represents the target loss function; L bbox represents the CIOU loss function; L cls represents the Focal Loss function; L aux represents the auxiliary loss value; β1 represents the loss weight corresponding to the CIOU loss function; β2 represents the loss weight corresponding to the Focal Loss function; β3 represents the loss weight corresponding to the auxiliary loss value.

[0086] It should be noted that the loss weights corresponding to the CIOU loss function, the Focal Loss function, and the auxiliary loss value are used to balance the contributions of each loss value to the final loss, and can be set as needed, and no specific restrictions are made here.

[0087] S104. Determine the region feature corresponding to the target prediction bounding box according to the target prediction bounding box and the feature map.

[0088] In a specific implementation, the target prediction bounding box and the feature map output by the target detection model are sent to a cropping module for feature cropping operation to determine the bounding box region feature corresponding to the target prediction bounding box; the dimensions of each bounding box region feature are aligned to determine the region feature corresponding to the target prediction bounding box.

[0089] Here, the process of feature cropping is the same as the processing of the sample training image during the training process of the target detection model, and will not be elaborated here.

[0090] S105. Determine the similarity between the text feature and the region feature, and use the target prediction bounding box corresponding to the region feature with the highest similarity as the target detection result.

[0091] In a specific implementation, a similarity calculation function such as the cosine similarity criterion is selected to calculate the similarity between the text feature of the class name and the region feature, and the bounding box corresponding to the region feature with the highest similarity is selected as the prediction result of the target class.

[0092] A target detection method provided by an embodiment of the present disclosure includes: obtaining a target image and target class information; determining the text feature corresponding to the target class information and the feature map corresponding to the target image; inputting the target image into a pre-trained target detection model to determine a target prediction bounding box; determining the region feature corresponding to the target prediction bounding box according to the target prediction bounding box and the feature map; determining the similarity between the text feature and the region feature, and using the target prediction bounding box corresponding to the region feature with the highest similarity as the target detection result. This can improve the consistency between class prediction and bounding box prediction in the target detection process, and further improve the accuracy and detection effect of target detection.

[0093] Those skilled in the art can understand that in the above method of the specific implementation manner, the writing order of each step does not mean a strict execution order that constitutes any limitation to the implementation process, and the specific execution order of each step should be determined according to its function and possible internal logic.

[0094] Based on the same inventive concept, an embodiment of the present disclosure also provides a target detection device corresponding to the target detection method. Since the principle of solving problems by the device in the embodiment of the present disclosure is similar to the above target detection method in the embodiment of the present disclosure, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.

[0095] Please refer to Figure 3 , Figure 3 which is a schematic diagram of a target detection device provided by an embodiment of the present disclosure. As shown in Figure 3 the target detection device 300 provided by an embodiment of the present disclosure includes:

[0096] An acquisition module 310, configured to acquire a target image and target class information.

[0097] A text feature and feature map determination module 320, configured to determine the text feature corresponding to the target class information and the feature map corresponding to the target image.

[0098] A target prediction module 330, configured to input the target image into a pre-trained target detection model to determine a target prediction bounding box.

[0099] The region feature determination module 340 is configured to determine the region feature corresponding to the target prediction bounding box according to the target prediction bounding box and the feature map.

[0100] The detection result determination module 350 is configured to determine the similarity between the text feature and the region feature, and use the target prediction bounding box corresponding to the region feature with the highest similarity as the target detection result.

[0101] Descriptions of the processing flows of the modules in the device and the interaction flows between the modules may refer to the relevant descriptions in the foregoing method embodiments, and will not be elaborated here.

[0102] A target detection device provided by an embodiment of the present disclosure obtains a target image and target category information; determines the text feature corresponding to the target category information and the feature map corresponding to the target image; inputs the target image into a pre-trained target detection model to determine a target prediction bounding box; determines the region feature corresponding to the target prediction bounding box according to the target prediction bounding box and the feature map; determines the similarity between the text feature and the region feature, and uses the target prediction bounding box corresponding to the region feature with the highest similarity as the target detection result. It can improve the consistency of category prediction and bounding box prediction in the target detection process, and further improve the accuracy and detection effect of target detection.

[0103] Corresponding to Figure 1 the target detection method in, an embodiment of the present disclosure further provides an electronic device 400, as Figure 4 shown, which is a schematic structural diagram of the electronic device 400 provided by an embodiment of the present disclosure, including:

[0104] A processor 41, a memory 42, and a bus 43; the memory 42 is used to store execution instructions, including an internal memory 421 and an external memory 422; the internal memory 421 here is also called the main memory, which is used to temporarily store the operation data in the processor 41 and the data exchanged with the external memory 422 such as a hard disk. The processor 41 exchanges data with the external memory 422 through the internal memory 421. When the electronic device 400 runs, the processor 41 communicates with the memory 42 through the bus 43, so that the processor 41 executes Figure 1 the steps of the target detection method in.

[0105] An embodiment of the present disclosure further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the target detection method described in the foregoing method embodiments. Wherein, the storage medium may be a volatile or non-volatile computer-readable storage medium.

[0106] The embodiments of the present disclosure also provide a computer program product, which includes computer instructions. When the computer instructions are executed by a processor, the steps of the object detection method described in the above method embodiments can be executed. For details, refer to the above method embodiments and will not be elaborated here.

[0107] Among them, the above computer program product can be specifically implemented in the form of hardware, software, or a combination thereof. In an alternative embodiment, the computer program product is specifically embodied as a computer storage medium. In another alternative embodiment, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.

[0108] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of the above-described device can refer to the corresponding process in the foregoing method embodiments and will not be elaborated here. In several embodiments provided by the present disclosure, it should be understood that the disclosed device and method can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some communication interfaces. The indirect couplings or communication connections of the devices or units can be electrical, mechanical, or other forms.

[0109] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0110] In addition, in each embodiment of the present disclosure, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0111] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such an understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present disclosure. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.

[0112] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present disclosure, used to illustrate the technical solutions of the present disclosure, rather than limiting them. The protection scope of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present disclosure can still modify the technical solutions recorded in the foregoing embodiments, or can easily conceive of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should all be covered by the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.

Claims

1. A target detection method, characterized in that, Including: Obtain a target image and target category information; Use a preset vision-language model to determine the text features corresponding to the target category information and the feature map corresponding to the target image; Input the target image and the target category information into a pre-trained target detection model to determine a target prediction bounding box; Determine the region features corresponding to the target prediction bounding box according to the target prediction bounding box and the feature map; Determine the similarity between the text features and the region features, and use the target prediction bounding box corresponding to the region feature with the highest similarity as the target detection result; Train the target detection model based on the following steps: Obtain a sample training image and the sample category label corresponding to the sample training image; Input the sample training image and preset sample target category information into a preset target detection model, and divide positive samples and negative samples according to the distribution result of the prediction label; Determine the positive sample category text features and positive sample region features corresponding to the positive samples; Determine an auxiliary loss value according to the similarity between the positive sample region features and the positive sample category text features; Construct a target loss function according to the auxiliary loss value, and supervise the training of the target detection model according to the target loss function; Construct the target loss function based on the following steps: For the bounding box prediction task, use the CIOU loss function as the loss metric standard; For the category classification task, use the Focal Loss function as the loss metric standard; Configure corresponding loss weights for the CIOU loss function, the Focal Loss function, and the auxiliary loss value respectively; Determine the target loss function according to the CIOU loss function, the Focal Loss function, the auxiliary loss value, and the corresponding loss weights.

2. The method according to claim 1, characterized in that, Determine the positive sample region features based on the following steps: Input the sample training image into a preset vision-language model to determine the sample feature map corresponding to the sample training image; Perform a cropping operation on the sample feature map according to the bounding box prediction result corresponding to the positive sample to determine the bounding box region features corresponding to the prediction bounding box in each positive sample; Align the dimensions of each of the bounding box region features to determine the positive sample region features.

3. The method according to claim 1, characterized in that Determine the positive sample category text features based on the following steps: Input the target category information into a preset vision-language model to extract the category name text features corresponding to each target category; Construct a text feature vector dictionary from all the category name text features; Query the category name text features corresponding to the category prediction result in the text feature vector dictionary according to the category prediction result corresponding to the positive sample as the positive sample category text features.

4. The method according to claim 1, wherein Construct the target loss function based on the following formula: Among them, L represents the target loss function; represents the CIOU loss function; represents the FocalLoss loss function; represents the auxiliary loss value; represents the loss weight corresponding to the CIOU loss function; represents the loss weight corresponding to the Focal Loss loss function; represents the loss weight corresponding to the auxiliary loss value.

5. The method according to claim 3, wherein: Extract the text features through the preset vision-language model; Among them, the preset vision-language model used to extract the category name text features is the same as the preset vision-language model used to extract the text features.

6. A target detection device, characterized in that, Including: An acquisition module, configured to acquire a target image and target category information; A text feature and feature map determination module, configured to determine the text features corresponding to the target category information and the feature map corresponding to the target image; A target prediction module, configured to input the target image and the target category information into a pre-trained target detection model to determine a target prediction bounding box; A region feature determination module, configured to determine the region features corresponding to the target prediction bounding box according to the target prediction bounding box and the feature map; A detection result determination module, configured to determine the similarity between the text features and the region features, and use the target prediction bounding box corresponding to the region feature with the highest similarity as the target detection result; The apparatus is further configured to train the target detection model based on the following steps: acquire a sample training image and a sample category label corresponding to the sample training image; Input the sample training image and preset sample target category information into a preset target detection model, and divide positive samples and negative samples according to the assignment result of the prediction label; determine the positive sample category text features and positive sample region features corresponding to the positive samples; Determine an auxiliary loss value according to the similarity between the positive sample region features and the positive sample category text features; construct a target loss function according to the auxiliary loss value, and supervise the training of the target detection model according to the target loss function; The apparatus is further configured to construct the target loss function based on the following steps: for the bounding box prediction task, use the CIOU loss function as the loss metric standard; for the category classification task, use the Focal Loss function as the loss metric standard; respectively configure corresponding loss weights for the CIOU loss function, the Focal Loss function, and the auxiliary loss value; determine the target loss function according to the CIOU loss function, the Focal Loss function, the auxiliary loss value, and the corresponding loss weights.

7. An electronic device, characterized in that, Including: A processor, a memory, and a bus, where the memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the steps of the target detection method according to any one of claims 1 to 5 are executed.

8. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is run by the processor, the steps of the target detection method according to any one of claims 1 to 5 are executed.

Citation Information

Patent Citations

  • Deep learning model training method, target object detection method and device

    CN114882321A

  • Vision and text cross-modal matching method based on consensus embedding space and similarity

    CN115935194A