Object Detection Method, Device, and Storage Medium

By mapping the encoding results of multiple target prompt information to a feature matrix, standardized prompt features are generated, and image features and these features are used for decoding processing to generate target prompt features that match image features, the problem of infinite expansion of the target feature pool in target detection is solved, and more efficient target detection is achieved.

CN119693634BActive Publication Date: 2025-05-27ZHEJIANG DAHUA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510204568.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-05-27
Estimated Expiration
2045-02-24

AI Technical Summary

Technical Problem

In object detection, due to the increase in time, infinitely exhaustive targets and different forms of targets are judged as multiple targets, resulting in infinite expansion of the target feature pool, increasing the time-consuming model training.

Method used

By obtaining the encoding results of multiple target prompt information, it is mapped into a feature matrix to obtain standardized prompt features; then the image features and standardized prompt features of the image to be detected are decoded to obtain weight parameters matching the image features, and these weight parameters are used to weight calculations for the standardized prompt features to generate target prompt features matching the image features, and finally object detection is performed based on these features.

Benefits of technology

Through this method, not only the generalization and stability of object detection are improved, but also the infinite increase in prompt features is avoided, storage resources and computing resources are saved, and model training costs are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119693634B_ABST
    Figure CN119693634B_ABST
Patent Text Reader

Abstract

The present application discloses an object detection method, device, and storage medium. The object detection method includes: obtaining target prompt features corresponding to a plurality of target prompt messages and mapping them into a feature matrix to obtain standardized prompt features corresponding to the unified target prompt features; obtaining an encoded result corresponding to the image to be detected to obtain image features, performing decoding processing on the image features and the standardized prompt features to obtain weight parameters matching the image features; using the weight parameters to perform weighted calculation on the standardized prompt features, and taking the calculation result as the target prompt features matching the image features; performing object detection on the image to be detected based on the image features and the target prompt features matching the image features to obtain an object detection result corresponding to the image to be detected. It can not only improve the generalization and stability of object detection through target prompt messages, but also avoid the infinite increase of prompt features, save storage resources and computing resources, and reduce the model training cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image processing, and in particular, to an object detection method, device, and storage medium. Background Art

[0002] With the rapid construction of intelligent systems, the reliability, scalability, and generalization of intelligent systems have become an increasingly important part. The application direction of intelligent algorithms has gradually expanded from solving closed-set problems to solving open-set problems.

[0003] In the related technologies of object detection, the open-set task mostly adopts the multi-modal large model solution. For example, align the multi-modal expressions of vision and language texts, and use the generalization of language texts to improve the scalability and generalization of vision algorithms to achieve object detection.

[0004] However, as the targets are infinitely enumerated over time, and different forms of the same type of target may be judged as multiple targets, it will cause the infinite expansion of the target feature pool and increase the time consumption of model training. Summary of the Invention

[0005] To solve the above technical problems, the present application provides at least one object detection method, device, and storage medium.

[0006] The first aspect of the present application provides an object detection method, including: obtaining encoded results corresponding to multiple object prompt messages respectively to obtain multiple object prompt features, mapping the multiple object prompt features into a feature matrix to obtain a standardized prompt feature corresponding to the multiple object prompt features unified; wherein, the object prompt message includes an image and / or text; obtaining an encoded result corresponding to the image to be detected to obtain an image feature, decoding the image feature and the standardized prompt feature to obtain a weight parameter matching the image feature; using the weight parameter to perform weighted calculation on the standardized prompt feature, and taking the calculation result as an object prompt feature matching the image feature; performing object detection on the image to be detected based on the image feature and the object prompt feature matching the image feature to obtain an object detection result corresponding to the image to be detected.

[0007] In one embodiment, decoding processing is performed on the image feature and the normalized prompt feature to obtain weight parameters matching the image feature, including: performing decoding processing on the image feature and the normalized prompt feature to obtain the weight parameter of each matrix element in the normalized prompt feature relative to each matrix element in the feature matrix to be solved; using the weight parameter to perform weighted calculation on the normalized prompt feature, and taking the calculation result as the target prompt feature matching the image feature, including: based on the weight parameter of each matrix element in the normalized prompt feature relative to the matrix element in the feature matrix to be solved, performing weighted summation calculation on the value of each matrix element in the normalized prompt feature, and taking the calculation result as the value of the matrix element in the feature matrix to be solved; combining the values of each matrix element in the feature matrix to be solved to obtain the target prompt feature matching the image feature.

[0008] In one embodiment, a feature normalization network and a normalization decoder are pre-trained; mapping a plurality of target prompt features to a feature matrix to obtain a normalized prompt feature uniformly corresponding to the plurality of target prompt features, including: inputting the plurality of target prompt features into the feature normalization network to obtain the normalized prompt feature output by the feature normalization network; performing decoding processing on the image feature and the normalized prompt feature to obtain weight parameters matching the image feature, including: inputting the image feature and the normalized prompt feature into the normalization decoder to obtain the weight parameters matching the image feature output by the normalization decoder.

[0009] In one embodiment, a plurality of prompt encoding networks are pre-trained; obtaining a plurality of target prompt features corresponding to the encoding results of a plurality of target prompt information respectively, including: inputting the target prompt information into the plurality of prompt encoding networks respectively to obtain the initial encoding results output by each prompt encoding network respectively; fusing each initial encoding result to obtain the target prompt feature corresponding to the target prompt information.

[0010] In one embodiment, the target prompt information includes a prompt image, and the prompt encoding network includes a multi-modal model and a text encoding model; inputting the target prompt information into the plurality of prompt encoding networks respectively to obtain the initial encoding results output by each prompt encoding network respectively, including: inputting the prompt image into the multi-modal model of the prompt encoding network to obtain the description text of the prompt image extracted by the multi-modal model; inputting the description text into the text encoding model of the prompt encoding network to obtain the text encoding feature output by the text encoding model, and taking the text encoding feature as the initial encoding result output by the encoding network.

[0011] In one embodiment, the method further includes: obtaining the hint accuracy rate corresponding to each target hint information; if the hint accuracy rate of the target hint information is less than or equal to a preset accuracy rate threshold, then taking the target hint information as special target hint information; obtaining the target-related image and / or text corresponding to the special target hint information to obtain additional hint content corresponding to the special target hint information; performing target detection on the image to be detected based on the image features and the target hint features matched with the image features, and obtaining the target detection result corresponding to the image to be detected, including: performing target detection on the image to be detected based on the image features, the target hint features matched with the image features, and the additional hint content corresponding to the special target hint information, and obtaining the target detection result corresponding to the image to be detected.

[0012] In one embodiment, performing target detection on the image to be detected based on the image features, the target hint features matched with the image features, and the additional hint content corresponding to the special target hint information, and obtaining the target detection result corresponding to the image to be detected, includes: obtaining an encoded result corresponding to the additional hint content to obtain additional hint features; calculating a matching degree between the image features corresponding to the image to be detected and the additional hint features; if the matching degree is greater than a preset matching degree threshold, then performing target detection on the image to be detected based on the image features, the target hint features matched with the image features, and the additional hint features, and obtaining the target detection result corresponding to the image to be detected.

[0013] In one embodiment, a target detection decoder is pre-trained; performing target detection on the image to be detected based on the image features and the target hint features matched with the image features, and obtaining the target detection result corresponding to the image to be detected, includes: obtaining an encoded result of the problem description text of the image to be detected to obtain problem text features; cascading the problem text features, the image features, and the target hint features matched with the image features to obtain cascaded features; inputting the image features and the cascaded features into the target detection decoder, and obtaining the target detection result output by the target detection decoder.

[0014] The second aspect of the present application provides an object detection device, which includes: a feature normalization module, configured to obtain encoded results corresponding to multiple object prompt messages respectively to obtain multiple object prompt features, map the multiple object prompt features into a feature matrix, and obtain a normalized prompt feature corresponding to the multiple object prompt features uniformly; wherein, the object prompt information includes an image and / or text; a feature decoding module, configured to obtain an encoded result corresponding to the image to be detected to obtain an image feature, perform decoding processing on the image feature and the normalized prompt feature, and obtain a weight parameter matching the image feature; a feature calculation module, configured to perform weighted calculation on the normalized prompt feature by using the weight parameter, and use the calculation result as an object prompt feature matching the image feature; an object detection module, configured to perform object detection on the image to be detected based on the image feature and the object prompt feature matching the image feature, and obtain an object detection result corresponding to the image to be detected.

[0015] The third aspect of the present application provides an electronic device, including a memory and a processor, and the processor is configured to execute program instructions stored in the memory to implement the above object detection method.

[0016] The fourth aspect of the present application provides a computer-readable storage medium, on which program instructions are stored, and when the program instructions are executed by a processor, the above object detection method is implemented.

[0017] In the above solution, by obtaining encoded results corresponding to multiple object prompt messages respectively to obtain multiple object prompt features, mapping the multiple object prompt features into a feature matrix, and obtaining a normalized prompt feature corresponding to the multiple object prompt features uniformly; obtaining an encoded result corresponding to the image to be detected to obtain an image feature, performing decoding processing on the image feature and the normalized prompt feature, and obtaining a weight parameter matching the image feature; performing weighted calculation on the normalized prompt feature by using the weight parameter, and using the calculation result as an object prompt feature matching the image feature; performing object detection on the image to be detected based on the image feature and the object prompt feature matching the image feature, and obtaining an object detection result corresponding to the image to be detected, not only can the generalization and stability of object detection be improved through object prompt information, but also the infinite increase of prompt features can be avoided, saving storage resources and computing resources, and reducing the model training cost.

[0018] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present application. Description of the Drawings

[0019] The drawings here are incorporated into the specification and constitute a part of this specification. These drawings show embodiments consistent with the present application and are used together with the specification to explain the technical solutions of the present application.

[0020] Figure 1It is a schematic diagram of the solution implementation environment shown in an exemplary embodiment of the present application;

[0021] Figure 2 It is a flowchart of the object detection method shown in an exemplary embodiment of the present application;

[0022] Figure 3 It is a schematic diagram of feature normalization and feature decoding shown in an exemplary embodiment of the present application;

[0023] Figure 4 It is a schematic diagram of encoding the object prompt information shown in an exemplary embodiment of the present application;

[0024] Figure 5 It is a schematic diagram of object detection shown in an exemplary embodiment of the present application;

[0025] Figure 6 It is a block diagram of the object detection device shown in an exemplary embodiment of the present application;

[0026] Figure 7 It is a schematic diagram of the structure of an electronic device shown in an exemplary embodiment of the present application;

[0027] Figure 8 It is a schematic diagram of the structure of a computer-readable storage medium shown in an exemplary embodiment of the present application. Detailed implementation manners

[0028] The solutions of the embodiments of the present application will be described in detail below with reference to the accompanying drawings of the specification.

[0029] In the following description, specific details such as specific system architectures, interfaces, and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the present application.

[0030] As used herein, the term "and / or" is merely an association information describing associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after. In addition, "multiple" herein means two or more than two. In addition, the term "at least one" herein means any one of multiple or any combination of at least two of multiple. For example, including at least one of A, B, and C can represent including any one or more elements selected from the set composed of A, B, and C.

[0031] The object detection method provided by the embodiments of the present application will be described below.

[0032] Please refer to Figure 1 , Figure 1It is a schematic diagram of the solution implementation environment shown in an exemplary embodiment of the present application. The solution implementation environment may include a terminal 110 and a server 120, and the terminal 110 and the server 120 are communicatively connected to each other.

[0033] The number of terminals 110 may be one or more. The terminal 110 may be a camera, a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart watch, etc., but is not limited thereto.

[0034] The server 120 may be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.

[0035] In one example, the server 120 may perform object detection processing on the to-be-detected image obtained from the terminal 110 to obtain an object detection result, and the server 120 may store the object detection result locally, return it to the terminal 110, or transmit it to other terminals.

[0036] In one example, a client of a target application is installed and run in the terminal 110. For example, the target application may be an application providing an object detection function, and object detection processing is performed on the to-be-detected image based on the target application to obtain an object detection result. The server 120 may be a background server of the target application for providing background services for the client of the target application.

[0037] For the object detection method provided by the embodiments of the present application, the execution subject of each step may be the terminal 110, such as the client of the target application installed and run in the terminal 110, or the server 120, or the terminal 110 and the server 120 cooperate with each other to execute, that is, part of the steps of the method are executed by the terminal 110 and the other part of the steps are executed by the server 120.

[0038] Please refer to Figure 2 , Figure 2 It is a flowchart of the object detection method shown in an exemplary embodiment of the present application. The object detection method may be applied to Figure 1 the implementation environment shown, and is specifically executed by the server in the implementation environment. It should be understood that the method may also be applicable to other exemplary implementation environments and is specifically executed by devices in other implementation environments. The embodiments of the present application do not limit the implementation environments applicable to the method.

[0039] Such asFigure 2 As shown in Figure 2 , the object detection method at least includes steps S210 to S240, which are introduced in detail as follows:

[0040] Step S210: Obtain the encoded results corresponding to multiple object prompt messages respectively to obtain multiple object prompt features, map the multiple object prompt features into a feature matrix, and obtain the standardized prompt features corresponding to the multiple object prompt features uniformly.

[0041] Among them, the object prompt message includes an image containing an object and / or text for describing the object.

[0042] For example, collect the images to be detected corresponding to the object detection results with anomalies in the object detection scenario, such as collecting the images to be detected corresponding to the object detection results with inaccurate recognition feedback by users, and use these images to be detected as the anomaly recognition result images, and obtain the object prompt message according to the anomaly recognition result images. For example, it can be that the feedback content of the user identifies the anomaly cause (such as misrecognition or missed recognition) and the anomaly area of the recognition anomaly, extract the anomaly area from the anomaly recognition result image, and combine the anomaly cause and the anomaly area to obtain the object prompt message; it can also convert the image content of the anomaly area into a text description to obtain the object prompt message; of course, it can also directly use the anomaly recognition result image as the object prompt message, and the present application does not limit this.

[0043] Perform encoding processing on the object prompt message to obtain the object prompt feature of the object prompt message.

[0044] In the actual application scenario, the object prompt messages will gradually increase. If the object prompt features are directly stored, it will cause the infinite expansion of the feature pool, infinitely increase the demand for device storage resources, and increase the time-consuming for subsequent comparison and encoding of object prompt messages, reducing the object recognition efficiency.

[0045] To solve the above problems, the present application maps multiple object prompt features into a feature matrix to obtain the standardized prompt features corresponding to the multiple object prompt features uniformly. The standardized prompt features can subsequently be decoded and estimated based on the image features corresponding to the image to be detected to obtain the object prompt features matching the image to be detected.

[0046] It should be noted that the relevant parameters used in the mapping process need to be determined by iterative training according to the training samples.

[0047] It should be noted that the target prompt feature can be a feature vector obtained by directly encoding the target prompt information; the target prompt feature can also be a standardized prompt feature mapped from a historical time period. For example, when a new target prompt information is received at the current time, the encoding result of the new target prompt information is obtained to get the new target prompt feature, and the previous standardized prompt feature is used as the historical standardized prompt feature. The new target prompt feature and the historical standardized prompt feature are encoded to be mapped into a feature matrix with a fixed size to obtain the current latest standardized prompt feature, which replaces the previous standardized prompt feature.

[0048] Step S220: Obtain the encoding result corresponding to the image to be detected to get the image feature, and perform decoding processing on the image feature and the standardized prompt feature to obtain the weight parameter matching the image feature.

[0049] Among them, the image to be detected can be an image pre-stored in the server or an image uploaded by the terminal.

[0050] For example, the image acquisition device is connected to the server through the network. The image acquisition device collects the image information of the environment in real time and uploads the collected environmental image to the server as the image to be detected.

[0051] Perform encoding processing on the image to be detected to obtain the image feature corresponding to the image to be detected.

[0052] Then, perform decoding processing on the standardized prompt feature according to the image feature to obtain the weight parameter matching the image feature.

[0053] That is, the weight parameter is jointly determined by the network parameter used in the decoding processing and the image feature corresponding to the image to be detected.

[0054] Step S230: Use the weight parameter to perform weighted calculation on the standardized prompt feature, and use the calculation result as the target prompt feature matching the image feature.

[0055] According to the weight parameter matching the image feature obtained by decoding, perform weighted calculation on the standardized prompt feature, which can convert the standardized prompt feature into the target prompt feature matching the image feature.

[0056] Exemplarily, by performing decoding processing on the image feature and the standardized prompt feature, the weight parameter of each matrix element in the standardized prompt feature relative to each matrix element in the feature matrix to be solved is obtained. Among them, the feature matrix to be solved is a matrix with a preset size, which contains multiple matrix elements to be solved.

[0057] For example, the feature matrix to be solved is expressed as , and the standardized prompt feature is expressed as , where n is the number of matrix elements. For the k-th matrix element to be solved in the feature matrix to be solved , the i-th matrix element in the decoded normalized hint feature is obtained Relative matrix element The weight parameter of is expressed as , and so on, decoding to obtain the weight parameter of each matrix element in the normalized hint feature relative to the matrix element .

[0058] Then, based on the weight parameters of each matrix element in the normalized hint feature relative to the matrix elements in the feature matrix to be solved, perform a weighted sum calculation on the values of each matrix element in the normalized hint feature, and use the calculation result as the value of the matrix element in the feature matrix to be solved; combine the values of each matrix element in the feature matrix to be solved to obtain the target hint feature that matches the image feature.

[0059] Specifically, according to the weight parameter of each matrix element in the normalized hint feature relative to the matrix element , the relevant formula for calculating the value of the matrix element can be expressed as the following formula 1:

[0060] (Formula 1)

[0061] Based on Formula 1, calculate the value of each matrix element in the feature matrix to be solved, so as to obtain the target hint feature that matches the image feature according to the calculated value of each matrix element in the feature matrix to be solved.

[0062] It should be noted that the above embodiments are only illustrative. In actual application scenarios, the feature sizes of the feature matrix to be solved and the normalized hint feature can be different, or it can be the weight parameter of the matrix element in the decoded normalized hint feature relative to the corresponding matrix element in the feature matrix to be solved, so as to directly perform a weighted calculation on each matrix element in the normalized hint feature through this weight parameter to obtain the target hint feature that matches the image feature. The specific implementation can be flexibly set according to the adopted network structure, the number of target hint information, the type of image to be detected, etc., and the present application does not limit this.

[0063] Exemplarily, please refer to Figure 3 , Figure 3 is a schematic diagram of feature normalization and feature decoding shown in an exemplary embodiment of the present application. As shown in Figure 3As shown, a feature normalization network and a normalization decoder are pre-trained; multiple target prompt features are input into the feature normalization network to obtain the normalized prompt features output by the feature normalization network; subsequently, by inputting the image features and the normalized prompt features into the normalization decoder, the target prompt features matching the image features output by the normalization decoder are obtained.

[0064] For example, the normalization decoder is represented as , the image features are represented as , and the normalized prompt features are represented as , then according to the image features and the normalized prompt features, the estimated target prompt features decoded are .

[0065] Optionally, the feature normalization network and the normalization decoder are trained using training samples. The training samples include sample target prompt information and sample input images. The model training loss can be obtained by calculating the difference between the actual output of the normalization decoder during training and the sample target prompt information, so as to adjust the parameters of the feature normalization network and the normalization decoder according to the model training loss until the loss converges.

[0066] Step S240: Perform target detection on the image to be detected based on the image features and the target prompt features matching the image features, and obtain the target detection result corresponding to the image to be detected.

[0067] Combining the image features and the target prompt features matching the image features to achieve target detection and obtain the target detection result.

[0068] In this application, multiple target prompt features are encoded into a normalized prompt feature, which can not only improve the generalization and stability of target detection through the target prompt information, but also avoid the infinite increase of prompt features, saving storage resources and computing resources. Moreover, if new target prompt information is added subsequently, only the network for generating the normalized prompt feature and the network for decoding the normalized prompt feature need to be trained and fine-tuned, reducing the model training cost.

[0069] Next, some embodiments of the present application will be described in detail.

[0070] In some embodiments, multiple prompt encoding networks are pre-trained; obtaining the encoding results corresponding to multiple target prompt information respectively in step S210 includes:

[0071] Step S211: Input the target prompt information into the multiple prompt encoding networks respectively to obtain the initial encoding results output by each prompt encoding network respectively.

[0072] It should be noted that the network structures of each prompt encoding network can be different and / or the network parameters can be different.

[0073] Each prompt encoding network respectively performs feature encoding on the input target prompt information, and takes the encoding results respectively output by each prompt encoding network as the initial encoding results.

[0074] Step S212: Fuse each initial encoding result to obtain the target prompt feature corresponding to the target prompt information.

[0075] Exemplarily, taking the target prompt information including a prompt image and the prompt encoding network including a multimodal model and a text encoding model as an example for illustration, please refer to Figure 4 , Figure 4 which is a schematic diagram of encoding the target prompt information shown in an exemplary embodiment of the present application. As Figure 4 shown, the prompt image Q is respectively input into the multimodal models of M prompt encoding networks to obtain the description text of the prompt image extracted by the multimodal models; the description text is respectively input into the text encoding models of M prompt encoding networks to obtain the text encoding features output by the text encoding models, and the finally obtained M text encoding features are used as the initial encoding results output by the encoding network, denoted as the initial encoding result set 。

[0076] Then, each initial encoding result in the initial encoding result set is fused.

[0077] For example, it can be directly adding each initial encoding result to obtain the target prompt feature corresponding to the prompt image Q, denoted as ; it can also be concatenating each initial encoding result, performing dimensionality reduction processing on the concatenated result, such as reducing the feature dimension of the concatenated result through a linear layer, so as to normalize the M initial encoding results into a feature matrix with a fixed length, and obtaining the target prompt feature finally corresponding to the prompt image.

[0078] Among them, the multimodal model and the text encoding model are open-source models or models fine-tuned and trained for specific tasks.

[0079] In the above embodiment, by respectively encoding and fusing the target prompt information through multiple prompt encoding networks, the generalization of the finally extracted target prompt feature can be improved, and the accuracy of subsequent target detection can be improved.

[0080] Optionally, the target prompt information includes misidentified target information and missed identified target information. The misidentified target information and the missed identified target information are divided to respectively obtain a misidentified target information set and a missed identified target information set, and each misidentified target information in the misidentified target information set is uniformly encoded to obtain a positive sample standardized prompt feature Each undetected target information in the undetected target information set is uniformly encoded to obtain a negative sample standardized prompt feature Combined with the image features corresponding to the image to be detected and the positive sample standardized prompt feature Decode to obtain a positive sample target prompt feature that matches the image features. Combine the image features corresponding to the image to be detected and the negative sample standardized prompt feature Decode to obtain a negative sample target prompt feature that matches the image features. Perform target detection based on the image features and the positive sample target prompt feature and negative sample target prompt feature corresponding to the image features to obtain the target detection result corresponding to the image to be detected.

[0081] Among them, in order to distinguish the positive sample target prompt feature and the negative sample target prompt feature during the target detection process, a marking bit can be added at the forefront or the end of the positive sample target prompt feature and the negative sample target prompt feature.

[0082] By constructing standardized prompt features for misdetected target information and undetected target information respectively, the accuracy of target recognition is improved.

[0083] In some embodiments, the method further includes: obtaining the prompt accuracy rate corresponding to each target prompt information; if the prompt accuracy rate of the target prompt information is less than or equal to a preset accuracy rate threshold, then the target prompt information is used as special target prompt information; obtaining the image and / or text associated with the target corresponding to the special target prompt information to obtain additional prompt content corresponding to the special target prompt information.

[0084] Exemplarily, it can be to collect the evaluation feedback of the user on the target detection result within a preset time period, analyze the evaluation feedback, and obtain the accuracy rate corresponding to the target prompt information used for each target detection result within the preset time period.

[0085] If the prompt accuracy rate of the target prompt information is less than or equal to the preset accuracy rate threshold, it indicates that it is still difficult to accurately identify the target corresponding to the target prompt information only relying on this target prompt information. Therefore, the target prompt information is used as special target prompt information, and the image and / or text associated with the target corresponding to the special target prompt information is obtained, such as an image containing the target or text describing the target, and these information are used as additional prompt content corresponding to the special target prompt information.

[0086] It can be to directly store the additional prompt content, or to perform encoding processing on the additional information and store the encoding result for subsequent comparison calculations.

[0087] Further, perform object detection on the image to be detected based on the image features, the target prompt features obtained by image feature matching, and the additional prompt content corresponding to the special target prompt information, to obtain the object detection result corresponding to the image to be detected.

[0088] Specifically, obtain the encoded result corresponding to the additional prompt content to obtain additional prompt features; calculate the matching degree between the image features corresponding to the image to be detected and the additional prompt features; if the matching degree is greater than the preset matching degree threshold, then perform object detection on the image to be detected based on the image features, the target prompt features obtained by image feature matching, and the additional prompt features, to obtain the object detection result corresponding to the image to be detected.

[0089] Exemplarily, it can be to calculate the cosine distance, Euclidean distance, etc. between feature vectors to obtain the matching degree between the image features corresponding to the image to be detected and the additional prompt features.

[0090] If the matching degree is greater than the preset matching degree threshold, then comprehensively use the image features, the target prompt features obtained by image feature matching, and the matching additional prompt features to perform object detection on the image to be detected, to obtain the object detection result corresponding to the image to be detected.

[0091] Among them, the preset matching degree threshold can be preset or flexibly calculated. For example, obtain the prompt accuracy rate of the target prompt information corresponding to each additional prompt feature, and based on the prompt accuracy rate of the target prompt information corresponding to each additional prompt feature, respectively set the preset matching degree threshold corresponding to each additional prompt feature. The prompt accuracy rate and the preset matching degree threshold are positively correlated, that is, the higher the prompt accuracy rate, the higher the preset matching degree threshold, and the lower the prompt accuracy rate, the lower the preset matching degree threshold, so as to improve the object detection accuracy.

[0092] In some embodiments, in addition to selecting the matching additional prompt features for object detection, it is also possible to directly perform object detection based on all the additional prompt features, that is, perform object detection by combining the image features, the target prompt features obtained by image feature matching, and all the additional prompt features. This application does not limit this.

[0093] In some real-time modes, an object detection decoder is pre-trained; perform object detection on the image to be detected based on the image features and the target prompt features obtained by image feature matching, to obtain the object detection result corresponding to the image to be detected, including: obtain the encoded result of the problem description text of the image to be detected to obtain the problem text features; cascade the problem text features, the image features, and the target prompt features obtained by image feature matching to obtain the cascaded features; input the image features and the cascaded features into the object detection decoder to obtain the object detection result output by the object detection decoder.

[0094] Specifically, please refer to Figure 5 ,Figure 5 A schematic diagram of object detection shown in an exemplary embodiment of the present application is as follows Figure 5 As shown, the object detection model includes an image encoder, an object detection decoder, and a prompt encoder. It receives an object detection request carrying the image to be detected and the problem description text. The image to be detected is input into the image encoder to obtain the image features corresponding to the image to be detected. Then, using Figure 3 the network structure shown, the target prompt features matching the image features are obtained. The problem description text, additional prompt content (image or text), and target prompt features are input into the prompt encoder to obtain the problem text features, additional prompt features, and encoded target prompt features. The image features are concatenated with the problem text features, additional prompt features, and encoded target prompt features to obtain concatenated features. The image features and the concatenated features are input into the object detection decoder to obtain the object detection result output by the object detection decoder, which is used to mark the position and type of each object in the image to be detected.

[0095] Among them, the above network structure can be implemented based on Convolutional Neural Networks (CNN), Recurrent Neural Network (RNN), Transformer network, etc., and the present application does not limit this.

[0096] The object detection method provided by the present application obtains multiple target prompt features by obtaining the encoding results corresponding to multiple target prompt information, maps the multiple target prompt features into a feature matrix to obtain the standardized prompt features corresponding to the multiple target prompt features unified; obtains the encoding result corresponding to the image to be detected to obtain the image features, decodes the image features and the standardized prompt features to obtain the weight parameters matching the image features; uses the weight parameters to perform weighted calculation on the standardized prompt features, and takes the calculation result as the target prompt features matching the image features; performs object detection on the image to be detected based on the image features and the target prompt features matching the image features to obtain the object detection result corresponding to the image to be detected, which can not only improve the generalization and stability of object detection through target prompt information, but also avoid the infinite increase of prompt features, save storage resources and computing resources, and reduce the model training cost.

[0097] Figure 6 The block diagram of the object detection device shown in an exemplary embodiment of the present application is as follows Figure 6 As shown, the exemplary object detection device 600 includes

[0098] A feature normalization module 610 is configured to obtain encoded results corresponding to multiple target prompt messages respectively to obtain multiple target prompt features, map the multiple target prompt features into a feature matrix, and obtain a normalized prompt feature corresponding to the multiple target prompt features uniformly; wherein, the target prompt messages include images and / or texts.

[0099] A feature decoding module 620 is configured to obtain an encoded result corresponding to an image to be detected to obtain an image feature, perform decoding processing on the image feature and the normalized prompt feature, and obtain a weight parameter matching the image feature.

[0100] A feature calculation module 630 is configured to perform weighted calculation on the normalized prompt feature by using the weight parameter, and use the calculation result as a target prompt feature matching the image feature.

[0101] A target detection module 640 is configured to perform target detection on the image to be detected based on the image feature and the target prompt feature matching the image feature, and obtain a target detection result corresponding to the image to be detected.

[0102] It should be noted that the target detection device provided in the above embodiment and the target detection method provided in the above embodiment belong to the same concept. The specific manners in which each module and unit perform operations have been described in detail in the method embodiment, and will not be repeated here. In practical applications, the target detection device provided in the above embodiment can, according to needs, allocate the above functions to different functional modules, that is, divide the internal structure of the device into different functional modules to complete all or part of the functions described above. This is not limited here.

[0103] Please refer to Figure 7 , Figure 7 which is a schematic structural diagram of an embodiment of an electronic device according to the present application. The electronic device 700 includes a memory 701 and a processor 702. The processor 702 is configured to execute program instructions stored in the memory 701 to implement the steps in any of the above target detection method embodiments. In a specific implementation scenario, the electronic device 700 may include, but is not limited to: a microcomputer, a server. In addition, the electronic device 700 may also include mobile devices such as a laptop computer, a tablet computer, etc., which are not limited here.

[0104] Specifically, the processor 702 is used to control itself and the memory 701 to implement the steps in any of the above-described target detection method embodiments. The processor 702 may also be referred to as a Central Processing Unit (CPU). The processor 702 may be an integrated circuit chip with signal processing capabilities. The processor 702 may also be a general-purpose processor, a Digital Signal Processor (DSP), an Application-Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 702 may be implemented jointly by integrated circuit chips.

[0105] Please refer to Figure 8 , Figure 8 which is a schematic structural diagram of an embodiment of the computer-readable storage medium of the present application. The computer-readable storage medium 800 stores program instructions 810 that can be run by a processor, and the program instructions 810 are used to implement the steps in any of the above-described target detection method embodiments.

[0106] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0107] The above descriptions of the various embodiments tend to emphasize the differences between the various embodiments. The similarities or similarities between them can be referred to each other. For the sake of brevity, they will not be repeated in this article.

[0108] In several embodiments provided by the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical, or other form.

[0109] In addition, in each embodiment of the present application, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

Claims

1. A target detection method, characterized in that: The method comprises: Obtaining encoding results corresponding to a plurality of target prompt information respectively to obtain a plurality of target prompt features, mapping the plurality of target prompt features to a feature matrix, and obtaining standardized prompt features uniformly corresponding to the plurality of target prompt features; wherein the target prompt information includes an image containing a target and / or text for describing the target; Obtaining the encoding result corresponding to the image to be detected to obtain image features, decoding the image features and the standardized prompt features, and obtaining weight parameters of each matrix element in the standardized prompt features relative to each matrix element in the feature matrix to be solved; Based on the weight parameter of each matrix element in the standardized prompt feature relative to the matrix element in the feature matrix to be solved, a weighted sum calculation is performed on the value of each matrix element in the standardized prompt feature, and the calculation result is used as the value of the matrix element in the feature matrix to be solved; Combining the value of each matrix element in the feature matrix to be solved, obtaining a target prompt feature that matches the image feature; Based on the image features and the target prompt features that match the image features, target detection is performed on the image to be detected to obtain a target detection result corresponding to the image to be detected.

2. The method according to claim 1, characterized in that A feature normalization network and a normalization decoder are pre-trained; the mapping of the plurality of target prompt features to a feature matrix to obtain normalized prompt features uniformly corresponding to the plurality of target prompt features comprises: Inputting the plurality of target prompt features into the feature normalization network to obtain normalized prompt features output by the feature normalization network; The decoding process of the image feature and the standardized prompt feature to obtain the weight parameter of each matrix element in the standardized prompt feature relative to each matrix element in the feature matrix to be solved includes: The image features and the standardized prompt features are input into the standardized decoder to obtain the weight parameters matched with the image features output by the standardized decoder.

3. The method according to claim 1, characterized in that: Pre-trained with multiple cue encoding networks; The step of obtaining the encoding results corresponding to the plurality of target prompt information to obtain a plurality of target prompt features includes: Inputting the target prompt information into a plurality of prompt encoding networks respectively, and obtaining initial encoding results output by each prompt encoding network respectively; Each initial encoding result is fused to obtain a target prompt feature corresponding to the target prompt information.

4. The method according to claim 3, characterized in that The target prompt information includes a prompt image, and the prompt encoding network includes a multimodal model and a text encoding model; the target prompt information is input into a plurality of prompt encoding networks respectively to obtain initial encoding results output by each prompt encoding network respectively, including: Inputting the prompt image into the multimodal model of the prompt encoding network to obtain the description text of the prompt image extracted by the multimodal model; The description text is input into the text encoding model of the prompt encoding network to obtain the text encoding features output by the text encoding model, and the text encoding features are used as the initial encoding results output by the encoding network.

5. The method according to claim 1, characterized in that The method further comprises: Obtain the prompt accuracy corresponding to each target prompt information; If the prompt accuracy of the target prompt information is less than or equal to a preset accuracy threshold, the target prompt information is used as special target prompt information; Acquire an image and / or text associated with the target corresponding to the special target prompt information, and obtain additional prompt content corresponding to the special target prompt information; The performing target detection on the image to be detected based on the image feature and the target prompt feature matched with the image feature to obtain the target detection result corresponding to the image to be detected includes: Based on the image features, the target prompt features matched by the image features and the additional prompt content corresponding to the special target prompt information, target detection is performed on the image to be detected to obtain a target detection result corresponding to the image to be detected.

6. The method according to claim 5, characterized in that The performing target detection on the image to be detected based on the image features, the target prompt features matched by the image features, and the additional prompt content corresponding to the special target prompt information to obtain the target detection result corresponding to the image to be detected includes: Obtaining the encoding result corresponding to the additional prompt content to obtain the additional prompt feature; Calculating the matching degree between the image feature corresponding to the image to be detected and the additional prompt feature; If the matching degree is greater than a preset matching degree threshold, target detection is performed on the image to be detected based on the image features, the target prompt features matched by the image features, and the additional prompt features to obtain a target detection result corresponding to the image to be detected.

7. The method according to claim 1, characterized in that A target detection decoder is pre-trained; the target detection is performed on the image to be detected based on the image feature and the target prompt feature matching the image feature to obtain the target detection result corresponding to the image to be detected, including: Obtaining the encoding result corresponding to the question description text of the image to be detected, and obtaining the question text feature; Cascading the question text feature, the image feature, and the target prompt feature matched with the image feature to obtain a cascade feature; The image features and the cascade features are input into an object detection decoder to obtain an object detection result output by the object detection decoder.

8. An electronic device, characterized in that: The electronic device comprises a memory and a processor, and the processor is used to execute program instructions stored in the memory to implement the steps in the method according to any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program instructions, and the program instructions can be executed by a processor to implement the steps in the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Meter image quality detection method and device and storage medium

    CN116977308A

  • Target detection method and device, electronic equipment and storage medium

    CN117437209A