Target Detection Model Training Method, Target Detection Method, Device and Equipment

By adopting the multi-model-combined object detection model training method to process the fusion of image and text features, the problems of low detection resolution and low accuracy in the prior art are solved, and accurate detection of objects to be detected in the image and identification of existing areas are achieved.

CN119851080BActive Publication Date: 2025-06-24PENG CHENG LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510320886.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-06-24
Estimated Expiration
2045-03-18

AI Technical Summary

Technical Problem

The artificial intelligence models in the prior art have problems with low resolution and limited model performance in image detection, resulting in the inability to accurately detect tiny objects in the image and low detection accuracy.

Method used

The object detection model training method is adopted, including an image encoding model, a feature encoding model and a feature decoding model. By obtaining the target sample area, category and description text of the object to be detected in the sample image, combining the image and text features for fusion processing, the target fusion feature is output to determine the predicted category probability and region.

Benefits of technology

The accurate detection of the object to be detected and its existing area in the image to be detected is realized, and the accuracy and resolution of the detection are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119851080B_ABST
    Figure CN119851080B_ABST
Patent Text Reader

Abstract

The embodiments of the present application disclose a method for training an object detection model, an object detection method, a device, and a device. The object detection model includes an image encoding model, a feature encoding model, and a feature decoding model. By obtaining a target sample region, a target sample category, and a sample description text corresponding to a sample object to be detected in a sample image; inputting the sample image into the image encoding model to output region image features corresponding to each sub-region in the sample image; inputting each region image feature and the sample description text feature into the feature encoding model to output a fusion feature corresponding to each sub-region; inputting the fusion feature corresponding to each sub-region into the feature decoding model to output a target fusion feature corresponding to each sub-region; determining a predicted class probability corresponding to the sample object to be detected and a predicted sample region in the sample image according to each target fusion feature and the sample description text; and iterating the object detection model based on the prediction result and the target result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and particularly relates to a method for training a target detection model, a target detection method, a device and a device. Background Art

[0002] With the development of artificial intelligence technology, target detection of some objects in an image can be performed using an artificial intelligence model.

[0003] In related technologies, often a single artificial intelligence model is used to implement target detection of some objects in an image. However, a single artificial intelligence model has certain limitations. For example, the resolution of the image that a single artificial intelligence model can input is relatively low, resulting in some small objects not being detectable. Additionally, the performance of a single artificial intelligence model is limited, leading to low detection accuracy for target detection of some objects in the image, or even unable to detect. Therefore, there is a technical problem that the artificial intelligence model in related technologies cannot accurately detect objects in the image. Summary of the Invention

[0004] Embodiments of the present application provide a method for training a target detection model, a target detection method, a device and a device. The trained target detection model can accurately detect the to-be-detected object and the existence region of the to-be-detected object in the to-be-detected image.

[0005] To achieve the above object, on the one hand, an embodiment of the present application provides a method for training a target detection model. The target detection model includes an image encoding model, a feature encoding model, and a feature decoding model. The method includes:

[0006] Obtain the target sample region, target sample category, and sample description text corresponding to the to-be-detected sample object in the sample image;

[0007] Input the sample image into the image encoding model, and output the region image feature corresponding to each sub-region in the sample image;

[0008] Determine the sample description text feature corresponding to the sample description text, and input each region image feature and the sample description text feature into the feature encoding model, and output the fusion feature corresponding to each sub-region;

[0009] Input the fusion feature corresponding to each sub-region into the feature decoding model, and output the target fusion feature corresponding to each sub-region;

[0010] Determine the predicted category probability corresponding to the to-be-detected sample object and the predicted sample region in the sample image according to each target fusion feature and the sample description text;

[0011] Determine the loss value corresponding to the target detection model according to the predicted class probability, the target sample class, the predicted sample region, and the target sample region. When the loss value does not meet the preset loss condition, iterate the target detection model, and return the target sample region, the target sample class, and the sample description text corresponding to the sample object to be detected in the acquired sample image until the loss value meets the preset loss condition, and obtain the trained image encoding model, the trained feature encoding model, and the trained feature decoding model.

[0012] To achieve the above object, an embodiment of the present application provides a target detection model training device on the one hand. The target detection model includes an image encoding model, a feature encoding model, and a feature decoding model. The device includes:

[0013] An acquisition module, configured to acquire a target sample region, a target sample class, and a sample description text corresponding to a sample object to be detected in a sample image;

[0014] An input module, configured to input the sample image into the image encoding model and output region image features corresponding to each sub-region in the sample image;

[0015] A determination module, configured to determine sample description text features corresponding to the sample description text, and input each region image feature and the sample description text features into the feature encoding model, and output fusion features corresponding to each sub-region;

[0016] A fusion module, configured to input the fusion features corresponding to each sub-region into the feature decoding model and output target fusion features corresponding to each sub-region;

[0017] An output module, configured to determine a predicted class probability corresponding to the sample object to be detected and a predicted sample region in the sample image according to each target fusion feature and the sample description text;

[0018] An iteration module, configured to determine the loss value corresponding to the target detection model according to the predicted class probability, the target sample class, the predicted sample region, and the target sample region. When the loss value does not meet the preset loss condition, iterate the target detection model, and return the target sample region, the target sample class, and the sample description text corresponding to the sample object to be detected in the acquired sample image until the loss value meets the preset loss condition, and obtain the trained image encoding model, the trained feature encoding model, and the trained feature decoding model.

[0019] In some embodiments, the image encoding model includes a first image encoding sub-model and a second image encoding sub-model; the input module is configured to:

[0020] Input the sample image into the first image encoding sub-model, and output the first regional image features corresponding to each sub-region in the sample image;

[0021] Input the sample image into the second image encoding sub-model, and output the second regional image features corresponding to each sub-region in the sample image;

[0022] Multiply the first regional image features corresponding to each sub-region by the first weight, multiply the second regional image features corresponding to each sub-region by the second weight, and then add them together to obtain the regional image features corresponding to each sub-region.

[0023] In some embodiments, the determination module is configured to:

[0024] Perform text segmentation processing on the sample description text to obtain a plurality of sub-texts;

[0025] Obtain a preset set of category names, and generate an input text sequence according to the preset set of category names and the plurality of sub-texts;

[0026] Input the input text sequence into the pre-trained text encoding model, and output the sample description text features corresponding to the sample description text.

[0027] In some embodiments, the output module is configured to:

[0028] Obtain the word vectors corresponding to each word in the sample description text;

[0029] Add the word vectors corresponding to each word and divide by the number of words in the sample description text to obtain the sentence text features;

[0030] Perform an inner product operation on the sentence text features and each target fusion feature to obtain the probability distribution corresponding to each sub-region belonging to each preset category;

[0031] Determine the predicted category probability corresponding to the sample object to be detected according to the probability distribution, and determine the predicted sample region corresponding to the sample object to be detected in each sub-region according to the predicted category probability.

[0032] In some embodiments, the iteration module includes a first loss sub-module, a second loss sub-module, and a third loss value module;

[0033] The first loss sub-module is configured to determine the classification loss value corresponding to the target detection model according to the predicted category probability and the target sample category;

[0034] The second loss sub-module is used to determine the detection loss value corresponding to the target detection model according to the predicted sample region and the target sample region;

[0035] The third loss sub-module is used to determine the intersection loss value corresponding to the target detection model according to the intersection over union between the predicted sample region and the target sample region;

[0036] Add the classification loss value, the detection loss value, and the intersection loss value to obtain the loss value corresponding to the target detection model.

[0037] In some embodiments, the first loss sub-module is used to:

[0038] Perform normalization processing on the predicted class probabilities to obtain normalized predicted class probabilities;

[0039] Determine the cross-entropy loss value corresponding to the predicted class probabilities and the target sample class according to the normalized predicted class probabilities, the target sample class, and a first preset hyperparameter;

[0040] Determine the first weight value corresponding to the cross-entropy loss value according to the target sample class, the normalized predicted class probabilities, and a second preset hyperparameter;

[0041] Determine the second weight value corresponding to the cross-entropy loss value according to the first preset hyperparameter and the target sample class;

[0042] Multiply the first weight value, the second weight value, and the cross-entropy loss value to obtain the target cross-entropy loss value;

[0043] Determine the classification loss value corresponding to the target detection model according to the total number of sub-regions in the sample image and the target cross-entropy loss value.

[0044] In some embodiments, the second loss sub-module is used to:

[0045] Determine a plurality of first coordinates corresponding to the predicted sample region and a plurality of second coordinates corresponding to the target sample region;

[0046] Perform difference calculation on the associated first coordinates and second coordinates among the plurality of first coordinates and the plurality of second coordinates to obtain a difference calculation result, and determine the absolute value corresponding to the difference calculation result as the detection loss value corresponding to the target detection model.

[0047] In some embodiments, the third loss sub-module is used to:

[0048] Determine the intersection value and the union value between the predicted sample region and the target sample region;

[0049] Divide the intersection value by the union value to obtain the intersection - union ratio, and determine the intersection loss value corresponding to the target detection model according to the intersection - union ratio.

[0050] To achieve the above object, on the one hand, an embodiment of the present application provides a target detection method, which is applied to a trained target detection model. The trained target detection model is a model trained by the target detection model training method provided by the embodiment of the present application. The trained target detection model includes a trained image encoding model, a trained feature encoding model, and a trained feature decoding model. The method includes:

[0051] Obtain a to - be - detected image and a task description text corresponding to a to - be - detected object in the to - be - detected image;

[0052] Input the to - be - detected image into the trained image encoding model, and output the image features corresponding to each sub - region in the to - be - detected image;

[0053] Determine the task description text features corresponding to the task description text, and input each image feature and the task description text features into the trained feature encoding model, and output the first fusion features corresponding to each sub - region;

[0054] Input the first fusion features corresponding to each sub - region into the trained feature decoding model, and output the second fusion features corresponding to each sub - region;

[0055] Determine the existence probability of the to - be - detected object in each sub - region according to each second fusion feature and the task description text;

[0056] Determine the sub - region corresponding to the highest existence probability in the to - be - detected image as the existence region of the to - be - detected object.

[0057] To achieve the above object, on the one hand, an embodiment of the present application provides a target detection device, which is applied to a trained target detection model. The trained target detection model is a model trained by the target detection model training method provided by the embodiment of the present application. The trained target detection model includes a trained image encoding model, a trained feature encoding model, and a trained feature decoding model. The device includes:

[0058] A first acquisition module, configured to obtain a to - be - detected image and a task description text corresponding to a to - be - detected object in the to - be - detected image;

[0059] A first input module, configured to input the to - be - detected image into the trained image encoding model, and output the image features corresponding to each sub - region in the to - be - detected image;

[0060] A second input module, configured to determine task description text features corresponding to the task description text, and input each image feature and the task description text features into a trained feature encoding model, and output first fusion features corresponding to each sub-region;

[0061] A third input module, configured to input the first fusion features corresponding to each sub-region into a trained feature decoding model, and output second fusion features corresponding to each sub-region;

[0062] A first determination module, configured to determine the presence probability of the object to be detected in each sub-region according to each second fusion feature and the task description text;

[0063] A second determination module, configured to determine the sub-region corresponding to the highest presence probability in the image to be detected as the presence region of the object to be detected.

[0064] To achieve the above object, an embodiment of the present application provides a computer-readable storage medium on the one hand. The computer-readable storage medium stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the object detection model training method provided by the embodiment of the present application or the object detection method provided by the embodiment of the present application.

[0065] To achieve the above object, an embodiment of the present application provides a computer device on the one hand, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the object detection model training method provided by the embodiment of the present application or the object detection method provided by the embodiment of the present application is implemented.

[0066] In an embodiment of the present application, the object detection model includes an image encoding model, a feature encoding model, and a feature decoding model. By obtaining a target sample region, a target sample category, and a sample description text corresponding to a sample object to be detected in a sample image; inputting the sample image into the image encoding model to output region image features corresponding to each sub-region in the sample image; determining sample description text features corresponding to the sample description text, and inputting each region image feature and the sample description text features into the feature encoding model to output fusion features corresponding to each sub-region; inputting the fusion features corresponding to each sub-region into the feature decoding model to output target fusion features corresponding to each sub-region; determining a predicted class probability corresponding to the sample object to be detected and a predicted sample region in the sample image according to each target fusion feature and the sample description text; determining a loss value corresponding to the object detection model according to the predicted class probability, the target sample category, the predicted sample region, and the target sample region. When the loss value does not meet the preset loss condition, the object detection model is iterated, and the target sample region, the target sample category, and the sample description text corresponding to the sample object to be detected in the sample image are obtained and returned until the loss value meets the preset loss condition, and a trained image encoding model, a trained feature encoding model, and a trained feature decoding model are obtained.

[0067] Therefore, first determine the target sample region, target sample category, and sample description text corresponding to the sample object to be detected in the sample image, then determine the sample description text features of the sample description text, then input the sample image into the image encoding model to output the region image features corresponding to each sub-region in the sample image, input each region image feature and the sample description text features into the feature encoding model to output the fusion features corresponding to each sub-region. The fusion features contain the correlation between the images of different sub-regions and the sample description text, and can associate and represent two different features in the same feature space. Then input the fusion features corresponding to each sub-region into the feature decoding model to output the target fusion features corresponding to each sub-region. The target fusion features can be understood as being generated for specific downstream tasks. For example, the fusion features are combined with more specific keyword texts of the sample object to generate. Finally, through the target fusion features and the sample description text, the predicted class probability corresponding to the sample object to be detected and the predicted sample region in the sample image can be determined. Then, according to the predicted class probability, target sample category, predicted sample region, and target sample region, the loss value corresponding to the target detection model is determined. When the loss value does not meet the preset loss condition, the target detection model is iterated, and the target sample region, target sample category, and sample description text corresponding to the sample object to be detected in the sample image are obtained and returned until the loss value meets the preset loss condition, and the trained image encoding model, trained feature encoding model, and trained feature decoding model are obtained. The trained target detection model can accurately detect the object to be detected and the existence region of the object to be detected in the image to be detected.

[0068] Other features and advantages of the present application will be described in the following specification, and partly will be obvious from the specification, or will be understood by implementing the present application. The objectives and other advantages of the present application can be achieved and obtained through the structures specifically pointed out in the specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.

[0070] Figure 1 It is a schematic diagram of the system framework corresponding to the target detection model training method and target detection method provided by the embodiments of the present application;

[0071] Figure 2 It is a schematic diagram of the application scenario of the trained target detection model provided by the embodiments of the present application;

[0072] Figure 3 is a schematic flowchart of a target detection model training method provided by an embodiment of the present application;

[0073] Figure 4 is a schematic structural diagram of a target detection model provided by an embodiment of the present application;

[0074] Figure 5 is another schematic flowchart of a target detection model training method provided by an embodiment of the present application;

[0075] Figure 6 is a schematic flowchart of a target detection method provided by an embodiment of the present application;

[0076] Figure 7 is a schematic structural diagram of a target detection model training device provided by an embodiment of the present application;

[0077] Figure 8 is a schematic structural diagram of a target detection model training device provided by an embodiment of the present application;

[0078] Figure 9 is a schematic structural diagram of a computer device provided by an embodiment of the present application. Specific Embodiments

[0079] To enable those skilled in the art of the present technology to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making creative efforts belong to the scope of protection of the present application.

[0080] It should be noted that in each specific embodiment of the present application, when it comes to performing relevant processing based on image data and visual event data, the user's permission or consent will be obtained first, and moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of the present application need to obtain sensitive personal information of the user, the user's separate permission or separate consent will be obtained through methods such as pop-up windows or jumping to a confirmation page. After clearly obtaining the user's separate permission or separate consent, the necessary user-related data for enabling the embodiments of the present application to operate normally will be obtained.

[0081] It should be noted that in some processes described in the specification, claims, and the above-mentioned drawings, a plurality of steps appear in a specific order. However, it should be clearly understood that these steps can be executed not in the order in which they appear herein or in parallel. The step numbers are only used to distinguish different steps, and the numbers themselves do not represent any execution order. In addition, descriptions such as "first", "second", or "target" in this article are used to distinguish similar objects and do not necessarily describe a specific order or sequence.

[0082] The object detection model training method and object detection method provided by the embodiments of the present application relate to the field of artificial intelligence technology. The object detection model training method and object detection method provided by the embodiments of the present application can be applied to a terminal, or to a server side, or can also be software running on a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or can be configured as a server cluster or distributed system composed of multiple physical servers, or can also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the object detection model training method and object detection method, etc., but is not limited to the above forms.

[0083] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer computer devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0084] Before further elaborating on the embodiments of the present application, the nouns and terms involved in the embodiments of the present application are described. The nouns and terms involved in the embodiments of the present application are applicable to the following explanations:

[0085] Attention Mechanism: A technique commonly used in computer science and machine learning that enables a model to be more accurate and efficient when processing sequential data. In traditional neural networks, the output of each neuron depends only on the outputs of all neurons in the previous layer. In the attention mechanism, however, the output of each neuron not only depends on the outputs of all neurons in the previous layer but can also be weighted according to different parts of the input data, that is, different weights are assigned to different parts. This allows the model to focus more on the key information in the input sequence, thereby improving the accuracy and efficiency of the model. In natural language processing, the attention mechanism is often used in tasks such as machine translation, speech recognition, and text summarization. Taking machine translation as an example, the attention mechanism can help the model better focus on the important parts in the source language and the target language during the translation process.

[0086] Normalized exponential function: That is, the Softmax function. In mathematics, especially in probability theory and related fields, the normalized exponential function is actually the gradient logarithm normalization of a finite-term discrete probability distribution. The softmax function can transform a set of numerical values into a probability distribution. In this process, larger numerical values will be transformed into larger probabilities, while smaller numerical values will be transformed into smaller probabilities.

[0087] The above are the explanations of relevant terms in this application. If other terms are involved in the subsequent content, they will be explained later.

[0088] First, describe the technical problems existing in the related technologies:

[0089] With the development of artificial intelligence technology, target detection of some objects in images can be performed using an artificial intelligence model.

[0090] In the related technologies, often a single artificial intelligence model is used to achieve target detection of some objects in images. However, a single artificial intelligence model has certain limitations. For example, the resolution of the images that a single artificial intelligence model can input is relatively low, which may result in some tiny objects not being detected. Additionally, the model performance of a single artificial intelligence model is limited, leading to a low detection accuracy for target detection of some objects in images, or even being unable to detect. Therefore, the artificial intelligence models in the related technologies have the technical problem of being unable to accurately detect objects in images.

[0091] To solve the above technical problems, an object detection model is provided in an embodiment of the present application. The object detection model includes an image encoding model, a feature encoding model, and a feature decoding model. First, the target sample region, target sample category, and sample description text corresponding to the sample object to be detected in the sample image are determined. Then, the sample description text features of the sample description text are determined. Next, the sample image is input into the image encoding model, and the region image features corresponding to each sub-region in the sample image are output. The region image features of each sub-region and the sample description text features are input into the feature encoding model, and the fusion features corresponding to each sub-region are output. The fusion features contain the correlation between the images of different sub-regions and the sample description text, and can associate and represent two different features in the same feature space. Then, the fusion features corresponding to each sub-region are input into the feature decoding model, and the target fusion features corresponding to each sub-region are output. The target fusion features can be understood as being generated for specific downstream tasks. For example, the fusion features are combined with more specific keyword texts of the sample object to generate. Finally, the prediction class probability corresponding to the sample object to be detected and the predicted sample region in the sample image can be determined through the target fusion features and the sample description text. Then, according to the prediction class probability, target sample category, predicted sample region, and target sample region, the loss value corresponding to the object detection model is determined. When the loss value does not meet the preset loss condition, the object detection model is iterated, and the target sample region, target sample category, and sample description text corresponding to the sample object to be detected in the sample image are obtained and returned until the loss value meets the preset loss condition, and the trained image encoding model, trained feature encoding model, and trained feature decoding model are obtained. The trained object detection model can accurately detect the object to be detected and the existence region of the object to be detected in the image to be detected.

[0092] The training process and application process of the object detection model will be explained later.

[0093] Please refer to Figure 1 , Figure 1 which is a schematic diagram of the system framework corresponding to the object detection model training method and object detection method provided in an embodiment of the present application. The object detection model training method and object detection method provided in an embodiment of the present application can be applied to this system framework.

[0094] The terminal 140 or the server 110 can be a device that executes the object detection model training method or the object detection method.

[0095] The terminal 140 includes, but is not limited to, mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, etc. The embodiments of the present application can be applied to various scenarios, including but not limited to object detection, image classification, etc. Additionally, it can be a single device or a collection of multiple devices combined. For example, multiple desktop computers are interconnected through a local area network, sharing a single monitor, etc. to work collaboratively, jointly constituting a terminal 140. The terminal 140 can communicate with the Internet 130 in a wired or wireless manner to exchange data.

[0096] The server 110 refers to a computer system that can provide certain services to the terminal 140. Compared with ordinary terminals 140, the server 110 has higher requirements in terms of stability, security, performance, etc. The server 110 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, as well as big data and artificial intelligence platforms.

[0097] The gateway 120 is also known as an internetwork connector and protocol converter. The gateway realizes network interconnection at the transport layer and is a computer system or device that acts as a conversion function. Between two systems using different communication protocols, data formats, or languages, and even with completely different architectures, the gateway is a translator. At the same time, the gateway can also provide filtering and security functions. Messages sent by the terminal 140 to the server 110 need to be sent to the corresponding server 110 through the gateway 120. Messages sent by the server 110 to the terminal 140 also need to be sent to the corresponding terminal 140 through the gateway 120.

[0098] In the embodiments of the present application, the object detection model trained by the object detection model training method can be applied to various scenarios such as object detection and image classification, without imposing limitations on the scenarios to which the trained object detection model in the present application is applied.

[0099] Please refer to Figure 2 , Figure 2 which is a schematic diagram of the application scenario of the trained object detection model provided by the embodiments of the present application.

[0100] Such as Figure 2As shown, the trained object detection model can be applied to object detection scenarios. For example, in the image to be detected, there are clouds and the sun, and then the task description text corresponding to the image to be detected is obtained. The task description text is: "Please find the sun in the image." Then, the image to be detected and the task description text are input into the trained object detection model. The trained object detection model can find the target object "sun" in the image to be detected according to the task description text, and can generate the bounding box corresponding to the "sun". The area contained within this bounding box is the existence area of the "sun".

[0101] Among them, the object detection model includes an image encoding model, a feature encoding model, and a feature decoding model. Regarding the training method corresponding to the trained object detection model, the following specific method can be adopted for training:

[0102] Obtain the target sample area, target sample category, and sample description text corresponding to the sample object to be detected in the sample image;

[0103] Input the sample image into the image encoding model, and output the regional image features corresponding to each sub-region in the sample image;

[0104] Determine the sample description text features corresponding to the sample description text, and input each regional image feature and the sample description text features into the feature encoding model to output the fusion features corresponding to each sub-region;

[0105] Input the fusion features corresponding to each sub-region into the feature decoding model to output the target fusion features corresponding to each sub-region;

[0106] Determine the predicted class probability corresponding to the sample object to be detected and the predicted sample area in the sample image according to each target fusion feature and the sample description text;

[0107] Determine the loss value corresponding to the object detection model according to the predicted class probability, target sample category, predicted sample area, and target sample area. When the loss value does not meet the preset loss condition, iterate the object detection model, and return to obtain the target sample area, target sample category, and sample description text corresponding to the sample object to be detected in the sample image until the loss value meets the preset loss condition, and obtain the trained image encoding model, trained feature encoding model, and trained feature decoding model.

[0108] To understand in more detail the object detection model training method provided by the embodiments of the present application, please refer to Figure 3 , Figure 3 is a schematic flowchart of the object detection model training method provided by the embodiments of the present application. The object detection model includes an image encoding model, a feature encoding model, and a feature decoding model. This object detection model training method may include the following steps:

[0109] Step 210: Obtain the target sample region, target sample category, and sample description text corresponding to the sample object to be detected in the sample image;

[0110] Step 220: Input the sample image into an image encoding model to output the region image features corresponding to each sub-region in the sample image;

[0111] Step 230: Determine the sample description text features corresponding to the sample description text, and input each region image feature and the sample description text features into a feature encoding model to output the fusion features corresponding to each sub-region;

[0112] Step 240: Input the fusion features corresponding to each sub-region into a feature decoding model to output the target fusion features corresponding to each sub-region;

[0113] Step 250: Determine the predicted class probability corresponding to the sample object to be detected and the predicted sample region in the sample image based on each target fusion feature and the sample description text;

[0114] Step 260: Determine the loss value corresponding to the target detection model based on the predicted class probability, target sample category, predicted sample region, and target sample region. When the loss value does not meet the preset loss condition, iterate the target detection model, and return to obtain the target sample region, target sample category, and sample description text corresponding to the sample object to be detected in the sample image until the loss value meets the preset loss condition, and obtain the trained image encoding model, trained feature encoding model, and trained feature decoding model.

[0115] The following will describe Steps 210 to 260 in detail.

[0116] In Step 210, obtain the target sample region, target sample category, and sample description text corresponding to the sample object to be detected in the sample image.

[0117] Among them, for the training of the target detection model, some sample images are required, and these sample images are used as training data. The sample image can be an image with a resolution of 1024x1024.

[0118] Taking a certain sample image as an example, the target sample region, target sample category, and sample description text corresponding to the sample object to be detected in the sample image can be obtained. In the sample image, a sample object to be detected can be determined in advance. This sample object to be detected occupies a region in the sample image, which is the target sample region. This sample object to be detected corresponds to a target sample category. For example, if the sample object to be detected is a cat, the target sample category is cat. This sample object to be detected corresponds to a sample description text. For example, the sample description text is a relevant description of this sample object to be detected in the sample image, such as the description text: "On a sunny lawn, there is a white cat lying in the sun."

[0119] It should be noted that the sample object to be detected in the sample image, the target sample region, target sample category, and sample description text corresponding to the sample object to be detected are all determined in advance. Subsequently, the target sample region and target sample category can be used as label data.

[0120] In step 220, the sample image is input into the image encoding model, and the region image features corresponding to each sub-region in the sample image are output.

[0121] Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of the object detection model provided by an embodiment of the present application. Among them, the object detection model includes a pre-trained text model, an image encoding model, a feature encoding model, and a feature decoding model. Among them, the image encoding model, feature encoding model, and feature decoding model are models that need to be iterated during the training process of the object detection model.

[0122] The sample image can be input into the image encoding model, and the region image features corresponding to each sub-region in the sample image are output. For example, in the sample image, there are multiple objects. The image encoding model can encode each sub-region where an object is located, thereby obtaining the region image features corresponding to each sub-region.

[0123] In some embodiments, the image encoding model includes a first image encoding sub-model and a second image encoding sub-model; inputting the sample image into the image encoding model and outputting the region image features corresponding to each sub-region in the sample image includes:

[0124] (1.1) Input the sample image into the first image encoding sub-model, and output the first region image features corresponding to each sub-region in the sample image;

[0125] (1.2) Input the sample image into the second image encoding sub-model, and output the second region image features corresponding to each sub-region in the sample image;

[0126] (1.3) Multiply the first regional image feature corresponding to each sub-region by the first weight, multiply the second regional image feature corresponding to each sub-region by the second weight, and then add them together to obtain the regional image feature corresponding to each sub-region.

[0127] Among them, the first image encoding sub-model and the second image encoding sub-model can be models of different types. For example, the first image encoding sub-model is the visual encoding model ViT of the APE model, and the second image encoding sub-model is the visual encoding model ViT-18B of the EVA-CLIP-18B model. The model structures and model parameters of these two models are different.

[0128] The sample image can be input into the first image encoding sub-model to output the first regional image feature corresponding to each sub-region in the sample image. For example, when the sample image is input into the visual encoding model ViT model, the first regional image feature Vape corresponding to each sub-region in the sample image is output.

[0129] The sample image can be input into the second image encoding sub-model to output the second regional image feature corresponding to each sub-region in the sample image. For example, when the sample image is input into the visual encoding model ViT-18B, the first regional image feature Vvit18b corresponding to each sub-region in the sample image is output.

[0130] Finally, multiply the first regional image feature corresponding to each sub-region by the first weight, multiply the second regional image feature corresponding to each sub-region by the second weight, and then add them together to obtain the regional image feature corresponding to each sub-region. For example, the calculation formula is as follows:

[0131] . Among them is the regional image feature, 1-W is the first weight, and W is the second weight. The first weight and the second weight are learnable, that is, during the iteration of the image encoding model, the first weight and the second weight will change with the iteration of the image encoding model.

[0132] As can be seen from the above, in this application, by combining the features respectively output by the first image encoding sub-model and the second image encoding sub-model, the regional image feature corresponding to each sub-region generated in this way can more accurately express the image information in the image corresponding to each sub-region. Compared with the related art that only uses a single visual encoding model to obtain the image features of different regions in the image, the regional image features in this application have richer feature information, which is beneficial to the detection of sample objects by the target detection model.

[0133] In step 230, determine the sample description text features corresponding to the sample description text, and input each regional image feature and the sample description text features into a feature encoding model to output the fusion features corresponding to each sub-region.

[0134] Among them, the sample description text needs to be transformed to generate sample description text features, and the sample description text features are vectors in a certain spatial dimension. Then, input each regional image feature and the sample description text features into the feature encoding model to output the fusion features corresponding to each sub-region.

[0135] In some embodiments, determining the sample description text features corresponding to the sample description text includes:

[0136] (1.1) Perform text segmentation processing on the sample description text to obtain multiple sub-texts;

[0137] (1.2) Obtain a preset set of category names, and generate an input text sequence according to the preset set of category names and the multiple sub-texts;

[0138] (1.3) Input the input text sequence into a pre-trained text encoding model to output the sample description text features corresponding to the sample description text.

[0139] Among them, the sample description text can be subjected to text segmentation processing to obtain multiple sub-texts. For example, sentences can be divided according to full stops, question marks, and exclamation marks. First, find the positions of these punctuation marks in the text, and then extract the content between two punctuation marks as a sub-text.

[0140] For another example, first determine some key words. For example, in the scene description, there are keywords such as "person", "vehicle", and "animal". Then find the positions of these keywords in the text, regard the content before each keyword as a sub-text, and regard the content from the keyword until before the next keyword as another sub-text.

[0141] Then obtain a preset set of category names, and generate an input text sequence according to the preset set of category names and the multiple sub-texts. Among them, the preset set of category names contains multiple preset categories, and the preset categories can be categories corresponding to different people and objects in the scene. For example, the preset categories include "cat", "dog", "child", "man", "woman", "old person", "tree", "flower", etc., and these preset categories can all be collected in the preset set of category names.

[0142] For each preset category name in the preset category name set, it can be used as a category text. These category texts and multiple sub-texts are arranged to generate an input text sequence. For example, each category text and each sub-text are concatenated to form a long text, and this long text is the input text sequence.

[0143] Finally, the input text sequence is input into the pre-trained text encoding model to output the sample description text features corresponding to the sample description text. For example, the pre-trained model can extract features from the input text sequence to generate sample description text features in a certain spatial dimension.

[0144] As can be seen from the above, in this application, by combining the preset category name set and the sub-texts in the sample description text, an input text sequence containing multiple categories is generated. The advantage of this is that the generated sample description text features can contain more information about the preset categories, which is beneficial to the learning of the preset categories by the target detection model.

[0145] After obtaining the sample description text features, each regional image feature and the sample description text features can be input into the feature encoding model to output the fusion features corresponding to each sub-region. Among them, the feature encoding model has a self-attention mechanism. The self-attention mechanism allows the feature encoding model to weigh the importance of each element in the sequence when processing sequence data. For the input sample description text features and each regional image feature, it can simultaneously focus on the information of these two different modalities.

[0146] Specifically, when the sample description text features and each regional image feature are input into the feature encoding model, they will first be converted into vector representations suitable for the feature encoding model to process. These vectors will be concatenated or combined in some way as the input sequence of the feature encoding model. Then the feature encoding model performs self-attention mechanism processing to obtain the corresponding fusion features of the two. This kind of fusion feature contains the semantic association and complementary information between text and image, and can be better used for subsequent tasks such as classification and generation.

[0147] In step 240, the fusion features corresponding to each sub-region are input into the feature decoding model to output the target fusion features corresponding to each sub-region.

[0148] Among them, after the fusion features output by the encoding model are input into the feature decoding model, the target fusion features are more targeted features generated through the decoding process on the basis of the fusion features. It is usually generated for specific downstream tasks. For example, in the target detection task, it can detect the image region of the sample object to be detected that matches the task description text in the sample image.

[0149] In some embodiments, object query features for querying a sample object to be detected can be obtained. The object query features can be task description text features generated based on the task description text. Then, the task description text features and the fusion features corresponding to each sub-region are input into a feature decoding model to output the target fusion features corresponding to each sub-region.

[0150] The object query features are like "clues" for the feature decoding model to search for the sample object to be detected in the sample image. They will guide the feature decoding model to focus on different regions in the image through the fusion features and try to find regions that match the object query features. For example, when faced with the prompt "find the cat in the image", the object query features will prompt the feature decoding model to search for regions with cat features in the sample image. By continuously adjusting the focus, the position and category of the cat can be accurately identified eventually.

[0151] The target fusion features can be understood as the finally output object embeddings. These object embeddings contain rich information, covering key elements such as the features, positions, and categories of various objects in the image, and are an important basis for subsequent alignment of regional images and text and prediction. When processing a complex scene image containing multiple objects, the object embeddings can effectively encode the unique attributes of each object, helping the object detection model distinguish different objects.

[0152] In step 250, the predicted class probability corresponding to the sample object to be detected and the predicted sample region in the sample image are determined according to each target fusion feature and the sample description text.

[0153] Among them, based on the generated target fusion features, the object detection model will further predict a series of key information, including scores, bounding boxes, and masks. The scores are used to evaluate the matching degree of each predicted object with the given prompt, reflecting the confidence of the model in the prediction accuracy; the bounding boxes accurately locate the positions of the target objects in the image, specifying their specific ranges in the image in coordinate form; the masks are used to more precisely segment the contours of the target objects. In instance segmentation and semantic segmentation tasks, the masks can clearly define the boundaries of the objects, which is crucial for processing image scenes where the foreground and background are complexly intertwined.

[0154] For example, when analyzing an urban street scene image containing people, vehicles, and buildings, the object detection model uses the scores obtained by prediction to judge the matching possibility of each detected object with the corresponding prompts (such as "pedestrian", "car", "building"), the bounding boxes accurately frame the positions of each object, and the masks precisely outline their contours.

[0155] In some embodiments, determining the predicted class probability corresponding to the sample object to be detected and the predicted sample region in the sample image according to each target fusion feature and the sample description text includes:

[0156] (1.1) Obtain the word vectors corresponding to each word in the sample description text;

[0157] (1.2) Add up the word vectors corresponding to each word and divide by the number of words in the sample description text to obtain the sentence text feature;

[0158] (1.3) Perform an inner product operation on the sentence text feature and each target fusion feature to obtain the probability distribution corresponding to each sub-region belonging to each preset category;

[0159] (1.4) Determine the predicted category probability corresponding to the sample object to be detected according to the probability distribution, and determine the predicted sample region corresponding to the sample object to be detected in each sub-region according to the predicted category probability.

[0160] Among them, the word vectors corresponding to each word in the sample description text can be obtained, and then the word vectors corresponding to each word are added up and divided by the number of words in the sample description text to obtain the sentence text feature. Among them, the sentence feature can be used to reduce the complexity of subsequent calculations and improve the calculation efficiency. By decomposing the sample description text into independent concepts and aggregating them into sentence-level embeddings, the model can receive hints from a large number of words and sentences in a single forward pass. This way greatly improves the scale of the model hints, enabling it to handle complex and diverse visual tasks. For example, in an image containing many objects and scenes, it can accurately identify various targets.

[0161] Then perform an inner product operation on the sentence text feature and each target fusion feature to obtain the probability distribution corresponding to each sub-region belonging to each preset category. This probability distribution can be understood as the score obtained through the target fusion feature above.

[0162] Determine the predicted category probability corresponding to the sample object to be detected according to the probability distribution, and determine the predicted sample region corresponding to the sample object to be detected in each sub-region according to the predicted category probability. For example, the highest probability distribution (score) can be determined as the predicted category probability, and the sub-region corresponding to the highest probability distribution can be determined as the predicted sample region corresponding to the sample object to be detected, believing that the sample object to be detected is most likely to exist in this predicted sample region. For example, if the sample object to be detected is a "cat", then it is considered that the category corresponding to the predicted sample region is "cat", and the "cat" exists in the predicted sample region. The predicted sample region can be represented by a bounding box. For example, the bounding box is a rectangle in the sample image, and the bounding box has certain coordinate information.

[0163] In step 260, the loss value corresponding to the object detection model is determined based on the predicted class probability, the target sample class, the predicted sample region, and the target sample region. When the loss value does not meet the preset loss condition, the object detection model is iterated, and the target sample region, the target sample class, and the sample description text corresponding to the object to be detected in the acquired sample image are returned until the loss value meets the preset loss condition, and the trained image encoding model, the trained feature encoding model, and the trained feature decoding model are obtained.

[0164] Among them, after obtaining the predicted class probability and the predicted sample region corresponding to the object to be detected, it is necessary to verify the predicted value to determine the difference between the predicted value and the true value. For example, this difference is the loss value, and the object detection model can be determined whether the training is completed by comparing the loss value with the preset loss condition.

[0165] When the loss value meets the preset loss condition, and after subsequent different sample images are input, the loss value corresponding to the object detection model still meets the preset loss condition. At this time, it indicates that the object detection model training is completed, and the trained object detection model is obtained. The trained object detection model includes the trained image encoding model, the trained feature encoding model, and the trained feature decoding model.

[0166] When the loss value does not meet the preset loss condition, the object detection model is iterated, and the target sample region, the target sample class, and the sample description text corresponding to the object to be detected in the acquired sample image are returned until the loss value meets the preset loss condition, and the trained image encoding model, the trained feature encoding model, and the trained feature decoding model are obtained.

[0167] In some embodiments, determining the loss value corresponding to the object detection model based on the predicted class probability, the target sample class, the predicted sample region, and the target sample region includes:

[0168] (1.1) Determine the classification loss value corresponding to the object detection model according to the predicted class probability and the target sample class;

[0169] (1.2) Determine the detection loss value corresponding to the object detection model according to the predicted sample region and the target sample region;

[0170] (1.3) Determine the intersection loss value corresponding to the object detection model according to the intersection over union between the predicted sample region and the target sample region;

[0171] (1.4) Add the classification loss value, the detection loss value, and the intersection loss value to obtain the loss value corresponding to the object detection model.

[0172] Among them, first determine the classification loss value corresponding to the object detection model according to the predicted class probability and the target sample class. The classification loss value is mainly obtained based on Focal Loss and is used to distinguish the foreground area from the background area.

[0173] Then determine the detection loss value corresponding to the object detection model according to the predicted sample area and the target sample area. The detection loss value is mainly used to measure the difference between the predicted sample area and the target sample area, so as to determine the prediction accuracy of the object detection model for the predicted sample area.

[0174] Then determine the intersection loss value corresponding to the object detection model according to the intersection over union (IoU) between the predicted sample area and the target sample area. Among them, the intersection loss value is mainly obtained based on GIOU (Generalized Intersection over Union)-loss and is mainly used to measure the similarity between two shapes (usually bounding boxes, such as in object detection tasks). The predicted sample area and the target sample area can be regarded as areas composed of bounding boxes.

[0175] Finally, add the classification loss value, the detection loss value, and the intersection loss value to obtain the loss value corresponding to the object detection model.

[0176] The advantage of doing this is that the loss value corresponding to the object detection model can be measured from multiple different types of loss dimensions, which can increase the accuracy of determining the loss value of the object detection model and is beneficial to improving the training efficiency of the object detection model.

[0177] In some embodiments, determining the classification loss value corresponding to the object detection model according to the predicted class probability and the target sample class includes:

[0178] (1.1.1) Normalize the predicted class probability to obtain the normalized predicted class probability;

[0179] (1.1.2) Determine the cross-entropy loss value corresponding to the predicted class probability and the target sample class according to the normalized predicted class probability, the target sample class, and the first preset hyperparameter;

[0180] (1.1.3) Determine the first weight value corresponding to the cross-entropy loss value according to the target sample class, the normalized predicted class probability, and the second preset hyperparameter;

[0181] (1.1.4) Determine the second weight value corresponding to the cross-entropy loss value according to the first preset hyperparameter and the target sample class;

[0182] (1.1.5) Multiply the first weight value, the second weight value, and the cross-entropy loss value to obtain the target cross-entropy loss value;

[0183] (1.1.6) Determine the classification loss value corresponding to the object detection model according to the total number of sub-regions in the sample image and the target cross-entropy loss value.

[0184] Among them, first perform normalization processing on the predicted class probabilities to obtain the normalized predicted class probabilities. For example, use the sigmoid function to perform normalization processing on the predicted class probabilities, and the specific calculation formula is as follows:

[0185] . Among them, pred_logits is the predicted class probability, which can be understood as the probability representation of each sub-region belonging to each preset class, and prob is the normalized predicted class probability.

[0186] Then determine the cross-entropy loss value corresponding to the predicted class probability and the target sample class according to the normalized predicted class probability, the target sample class, and the first preset hyperparameter. Specifically, the calculation formula is as follows:

[0187] . Among them is the cross-entropy loss value, targets is the representation corresponding to the target sample class, which can specifically be obtained through the one-hot encoding of the target sample class, is the first preset hyperparameter, and the corresponding value is 0.25.

[0188] Then determine the first weight value corresponding to the cross-entropy loss value according to the target sample class, the normalized predicted class probability, and the second preset hyperparameter. Among them, the first weight value can specifically be expressed as , where is the second preset hyperparameter, and the value can be 2.0. Among them, the intermediate variable The specific calculation method is: .

[0189] Then determine the second weight value corresponding to the cross-entropy loss value according to the first preset hyperparameter and the target sample class. The specific calculation method is: , where is the second weight value.

[0190] Then multiply the first weight value, the second weight value, and the cross-entropy loss value to obtain the target cross-entropy loss value. The specific calculation method is: . loss_l is the target cross-entropy loss value. By the first weight value and the second weight value, the adjustment of the cross-entropy loss can be realized, so as to obtain a more accurate target cross-entropy loss value.

[0191] Finally, determine the classification loss value corresponding to the object detection model according to the total number of sub-regions in the sample image and the target cross-entropy loss value. The specific calculation formula is:

[0192] Among them, num_boxes is the total number of sub-regions in the sample image, and i is the i-th sub-region. is the target cross-entropy loss value corresponding to the i-th sub-region. is the classification loss value.

[0193] The advantage of doing this is that it can provide necessary supervision signals for multi-task learning through the classification loss value, so as to accurately guide the training of the object detection model.

[0194] In some embodiments, the detection loss value corresponding to the object detection model is determined according to the predicted sample region and the target sample region, including:

[0195] (1.2.1) Determine multiple first coordinates corresponding to the predicted sample region and multiple second coordinates corresponding to the target sample region;

[0196] (1.2.2) Calculate the difference between the associated first coordinates and second coordinates among the multiple first coordinates and multiple second coordinates to obtain a difference calculation result, and determine the absolute value corresponding to the difference calculation result as the detection loss value corresponding to the object detection model.

[0197] Among them, the predicted sample region and the target sample region can be understood as the regions formed by two bounding boxes respectively. In the bounding box corresponding to the predicted sample region, multiple first coordinates are obtained. In the bounding box corresponding to the target sample region, multiple second coordinates are obtained.

[0198] Calculate the difference between the associated first coordinates and second coordinates among the multiple first coordinates and multiple second coordinates to obtain a difference calculation result, and determine the absolute value corresponding to the difference calculation result as the detection loss value corresponding to the object detection model.

[0199] For example, the bounding box is a rectangle, which contains four vertices: upper left, lower left, upper right, and lower right. The difference between the first coordinates and the second coordinates corresponding to each vertex can be calculated to obtain a difference calculation result. And the absolute value corresponding to the difference calculation result is determined as the detection loss value corresponding to the object detection model. The specific calculation formula is:

[0200] . Among them, is the detection loss value. is the target sample region. is the predicted sample region.

[0201] The advantage of doing this is that it can determine the detection loss between the predicted sample region and the target sample region to measure the loss value of the object detection model, and improve the accuracy of the determined loss value.

[0202] In some embodiments, determining the intersection loss value corresponding to the object detection model according to the intersection over union (IoU) between the predicted sample region and the target sample region includes:

[0203] (1.3.1) Determine the intersection value and the union value between the predicted sample region and the target sample region;

[0204] (1.3.2) Divide the intersection value by the union value to obtain the IoU, and determine the intersection loss value corresponding to the object detection model according to the IoU.

[0205] Among them, by determining the intersection value and the union value between the predicted sample region and the target sample region, the size of the intersection region between the predicted sample region and the target sample region can be determined, so as to obtain the intersection value, denoted as IntersectionArea. The size of the union region between the predicted sample region and the target sample region can be determined, so as to obtain the intersection value, denoted as UnionArea.

[0206] Then divide the intersection value by the union value to obtain the IoU, and the calculation method is: , where is the IoU.

[0207] Then determine the intersection loss value corresponding to the object detection model according to the IoU, and the specific calculation method is: . Among them, is the intersection loss value.

[0208] Finally, add the classification loss value, the detection loss value and the intersection loss value to obtain the loss value corresponding to the object detection model.

[0209] Among them, when the loss value meets the preset loss condition, and after different sample images are input subsequently, the loss value corresponding to the object detection model still meets the preset loss condition. At this time, it indicates that the object detection model training is completed, and the trained object detection model is obtained. The trained object detection model includes a trained image encoding model, a trained feature encoding model and a trained feature decoding model.

[0210] When the loss value does not meet the preset loss condition, the object detection model is iterated. For example, the parameters of the object detection model are adjusted, and the target sample region, the target sample category and the sample description text corresponding to the object to be detected in the acquired sample image are returned until the loss value meets the preset loss condition, and the trained image encoding model, the trained feature encoding model and the trained feature decoding model are obtained. The trained image encoding model includes a trained first image encoding sub-model and a trained second image encoding sub-model.

[0211] The trained object detection model can accurately detect the object to be detected and the existence region of the object to be detected in the image to be detected.

[0212] In the embodiment of the present application, the object detection model includes an image encoding model, a feature encoding model, and a feature decoding model. By obtaining the target sample region, target sample category, and sample description text corresponding to the sample object to be detected in the sample image; inputting the sample image into the image encoding model to output the region image features corresponding to each sub-region in the sample image; determining the sample description text features corresponding to the sample description text, and inputting each region image feature and the sample description text features into the feature encoding model to output the fusion features corresponding to each sub-region; inputting the fusion features corresponding to each sub-region into the feature decoding model to output the target fusion features corresponding to each sub-region; determining the predicted category probability corresponding to the sample object to be detected and the predicted sample region in the sample image according to each target fusion feature and the sample description text; determining the loss value corresponding to the object detection model according to the predicted category probability, target sample category, predicted sample region, and target sample region. When the loss value does not meet the preset loss condition, the object detection model is iterated, and the target sample region, target sample category, and sample description text corresponding to the sample object to be detected in the sample image are obtained and returned until the loss value meets the preset loss condition, and the trained image encoding model, trained feature encoding model, and trained feature decoding model are obtained.

[0213] Therefore, first determine the target sample region, target sample category, and sample description text corresponding to the sample object to be detected in the sample image. Then, determine the sample description text features of the sample description text. Next, input the sample image into an image encoding model to output the region image features corresponding to each sub-region in the sample image. Input each region image feature and the sample description text features into a feature encoding model to output the fusion features corresponding to each sub-region. The fusion features contain the correlation between the images of different sub-regions and the sample description text, and can associate and represent the two different features in the same feature space. Then, input the fusion features corresponding to each sub-region into a feature decoding model to output the target fusion features corresponding to each sub-region. The target fusion features can be understood as being generated for a specific downstream task. For example, the fusion features are combined with more specific keyword texts of the sample object to generate. Finally, the prediction class probability corresponding to the sample object to be detected and the predicted sample region in the sample image can be determined through the target fusion features and the sample description text. Then, determine the loss value corresponding to the target detection model according to the prediction class probability, target sample category, predicted sample region, and target sample region. When the loss value does not meet the preset loss condition, iterate the target detection model, and return to obtain the target sample region, target sample category, and sample description text corresponding to the sample object to be detected in the sample image until the loss value meets the preset loss condition, and obtain the trained image encoding model, trained feature encoding model, and trained feature decoding model. The trained target detection model can accurately detect the object to be detected and the existence region of the object to be detected in the image to be detected.

[0214] Please refer to Figure 5 , Figure 5 which is another flowchart of the target detection model training method provided by the embodiments of the present application. The target detection model training method may include:

[0215] Step 301, obtain the target sample region, target sample category, and sample description text corresponding to the sample object to be detected in the sample image;

[0216] Step 302, input the sample image into the first image encoding sub-model to output the first region image features corresponding to each sub-region in the sample image;

[0217] Step 303, input the sample image into the second image encoding sub-model to output the second region image features corresponding to each sub-region in the sample image;

[0218] Step 304, multiply the first region image features corresponding to each sub-region by the first weight, multiply the second region image features corresponding to each sub-region by the second weight, and then add them to obtain the region image features corresponding to each sub-region;

[0219] Step 305: Determine the sample description text features corresponding to the sample description text, and input each region image feature and sample description text feature into the feature encoding model to output the fusion feature corresponding to each sub-region;

[0220] Step 306: Input the fusion feature corresponding to each sub-region into the feature decoding model to output the target fusion feature corresponding to each sub-region;

[0221] Step 307: Determine the predicted class probability corresponding to the sample object to be detected and the predicted sample region in the sample image according to each target fusion feature and the sample description text;

[0222] Step 308: Determine the classification loss value corresponding to the target detection model according to the predicted class probability and the target sample class;

[0223] Step 309: Determine the detection loss value corresponding to the target detection model according to the predicted sample region and the target sample region;

[0224] Step 310: Determine the intersection loss value corresponding to the target detection model according to the intersection over union between the predicted sample region and the target sample region;

[0225] Step 311: Add the classification loss value, the detection loss value and the intersection loss value to obtain the loss value corresponding to the target detection model. When the loss value does not meet the preset loss condition, iterate the target detection model, and return to obtain the target sample region, the target sample class and the sample description text corresponding to the sample object to be detected in the sample image until the loss value meets the preset loss condition, and obtain the trained image encoding model, the trained feature encoding model and the trained feature decoding model.

[0226] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the detailed description of the above target detection model training method, which will not be elaborated here.

[0227] Please refer to Figure 6 , Figure 6 which is a schematic flowchart of the target detection method provided by the embodiment of the present application. The target detection method is applied to the trained target detection model, and the trained target detection model is a model trained by the target detection model training method. The trained model includes a trained image encoding model, a trained feature encoding model and a trained feature decoding model.

[0228] Among them, the target detection method includes:

[0229] Step 410: Obtain the image to be detected and the task description text corresponding to the object to be detected in the image to be detected;

[0230] Step 420: Input the image to be detected into the trained image encoding model, and output the image features corresponding to each sub-region in the image to be detected;

[0231] Step 430: Determine the task description text features corresponding to the task description text, and input each image feature and the task description text features into the trained feature encoding model, and output the first fusion features corresponding to each sub-region;

[0232] Step 440: Input the first fusion features corresponding to each sub-region into the trained feature decoding model, and output the second fusion features corresponding to each sub-region;

[0233] Step 450: Determine the presence probability of the object to be detected in each sub-region according to each second fusion feature and the task description text;

[0234] Step 460: Determine the sub-region corresponding to the highest presence probability in the image to be detected as the presence region of the object to be detected.

[0235] Steps 410 to 460 will be described in detail below.

[0236] In step 410, obtain the image to be detected and the task description text corresponding to the object to be detected in the image to be detected.

[0237] The image to be detected can be an image with a resolution of 1024x1024. There is an object to be detected in the image to be detected, for example, the object to be detected is a person. The task description text can be obtained, and the task description text can be: "Please determine whether there is a person present".

[0238] In step 420, input the image to be detected into the trained image encoding model, and output the image features corresponding to each sub-region in the image to be detected.

[0239] Then, the image to be detected can be input into the trained image encoding model, and the image features corresponding to each sub-region in the image to be detected are output. Among them, the trained image encoding model includes a trained first image encoding sub-model and a trained second image encoding sub-model.

[0240] The image to be detected can be input into the trained first image encoding sub-model and the trained second image encoding sub-model respectively, and the target first image features corresponding to each sub-region output by the trained first image encoding sub-model and the target second image features corresponding to each sub-region output by the trained second image encoding sub-model are obtained.

[0241] Finally, multiply the target first image feature by the target first weight to obtain the first result. Multiply the target second image feature by the second weight to obtain the second result. Finally, add the first result and the second result to obtain the image feature corresponding to each sub-region in the detected image.

[0242] In step 430, determine the task description text feature corresponding to the task description text, and input each image feature and the task description text feature into the trained feature encoding model, and output the first fusion feature corresponding to each sub-region.

[0243] It is possible to determine the task description text feature corresponding to the task description text, and input each image feature and the task description text feature into the trained feature encoding model, and output the first fusion feature corresponding to each sub-region. The first fusion feature combines the image information in the image to be detected and the text information in the task description text, realizing the feature enhancement of each sub-region.

[0244] In step 440, input the first fusion feature corresponding to each sub-region into the trained feature decoding model, and output the second fusion feature corresponding to each sub-region.

[0245] Input the first fusion feature corresponding to each sub-region into the trained feature decoding model, and output the second fusion feature corresponding to each sub-region. The second fusion feature can be understood as the finally output object embedding. These object embeddings contain rich information, covering key elements such as the features, positions, and categories of various objects in the image, and are an important basis for subsequent alignment of regional images and texts and prediction. When processing complex scene images containing multiple objects, object embeddings can effectively encode the unique attributes of each object, helping the target detection model distinguish different objects.

[0246] In step 450, determine the existence probability of the object to be detected in each sub-region according to each second fusion feature and the task description text.

[0247] Among them, the task description text can be input into the pre-trained text model to generate sentence-level features. Finally, align each second fusion feature with the sentence-level features, and then multiply the aligned sentence-level features by each second fusion feature to obtain the existence probability of the object to be detected in each sub-region.

[0248] In step 460, determine the sub-region corresponding to the highest existence probability in the image to be detected as the existence region of the object to be detected.

[0249] It is possible to determine the sub-region corresponding to the highest existence probability in the image to be detected as the existence region of the object to be detected. Finally, determine the object within the existence region as the finally detected object.

[0250] As can be seen from the above, the trained object detection model in this application can accurately detect the object to be detected and the existence area of the object to be detected in the image to be detected.

[0251] Please refer to Figure 7 , Figure 7 which is a schematic structural diagram of the object detection model training device provided by an embodiment of this application. The object detection model training device can execute the above object detection model training method.

[0252] In the embodiments of this application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the functions of the module or unit.

[0253] The object detection model includes an image encoding model, a feature encoding model, and a feature decoding model. The object detection model training device 500 includes:

[0254] An acquisition module 510, configured to acquire a target sample region, a target sample category, and a sample description text corresponding to the object to be detected in the sample image;

[0255] An input module 520, configured to input the sample image into the image encoding model and output region image features corresponding to each sub-region in the sample image;

[0256] A determination module 530, configured to determine sample description text features corresponding to the sample description text, and input each region image feature and the sample description text features into the feature encoding model to output fusion features corresponding to each sub-region;

[0257] A fusion module 540, configured to input the fusion features corresponding to each sub-region into the feature decoding model to output target fusion features corresponding to each sub-region;

[0258] An output module 550, configured to determine the predicted class probability corresponding to the object to be detected in the sample image and the predicted sample region in the sample image according to each target fusion feature and the sample description text;

[0259] An iterative module 560 is used to determine the loss value corresponding to the object detection model according to the predicted class probability, the target sample class, the predicted sample region, and the target sample region. When the loss value does not meet the preset loss condition, the object detection model is iterated, and the target sample region, the target sample class, and the sample description text corresponding to the object to be detected in the acquired sample image are returned until the loss value meets the preset loss condition, and the trained image encoding model, the trained feature encoding model, and the trained feature decoding model are obtained.

[0260] In some embodiments, the image encoding model includes a first image encoding sub-model and a second image encoding sub-model; an input module 520 is configured to:

[0261] Input the sample image into the first image encoding sub-model, and output the first regional image feature corresponding to each sub-region in the sample image;

[0262] Input the sample image into the second image encoding sub-model, and output the second regional image feature corresponding to each sub-region in the sample image;

[0263] Multiply the first regional image feature corresponding to each sub-region by the first weight, multiply the second regional image feature corresponding to each sub-region by the second weight, and then add them to obtain the regional image feature corresponding to each sub-region.

[0264] In some embodiments, a determination module 530 is configured to:

[0265] Perform text segmentation processing on the sample description text to obtain multiple sub-texts;

[0266] Obtain a preset set of category names, and generate an input text sequence according to the preset set of category names and the multiple sub-texts;

[0267] Input the input text sequence into the pre-trained text encoding model, and output the sample description text feature corresponding to the sample description text.

[0268] In some embodiments, an output module 550 is configured to:

[0269] Obtain the word vector corresponding to each word in the sample description text;

[0270] Add the word vectors corresponding to each word and then divide by the number of words in the sample description text to obtain the sentence text feature;

[0271] Perform an inner product operation on the sentence text feature and each target fusion feature to obtain the probability distribution corresponding to each sub-region belonging to each preset category;

[0272] Determine the predicted class probability corresponding to the sample object to be detected according to the probability distribution, and determine the predicted sample region corresponding to the sample object to be detected in each sub-region according to the predicted class probability.

[0273] In some embodiments, the iterative module 560 includes a first loss sub-module, a second loss sub-module, and a third loss value module;

[0274] The first loss sub-module is used to determine the classification loss value corresponding to the target detection model according to the predicted class probability and the target sample class;

[0275] The second loss sub-module is used to determine the detection loss value corresponding to the target detection model according to the predicted sample region and the target sample region;

[0276] The third loss sub-module is used to determine the intersection loss value corresponding to the target detection model according to the intersection over union between the predicted sample region and the target sample region;

[0277] Add the classification loss value, the detection loss value, and the intersection loss value to obtain the loss value corresponding to the target detection model.

[0278] In some embodiments, the first loss sub-module is used for:

[0279] Perform normalization processing on the predicted class probability to obtain the normalized predicted class probability;

[0280] Determine the cross-entropy loss value corresponding between the predicted class probability and the target sample class according to the normalized predicted class probability, the target sample class, and the first preset hyperparameter;

[0281] Determine the first weight value corresponding to the cross-entropy loss value according to the target sample class, the normalized predicted class probability, and the second preset hyperparameter;

[0282] Determine the second weight value corresponding to the cross-entropy loss value according to the first preset hyperparameter and the target sample class;

[0283] Multiply the first weight value, the second weight value, and the cross-entropy loss value to obtain the target cross-entropy loss value;

[0284] Determine the classification loss value corresponding to the target detection model according to the total number of sub-regions in the sample image and the target cross-entropy loss value.

[0285] In some embodiments, the second loss sub-module is used for:

[0286] Determine multiple first coordinates corresponding to the predicted sample region and multiple second coordinates corresponding to the target sample region;

[0287] Calculate the difference between the associated first coordinates and second coordinates among multiple first coordinates and multiple second coordinates to obtain a difference calculation result, and determine the absolute value corresponding to the difference calculation result as the detection loss value corresponding to the target detection model.

[0288] In some embodiments, the third loss sub-module is configured to:

[0289] Determine the intersection value and union value between the predicted sample region and the target sample region;

[0290] Divide the intersection value by the union value to obtain the intersection over union (IoU), and determine the intersection loss value corresponding to the target detection model according to the IoU.

[0291] In the above embodiments, the descriptions of the respective embodiments have their own focuses. For the parts not described in detail in a certain embodiment, reference may be made to the detailed description of the above target detection model training method, which will not be elaborated here.

[0292] In the embodiments of the present application, the target detection model includes an image encoding model, a feature encoding model, and a feature decoding model. The acquisition module 510 acquires the target sample region, the target sample category, and the sample description text corresponding to the sample object to be detected in the sample image; the input module 520 inputs the sample image into the image encoding model and outputs the region image features corresponding to each sub-region in the sample image; the determination module 530 determines the sample description text features corresponding to the sample description text, and inputs each region image feature and the sample description text features into the feature encoding model to output the fusion features corresponding to each sub-region; the fusion module 540 inputs the fusion features corresponding to each sub-region into the feature decoding model to output the target fusion features corresponding to each sub-region; the output module 550 determines the predicted class probability corresponding to the sample object to be detected and the predicted sample region in the sample image according to each target fusion feature and the sample description text; the iteration module 560 determines the loss value corresponding to the target detection model according to the predicted class probability, the target sample category, the predicted sample region, and the target sample region. When the loss value does not meet the preset loss condition, the target detection model is iterated, and the target sample region, the target sample category, and the sample description text corresponding to the sample object to be detected in the sample image are returned until the loss value meets the preset loss condition, obtaining the trained image encoding model, the trained feature encoding model, and the trained feature decoding model.

[0293] Therefore, first determine the target sample region, target sample category, and sample description text corresponding to the sample object to be detected in the sample image. Then, determine the sample description text features of the sample description text. Next, input the sample image into the image encoding model to output the region image features corresponding to each sub-region in the sample image. Input each region image feature and the sample description text features into the feature encoding model to output the fusion features corresponding to each sub-region. The fusion features contain the correlation between the images of different sub-regions and the sample description text, and can represent the association of two different features in the same feature space. Then, input the fusion features corresponding to each sub-region into the feature decoding model to output the target fusion features corresponding to each sub-region. The target fusion features can be understood as being generated for a specific downstream task. For example, the fusion features are combined with more specific keyword text of the sample object to generate. Finally, the prediction class probability corresponding to the sample object to be detected and the predicted sample region in the sample image can be determined through the target fusion features and the sample description text. Then, determine the loss value corresponding to the target detection model according to the prediction class probability, target sample category, predicted sample region, and target sample region. When the loss value does not meet the preset loss condition, iterate the target detection model, and return to obtain the target sample region, target sample category, and sample description text corresponding to the sample object to be detected in the sample image until the loss value meets the preset loss condition, and obtain the trained image encoding model, trained feature encoding model, and trained feature decoding model. The trained target detection model can accurately detect the object to be detected and the existence region of the object to be detected in the image to be detected.

[0294] Please refer to Figure 8 , Figure 8 which is a schematic structural diagram of the target detection device provided by the embodiment of the present application. The target detection device can execute the above target detection method.

[0295] The target detection device 600 is applied to the trained target detection model. The trained target detection model is a model trained by the target detection model training method provided by the embodiment of the present application. The trained model includes a trained image encoding model, a trained feature encoding model, and a trained feature decoding model. The device includes:

[0296] The first acquisition module 610 is configured to acquire the image to be detected and the task description text corresponding to the object to be detected in the image to be detected;

[0297] The first input module 620 is configured to input the image to be detected into the trained image encoding model to output the image features corresponding to each sub-region in the image to be detected;

[0298] A second input module 630, configured to determine task description text features corresponding to task description text, and input each image feature and the task description text features into a trained feature encoding model, and output first fusion features corresponding to each sub-region;

[0299] A third input module 640, configured to input the first fusion features corresponding to each sub-region into a trained feature decoding model, and output second fusion features corresponding to each sub-region;

[0300] A first determination module 650, configured to determine the existence probability of the object to be detected in each sub-region according to each second fusion feature and the task description text;

[0301] A second determination module 660, configured to determine the sub-region corresponding to the highest existence probability in the image to be detected as the existence region of the object to be detected.

[0302] In the above embodiments, the descriptions of the various embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the detailed description of the above object detection method, which will not be elaborated here.

[0303] An embodiment of the present application further provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above object detection model training method or object detection method is implemented. The computer device may be any intelligent terminal including a tablet computer, a vehicle-mounted computer, etc.

[0304] Please refer to Figure 9 , Figure 9 which schematically shows the hardware structure of a computer device according to another embodiment. The computer device includes:

[0305] A processor 701, which may be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is configured to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;

[0306] The memory 702 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 702 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 702, and are called by the processor 701 to execute the target detection model training method or the target detection method of the embodiments of this application;

[0307] The input / output interface 703 is used to implement information input and output;

[0308] The communication interface 704 is used to implement communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or through wireless means (such as mobile network, WIFI, Bluetooth, etc.);

[0309] The bus 705 transmits information between various components of the device (such as the processor 701, the memory 702, the input / output interface 703, and the communication interface 704);

[0310] Among them, the processor 701, the memory 702, the input / output interface 703, and the communication interface 704 are communicatively connected to each other inside the device through the bus 705.

[0311] The embodiments of this application also provide a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the above-mentioned target detection model training method or target detection method.

[0312] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0313] In the embodiment of the present application, the object detection model includes an image encoding model, a feature encoding model, and a feature decoding model. By obtaining the target sample region, target sample category, and sample description text corresponding to the sample object to be detected in the sample image; inputting the sample image into the image encoding model to output the region image features corresponding to each sub-region in the sample image; determining the sample description text features corresponding to the sample description text, and inputting each region image feature and the sample description text features into the feature encoding model to output the fusion features corresponding to each sub-region; inputting the fusion features corresponding to each sub-region into the feature decoding model to output the target fusion features corresponding to each sub-region; determining the predicted class probability corresponding to the sample object to be detected and the predicted sample region in the sample image according to each target fusion feature and the sample description text; determining the loss value corresponding to the object detection model according to the predicted class probability, target sample category, predicted sample region, and target sample region. When the loss value does not meet the preset loss condition, the object detection model is iterated, and the target sample region, target sample category, and sample description text corresponding to the sample object to be detected in the sample image are returned until the loss value meets the preset loss condition, and the trained image encoding model, trained feature encoding model, and trained feature decoding model are obtained.

[0314] Therefore, first determine the target sample region, target sample category, and sample description text corresponding to the sample object to be detected in the sample image. Then, determine the sample description text features of the sample description text. Next, input the sample image into the image encoding model to output the region image features corresponding to each sub-region in the sample image. Input each region image feature and the sample description text features into the feature encoding model to output the fusion features corresponding to each sub-region. The fusion features contain the correlation between the images of different sub-regions and the sample description text, and can represent the association of two different features in the same feature space. Then, input the fusion features corresponding to each sub-region into the feature decoding model to output the target fusion features corresponding to each sub-region. The target fusion features can be understood as being generated for specific downstream tasks. For example, the fusion features are combined with more specific keyword texts of the sample object to generate. Finally, the prediction class probability corresponding to the sample object to be detected and the predicted sample region in the sample image can be determined through the target fusion features and the sample description text. Then, determine the loss value corresponding to the target detection model according to the prediction class probability, target sample category, predicted sample region, and target sample region. When the loss value does not meet the preset loss condition, iterate the target detection model, and return to obtain the target sample region, target sample category, and sample description text corresponding to the sample object to be detected in the sample image until the loss value meets the preset loss condition, and obtain the trained image encoding model, trained feature encoding model, and trained feature decoding model. The trained target detection model can accurately detect the object to be detected and the existence region of the object to be detected in the image to be detected.

[0315] The embodiments described in the embodiments of the present application are for more clearly explaining the technical solutions of the embodiments of the present application, and do not constitute a limitation to the technical solutions provided by the embodiments of the present application. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0316] Those skilled in the art can understand that the technical solutions shown in the figure do not constitute a limitation to the embodiments of the present application, and may include more or fewer steps than shown in the figure, or combine some steps, or different steps.

[0317] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0318] Those of ordinary skill in the art will understand that all or some of the steps in the methods disclosed above, and the functional modules / units in systems and devices, can be implemented as software, firmware, hardware, or a suitable combination thereof.

[0319] As used in the specification of this application and the above-mentioned drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0320] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Here, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or a similar expression means any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0321] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above-mentioned unit division is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.

[0322] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0323] In addition, the functional units in various embodiments of the present application may be integrated in a processing unit, may exist separately as individual physical units, or two or more units may be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0324] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs and other various media that can store programs.

[0325] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings. However, this does not limit the scope of the rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall fall within the scope of the rights of the embodiments of the present application.

Claims

1. A target detection model training method, characterized in that: The target detection model includes an image encoding model, a feature encoding model and a feature decoding model, and the method includes: Obtain the target sample area, target sample category and sample description text corresponding to the sample object to be detected in the sample image; Inputting the sample image into an image coding model, and outputting regional image features corresponding to each sub-region in the sample image; Determine the sample description text features corresponding to the sample description text, input each region image feature and the sample description text features into a feature encoding model, and output the fusion features corresponding to each sub-region; Inputting the fused features corresponding to each sub-region into the feature decoding model, and outputting the target fused features corresponding to each sub-region; Determine the predicted category probability corresponding to the sample object to be detected and the predicted sample area in the sample image according to each target fusion feature and the sample description text; A loss value corresponding to the target detection model is determined according to the predicted category probability, the target sample category, the predicted sample area and the target sample area. When the loss value does not meet the preset loss condition, the target detection model is iterated, and the target sample area, target sample category and sample description text corresponding to the sample object to be detected in the sample image are returned until the loss value meets the preset loss condition, thereby obtaining a trained image coding model, a trained feature coding model and a trained feature decoding model.

2. The target detection model training method according to claim 1, characterized in that: The image coding model comprises a first image coding sub-model and a second image coding sub-model; The step of inputting the sample image into an image coding model and outputting a regional image feature corresponding to each sub-region in the sample image comprises: Inputting the sample image into a first image coding sub-model, and outputting first region image features corresponding to each sub-region in the sample image; Inputting the sample image into a second image coding sub-model, and outputting second region image features corresponding to each sub-region in the sample image; The first regional image feature corresponding to each sub-region is multiplied by a first weight, and the second regional image feature corresponding to each sub-region is multiplied by a second weight, and then added together to obtain the regional image feature corresponding to each sub-region.

3. The target detection model training method according to claim 1, characterized in that: The determining of the sample description text feature corresponding to the sample description text includes: Performing text segmentation processing on the sample description text to obtain multiple sub-texts; Acquire a preset category name set, and generate an input text sequence by arranging the preset category name set and the multiple subtexts; The input text sequence is input into a pre-trained text encoding model, and the sample description text features corresponding to the sample description text are output.

4. The target detection model training method according to claim 1, characterized in that: The step of determining the predicted category probability corresponding to the sample object to be detected and the predicted sample area in the sample image according to each target fusion feature and the sample description text includes: Obtain the word vector corresponding to each word in the sample description text; Add the word vectors corresponding to each word and divide the sum by the number of words in the sample description text to obtain the sentence text feature; Perform an inner product operation on the sentence text feature and each target fusion feature to obtain a probability distribution corresponding to each sub-region belonging to each preset category; The predicted category probability corresponding to the sample object to be detected is determined according to the probability distribution, and the predicted sample area corresponding to the sample object to be detected is determined in each sub-area according to the predicted category probability.

5. The target detection model training method according to claim 1, characterized in that: The determining the loss value corresponding to the target detection model according to the predicted category probability, the target sample category, the predicted sample area and the target sample area includes: Determine a classification loss value corresponding to the target detection model according to the predicted category probability and the target sample category; Determine a detection loss value corresponding to the target detection model according to the predicted sample area and the target sample area; Determining an intersection loss value corresponding to the target detection model according to an intersection-over-union ratio between the predicted sample area and the target sample area; The classification loss value, the detection loss value and the intersection loss value are added to obtain a loss value corresponding to the target detection model.

6. The target detection model training method according to claim 5, characterized in that: The step of determining the classification loss value corresponding to the target detection model according to the predicted category probability and the target sample category includes: Normalizing the predicted category probabilities to obtain normalized predicted category probabilities; Determining a cross entropy loss value corresponding to the predicted category probability and the target sample category according to the normalized predicted category probability, the target sample category and a first preset hyperparameter; Determine a first weight value corresponding to the cross entropy loss value according to the target sample category, the normalized predicted category probability and a second preset hyperparameter; Determine a second weight value corresponding to the cross entropy loss value according to the first preset hyperparameter and the target sample category; Multiplying the first weight value, the second weight value, and the cross entropy loss value to obtain a target cross entropy loss value; A classification loss value corresponding to the target detection model is determined according to the total number of sub-regions in the sample image and the target cross entropy loss value.

7. The target detection model training method according to claim 5, characterized in that: The determining, according to the predicted sample area and the target sample area, a detection loss value corresponding to the target detection model includes: Determine a plurality of first coordinates corresponding to the predicted sample area and a plurality of second coordinates corresponding to the target sample area; A difference calculation is performed on the first coordinates and the second coordinates associated with the multiple first coordinates and the multiple second coordinates to obtain a difference calculation result, and an absolute value corresponding to the difference calculation result is determined as a detection loss value corresponding to the target detection model.

8. The target detection model training method according to claim 5, characterized in that: The determining, according to the intersection-over-union ratio between the predicted sample region and the target sample region, an intersection loss value corresponding to the target detection model comprises: Determine an intersection value and a union value between the prediction sample area and the target sample area; The intersection value is divided by the union value to obtain an intersection-and-union ratio, and an intersection loss value corresponding to the target detection model is determined according to the intersection-and-union ratio.

9. A target detection model training device, characterized in that: The target detection model includes an image coding model, a feature coding model and a feature decoding model, and the device includes: An acquisition module is used to acquire a target sample area, a target sample category and a sample description text corresponding to a sample object to be detected in a sample image; An input module, used for inputting the sample image into an image coding model, and outputting regional image features corresponding to each sub-region in the sample image; A determination module, used to determine the sample description text features corresponding to the sample description text, and input each region image feature and the sample description text features into a feature coding model, and output a fusion feature corresponding to each sub-region; A fusion module, used for inputting the fusion features corresponding to each sub-region into the feature decoding model, and outputting the target fusion features corresponding to each sub-region; An output module, used to determine the predicted category probability corresponding to the sample object to be detected and the predicted sample area in the sample image according to each target fusion feature and the sample description text; An iterative module is used to determine the loss value corresponding to the target detection model according to the predicted category probability, the target sample category, the predicted sample area and the target sample area. When the loss value does not meet the preset loss condition, the target detection model is iterated to return to obtain the target sample area, target sample category and sample description text corresponding to the sample object to be detected in the sample image until the loss value meets the preset loss condition, thereby obtaining a trained image coding model, a trained feature coding model and a trained feature decoding model.

10. A target detection method, characterized in that: The trained target detection model is applied to a trained target detection model, wherein the trained target detection model is a model trained by the target detection model training method according to any one of claims 1 to 8, wherein the trained target detection model includes a trained image coding model, a trained feature coding model and a trained feature decoding model, and the method includes: Obtaining a task description text corresponding to an image to be detected and an object to be detected in the image to be detected; Inputting the image to be detected into the trained image coding model, and outputting the image features corresponding to each sub-region in the image to be detected; Determine the task description text feature corresponding to the task description text, input each image feature and the task description text feature into the trained feature encoding model, and output the first fusion feature corresponding to each sub-region; Inputting the first fused feature corresponding to each sub-region into the trained feature decoding model, and outputting the second fused feature corresponding to each sub-region; Determine the existence probability of the object to be detected in each sub-region according to each second fusion feature and the task description text; The sub-region corresponding to the highest existence probability in the image to be detected is determined as the existence region of the object to be detected.

11. A target detection device, characterized in that: The trained target detection model is applied to a trained target detection model, wherein the trained target detection model is a model trained by the target detection model training method according to any one of claims 1 to 8, wherein the trained target detection model includes a trained image coding model, a trained feature coding model and a trained feature decoding model, and the device includes: A first acquisition module is used to acquire the image to be detected and the task description text corresponding to the object to be detected in the image to be detected; A first input module, used for inputting the image to be detected into the trained image coding model, and outputting the image features corresponding to each sub-region in the image to be detected; A second input module is used to determine the task description text feature corresponding to the task description text, and input each image feature and the task description text feature into the trained feature encoding model, and output the first fusion feature corresponding to each sub-region; A third input module, used for inputting the first fused feature corresponding to each sub-region into the trained feature decoding model, and outputting the second fused feature corresponding to each sub-region; A first determination module, used to determine the existence probability of the object to be detected in each sub-region according to each second fusion feature and the task description text; The second determination module is used to determine the sub-region corresponding to the highest existence probability in the image to be detected as the existence region of the object to be detected.

12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a plurality of instructions, which are suitable for loading by a processor to execute the target detection model training method described in any one of claims 1 to 8 or the target detection method described in claim 10.

13. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the target detection model training method described in any one of claims 1 to 8 or the target detection method described in claim 10 is implemented.

Citation Information

Patent Citations

  • Semantic segmentation model training method and apparatus, and electronic device and storage medium

    WO2024012251A1

  • Method and apparatus for training text-graph model, device, and storage medium

    WO2025035924A1