Target detection method and device, equipment and storage medium
By using the unlimited object detection model of detectable target categories, and using the limited and unlimited training image training, the problems of poor flexibility and high cost of newly added target categories in the prior art are solved, and fast and low-cost object detection are achieved.
Patent Information
- Application Number
- CN202410016257.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-04
- Publication Date
- 2025-07-04
AI Technical Summary
Existing object detection methods require modification of model structure and retraining when adding new target categories of interest, resulting in poor flexibility and high cost.
By using an object detection model with an unlimited target category, training is performed using the first training image defined by the detected target category and the unlimited second training image to be trained, and rapid detection of any new category target is achieved.
There is no need to retrain the model and add new object detectors, which reduces the cost of object detection and improves detection efficiency and flexibility.
Smart Images

Figure CN120259710A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of computer technologies, and in particular, to an object detection method, apparatus, device, and storage medium. Background Art
[0002] With the rapid development of Artificial Intelligence (AI) technology, AI technology has been used in various fields, such as object detection. Current object detection methods are mainly divided into a localization task and a classification task. The localization task is to frame the location where the object is located, while the classification task is to determine the category of the object.
[0003] Current object detection methods can only detect categories in a pre-determined set of categories. If a new category of object of interest is added, the number of categories output by the classifier of the object detection model needs to be changed, and a dataset containing the new category of object needs to be prepared to retrain the model. This results in poor flexibility and high cost of object detection. Summary of the Invention
[0004] The present application provides an object detection method, apparatus, device, and storage medium, which can achieve rapid detection of objects of new categories, improve the flexibility of object detection, and reduce the cost of object detection.
[0005] In a first aspect, the present application provides an object detection method, including:
[0006] Obtaining a target image to be detected and N target categories, where N is a positive integer;
[0007] Processing the target image and the category names of the N target categories through an object detection model whose detectable target categories are not limited, and detecting the objects in the target image that belong to the N target categories;
[0008] Wherein, the object detection model is trained based on a first training image and a second training image. The first training image is an image in which the detectable target categories are limited, and the second training image is an image in which the detectable target categories are not limited.
[0009] In a second aspect, the present application provides an object detection apparatus, including:
[0010] An obtaining unit, configured to obtain a target image to be detected and N target categories, where N is a positive integer;
[0011] A detecting unit, configured to process the target image and the category names of the N target categories through an object detection model whose detectable target categories are not limited, and detect the objects in the target image that belong to the N target categories;
[0012] Among them, the target detection model is trained based on a first training image and a second training image. The first training image is an image in which the target categories to be detected are defined, and the second training image is an image in which the target categories to be detected are not defined.
[0013] In a third aspect, a computing device is provided, including a processor and a memory. The memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to execute the method in the first aspect or its various implementation manners above.
[0014] In a fourth aspect, a chip is provided for implementing the method in any one of the first aspects or its various implementation manners above. Specifically, the chip includes: a processor, which is used to call and run a computer program from a memory, so that a device installed with the chip executes the method in any one of the first aspects or its various implementation manners above.
[0015] In a fifth aspect, a computer-readable storage medium is provided for storing a computer program, and the computer program enables a computer to execute the method in any one of the first aspects or its various implementation manners above.
[0016] In a sixth aspect, a computer program product is provided, including computer program instructions, and the computer program instructions enable a computer to execute the method in any one of the first aspects or its various implementation manners above.
[0017] In a seventh aspect, a computer program is provided, which when running on a computer, enables the computer to execute the method in any one of the first aspects or its various implementation manners above.
[0018] In summary, the present application obtains a target image to be detected and N target categories, and processes the target image and the category names of the N target categories through a target detection model with no limit on the detectable target categories, so as to detect the objects belonging to the N target categories in the target image. That is to say, the embodiment of the present application trains the target detection model through a first training image with limited detectable target categories and a second training image with unlimited detectable target categories, so that the trained target detection model can detect targets of any newly added category. In this way, in an actual application scenario, when new targets of interest need to be added, only the N target categories of interest and the target image to be detected need to be input into the trained target detection model, and then the target detection model with no limit on the detectable target categories processes the target image and the category names of the N target categories, so as to detect the objects belonging to the N target categories in the target image, without the need to update the target detection model again or add a new target detector, thereby reducing the cost of target detection and improving the efficiency and flexibility of target detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0020] Figure 1 It is a schematic diagram of an implementation environment related to the embodiment of the present application;
[0021] Figure 2 It is a schematic flowchart of a model training method provided by an embodiment of the present application;
[0022] Figure 3 It is a schematic diagram of model training related to the embodiment of the present application;
[0023] Figure 4 It is a schematic structural diagram of a target detection model proposed by the embodiment of the present application;
[0024] Figure 5 It is a schematic diagram when training the target detection model with the first training data;
[0025] Figure 6 It is a schematic diagram when training the target detection model with the second training data;
[0026] Figure 7 It is a schematic flowchart of a target detection method provided by an embodiment of the present application;
[0027] Figure 8 Schematic diagram for performing object detection using a trained object detection model;
[0028] Figure 9 Schematic block diagram of an object detection device provided by an embodiment of the present application;
[0029] Figure 10 Schematic block diagram of a computing device provided by an embodiment of the present application. Detailed implementation manners
[0030] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0031] It should be noted that the terms "first", "second", etc. in the description and claims of the present application and the above accompanying drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In the embodiments of the present invention, "B corresponding to A" means that B is associated with A. In one implementation, B can be determined according to A. However, it should also be understood that determining B according to A does not mean determining B only according to A, and B can also be determined according to A and / or other information. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server including a series of steps or units does not necessarily need to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices. In the description of the present application, unless otherwise specified, "a plurality of" means two or more than two.
[0032] The technical solutions proposed in the present application can be applied to technical fields such as artificial intelligence and object detection, and are used to reduce the cost of object detection and improve the flexibility of object detection.
[0033] Next, the related concepts involved in the embodiments of the present application will be introduced.
[0034] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.
[0035] Artificial intelligence technology is an interdisciplinary subject that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0036] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0037] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields. For example, common ones include smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, robots, smart healthcare, smart customer service, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0038] In the embodiments of this application, artificial intelligence technology is applied to object detection to provide an open-category object detection model, which can achieve fast detection of newly added category objects without retraining the model, making the cost of object detection low, and can also achieve detection of any newly added category objects, making the flexibility of object detection high.
[0039] Object detection methods are mainly divided into a localization task and a classification task. Among them, the localization task is to frame the location where the object is located, while the classification task is to determine the category of the object.
[0040] Current object detection models can only detect object categories within a pre-determined set of categories. When the number of object categories of interest increases in an actual application scenario, the first solution is to modify the network structure and retrain. The second solution is to add a new object detector. The disadvantage of the first solution is that the iteration of the model takes time and requires redeployment, while the disadvantage of the second solution is that the cost is high and shows a gradually increasing trend. For scenarios where the objects change greatly, such as shopping and recommendation, the update or addition of the model limits the application update efficiency.
[0041] To solve the above technical problems, an embodiment of the present application provides a general object detection model with an open category set. Through this object detection model, accurate and fast detection of objects of any newly added category can be achieved without retraining the model and adding a new object detector. Specifically, the object detection model is trained with a first training image in which the object categories to be detected are limited, and a second training image in which the object categories to be detected are not limited, so that the trained object detection model can detect objects of any newly added category. In such an actual application scenario, when new objects of interest need to be added, only the N object categories of interest and the object image to be detected are input into the trained object detection model. Then, through this object detection model with unlimited object categories to be detected, the object image and the category names of the N object categories are processed, and the objects belonging to the N object categories in the object image can be detected without re-updating the object detection model or adding a new object detector, thereby reducing the cost of object detection and improving the efficiency and flexibility of object detection.
[0042] The implementation environment of the embodiment of the present application is introduced below.
[0043] Figure 1 It is a schematic diagram of an implementation environment related to an embodiment of the present application, including a terminal device 101 and a computing device 102.
[0044] As Figure 1 shown, the computing device 102 in the embodiment of the present application includes an object detection model. Among them, the terminal device 101 obtains a training set, for example, obtains a plurality of first training images and a plurality of second training images, and sends the training set to the computing device 102. The computing device 102 uses the training set and adopts the model training method provided by the embodiment of the present application to train the object detection model.
[0045] Exemplarily, the computing device 102 obtains a plurality of first training images and a plurality of second training images, and uses the plurality of first training images and the plurality of second training images to train the object detection model. Among them, the first training image is an image in which the object categories to be detected are defined, and the second training image is an image in which the object categories to be detected are not defined. In this way, the computing device 102 can initially train the object detection model using the first training images to obtain the object detection model after initial training. Then, the computing device 102 uses the second training images to retrain the object detection model after initial training. Since the object categories to be detected in the second training images are not defined, when using the second training images to train the object detection model, the object detection model can detect objects of various categories. That is to say, in the embodiment of the present application, a general object detection model for an open category set is trained through the first training images and the second training images. In actual detection, the object image to be detected and N object categories of interest are obtained, and the object detection model is used to process the object image and the N object categories to detect the objects belonging to the N object categories in the object image.
[0046] In some embodiments, as Figure 1 shown, the application scenario may further include a database 103, which includes historical data, such as historical object detection data. In the embodiment of the present application, the terminal device 101 is communicatively connected to the database 103 and can write data into the database 103, and the computing device 102 is also communicatively connected to the database 102 and can read data from the database 103. In one example, during the model training process of the embodiment of the present application, when training the model, the computing device 102 obtains historical data from the database 103 as training samples. Then, the computing device 102 uses the training samples to train the object detection model to obtain the object detection model after training. Optionally, the computing device 102 can save the object detection model after training in the computing device 102. Optionally, the computing device 102 can send the object detection model after training to the terminal device 101 for saving.
[0047] In the embodiment of the present application, the object detection method can be implemented by the terminal device 101, or by the computing device 102, or by a system composed of the terminal device 101 and the computing device 102.
[0048] In some embodiments, when the target detection method of the embodiments of the present application is executed by the terminal device 101, the terminal device 101 obtains the trained target detection model from the computing device 102. In this way, when performing target detection, the terminal device 101 obtains the target image to be detected and N target categories of interest, and processes the target image and the N target categories of interest through the trained target detection model to detect the objects belonging to the N target categories in the target image.
[0049] In some embodiments, when the target detection method of the embodiments of the present application is executed by the computing device 102, the terminal device 101 sends the obtained target image to be detected and N target categories of interest to the computing device 102. The computing device 102 uses the target detection model stored by itself to process the target image and the N target categories of interest to detect the objects belonging to the N target categories in the target image.
[0050] The embodiments of the present application do not limit the specific type of the terminal device 101. In some embodiments, the terminal device 101 may include but is not limited to: mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, wearable intelligent devices, medical devices, etc. The device is often configured with a display device, and the display device may also be a display, a display screen, a touch screen, etc. The touch screen may also be a touch panel, a touch screen panel, etc.
[0051] In some embodiments, the computing device 102 is a terminal device with data processing functions, such as mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, wearable intelligent devices, medical devices, etc.
[0052] In some embodiments, the computing device 102 is a server. The server can be one or more. When there are multiple servers, at least two servers are used to provide different services, and / or at least two servers are used to provide the same service, such as providing the same service in a load balancing manner. The embodiments of the present application do not limit this. Among them, the above server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The server 102 can also become a node of the blockchain.
[0053] In the embodiments of the present application, the terminal device 101 and the computing device 102 can be directly or indirectly connected through wired communication or wireless communication, and the present application does not limit this here.
[0054] It should be noted that the implementation environment of the embodiments of the present application includes but is not limited to Figure 1 as shown.
[0055] The technical solutions of the embodiments of the present application will be described in detail below through some embodiments. These several embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0056] First, the model training process of the embodiments of the present application will be introduced.
[0057] Figure 2 is a schematic flowchart of a target detection model training method provided by an embodiment of the present application. The execution subject of the embodiments of the present application is a device with the function of training a model, such as a model training device. In some embodiments, the model training device can be Figure 1 the computing device in Figure 1 or the terminal device in Figure 1 or a system composed of a computing device and a terminal device in
[0058] For the sake of description, the embodiments of the present application will be described by taking the execution subject as a computing device as an example. Figure 2 The training process of the target detection model will be introduced below in combination with
[0059] As Figure 2 shown, the target detection model training process of the embodiments of the present application includes:
[0060] S101. Obtain first training data and second training data.
[0061] Among them, the first training data includes a plurality of first training images and the target detection results of each first training image under a limited number of target categories.
[0062] In the embodiments of the present application, the first training image can be understood as an image in which the target categories to be detected are limited. For example, the first training image includes 10 objects, but only 5 types of target objects in the first training image can be detected. That is to say, the target categories that can be detected in the first training image are predetermined.
[0063] In one example, the above first training data is data collected in an actual application scenario.
[0064] For example, in an actual application scenario, there is usually already one or several traditional object detectors. Each object detector can detect object targets in a fixed category and regard other areas as the background. Based on different service deployment and usage methods, the embodiments of the present application may obtain data in the following two cases as the first training data.
[0065] Case 1: Each service application is in a different scenario. At this time, only the object detection results of the first training images under the category set corresponding to the service can be obtained. For example, assume that the services include Service 1, Service 2, and Service 3. The category set corresponding to Service 1 includes: Category A1, Category A2, Category A3... Category A20. The category set corresponding to Service 2 includes: Category B1, Category B2, Category B3... Category B15. The category set corresponding to Service 3 includes: Category C1, Category C2, Category C3... Category C30. In this way, the object detection results of at least one first training image in the Service 1 scenario can be obtained. The target categories that each training image in these first training images can be detected are limited to these 20 categories, namely Category A1, Category A2, Category A3... Category A20. At the same time, the computing device can obtain the object detection results of at least one first training image in the Service 2 scenario. The target categories that each training image in these first training images can be detected are limited to these 15 categories, namely Category B1, Category B2, Category B3... Category B15. The computing device can obtain at least one first training image in the Service 3 scenario. The target categories that each training image in these first training images can be detected are limited to these 30 categories, namely Category C1, Category C2, Category C3... Category C30. That is to say, in Case 1, the computing device can obtain the object detection results of multiple first training images under different category sets. Exemplarily, the computing device can obtain the first training data as shown in Table 1:
[0066] Table 1
[0067]
[0068] As shown in Table 1, the computing device can obtain n1 first training images and the object detection results of these n1 first training images under the category set 1 (such as various limited category sets such as the above-mentioned Category A1, Category A2, Category A3... Category A20). And the computing device can obtain n2 first training images and the object detection results of these n2 first training images under the category set 2 (such as various limited category sets such as the above-mentioned Category B1, Category B2, Category B3... Category B15).
[0069] Among them, the object detection result of the first training image includes the detection bounding box of the object belonging to the specified category in the first training image (i.e., the position information of the object), and the category of the object in the detection bounding box.
[0070] Case 2: All service applications are in the same scenario. At this time, the object detection results of the first training image under all services can be obtained. For example, assume that the services include Service 1, Service 2, and Service 3. The category set corresponding to Service 1 includes: Category A1, Category A2, Category A3... Category A20. The category set corresponding to Service 2 includes: Category B1, Category B2, Category B3... Category B15. The category set corresponding to Service 3 includes: Category C1, Category C2, Category C3... Category C30. In this way, the object detection results of the first training image under the three services of Service 1, Service 2, and Service 3 can be obtained. At this time, the object categories that can be detected in the first training image are limited to Category A1, Category A2, Category A3... Category A20, Category B1, Category B2, Category B3... Category B15, and Category C1, Category C2, Category C3... Category C30, a total of 65 categories. That is to say, in Case 2, the computing device can obtain the object detection results of multiple first training images under all preset category sets. Exemplarily, the computing device can obtain the first training data as shown in Table 2:
[0071] Table 2
[0072]
[0073] As shown in Table 2, the computing device can obtain multiple first training images such as n3 + n4, and the object detection results of these multiple first training images under all category sets (for example, all category sets include: the above-mentioned Category A1, Category A2, Category A3... Category A20, Category B1, Category B2, Category B3... Category B15, and Category C1, Category C2, Category C3... Category C30, a total of 65 categories).
[0074] In some embodiments, the first training data of the embodiments of the present application may only include obtaining multiple first training images and the object detection results of each first training image through the method of Case 1.
[0075] In some embodiments, the first training data of the embodiments of the present application may only include obtaining multiple first training images and the object detection results of each first training image through the method of Case 2.
[0076] In some embodiments, the first training data of the embodiments of the present application may include obtaining multiple first training images and the object detection results of each first training image through the methods of Case 1 and Case 2.
[0077] The embodiments of the present application do not limit the specific manner in which the computing device obtains the object detection results of each first training image in the first training data.
[0078] In a possible implementation manner, the object detection results of each first training image in the first training data are determined by manual annotation. For example, the object detection results of the first training image under categories A1, A2, A3... A20 are manually annotated, that is, the detection frames (i.e., position information) of the objects belonging to the 20 categories of A1, A2, A3... A20 in the first training image are annotated, as well as the category of the object in each detection frame.
[0079] In a possible implementation manner, the computing device performs object detection on the first training image through at least one trained object detector, and the object detection results of the first training image under a limited number of object categories can be obtained.
[0080] Illustrating with an example, assume there are 3 object detectors, and the object categories that these 3 object detectors can detect are limited and not completely the same. For example, the object categories that the first object detector can detect include categories A1, A2, A3... A20, the object categories that the second object detector can detect include categories B1, B2, B3... B15, and the object categories that the third object detector can detect can include categories C1, C2, C3... C30. In this way, in case 1, the computing device can perform object detection on a preset number (such as 100,000, or other values) of first training images through the first object detector, and the object detection results of these 100,000 first training images under categories A1, A2, A3... A20 can be obtained. The computing device can perform object detection on a preset number (such as 100,000, or other values) of first training images through the second object detector, and the object detection results of these 100,000 first training images under categories B1, B2, B3... B15 can be obtained. And the computing device can perform object detection on a preset number (such as 100,000, or other values) of first training images through the third object detector, and the object detection results of these 100,000 first training images under categories C1, C2, C3... C30 can be obtained. In case 2, the computing device can perform object detection on a certain number (such as 200,000) of first training images through the first object detector, the second object detector, and the third object detector respectively, and the object detection results of these 200,000 first training images under 65 categories including categories A1, A2, A3... A20, categories B1, B2, B3... B15, and categories C1, C2, C3... C30 can be obtained.
[0081] That is to say, in this embodiment, the computing device obtains the first training data based on several existing object detectors of limited categories (also known as object detection models), and uses the first training data to train a general object detection model for an open category set.
[0082] The second training data of the embodiments of the present application includes a plurality of second training images and description information of each second training image.
[0083] The second training images of the embodiments of the present application can be understood as images whose detectable target categories are not limited. For example, if the second training image includes 10 objects, then the 10 types of target objects in the second training image can be detected. That is to say, the detectable target categories in the second training image are not predetermined in advance, and it is possible to detect each type of object included in the second training image.
[0084] In one example, the above second training data is data collected in a general scenario.
[0085] For example, in a general scenario, a certain number (such as more than 1 million, or other numbers) of second training images and description information of each second training image are collected. Exemplarily, the computing device can obtain the second training data from open-source text-image multi-modal training data, and this second training data is used to train the object detection model's perception ability for general objects.
[0086] After the computing device obtains the first training data and the second training data based on the above steps, it executes the following step S102.
[0087] S102: Initially train the object detection model with the first training data.
[0088] In the embodiments of the present application, as Figure 3 shown, the training process of the computing device using the first training data and the second training data for the object detection model includes two parts. The first part is to initially train the object detection model with the first training data to obtain the object detection model after initial training. The second part is to retrain the object detection model after initial training with the second training data to obtain the trained object detection model.
[0089] The embodiments of the present application do not limit the specific process of the computing device initially training the object detection model with the first training data.
[0090] In some embodiments, for each first training image in the first training data, the computing device inputs the first training image into the object detection model to obtain the detection result of the first training image output by the object detection model. Then, the detection result output by the object detection model is compared with the object detection result of the first training image to determine the loss of the object detection model. Furthermore, based on this loss, the parameters in the object detection model are adjusted. Through repeated iteration, the initially trained object detection model is obtained.
[0091] In some embodiments, the above S102 includes the following steps S102-A to S102-C:
[0092] S102-A. For each first training image in the first training data, through the object detection model, process the first training image and the class names of the limited object classes corresponding to the first training image to obtain the first detection result of the first training image under the limited object classes;
[0093] S102-B. Based on the object detection result and the first detection result of the first training image under the limited object classes, determine the first loss of the object detection model;
[0094] S102-C. Based on the first loss, perform initial training on the object detection model.
[0095] In this implementation manner, for each first training image in the first training data, the computing device processes the first training image and the class names of the limited object classes corresponding to the first training image to obtain the detection result of the first training image under the effective classes. For the convenience of description, this detection result is denoted as the first detection result.
[0096] As can be seen from the above, in this implementation manner, the data input by the computing device into the object detection model, in addition to the first training image, also includes the class names of the limited object classes corresponding to the first training image. Assume that the limited object classes corresponding to the first training image are class A1, class A2, class A3... class A20, these 20 classes. Then the computing device inputs the first training image and the names of these 20 classes, namely class A1, class A2, class A3... class A20, into the object detection model for learning, so that the object detection model aligns the image information of the first training image with the text information of these 20 class names.
[0097] The embodiments of the present application do not limit the specific network structure of the object detection model.
[0098] In a possible implementation manner, such as Figure 4As shown in the figure, the object detection model of the embodiment of the present application includes a visual module and a text module, where the visual module is used to process the input training images, and the text module is used to process the input text data. Based on this, the above S102-A includes the following steps of S102-A1 to S102-A3:
[0099] S102-A1. Through the visual module, perform object detection on the first training image to obtain candidate boxes of the objects in the first training image, and extract the image feature information of each candidate box;
[0100] S102-A2. Through the text module, extract the text feature information of each category name in the valid target categories corresponding to the first training image;
[0101] S102-A3. Based on the image feature information of each candidate box and the text feature information of each category name in the limited target categories, determine the first detection result of the first training image under the limited target categories.
[0102] As can be seen from the above, the first training data of the embodiment of the present application includes multiple first training images, and the object detection results of each first training image under the corresponding limited target categories. As Figure 4 shown, the object detection model includes a visual module (also called a detection module in some embodiments) and a text module. In this way, for each first training image in the first training data, the computing device can input the first training image into the visual module for object detection and feature extraction, and can obtain the candidate boxes of the objects in the first training image, and the image feature information of each candidate box. At the same time, the computing device inputs the category names of the limited target categories corresponding to the first training image into the text module for processing, and obtains the text feature information of each category name in the limited target categories.
[0103] For example, assume that the limited target categories corresponding to the first training image are category A1, category A2, category A3... category A20, these 20 categories. Then the computing device inputs the first training image into the visual module for processing, and obtains the candidate boxes of the objects in the first training image and the image feature information of each candidate box. At the same time, the computing device inputs the category names of category A1, category A2, category A3... category A20, these 20 categories into the text module for processing, and obtains the text feature information of each category name in these 20 categories.
[0104] Next, the computing device determines the first detection result of the first training image under the limited target categories based on the image feature information of each candidate box in the first training image and the text feature information of each category name in the limited target categories corresponding to the first training image.
[0105] The embodiments of this application do not limit the specific manner in which the computing device determines the first detection result of the first training image under the limited target categories based on the image feature information of each candidate box in the first training image and the text feature information of each category name in the limited target categories corresponding to the first training image.
[0106] In a possible implementation, the computing device matches and aligns the image feature information of each candidate box in the first training image with the text feature information of each category name in the limited target categories corresponding to the first training image to determine the category of the object in each candidate box in the first training image. For example, for the a-th candidate box in the first training image and the b-th category name in the limited target categories corresponding to the first training image, the computing device determines the similarity (such as cosine similarity) between the image feature information of the a-th candidate box and the text feature information of the b-th category name. Furthermore, the similarity between the a-th candidate box and the b-th category name can be determined. Based on this method, the similarity between each candidate box in the first training image and each category name in the limited target categories can be determined. Then, the category corresponding to the category name with the highest similarity to the candidate box is determined as the category of the object in the candidate box.
[0107] In a possible implementation, the computing device multiplies the image feature information of each candidate box by the text feature information of each category name in the limited target categories to determine the probability value that the object in each candidate box belongs to each category in the limited target categories. Furthermore, based on this probability value, it can be determined which category in the limited target categories the object in each candidate box belongs to, and then the first detection result of the first training image is obtained.
[0108] Illustrated by way of example, assume that the limited target categories corresponding to the first training image include m categories, such as Figure 5 as shown, including: person, hat, tree, and street lamp, these 4 categories. The first training image is input into the vision module for image feature extraction and n candidate boxes are generated. For the i-th (i = 1, 2, 3..n) candidate box among the n candidate boxes, the image feature information corresponding to the i-th candidate box is denoted as V i . At the same time, the computing device inputs the m target category names corresponding to the first training image, such as the category names of person, hat, tree, and street lamp, into the text module for category name feature extraction, and obtains the text feature information of each category name in the m target categories. For example, for the j-th (j = 1, 2, 3..m) target category among the m target categories, the text feature information of the category name of the j-th target category is denoted as T j . Then, the computing device takes the dot product of the image feature information of the n candidate boxes and the text feature information of the m category names S = VT *T.S∈R n×m For each candidate box among the n candidate boxes, the probability that the object in each candidate box belongs to each of the m target categories can be obtained.
[0109] The embodiments of the present application do not limit the specific network structure of the vision module. Exemplarily, the vision module may be the backbone network of any detection network (backbone), such as Fast RCNN, YOLO, etc.
[0110] The embodiments of the present application do not limit the specific network structure of the text module. Exemplarily, the text module may be the backbone network of any text recognition network (backbone), such as Bert, etc.
[0111] By performing the above processing on each first training image in the first training data, the computing device can obtain the first detection result of each first training image under the limited target categories.
[0112] Next, the computing device executes S102-B above, and determines the first loss of the object detection model based on the object detection result and the first detection result of the first training image under the limited target categories.
[0113] The embodiments of the present application do not limit the specific manner in which the computing device determines the first loss of the object detection model based on the object detection result and the first detection result of the first training image under the limited target categories.
[0114] In a possible implementation manner, based on the above steps, the computing device can determine the first detection result of each first training image in the first training data under the limited target categories. At the same time, the computing device obtains the object detection result of each first training image in the first training data under the limited target categories. In this way, the computing device can determine the first loss of the object detection model based on the object detection result and the first detection result of each first training image under the limited target categories. For example, the computing device determines the first loss of the object detection model based on the difference between the object detection result and the first detection result of each first training image under the limited target categories.
[0115] In another possible implementation manner, S102-B above includes the following steps S102-B1 to S102-B3:
[0116] S102-B1. Determine the category prediction loss of the object detection model based on the first detection result and the object detection result of the first training image under the limited target categories;
[0117] S102-B2. Determine the position prediction loss of the target detection model based on the position information of each candidate box in the first detection result and the position information of each object detection box in the target detection result;
[0118] S102-B3. Determine the first loss based on the class prediction loss and the position prediction loss.
[0119] In this implementation manner, the computing device determines the class prediction loss of the target detection model based on the first detection result and the target detection result of the first training image under limited target categories. At the same time, the computing device determines the position prediction loss of the target detection model based on the position information of each candidate box in the first detection result and the position information of each object detection box in the target detection result, and then determines the first loss based on the class prediction loss and the position prediction loss. That is to say, in the embodiments of the present application, the first loss of the target detection model is determined by two parts. The first part is the class prediction loss of the target detection model, and the second part is the position prediction loss of the target detection model.
[0120] The embodiments of the present application do not limit the specific manner in which the computing device determines the first detection result and the target detection result of the first training image under limited target categories and determines the class prediction loss of the target detection model.
[0121] In one example, the computing device compares the class of the object in each candidate box in the first detection result of the first training image under limited target categories with the class of the object in the corresponding detection box in the target detection result to determine the class prediction loss of the target detection model. For example, for the i-th candidate box in the first detection result, based on the position information, it is determined that the corresponding detection box of the i-th candidate box in the target detection result is the j-th detection box. In this way, the class prediction loss of the target detection model for the i-th candidate box can be determined by comparing the class of the object predicted by the target detection model in the i-th candidate box with the class of the object in the j-th detection box. Based on the class prediction loss corresponding to each candidate box in each first training image, the class prediction loss of the target detection model can be determined.
[0122] In one example, the above S102-B1 includes the following steps of S102-B11 and S102-B12:
[0123] S102-B11. Match each candidate box included in the first detection result with each object detection box included in the target detection result to determine the supervision label corresponding to each candidate box;
[0124] S102 - B12. Determine the class prediction loss based on the probability that each object in each candidate box included in the first detection result belongs to each category in the target category set, and the supervision label corresponding to each candidate box.
[0125] In this implementation manner, for each first training image in the first training data, the computing device matches each candidate box included in the first detection result of the first training image with each object detection box included in the object detection result of the first training image, and determines the supervision label corresponding to each candidate box included in the first detection result of the first training image.
[0126] For example, for each candidate box included in the first detection result of the first training image, such as the i-th candidate box, based on the position information of the i-th candidate box, perform size and position matching with the position information of each object detection box included in the object detection result of the first training image. If the size and position matching degree between the i-th candidate box and a certain detection box is greater than or equal to the preset value, it is determined that the supervision label of the i-th candidate box and the detection box is 1. If the size and position matching degree between the i-th candidate box and a certain detection box is less than the preset value, it is determined that the supervision label of the i-th candidate box and the detection box is 0. Repeat the execution to determine the supervision label between the i-th candidate box and each object detection box included in the object detection result of the first training image.
[0127] For another example, for the j-th candidate box included in the first detection result of the first training image, determine the intersection over union (IoU) between the j-th candidate box and each object detection box included in the object detection result of the first training image, where j is a positive integer. If the IoU between the j-th candidate box and the p-th object detection box among each object detection box included in the first object detection result is greater than or equal to the preset value (for example, 0.7, or other values), then set the supervision label Y j,o of the j-th candidate box under the p-th object detection box to 1, where p is a positive integer. If the IoU between the j-th candidate box and the p-th object detection box is less than the preset value, then set the supervision label Y j,p of the j-th candidate box under the p-th object detection box to 0. In this way, the supervision label between the j-th candidate box and each object detection box included in the object detection result of the first training image can be determined.
[0128] Referring to the above method, the computing device can determine the supervision label corresponding to each candidate box included in the first detection result of the first training image, and further determine the class prediction loss of the object detection model based on the probability that each object in each candidate box included in the first detection result belongs to each category in the finite target category, and the supervision label corresponding to each candidate box.
[0129] The embodiments of the present application do not limit the specific manner in which the computing device determines the class prediction loss of the object detection model based on the probability that each object in each candidate box included in the first detection result belongs to each category in the finite target categories, and the supervision label corresponding to each candidate box.
[0130] In a possible implementation, the computing device determines the cross-entropy loss between the probability that each object in each candidate box included in the first detection result of the first training image belongs to each category in the finite target categories corresponding to the first training image, and the supervision label corresponding to each candidate box, as the class prediction loss.
[0131] Exemplarily, the computing device determines the class prediction loss of the object detection model through the following formula (1):
[0132] L cls = CrossEntropy(S, Y) (1)
[0133] where L cls is the class prediction loss of the object detection model, CrossEntropy is the cross-entropy loss function, S is the probability that each object in each candidate box included in the first detection result of the first training image belongs to each category in the finite target categories corresponding to the first training image, and Y is the supervision label corresponding to each candidate box included in the first detection result of the first training image.
[0134] Next, the specific process of S102-B2 above, where the computing device determines the position prediction loss of the object detection model based on the position information of each candidate box in the first detection result and the position information of each object detection box in the object detection result, will be introduced.
[0135] It should be noted that there is no order of execution between S102-B2 and S102-B1 above. That is to say, S102-B2 can be executed before S102-B1, or after S102-B1, or executed synchronously with S102-B1.
[0136] The embodiments of the present application do not limit the specific manner in which the computing device determines the position prediction loss of the object detection model based on the position information of each candidate box in the first detection result and the position information of each object detection box in the object detection result.
[0137] In a possible implementation, for each candidate box, based on the position information of the candidate box, the object detection box closest to the position of the candidate box is determined, and based on the position information of the closest object detection box and the position information of the candidate box, the position deviation of the candidate box is determined. Further, based on the position deviation of each candidate box, the position prediction loss of the target detection model is determined.
[0138] In a possible implementation, the computing device determines the position prediction loss of the target detection model through a coordinate supervision loss function.
[0139] Exemplarily, the computing device determines the position prediction loss of the target detection model through the following formula (2):
[0140]
[0141]
[0142] where L box is the position prediction loss of the target detection model, (x u , y u , w u , h u ) is the position information of the candidate box, and (x, y, w, h) is the position information of the object detection box.
[0143] It should be noted that the computing device can also determine the position prediction loss of the target detection model based on the position information of each candidate box in the reference detection result and the position information of each object detection box in the target detection result in other ways, and the embodiments of the present application do not limit this.
[0144] After the computing device determines the class prediction loss and the position prediction loss of the target detection model based on the above steps, based on the class prediction loss and the position prediction loss, the first loss of the target detection model is determined. For example, the computing device determines the sum of the class prediction loss and the position prediction loss as the first loss.
[0145] Exemplarily, the computing device determines the first loss L1 through the following formula (3):
[0146] L1 = l cls + L box (3)
[0147] After the computing device determines the first loss of the target detection model through the above steps, it performs initial training on the target detection model based on this first loss. For example, based on this first loss, the parameters in the target detection model are updated to obtain the target detection model after initial training. As can be seen from the above, the first training data includes multiple first training images. During each initial training, a batch of first training images and their corresponding target detection results can be selected to train the target detection model. Then, another batch of first training images and their corresponding target detection results are used to adjust the parameters of the target detection model with adjusted parameters again. This process is repeated iteratively for multiple times, and finally, the target detection model after initial training is obtained.
[0148] S103. Use the second training data to retrain the target detection model after initial training to obtain the trained target detection model.
[0149] As can be seen from the above, the computing device uses the first training data with the detectable target categories limited to perform initial training on the target detection model to obtain the target detection model after initial training. Then, the second training data with the detectable target categories not limited is used to retrain the target detection model after initial training to obtain the trained target detection model.
[0150] In the embodiments of the present application, the second training data includes multiple second training images and the description information of each second training image. In this way, the computing device can process the second training images and the description information of the second training images based on the target detection model after initial training to determine the target detection results of the second training images. Finally, based on the second training images and the target detection results of the second training images, the target detection model after initial training is retrained to obtain the finally trained target detection model.
[0151] The embodiments of the present application do not limit the specific manner in which the computing device uses the second training data to retrain the target detection model after initial training to obtain the trained target detection model.
[0152] In some embodiments, the computing device uses the description information of the second training images in the second training data as the target classification set corresponding to the second training images, and then inputs the second training images and the description information into the initially trained object detection model to determine the object detection frames in the second training images, the image feature information of each object detection frame, and the text feature information of the description information. Then, the computing device matches the image feature information of each object detection frame in the second training images with the text feature information of the description information to determine the category of the object in each object detection frame in the second training images, thereby obtaining the object detection result of the second training images. Next, the computing device uses the initially trained object detection model to perform object detection on the second training images to obtain the detection result of the second training images predicted by the initially trained object detection model. For the sake of convenience of description, this detection result is denoted as the second detection result. Finally, the computing device determines the loss of the initially trained object detection model based on the second detection result and the object detection result of the second training images, and adjusts the parameters in the initially trained object detection model based on this loss. By repeating the iteration multiple times, the finally trained object detection model can be obtained.
[0153] In some embodiments, the above S103 includes the following steps S103-A to S103-E:
[0154] S103-A. For each second training image, extract at least one keyword from the description information of the second training image;
[0155] S103-B. Determine the object detection result of the second training image under at least one keyword;
[0156] S103-C. Process the second training image and at least one keyword through the initially trained object detection model to obtain the second detection result of the second training image under at least one keyword;
[0157] S103-D. Determine the second loss of the object detection model based on the object detection result and the second detection result of the second training image under at least one keyword;
[0158] S103-E. Based on the second loss, retrain the initially trained object detection model to obtain the trained object detection model.
[0159] In this implementation, when the computing device uses the second training data to retrain the initially trained object detection model, since the second training images included in the second training data do not limit the target categories, the computing device can determine the target category set corresponding to the second training images based on the description information of the second training images.
[0160] The embodiments of the present application do not limit the specific manner in which the computing device determines the target category set corresponding to the second training image based on the description information of the second training image.
[0161] In a possible implementation, by manually analyzing the description information of the second training image, the key information included in the description information of the second training image can be obtained, and then based on this key information, the target category set corresponding to the second training image can be determined. For example, the description information of the second training image is "a person wearing a hat, with a tree and a street lamp behind", so that a person can analyze this description information to obtain the key information included in this description information as: person, hat, tree, and street lamp. In this way, the four target categories of "person, hat, tree, and street lamp" can be determined as the target category set corresponding to the second training image.
[0162] In a possible implementation, the computing device extracts at least one keyword (or at least one noun) from the description information of the second training image through the Name Entity Recognition (NER) method, and then determines the at least one extracted keyword as the target category set corresponding to the second training image. For example, as Figure 6 shown, the description information of the second training image is "a person wearing a hat, with a tree and a street lamp behind", so that the computing device extracts at least one keyword from this description information through the NER method as: person, hat, tree, and street lamp. In this way, the computing device can determine the four target categories of "person, hat, tree, and street lamp" as the target category set corresponding to the second training image.
[0163] After the computing device extracts at least one keyword from the description information of the second training image, it uses the at least one keyword as the target category set of the second training image to determine the target detection result of the second training image under the at least one keyword.
[0164] In the embodiments of the present application, the ways for the computing device to determine the target detection result of the second training image under the at least one keyword at least include the following several kinds:
[0165] Method 1: Determine the target detection result of the second training image under the at least one keyword through manual annotation. For example, at least one keyword of the second training image includes "person, hat, tree, and street lamp", so that the target detection results of the second training image in the four categories of "person, hat, tree, and street lamp" are manually annotated, that is, the detection frames (i.e., position information) of the objects belonging to the four categories of "person, hat, tree, and street lamp" in the second training image are annotated, as well as the category of the object in each detection frame.
[0166] In the second method, the computing device processes the second training image and at least one keyword corresponding to the second training image through the object detection model after initial training, and obtains the object detection result of the second training image under the at least one keyword.
[0167] The object detection model in the embodiments of the present application is an open-category detection model that can detect objects of any category. After the computing device uses the first training data to perform initial training on the object detection model, the object detection model after the initial training has a certain detection ability for objects of any category. Based on this, the computing device can process the second training image and at least one keyword corresponding to the second training image through the object detection model after initial training, and obtain the object detection result of the second training image under the at least one keyword.
[0168] The embodiments of the present application do not limit the specific method for the computing device to process the second training image and at least one keyword corresponding to the second training image through the object detection model after initial training, and obtain the object detection result of the second training image under the at least one keyword.
[0169] In a possible implementation manner, the data input by the computing device into the object detection model after initial training, in addition to the second training image, further includes at least one keyword corresponding to the second training image. In this way, the object detection model after initial training detects each object detection box in the second training image, determines the keyword corresponding to each object detection box from the at least one keyword, and then determines the keyword corresponding to the object detection box as the category of the object detection box, so as to determine the object detection result of the second training image.
[0170] In a possible implementation manner, such as Figure 4As shown, the object detection model of the embodiment of the present application includes a visual module and a text module, and the corresponding initially trained object detection model includes an initially trained visual module and an initially trained text module. Based on this, the computing device performs object detection on the second training image through the initially trained visual module, obtains the object detection boxes of the second training image, and extracts the image feature information of each object detection box in the second training image. At the same time, the computing device extracts the text feature information of each keyword in at least one keyword through the initially trained text module. For each object detection box in the second training image, for example, the i-th object detection box, the image feature information of the i-th object detection box is matched with the text feature information of each keyword to determine the class label of the object in the i-th object detection box, where i is a positive integer. For example, the similarity (such as cosine similarity, etc.) between the image feature information of the i-th object detection box and the text feature information of each keyword is determined. If the similarity between the image feature information of the i-th object detection box and the text feature information of a certain keyword is greater than or equal to the preset threshold, then this keyword is determined as the class of the object in the i-th object detection box. In this way, the computing device can obtain the position information of each object detection box in the second training image and the class of the object in each object detection box, and obtain the object detection result of the second training image.
[0171] That is to say, the general data (i.e., the second training data) includes the second training image and the corresponding text description information. Since the general data is used to train the text-image multi-modal alignment, most of the text description information can correspond to the image visual content. However, for the detection task, the object detection box information is still lacking. Therefore, the embodiment of the present application uses the above initially trained object detection model to perform inference on the general data set, and the object detection boxes in the second training image can be obtained. In order to determine the corresponding object class in the object detection box in the second training image, it is necessary to match the visual features in the object detection box with the possible class name features. Therefore, for the text description information, as Figure 6 shown, the computing device first uses the NER method to extract at least one keyword from the description information of the second training image. Then, using the initially trained object detection model, the text feature information of these keywords is extracted, and the image feature information of the object detection box is matched with these text feature information one by one. The keyword with a matching similarity greater than the threshold is determined as the class of the object in the object detection box. In this way, on the general data set, the object class corresponding to each second training image and the corresponding detection box can be obtained, that is, the object detection result of each second training image is obtained, which is used to retrain the initially trained object detection model.
[0172] After the computing device determines the object detection result of the second training image based on the above steps, it executes the steps of S103-C above, and processes the second training image and at least one keyword through the object detection model after initial training to obtain the second detection result of the second training image under at least one keyword.
[0173] In the embodiments of the present application, for each second training image in the second training data, the computing device processes the second training image and at least one keyword corresponding to the second training image to obtain the detection result of the second training image under the at least one keyword. For the sake of description, this detection result is denoted as the second detection result.
[0174] The embodiments of the present application do not limit the specific network structure of the object detection model.
[0175] In a possible implementation manner, as Figure 4 shown, the object detection model of the embodiments of the present application includes a visual module and a text module, where the visual module is used to process the input training image, and the text module is used to process the input text data. Based on this, the above S103-C includes the following steps of S103-C1 to S103-C3:
[0176] S103-C1. Through the visual module after initial training, perform object detection on the second training image to obtain the candidate boxes of the objects in the second training image, and extract the image feature information of each candidate box;
[0177] S103-C2. Through the text module after initial training, extract the text feature information of each keyword in at least one keyword corresponding to the second training image;
[0178] S103-C3. Based on the image feature information of each candidate box and the text feature information of each keyword in at least one keyword, determine the second detection result of the second training image under the at least one keyword.
[0179] As can be seen from the above, the second training data of the embodiments of the present application includes multiple second training images, and the object detection results of each second training image under the corresponding at least one keyword. As Figure 4 shown, the object detection model includes a visual module and a text module. In this way, for each second training image in the second training data, the computing device can input the second training image into the visual module for object detection and feature extraction, and can obtain the candidate boxes of the objects in the second training image, and the image feature information of each candidate box. At the same time, the computing device inputs at least one keyword corresponding to the second training image into the text module for processing to obtain the text feature information of each keyword in the at least one keyword.
[0180] Next, the computing device determines a second detection result of the second training image under the limited target category based on the image feature information of each candidate box in the second training image and the text feature information of each keyword among at least one keyword corresponding to the second training image.
[0181] The embodiments of the present application do not limit the specific manner in which the computing device determines a second detection result of the second training image under the limited target category based on the image feature information of each candidate box in the second training image and the text feature information of each keyword among at least one keyword corresponding to the second training image.
[0182] In a possible implementation, the computing device matches and aligns the image feature information of each candidate box in the second training image with the text feature information of each keyword among at least one keyword corresponding to the second training image to determine the category of the object in each candidate box in the second training image. For example, for the a-th candidate box in the second training image and the b-th keyword among at least one keyword corresponding to the second training image, the computing device determines the similarity (such as cosine similarity) between the image feature information of the a-th candidate box and the text feature information of the b-th keyword, and then can determine the similarity between the a-th candidate box and the b-th keyword. Based on this method, the similarity between each candidate box in the second training image and each keyword among at least one keyword can be determined, and then the keyword with a similarity greater than or equal to a preset threshold to the candidate box is determined as the category of the object in the candidate box.
[0183] In a possible implementation, the computing device multiplies the image feature information of each candidate box and the text feature information of each keyword among at least one keyword to determine the probability value that the object in each candidate box belongs to each keyword among at least one keyword. Then, based on this probability value, it can be determined which keyword among at least one keyword the object in each candidate box belongs to, and then the second detection result of the second training image is obtained.
[0184] By performing the above processing on each second training image in the second training data, the computing device can obtain the second detection result of each second training image under the corresponding at least one keyword.
[0185] Next, the computing device executes the above S103-D, and determines a second loss of the target detection model based on the target detection result and the second detection result of the second training image under at least one keyword.
[0186] The embodiments of the present application do not limit the specific manner in which the computing device determines the second loss of the object detection model based on the object detection results and the second detection results of the second training images under the corresponding at least one keyword.
[0187] In a possible implementation manner, based on the above steps, the computing device can determine the second detection results of each second training image in the second training data under the corresponding at least one keyword. At the same time, the computing device obtains the object detection results of each second training image in the second training data under the corresponding at least one keyword. In this way, the computing device can determine the second loss of the object detection model based on the object detection results and the second detection results of each second training image under the corresponding at least one keyword. For example, the computing device determines the second loss of the object detection model based on the difference between the object detection results and the second detection results of each second training image under the corresponding at least one keyword.
[0188] In another possible implementation manner, the above S103-D includes the following steps S103-D1 to S103-D3:
[0189] S103-D1. Determine the class prediction loss of the object detection model based on the second detection results and the object detection results of the second training images under the corresponding at least one keyword;
[0190] S103-D2. Determine the position prediction loss of the object detection model based on the position information of each candidate box in the second detection results and the position information of each object detection box in the object detection results;
[0191] S103-D3. Determine the second loss based on the class prediction loss and the position prediction loss.
[0192] In this implementation manner, the computing device determines the class prediction loss of the object detection model based on the second detection results and the object detection results of the second training images under the corresponding at least one keyword. At the same time, the computing device determines the position prediction loss of the object detection model based on the position information of each candidate box in the second detection results and the position information of each object detection box in the object detection results, and then determines the second loss based on the class prediction loss and the position prediction loss. That is to say, in the embodiments of the present application, the second loss of the object detection model is determined by two parts. The first part is the class prediction loss of the object detection model, and the second part is the position prediction loss of the object detection model.
[0193] The embodiments of the present application do not limit the specific manner in which the computing device determines the second detection results and the object detection results of the second training images under the corresponding at least one keyword and determines the class prediction loss of the object detection model.
[0194] In one example, the computing device compares the category of the object in each candidate box in the second detection result of the second training image under at least one corresponding keyword with the category of the object in the corresponding detection box in the target detection result to determine the category prediction loss of the target detection model. For example, for the i-th candidate box in the second detection result, based on the position information, it is determined that the detection box corresponding to the i-th candidate box in the target detection result is the j-th detection box. In this way, the category prediction loss of the target detection model for the i-th candidate box can be determined by comparing the category of the object in the i-th candidate box predicted by the target detection model with the category of the object in the j-th detection box. Based on the category prediction loss corresponding to each candidate box in each second training image, the category prediction loss of the target detection model can be determined.
[0195] In one example, the above S103-D1 includes the following steps of S103-D11 and S103-D12:
[0196] S103-D11: Match each candidate box included in the second detection result with each object detection box included in the target detection result to determine the supervision label corresponding to each candidate box;
[0197] S103-D12: Determine the category prediction loss based on the probability that the object in each candidate box included in the second detection result belongs to each category in the target category set and the supervision label corresponding to each candidate box.
[0198] In this implementation, for each second training image in the second training data, the computing device matches each candidate box included in the second detection result of the second training image with each object detection box included in the target detection result of the second training image to determine the supervision label corresponding to each candidate box included in the second detection result of the second training image.
[0199] For example, for each candidate box included in the second detection result of the second training image, such as the i-th candidate box, based on the position information of the i-th candidate box, a size and position match is performed with the position information of each object detection box included in the target detection result of the second training image. If the size and position match degree between the i-th candidate box and a certain detection box is greater than or equal to the preset value, it is determined that the supervision label of the i-th candidate box and the detection box is 1. If the size and position match degree between the i-th candidate box and a certain detection box is less than the preset value, it is determined that the supervision label of the i-th candidate box and the detection box is 0. By repeating the execution, the supervision label between the i-th candidate box and each object detection box included in the target detection result of the second training image can be determined.
[0200] For another example, for the j-th candidate box included in the second detection result of the second training image, determine the intersection over union (IoU) between the j-th candidate box and each object detection box included in the target detection result of the second training image, where j is a positive integer. If the IoU between the j-th candidate box and the p-th object detection box among each object detection box included in the second target detection result is greater than or equal to a preset value (e.g., 0.7, or other values), then set the supervision label Y of the j-th candidate box under the p-th object detection box j,p to 1, where p is a positive integer. If the IoU between the j-th candidate box and the p-th object detection box is less than the preset value, then set the supervision label Y of the j-th candidate box under the p-th object detection box j,p to 0. In this way, the supervision label between the j-th candidate box and each object detection box included in the target detection result of the second training image can be determined.
[0201] Referring to the above method, the computing device can determine the supervision label corresponding to each candidate box included in the second detection result of the second training image, and then, based on the probability that the object in each candidate box included in the second detection result belongs to each keyword of at least one keyword, and the supervision label corresponding to each candidate box, determine the class prediction loss of the target detection model.
[0202] The embodiments of the present application do not limit the specific manner in which the computing device determines the class prediction loss of the target detection model based on the probability that the object in each candidate box included in the second detection result belongs to each keyword of at least one keyword, and the supervision label corresponding to each candidate box.
[0203] In a possible implementation manner, the computing device determines the cross-entropy loss between the probability that the object in each candidate box included in the second detection result of the second training image belongs to each keyword of at least one keyword corresponding to the second training image, and the supervision label corresponding to each candidate box, as the class prediction loss.
[0204] Next, the specific process of S103-D2 above, where the computing device determines the position prediction loss of the target detection model based on the position information of each candidate box in the second detection result and the position information of each object detection box in the target detection result, will be introduced.
[0205] It should be noted that when specifically executing S103-D2 and S103-D1 above, there is no order requirement. That is to say, S103-D2 can be executed before S103-D1, or after S103-D1, or executed synchronously with S103-D1.
[0206] The embodiments of the present application do not limit the specific manner in which the computing device determines the position prediction loss of the target detection model based on the position information of each candidate box in the second detection result and the position information of each object detection box in the target detection result.
[0207] In a possible implementation, for each candidate box, based on the position information of the candidate box, the object detection box closest to the position of the candidate box is determined, and based on the position information of the closest object detection box and the position information of the candidate box, the position deviation of the candidate box is determined. Further, based on the position deviation of each candidate box, the position prediction loss of the target detection model is determined.
[0208] In a possible implementation, the computing device determines the position prediction loss of the target detection model through a coordinate supervised loss function. Specifically, reference can be made to the above formula (2), which will not be elaborated here.
[0209] After the computing device determines the class prediction loss and the position prediction loss of the target detection model with respect to the second training image based on the above steps, based on the class prediction loss and the position prediction loss, the computing device determines the second loss of the target detection model with respect to the second training image. For example, the computing device determines the sum of the class prediction loss and the position prediction loss as the second loss.
[0210] After the computing device determines the second loss of the target detection model through the above steps, the computing device retrains the initially trained target detection model based on the second loss. For example, based on the second loss, the parameters in the initially trained target detection model are updated to obtain the trained target detection model. As can be seen from the above, the second training data includes multiple second training images. During each training, a batch of second training images and their corresponding target detection results can be selected to train the initially trained target detection model, and then another batch of second training images and their corresponding target detection results can be used to adjust the parameters of the target detection model with adjusted parameters again. This process is repeated multiple times, and finally the trained target detection model is obtained.
[0211] In some embodiments, the computing device can use the first training data and the second training data to retrain the initially trained target detection model to obtain the trained target detection model. Among them, the specific manner of using the first training data and the second training data to retrain the initially trained target detection model is basically the same as the specific process of using the second training data to retrain the initially trained target detection model described above, and reference can be made to the above description.
[0212] The model training method provided by the embodiments of the present application obtains first training data and second training data, where the first training data includes a plurality of first training images and the object detection results of each first training image under limited object categories, and the second training data includes a plurality of second training images and the description information of each second training image. Then, the object detection model is initially trained with the first training data, and the initially trained object detection model is retrained with the second training data to obtain the trained object detection model. That is to say, the embodiments of the present application train the object detection model with the first training images whose detectable object categories are limited and the second training images whose detectable object categories are not limited, so that the trained object detection model can detect objects of any newly added category. In this way, in the actual application scenario, when new objects of interest need to be added, only the N object categories of interest and the object images to be detected are input into the trained object detection model, and then the object detection model with unlimited detectable object categories processes the object images and the category names of the N object categories, so as to detect the objects belonging to the N object categories in the object images, without the need to update the object detection model again or add new object detectors, thereby reducing the cost of object detection and improving the efficiency and flexibility of object detection.
[0213] As described above in conjunction with Figures 2 to 6 , the embodiments of the model training method of the present application have been described in detail. Next, the object detection method provided by the embodiments of the present application will be introduced.
[0214] Figure 7 FIG. is a schematic flowchart of an object detection method provided by an embodiment of the present application. The execution subject of the embodiment of the present application is a device with an object detection function, such as an object detection device. In some embodiments, the object detection device may be a Figure 1 computing device in Figure 1 , or a Figure 1 terminal device in
[0215] Figure 7
[0216]
[0217] S201. Obtain an object image to be detected and N object categories.
[0218] where N is a positive integer.
[0218] In the embodiments of the present application, the target detection model is a detection model whose detectable target categories are not limited. Thus, when target detection needs to be performed using this target detection model, not only the target image to be detected needs to be input, but also N target categories of interest need to be input. That is to say, in the embodiments of the present application, when different target categories need to be detected, only the target categories of interest need to be input into the target detection model, and this target detection model can achieve the detection of different target categories without adjusting the network output structure or retraining the model, greatly reducing the cost of target detection and enhancing the flexibility of target detection.
[0219] S202. Process the target image and the class names of N target categories using a target detection model whose detectable target categories are not limited, and detect the objects in the target image that belong to the N target categories.
[0220] Among them, the target detection model is trained based on a first training image and a second training image. The first training image is an image in which the detectable target categories are limited, and the second training image is an image in which the detectable target categories are not limited.
[0221] Among them, the specific training process of the target detection model refers to the specific description of the above embodiments.
[0222] In the embodiments of the present application, when the computing device performs target detection, it first needs to define a set N = {N1, N2,... N n} of the class names of the target categories of interest. The set of class names and the target image are used as the input of the target detection model, and the detection of various target categories of interest can be achieved. That is to say, in the embodiments of the present application, if the target categories of interest need to be modified, only the set of class names of the input target categories needs to be adjusted, without adjusting the network output structure of the target detection model or retraining the model.
[0223] For example, if the user needs to detect the detection results of target image 1 under three categories, namely category A, category B, and category C, and needs to detect the detection results of target image 2 under four categories, namely category D, category E, category F, and category B. During actual detection, the computing device only needs to obtain target image 1 and the three categories of category A, category B, and category C, and input target image 1 and the three categories of category A, category B, and category C into the above-mentioned trained target detection model for processing. The target detection model can then output the detection frames and categories of the objects in target image 1 that belong to the three categories of category A, category B, and category C, that is, the target detection of target image 1 under the three categories of category A, category B, and category C is achieved.
[0224] Similarly, the computing device can obtain the target image 2, as well as four categories, namely category D, category E, category F, and category B, and input the target image 2 and these four categories into the above-mentioned trained object detection model for processing. The object detection model can then output the detection frames and categories of the objects in the target image 2 that belong to the four categories of category D, category E, category F, and category B, that is, the object detection of the target image 2 in the four categories of category D, category E, category F, and category B is achieved.
[0225] The embodiments of the present application do not limit the specific network structure of the object detection model.
[0226] In a possible implementation manner, as Figure 8 shown, the object detection model includes a visual module and a text module. In this way, the computing device can perform object detection on the target image through the visual module, obtain the detection frames of the objects in the target image, and extract the image feature information of each detection frame in the target image. At the same time, the computing device extracts the text feature information of each category name in the above-mentioned N target categories through the text module. Finally, for the k-th detection frame in the target image, the image feature information of the k-th detection frame is matched with the text feature information of each category name in the N target categories to determine the category of the object in the k-th detection frame, where k is a positive integer. For example, determine the similarity between the image feature information of the k-th detection frame and the text feature information of each category name in the N target categories; if the similarity between the image feature information of the k-th detection frame and the text feature information of the first category name in the N target categories is greater than the preset threshold, then determine that the category of the object in the k-th detection frame is the first category. Another example is to multiply the image feature information of the k-th detection frame by the text feature information of each target category in the N target categories to determine the probability values of the object in the k-th detection frame belonging to each target category in the N target categories, and then based on these probability values, determine the category of the object in the k-th detection frame. Repeating the above steps can determine the category of the object in each detection frame included in the target image, and thus achieve the object detection of the target image.
[0227] The target detection method provided by the embodiments of the present application obtains a target image to be detected and N target categories of interest, and processes the target image and the category names of the N target categories through a target detection model with unlimited detectable target categories, so as to detect the objects belonging to the N target categories in the target image. That is to say, in the embodiments of the present application, during application, only the set of category names of the target categories of interest and the target image need to be used as the input of the target detection model, and the target detection model can match and detect the categories of the objects in the target image, thereby achieving flexible adjustment of the set of target categories of interest without adjusting the model structure of the target detection model or retraining the target detection model, and can quickly implement the detection ability for newly added object categories in practical applications.
[0228] As described above in conjunction with Figures 2 to 8 , the embodiments of the model training and target detection methods of the present application have been described in detail. Below in conjunction with Figure 9 , the device embodiments of the present application will be described in detail.
[0229] Figure 9 FIG. is a schematic block diagram of a target detection device provided by an embodiment of the present application. The device 10 can be applied to a computing device.
[0230] As Figure 9 shown, the target detection device 10 includes:
[0231] An acquisition unit 11, configured to acquire a target image to be detected and N target categories, where N is a positive integer;
[0232] A detection unit 12, configured to process the target image and the category names of the N target categories through a target detection model with unlimited detectable target categories, so as to detect the objects belonging to the N target categories in the target image;
[0233] Wherein, the target detection model is trained based on a first training image and a second training image, the first training image is an image with limited detectable target categories, and the second training image is an image with unlimited detectable target categories.
[0234] In some embodiments, the object detection model includes a visual module and a text module; the detection unit 12 is specifically configured to perform object detection on the target image through the visual module to obtain detection frames of each object in the target image, and extract image feature information of each detection frame in the target image; through the text module, extract text feature information of each class name in the N target classes; for the k-th detection frame in the target image, match the image feature information of the k-th detection frame with the text feature information of each class name in the N target classes to determine the class of the object in the k-th detection frame, where k is a positive integer.
[0235] In some embodiments, the detection unit 12 is specifically configured to determine the similarity between the image feature information of the k-th detection frame and the text feature information of each class name in the N target classes; if the similarity between the image feature information of the k-th detection frame and the text feature information of the first class name in the N target classes is greater than a preset threshold, then determine that the class of the object in the k-th detection frame is the first class.
[0236] In some embodiments, the training process of the object detection model includes: obtaining first training data and second training data, where the first training data includes a plurality of first training images and the object detection results of each first training image under a limited number of target classes, and the second training data includes a plurality of second training images and the description information of each second training image; performing initial training on the object detection model through the first training data; performing re-training on the object detection model after initial training through the second training data to obtain the trained object detection model.
[0237] In some embodiments, the performing initial training on the object detection model through the first training data includes: for each first training image in the first training data, processing the first training image and the class names of the limited number of target classes corresponding to the first training image through the object detection model to obtain the first detection result of the first training image under the limited number of target classes; based on the object detection result and the first detection result of the first training image under the limited number of target classes, determining the first loss of the object detection model; and performing initial training on the object detection model based on the first loss.
[0238] In some embodiments, the object detection result of the first training image under a limited number of target classes is obtained by performing object detection on the first training image through at least one trained object detector.
[0239] In some embodiments, re-training the initially trained object detection model with the second training data to obtain the trained object detection model includes: for each second training image, extracting at least one keyword from the description information of the second training image; determining the object detection result of the second training image under the at least one keyword; processing the second training image and the at least one keyword through the initially trained object detection model to obtain a second detection result of the second training image under the at least one keyword; determining a second loss of the object detection model based on the object detection result and the second detection result of the second training image under the at least one keyword; and re-training the initially trained object detection model based on the second loss to obtain the trained object detection model.
[0240] In some embodiments, determining the object detection result of the second training image under the at least one keyword includes: processing the second training image and the at least one keyword through the initially trained object detection model to obtain the object detection result of the second training image under the at least one keyword.
[0241] In some embodiments, the object detection model includes a visual module and a text module; processing the second training image and the at least one keyword through the initially trained object detection model to obtain the object detection result of the second training image under the at least one keyword includes: performing object detection on the second training image through the initially trained visual module to obtain the object detection bounding boxes of the second training image, and extracting the image feature information of each object detection bounding box in the second training image; extracting the text feature information of each keyword in the at least one keyword through the initially trained text module; for the i-th object detection bounding box in the second training image, matching the image feature information of the i-th object detection bounding box with the text feature information of each keyword to determine the class label of the object in the i-th object detection bounding box, where i is a positive integer.
[0242] In some embodiments, the target detection model includes a visual module and a text module, and the reference detection result includes any one of a first detection result and a second detection result. Determining the reference detection result includes the following steps: Through the visual module, perform object detection on the target training image to obtain candidate boxes of objects in the target training image, and extract image feature information of each candidate box. The target training image is the first training image or the second training image; Through the text module, extract text feature information of each class name in the target class set corresponding to the target training image. If the target training image is the first training image, the target class set includes a finite number of target classes corresponding to the first training image. If the target training image is the second training image, the target class set includes at least one keyword corresponding to the second training image; Based on the image feature information of each candidate box and the text feature information of each class name in the target class set, determine the reference detection result.
[0243] In some embodiments, the determining the reference detection result based on the image feature information of each candidate box and the text feature information of each class name in the target class set includes: Multiply the image feature information of each candidate box and the text feature information of each class name in the target class set to determine the probability value that the object in each candidate box belongs to each class in the target class set.
[0244] In some embodiments, the target loss includes a first loss and a second loss. Determining the target loss includes the following steps: Based on the reference detection result of the target image under the target class set and the first target detection result, determine the class prediction loss of the target detection model. If the reference detection result is the first detection result, the first target detection result is the target detection result of the first training image under the finite target classes. If the reference detection result is the second detection result, the first target detection result is the target detection result of the second training image under the at least one keyword; Based on the position information of each candidate box in the reference detection result and the position information of each object detection box in the target detection result, determine the position prediction loss of the target detection model; Based on the class prediction loss and the position prediction loss, determine the target loss.
[0245] In some embodiments, determining the class prediction loss of the target detection model based on the reference detection result and the first target detection result of the target image under the target class set includes: matching each candidate box included in the reference detection result with each object detection box included in the first target detection result to determine the supervision label corresponding to each candidate box; determining the class prediction loss based on the probability that the objects in each candidate box included in the reference detection result belong to each class in the target class set and the supervision label corresponding to each candidate box.
[0246] In some embodiments, the matching each candidate box included in the reference detection result with each object detection box included in the first target detection result to determine the supervision label corresponding to each candidate box includes: for the j-th candidate box, determining the intersection over union (IoU) between the j-th candidate box and each object detection box included in the first target detection result, where j is a positive integer; if the IoU between the j-th candidate box and the p-th object detection box among the object detection boxes included in the first target detection result is greater than or equal to a preset value, setting the supervision label of the j-th candidate box under the p-th object detection box to 1, where p is a positive integer; if the IoU between the j-th candidate box and the p-th object detection box is less than the preset value, setting the supervision label of the j-th candidate box under the p-th object detection box to 0.
[0247] In some embodiments, the determining the class prediction loss based on the probability that the objects in each candidate box included in the reference detection result belong to each class in the target class set and the supervision label corresponding to each candidate box includes: determining the cross-entropy loss between the probability that the objects in each candidate box belong to each class in the target class set and the supervision label corresponding to each candidate box as the class prediction loss.
[0248] In some embodiments, the determining the target loss based on the class prediction loss and the location prediction loss includes: determining the sum of the class prediction loss and the location prediction loss as the target loss.
[0249] It should be understood that the apparatus embodiments and the method embodiments can correspond to each other, and similar descriptions can refer to the method embodiments. To avoid repetition, it will not be elaborated here. Specifically, Figure 9 The illustrated apparatus can execute the embodiments of the above-mentioned target detection method, and the foregoing and other operations and / or functions of each module in the apparatus respectively implement the corresponding method embodiments of the computing device. For the sake of brevity, it will not be elaborated here.
[0250] In the foregoing, the apparatus according to the embodiments of the present application has been described from the perspective of functional modules. It should be understood that the functional modules can be implemented in the form of hardware, can also be implemented by instructions in the form of software, or can be implemented by a combination of hardware and software modules. Specifically, the steps of the method embodiments in the present application can be completed by the integrated logic circuit of the hardware in the processor and / or instructions in the form of software. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or can be executed and completed by a combination of the hardware and software modules in the decoding processor. Optionally, the software module can be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps in the foregoing method embodiments.
[0251] Figure 10 is a schematic block diagram of a computing device provided by an embodiment of the present application, Figure 10 and the computing device can be used to execute the foregoing model training method and / or object detection method.
[0252] As Figure 10 shown, the computing device 30 may include:
[0253] A memory 31 and a processor 32. The memory 31 is used to store a computer program 33 and transmit the program code 33 to the processor 32. In other words, the processor 32 can call and run the computer program 33 from the memory 31 to implement the method in the embodiments of the present application.
[0254] For example, the processor 32 can be used to execute the steps in the foregoing method according to the instructions in the computer program 33.
[0255] In some embodiments of the present application, the processor 32 may include, but is not limited to:
[0256] A general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and the like.
[0257] In some embodiments of the present application, the memory 31 includes, but is not limited to:
[0258] Volatile memory and / or non-volatile memory. Among them, the non-volatile memory can be Read-Only Memory (ROM), Programmable ROM (PROM), Erasable PROM (EPROM), Electrically Erasable PROM (EEPROM), or flash memory. The volatile memory can be Random Access Memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double DataRate SDRAM (DDR SDRAM), Enhanced SDRAM (ESDRAM), synch link DRAM (SLDRAM), and Direct Rambus RAM (DR RAM).
[0259] In some embodiments of the present application, the computer program 33 can be divided into one or more modules, which are stored in the memory 31 and executed by the processor 32 to complete the method for recording a page provided by the present application. The one or more modules can be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program 33 in the computing device.
[0260] As Figure 10 shown, the computing device 30 may further include:
[0261] A transceiver 34, which can be connected to the processor 32 or the memory 31.
[0262] Among them, the processor 32 can control the transceiver 34 to communicate with other devices. Specifically, it can send information or data to other devices, or receive information or data sent by other devices. The transceiver 34 can include a transmitter and a receiver. The transceiver 34 may further include an antenna, and the number of antennas can be one or more.
[0263] It should be understood that the various components in the computing device 30 are connected through a bus system, where the bus system includes, in addition to the data bus, a power bus, a control bus, and a status signal bus.
[0264] According to one aspect of the present application, there is provided a computer storage medium having a computer program stored thereon, and when the computer program is executed by a computer, the computer is enabled to execute the method of the above method embodiment. Or rather, the embodiment of the present application further provides a computer program product containing instructions, and when the instructions are executed by a computer, the computer is enabled to execute the method of the above method embodiment.
[0265] According to another aspect of the present application, there is provided a computer program product or a computer program, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, enabling the computer device to execute the method of the above method embodiment.
[0266] In other words, when implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server, a data center, etc. that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0267] Those of ordinary skill in the art will realize that the modules and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0268] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or modules can be electrical, mechanical, or other forms.
[0269] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. For example, in each embodiment of this application, the functional modules can be integrated into a processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.
[0270] The above content is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A target detection method, characterized in that, Including: Obtaining a target image to be detected and N target categories, where N is a positive integer; Processing the target image and the category names of the N target categories through a target detection model with no limitation on detectable target categories to detect the objects in the target image that belong to the N target categories; Wherein, the target detection model is trained based on a first training image and a second training image, the first training image is an image with limited detectable target categories, and the second training image is an image with no limitation on detectable target categories.
2. The method according to claim 1, characterized in that, The target detection model includes a visual module and a text module; The processing the target image and the category names of the N target categories through a target detection model with no limitation on detectable target categories to detect the objects in the target image that belong to the N target categories includes: Performing object detection on the target image through the visual module to obtain the detection frames of each object in the target image, and extracting the image feature information of each detection frame in the target image; Extracting the text feature information of each category name in the N target categories through the text module; For the k-th detection frame in the target image, matching the image feature information of the k-th detection frame with the text feature information of each category name in the N target categories to determine the category of the object in the k-th detection frame, where k is a positive integer.
3. The method according to claim 2, wherein The matching the image feature information of the k-th detection frame with the text feature information of each category name in the N target categories to determine the category of the object in the k-th detection frame includes: Determining the similarity between the image feature information of the k-th detection frame and the text feature information of each category name in the N target categories; If the similarity between the image feature information of the k-th detection frame and the text feature information of the first category name in the N target categories is greater than a preset threshold, then determining that the category of the object in the k-th detection frame is the first category.
4. The method according to any one of claims 1 to 3, characterized in that, The training process of the target detection model includes: Obtaining first training data and second training data, the first training data includes a plurality of first training images and the target detection results of each first training image under limited target categories, and the second training data includes a plurality of second training images and the description information of each second training image; Performing initial training on the target detection model through the first training data; Performing retraining on the initially trained target detection model through the second training data to obtain the trained target detection model.
5. The method according to claim 4, wherein The performing initial training on the target detection model through the first training data includes: For each first training image in the first training data, processing the first training image and the category names of the limited target categories corresponding to the first training image through the target detection model to obtain the first detection results of the first training image under the limited target categories; Determine a first loss of the object detection model based on the object detection result and the first detection result of the first training image under the limited object categories; Based on the first loss, perform initial training on the object detection model.
6. The method according to claim 5, wherein The object detection result of the first training image under the limited object categories is obtained by performing object detection on the first training image using at least one trained object detector.
7. The method according to claim 4, wherein The retraining of the initially trained object detection model using the second training data to obtain a trained object detection model includes: For each second training image, extract at least one keyword from the description information of the second training image; Determine the object detection result of the second training image under the at least one keyword; Process the second training image and the at least one keyword through the initially trained object detection model to obtain a second detection result of the second training image under the at least one keyword; Determine a second loss of the object detection model based on the object detection result and the second detection result of the second training image under the at least one keyword; Based on the second loss, perform retraining on the initially trained object detection model to obtain the trained object detection model.
8. The method according to claim 7, wherein The determination of the object detection result of the second training image under the at least one keyword includes: Process the second training image and the at least one keyword through the initially trained object detection model to obtain the object detection result of the second training image under the at least one keyword.
9. The method according to claim 8, characterized in that, The object detection model includes a visual module and a text module; The process of processing the second training image and the at least one keyword through the initially trained object detection model to obtain the object detection result of the second training image under the at least one keyword includes: Perform object detection on the second training image through the initially trained visual module to obtain the object detection boxes of the second training image, and extract the image feature information of each object detection box in the second training image; Extract the text feature information of each keyword in the at least one keyword through the initially trained text module; For the i-th object detection box in the second training image, match the image feature information of the i-th object detection box with the text feature information of each keyword to determine the class label of the object in the i-th object detection box, where i is a positive integer.
10. The method according to claim 5 or 7, characterized in that, The object detection model includes a visual module and a text module, and the reference detection result includes any one of the first detection result and the second detection result. The determination of the reference detection result includes the following steps: Perform object detection on the target training image through the visual module to obtain the candidate boxes of the objects in the target training image, and extract the image feature information of each candidate box, where the target training image is the first training image or the second training image; Through the text module, extract the text feature information of each class name in the target class set corresponding to the target training image. If the target training image is the first training image, the target class set includes the limited target classes corresponding to the first training image. If the target training image is the second training image, the target class set includes at least one keyword corresponding to the second training image; Based on the image feature information of each candidate box and the text feature information of each class name in the target class set, determine the reference detection result.
11. The method according to claim 10, wherein The determining the reference detection result based on the image feature information of each candidate box and the text feature information of each class name in the target class set includes: Multiply the image feature information of each candidate box by the text feature information of each class name in the target class set to determine the probability value that the object in each candidate box belongs to each class in the target class set.
12. The method according to claim 10, characterized in that, The target loss includes a first loss and a second loss. Determining the target loss includes the following steps: Based on the reference detection result of the target image under the target class set and the first target detection result, determine the class prediction loss of the target detection model. If the reference detection result is the first detection result, the first target detection result is the target detection result of the first training image under the limited target classes. If the reference detection result is the second detection result, the first target detection result is the target detection result of the second training image under the at least one keyword; Based on the position information of each candidate box in the reference detection result and the position information of each object detection box in the target detection result, determine the position prediction loss of the target detection model; Based on the class prediction loss and the position prediction loss, determine the target loss.
13. The method according to claim 12, characterized in that, The determining the class prediction loss of the target detection model based on the reference detection result of the target image under the target class set and the first target detection result includes: Match each candidate box included in the reference detection result with each object detection box included in the first target detection result to determine the supervision label corresponding to each candidate box; Based on the probability that the object in each candidate box included in the reference detection result belongs to each class in the target class set and the supervision label corresponding to each candidate box, determine the class prediction loss.
14. The method according to claim 13, characterized in that, The matching each candidate box included in the reference detection result with each object detection box included in the first target detection result to determine the supervision label corresponding to each candidate box includes: For the j-th candidate box, determine the intersection over union of the j-th candidate box and each object detection box included in the first target detection result, where j is a positive integer; If the intersection over union (IoU) of the j-th candidate box and the p-th object detection box among the object detection boxes included in the first target detection result is greater than or equal to a preset value, then set the supervision label of the j-th candidate box under the p-th object detection box to 1, where p is a positive integer; If the IoU of the j-th candidate box and the p-th object detection box is less than the preset value, then set the supervision label of the j-th candidate box under the p-th object detection box to 0.
15. The method according to claim 13, characterized in that Determining the class prediction loss based on the probability that the object in each candidate box included in the reference detection result belongs to each category in the target category set, and the supervision label corresponding to each candidate box, includes: Determine the cross-entropy loss between the probability that the object in each candidate box belongs to each category in the target category set and the supervision label corresponding to each candidate box, as the class prediction loss.
16. The method according to claim 12, characterized in that Determining the target loss based on the class prediction loss and the location prediction loss, includes: Determine the sum of the class prediction loss and the location prediction loss as the target loss.
17. A target detection device, characterized in that, Includes: An acquisition unit, configured to acquire the speech recognition text after performing speech recognition on the speech to be recognized; A target detection unit, configured to perform target detection on the speech recognition text of the speech to be recognized through a target detection model to obtain the target detection result of the speech to be recognized; Wherein, the training process of the target detection model includes the following steps: acquiring the speech recognition text and the manual recognition text of each training sample in N training samples, and extracting the sentence vector representation of the speech recognition text and the sentence vector representation of the manual recognition text of each training sample through the target detection model, where N is a positive integer; based on a preset additional margin, and the sentence vector representation of the speech recognition text and the sentence vector representation of the manual recognition text of each training sample, enhancing the decision boundary between training samples with different intents among the N training samples to obtain the model loss of the target detection model; training the target detection model based on the model loss.
18. A computer device, including a processor and a memory; The memory is used to store a computer program; The processor is configured to execute the computer program to implement the method according to any one of claims 1 to 16 above.
19. A computer-readable storage medium, characterized in that, For storing a computer program; The computer program causes the computer to execute the method according to any one of claims 1 to 16 above.