Pseudo-label-based model training method, target detection method, device and equipment

By training an object detection model using a pseudo-label-based method, and utilizing a small number of manually labeled sample images and a large number of general images, the training cost is reduced and the detection efficiency is improved, thus solving the problem of high training cost of deep learning models.

CN120833531APending Publication Date: 2025-10-24TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410464789.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-16
Publication Date
2025-10-24

Smart Images

  • Figure CN120833531A_ABST
    Figure CN120833531A_ABST
Patent Text Reader

Abstract

The invention provides a pseudo-label-based model training method, a target detection method, a target detection device and target detection equipment, which can be applied to various scenes such as cloud technology, artificial intelligence, intelligent traffic, auxiliary driving and the like. The method comprises the following steps of: training a classification model and a first target detection model by utilizing a first sample image set in a target scene, a label of the first sample image set, a second sample image set in a non-target scene and a label of the second sample image set; and constructing a label of a third sample image set with a label by using a classification model and the first target detection model, and further training the second target detection model directly based on the third sample image set and the label of the third sample image set. According to the method, the number of sample images needing to be manually annotated can be reduced, and then the training cost of the model can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of machine learning, and more particularly, to a pseudo-label-based model training method, a target detection method, an apparatus and a device. BACKGROUND

[0002] Target detection is a core task in the field of computer vision, which aims to identify and locate one or more specific targets in a given image or video. Target detection not only needs to determine whether there is a target of interest in the image, but also needs to determine the location of the target, which is usually represented by a bounding box.

[0003] In related technologies, target detection can be implemented based on deep learning methods, that is, by training a detection model based on deep learning, such as U-net or Mask R-CNN, and then detecting targets in an image based on the detection model.

[0004] However, although deep learning-based methods perform well in many scenarios, a large amount of sample data needs to be manually labeled when training the detection model, which results in excessive training costs. SUMMARY

[0005] The embodiments of the present application provide a pseudo-label-based model training method, a target detection method, an apparatus and a device, which can reduce the training cost of the model while ensuring the performance of the model.

[0006] In a first aspect, the embodiments of the present application provide a pseudo-label-based model training method, comprising:

[0007] obtaining a first sample image set in a target scene and a second sample image set in a non-target scene;

[0008] training a classification model based on the first sample image set with the target scene as a label and based on the second sample image set with the non-target scene as a label;

[0009] training a first target detection model based on each first sample image in the first sample image set with the class of the target in the each first sample image and the detection box where the target is located as a label, and based on each second sample image in the second sample image set with the class of the target in the each second sample image and the detection box where the target is located as a label;

[0010] Based on the unlabeled sample image set, the classification model is used to predict a classification result of each unlabeled sample image in the unlabeled sample image set;

[0011] Based on the classification result of each unlabeled sample image, a third sample image set with a scene being the target scene is selected from the unlabeled sample image set, and based on the third sample image set, the first target detection model is used to predict a category of a target in each third sample image in the third sample image set and a detection box in which the target is located;

[0012] Based on the third sample image set, a second target detection model is trained by taking the category of the target in each third sample image and the detection box in which the target is located as a label.

[0013] In a second aspect, an embodiment of the present application provides a target detection method, and the method has the characteristics that the method comprises:

[0014] An image to be detected is acquired;

[0015] A second target detection model deployed on a user terminal is used to predict a category of a target in the image to be detected and a detection box in which the target is located;

[0016] The second target detection model is obtained by training in the following manner:

[0017] A first sample image set under a target scene and a second sample image set under a non-target scene are acquired;

[0018] A classification model is trained based on the first sample image set by taking the target scene as a label and based on the second sample image set by taking the non-target scene as a label;

[0019] A first target detection model is trained based on each first sample image in the first sample image set by taking a category of a target in each first sample image and a detection box in which the target is located as a label and based on each second sample image in the second sample image set by taking a category of a target in each second sample image and a detection box in which the target is located as a label;

[0020] Based on the unlabeled sample image set, the classification model is used to predict a classification result of each unlabeled sample image in the unlabeled sample image set;

[0021] Based on the classification result of each unlabeled sample image, a third sample image set with a scene being the target scene is selected from the unlabeled sample image set, and based on the third sample image set, the first target detection model is used to predict a category of a target in each third sample image in the third sample image set and a detection box in which the target is located;

[0022] Based on the third sample image set, the type of the target in each third sample image and the detection frame in which the target is located are taken as labels to train the second target detection model.

[0023] In a third aspect, an embodiment of the present application provides a model training device based on pseudo labels, including:

[0024] An acquisition unit is configured to acquire a first sample image set in a target scene and a second sample image set in a non-target scene.

[0025] A first training unit is configured to train a classification model based on the first sample image set and taking the target scene as a label, and based on the second sample image set and taking the non-target scene as a label.

[0026] A second training unit is configured to train a first target detection model based on each first sample image in the first sample image set and taking the category of the target in the first sample image and the detection frame in which the target is located as labels, and based on each second sample image in the second sample image set and taking the category of the target in the second sample image and the detection frame in which the target is located as labels.

[0027] A first prediction unit is configured to predict a classification result of each unlabeled sample image in an unlabeled sample image set based on the unlabeled sample image set and using the classification model.

[0028] A second prediction unit is configured to select a third sample image set in which the scene is the target scene from the unlabeled sample image set based on the classification result of each unlabeled sample image, and predict the category of the target in each third sample image in the third sample image set and the detection frame in which the target is located based on the third sample image set and using the first target detection model.

[0029] A third training unit is configured to train a second target detection model based on the third sample image set and taking the type of the target in each third sample image and the detection frame in which the target is located as labels.

[0030] In a fourth aspect, an embodiment of the present application provides a target detection device, including:

[0031] An acquisition unit is configured to acquire a to-be-detected image.

[0032] A prediction unit is configured to predict the type of the target in the to-be-detected image and the detection frame in which the target is located by using a second target detection model deployed on a user terminal.

[0033] The second target detection model is trained in the following manner:

[0034] obtain a first sample image set in a target scene and a second sample image set in a non-target scene;

[0035] train a classification model based on the first sample image set and the second sample image set;

[0036] train a first target detection model based on each first sample image in the first sample image set and each second sample image in the second sample image set;

[0037] predict a classification result of each unlabeled sample image in the unlabeled sample image set based on the classification model;

[0038] select a third sample image set in the unlabeled sample image set based on the classification result of each unlabeled sample image, and predict a target type and a detection frame of a target in each third sample image in the third sample image set based on the first target detection model and the third sample image set;

[0039] train the second target detection model based on the third sample image set and the target type and the detection frame of the target in each third sample image.

[0040] In a fifth aspect, an electronic device is provided, and the electronic device comprises:

[0041] a processor adapted to implement computer instructions; and

[0042] a computer readable storage medium, the computer readable storage medium storing computer instructions, the computer instructions being adapted to be loaded by the processor and to execute the method provided in the first aspect or the second aspect.

[0043] In a sixth aspect, a computer readable storage medium is provided, the computer readable storage medium storing computer instructions, the computer instructions being read and executed by a processor of a computer device, so that the computer device executes the method provided in the first aspect or the second aspect.

[0044] In a seventh aspect, embodiments of the present application provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the method provided in the first or second aspect mentioned above.

[0045] In an embodiment of the present application, a first sample image set in a target scene and a second sample image set in a non-target scene are obtained; a classification model is trained based on the first sample image set with the target scene as a label, and based on the second sample image set with the non-target scene as a label; a first target detection model is trained based on each first sample image in the first sample image set with the category of the target and the detection frame where the target is located in each first sample image as a label, and based on each second sample image in the second sample image set with the category of the target and the detection frame where the target is located in each second sample image as a label; Based on the unlabeled sample image set, the classification model is used to predict the classification result of each unlabeled sample image in the unlabeled sample image set; based on the classification result of each unlabeled sample image, a third sample image set with the scene as the target scene is selected from the unlabeled sample image set, and based on the third sample image set, the first target detection model is used to predict the category of the target and the detection box where the target is located in each third sample image in the third sample image set; based on the third sample image set, the second target detection model is trained with the type of the target and the detection box where the target is located in each third sample image as a label.

[0046] That is, for the target scene, when training the second target detection model, since the scene requirement of the second sample image set is low, a sample image set that is easy to obtain and has labels can be obtained as the second sample image set, that is, only the first sample image set needs to be labeled, after the classification model and the first target detection model are trained, a large number of unlabeled sample image sets can be converted into labeled third sample image sets by the classification model and the first target detection model, and then the second target detection model is trained based on the labeled third sample image set, therefore, the scheme of the present application can not only ensure sufficient sample images to train the second target detection model, but also reduce the number of manually labeled sample images, and thus the training cost of the model can be reduced on the basis of ensuring the performance of the model.

[0047] In short, the scheme of the present application trains a classification model and a first target detection model using a small amount of first sample image set and a large amount of general second sample image set, then selects a third sample image set under the target scene from the unlabeled sample image set by using the classification model, and obtains the label of the third sample image set by using the first target detection model, and then trains the second target detection model based on the labeled third sample image set, therefore, the scheme of the present application can not only ensure sufficient sample images to train the second target detection model, but also reduce the number of manually labeled sample images, and thus the training cost of the model can be reduced on the basis of ensuring the performance of the model.

[0048] In addition, since the second target detection model can be based on a single image recognition, it can reduce the computational amount of the second target detection model under the premise of ensuring the accuracy, and thus the detection efficiency can be improved. BRIEF DESCRIPTION OF DRAWINGS

[0049] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed to be used in the description of the embodiments of the present application will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor under the premise of the drawings.

[0050] Figure 1 is an example of the system framework provided by the embodiments of the present application.

[0051] Figure 2 is a schematic flow chart of the model training method based on pseudo labels provided by the embodiments of the present application.

[0052] Figure 3 is a schematic structural diagram of the classification model, the first target detection model or the first image segmentation model provided by the embodiments of the present application.

[0053] Figure 4 is a schematic structural diagram of the second target detection model or the second image segmentation model provided by the embodiments of the present application.

[0054] Figure 5 is an example of the training principle of the second target detection model or the second image segmentation model provided by the embodiments of the present application.

[0055] Figure 6 is an example of the training principle of the second target detection model or the second image segmentation model provided by the embodiments of the present application.

[0056] Figure 7 is a schematic flow chart of the target detection method provided by the embodiments of the present application.

[0057] Figure 8 is another schematic flow chart of the target detection method provided by the embodiments of the present application.

[0058] Figure 9 is a schematic block diagram of the model training device based on pseudo labels provided by the embodiments of the present application.

[0059] Figure 10 is a schematic block diagram of the target detection device provided by the embodiments of the present application.

[0060] Figure 11 is a schematic block diagram of the electronic device provided by the embodiments of the present application. DETAILED DESCRIPTION

[0061] In the following, the technical solutions in the embodiments of the present application will be described clearly with the drawings in the embodiments of the present application. Obviously, the described embodiments are some of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present application.

[0062] To facilitate understanding of the technical solutions provided in the present application, the related terms are explained as follows.

[0063] Pre-training model (Pre-training model): The pre-training model is also called cornerstone model, large model, which refers to a deep neural network (Deep neural network, DNN) with large parameters. The pre-training model is trained on a large amount of unlabeled data, and the function approximation capability of the large parameter DNN is used to extract common features from the data. Through fine tuning, parameter efficient fine tuning (PEFT), prompt-tuning and other technologies, the pre-training model is suitable for downstream tasks. Therefore, the pre-training model can achieve ideal results in the few-shot or zero-shot scene. The PTM can be divided into language models (ELMO, BERT, GPT), visual models (swin-transformer, ViT, V-MOE), speech models (VALL-E), multi-modal models (ViBERT, CLIP, Flamingo, Gato) according to the data modalities processed, wherein the multi-modal model refers to a model that establishes the feature representation of two or more data modalities. The pre-training model is an important tool for output artificial intelligence generated content (AIGC), and can also be used as a general interface connecting multiple specific task models.

[0064] Distributed training: refers to splitting and sharing the workload of training the model to multiple microprocessors. The parameters of the large model are large, and the training data are large, which exceeds the capacity of a single machine, so distributed parallel speedup is needed. Parallel mechanisms include data parallelism (Data Parallel, DP), model parallelism (Model Parallel, MP), pipeline parallelism (Pipeline Parallel, PP), and hybrid parallelism (Hybrid parallel, HP). Structural design includes parameter server (Parameter Server), reduce (Reduce), MPI, etc.

[0065] Model compression and quantization: refers to the technology of compression and quantization to help reduce the model size and accelerate the model inference, so as to reduce the cost of model in storage and calculation. Model compression usually includes pruning, low-rank decomposition, knowledge distillation, etc. Model quantization refers to converting floating-point number parameters in the model to fixed-point number or integer parameters, thereby reducing the model size and accelerating the model inference.

[0066] Adaptive computing: refers to automatically adjusting the amount of calculation and accuracy of the model according to different input data, in order to achieve the purpose of improving the calculation efficiency of the model while maintaining the accuracy of the model. Adaptive computing can flexibly adjust the amount of calculation and accuracy of the model on different input data, so as to better balance the calculation efficiency and accuracy of the model.

[0067] Model parallel computing: refers to distributing the calculation tasks of the model to multiple computing devices (such as CPU, GPU, TPU, etc.) to perform calculation at the same time, so as to accelerate the training and inference of the model. Model parallel computing can effectively utilize computing resources, improve the calculation efficiency and training speed of the model.

[0068] It should be noted that the terms used in the embodiment part of the present application are only used to explain the embodiments of the present application, and are not intended to limit the present application.

[0069] For example, the term "and / or" in this paper is only a description of the association relationship between the associated objects, which means that there may be three relationships, for example, A and / or B, which can represent the following three cases: A exists alone, A and B exist at the same time, and B exists alone. The term "at least one" is only a description of the combination relationship of the listed objects, which means that one or more items may exist, for example, at least one of the following: A, B, C, which can represent the following combination cases: A exists alone, B exists alone, C exists alone, A and B exist at the same time, A and C exist at the same time, B and C exist at the same time, A, B and C exist at the same time. The term "multiple" means two or more. The character " / " generally represents that the front and rear associated objects are a "or" relationship.

[0070] For example, the term “corresponding” can represent a direct or indirect relationship between the two, can also represent an associated relationship between the two, and can also indicate a relationship with the indicated, configured, and the like. The term “indicates” can be direct indication, can also be indirect indication, and can also represent an associated relationship. For example, A indicates B, which can represent that A directly indicates B, for example, B can be obtained through A; can also represent that A indirectly indicates B, for example, A indicates C, and B can be obtained through C; and can also represent that A and B have an associated relationship. The term “predefined” or “preconfigured” can pre-store corresponding codes, tables or other related information that can be used for indication in the device, or can be indicated by the protocol. “Protocol” can refer to the standard protocol in the art. The term “at” can be interpreted as “if” or “if” or “when” or “in response to” and the like. Similarly, depending on the context, the phrase “if determined” or “if detected (stated condition or event)” can be interpreted as “when determined” or “in response to determination” or “when detected (stated condition or event)” or “in response to detection (stated condition or event)” and the like. The terms “first”, “second”, “third”, “fourth”, “A”, “B” and the like are used to distinguish different objects, not to describe a specific order. The terms “include” and “have” and any variations thereof are intended to cover non-exclusive (or non-exclusive) inclusion. Among them, the digital video compression technology is mainly to compress the huge digital image video data, so as to facilitate transmission and storage.

[0071] The following describes the scenarios to which the model training method provided in the present application is applicable.

[0072] (1) Artificial intelligence (Artificial Intelligence, AI) scene.

[0073] Among them, AI is to use digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science, which tries to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines, so that machines have the functions of perception, reasoning and decision-making.

[0074] Artificial intelligence technology is a comprehensive discipline involving a wide range of fields, both hardware and software technologies. Artificial intelligence basic technologies generally include, for example, sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-training model technology, operation / interaction system, mechatronics, etc. Among them, the pre-training model, also known as a large model or a basic model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. Artificial intelligence software technologies mainly include computer vision (CV) technology, speech technology, nature language processing (NLP) technology, machine learning (ML), and autonomous driving technology, etc.

[0075] As an implementation manner, the model training method based on pseudo labels provided in the application can also be referred to as a model training method based on CV pseudo labels, which refers to pseudo labels determined or obtained based on the recognition result of image recognition.

[0076] As another implementation manner, the model trained based on the model training method based on pseudo labels provided in the application can be applied to CV technology. For example, the model trained based on the model training method based on pseudo labels provided in the application can be used for image processing and other operations.

[0077] Among them, computer vision technology is a science that studies how to make machines "see". Further, it refers to using cameras and computers to replace human eyes to identify, detect and measure targets, and further process graphics so that computer processing becomes images more suitable for human observation or transmission to instruments for detection. As a scientific discipline, computer vision researches related theories and technologies, trying to establish artificial intelligence systems that can obtain information from images or multidimensional data. Large model technology brings important changes to the development of computer vision technology. Swin-transformer, Vision Transformer (ViT), Vision Mixture of Experts (V-MOE), Masked Auto Encoder (MAE) and other pre-training models in the field of vision can be quickly and widely applied to downstream specific tasks after fine tuning. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, and other technologies. It also includes common face recognition, fingerprint recognition and other biometric identification technologies.

[0078] As an implementation manner, the model training method based on pseudo-label provided in the application can also be referred to as a model training method based on ML pseudo-label. The ML-based pseudo-label can be a pseudo-label predicted or inferred based on an ML-based model.

[0079] ML is a multi-disciplinary subject, involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity and other disciplines. It is a specialized study of how computers simulate or implement human learning behavior to acquire new knowledge or skills, reorganize existing knowledge structure to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental approach to making computers intelligent. Its applications are widespread in various fields of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rule-based learning. Pre-training models are the latest developments in deep learning, integrating the above technologies.

[0080] As an implementation manner, the model trained by the model training method based on pseudo-label provided in the application can be applied to automatic driving technology. For example, the model trained by the model training method based on pseudo-label provided in the application can be used to predict information related to automatic driving technology.

[0081] Autonomous driving technology refers to the technology that enables a vehicle to drive itself without the need for a human driver. It typically includes technologies such as high-precision maps, environmental perception, computer vision, behavior decision-making, path planning, and motion control. Autonomous driving includes single-vehicle intelligence, vehicle-road coordination, and networked cloud control, among other development paths. Autonomous driving technology has a wide range of applications, and is currently being used in areas such as logistics, public transportation, taxis, and intelligent transportation. In the future, it will continue to develop and expand into new areas.

[0082] With the research and development of artificial intelligence technology, it has been applied in various fields, such as smart home, smart wearable devices, virtual assistants, smart speakers, intelligent marketing, autonomous driving, unmanned vehicles, drones, digital twins, virtual humans, robots, artificial intelligence generated content (AIGC), conversational interactions, intelligent medical care, intelligent customer service, and game AI. With the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0083] (2) Cloud technology scenarios.

[0084] Cloud technology refers to a series of technologies that enable computing resources such as servers, storage, databases, networks, software, and analytics to be provided, managed, and delivered over the internet. These resources are centrally managed in the cloud, and users can access and use these services flexibly according to their needs, without the need to own or maintain physical hardware on-site.

[0085] As an implementation manner, the computing involved in the present application can be implemented as cloud computing.

[0086] For example, the model training method provided by the present application can be a cloud computing-based model training method.

[0087] For another example, the execution subject of the model training method provided by the present application can be a device providing cloud computing services, such as a cloud server or a user terminal.

[0088] Cloud computing is a computing model that distributes computing tasks across a pool of resources composed of a large number of computers, enabling various application systems to obtain computing power, storage space, and information services as needed. The network that provides resources is called a "cloud". The resources in the "cloud" can be infinitely expanded and can be accessed at any time, used on demand, expanded at any time, and paid for according to usage.

[0089] As a basic capability provider of cloud computing, a cloud computing resource pool, which can also be referred to as a cloud platform or an infrastructure as a service (IaaS) platform, is established, and various types of virtual resources are deployed in the resource pool for external customers to select and use. The cloud computing resource pool mainly includes computing devices (virtualized machines containing operating systems), storage devices, and network devices.

[0090] According to logical functions, a platform as a service (PaaS) layer can be deployed on an infrastructure as a service (IaaS) layer, and a software as a service (SaaS) layer is deployed on the PaaS layer, or the SaaS is directly deployed on the IaaS. The PaaS is a platform for software running, such as a database and a web container. The SaaS is various business software, such as a web portal and a short message massager. Generally, the SaaS and the PaaS are upper layers relative to the IaaS.

[0091] As another implementation manner, the model trained based on the pseudo-label-based model training method provided in the present application can be implemented as an artificial intelligence cloud service.

[0092] The artificial intelligence cloud service is also referred to as AI as a Service (AIaaS). This is a mainstream service mode of an artificial intelligence platform. Specifically, the AIaaS platform splits several common AI services and provides independent or packaged services in the cloud. This service mode is similar to opening an AI theme mall: all developers can access and use one or more artificial intelligence services provided by the platform through an Application Programming Interface (API), and some experienced developers can also use the AI framework and AI infrastructure provided by the platform to deploy and maintain their own cloud artificial intelligence services.

[0093] (3) Map connected vehicle scenario.

[0094] Map connected vehicle refers to a scenario in which, in an intelligent transportation system, vehicles are connected to the Internet through vehicle-mounted sensors and communication technologies, realize real-time data exchange with the external environment, and perform data analysis and processing through a map service to provide more intelligent driving experience and traffic management. This scenario usually involves the following technical and application directions: vehicle navigation system, real-time traffic detection and control system, intelligent driving assistance system, vehicle remote control and management, shared travel service, and smart city traffic management.

[0095] The development of map car networking not only improves the travel experience of drivers, but also improves the transportation efficiency of roads, reduces the incidence of traffic accidents, and has a positive impact on environmental protection and energy conservation. With the continuous progress of 5G communication technology, artificial intelligence, big data analysis and other technologies, map car networking scenarios will play an increasingly important role in the future of intelligent transportation and autonomous driving.

[0096] As another implementation manner, the model trained based on the pseudo-label-based model training method provided in the present application can be deployed as a lightweight model on a user device, which includes but is not limited to a mobile phone, a computer, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, etc. The present embodiment can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, intelligent transportation, and assisted driving.

[0097] (4) Short video creation scenario.

[0098] The short video creation scenario can include multiple sub-scenarios such as video shooting, video editing, and video packaging.

[0099] Taking a skiing short video as an example, the model trained based on the pseudo-label-based model training method provided in the present application can be used to detect targets in a skiing image, such as a skiing subject (e.g., including a person and skiing equipment), a ski, and ski goggles, and the detected targets can be widely applied to multiple sub-scenarios of short video creation, such as video shooting, video editing, and video packaging.

[0100] As an implementation manner, in the video shooting scenario, the model trained based on the pseudo-label-based model training method provided in the present application can be applied to realize a scenario-based shooting special effect. For example, when shooting a skiing video, the model can be used to segment a skiing subject (e.g., including a person and skiing equipment), a ski, or ski goggles in the picture in real time, and then the skiing subject (e.g., including a person and skiing equipment), the ski, or the ski goggles can be highlighted in real time through a special effect, or decorative elements can be added to the background to create a skiing-themed shooting scene, increasing the creativity and interest of the shooting video.

[0101] As another implementation manner, in the video editing scenario, the model trained based on the pseudo-label-based model training method provided in the present application can segment a skiing subject (e.g., including a person and skiing equipment) from the background, and then realize virtual background replacement to provide users with rich creative material selection.

[0102] As another implementation manner, in a video packaging scenario, the model trained based on the pseudo-label-based model training method provided in the present application can be applied to creative special effect production of a skiing theme. For example, by separately segmenting a person and a ski, different visual special effects can be applied to them in the editing process, such as person contour outlining and light emission, ski contour outlining and light emission, person background processing, person masking, person duplication, person trailing, and the like, to enhance the visual impact and aesthetic sense of the video.

[0103] It should be noted that the application scenarios described above are only examples, and the model training method provided in the present application can also be applied to other scenarios, which are not limited herein.

[0104] It should be noted that the acquisition and processing of the related data set (for example, the sample image set used for model training) in the present application should strictly comply with the requirements of the relevant national laws and regulations, obtain the informed consent or separate consent of the personal information subject, and within the scope of authorization of the laws and regulations and the personal information subject, carry out subsequent data use and processing behavior. Among them, the acquisition of the data set can be realized in the form of collecting logs, for example, the relevant data can be acquired by acquiring the logs of the target memory or the target processor. When collecting logs, the collection granularity can be in units of days or weeks, and the present application does not make a specific limitation.

[0105] Figure 1 is an example of the system framework 100 provided by the embodiments of the present application.

[0106] As shown in Figure 1 , the system framework 100 can include a server 11 and a terminal 12. The terminal 12 and the server 11 can be directly or indirectly connected through wired or wireless communication, which is not limited by the present application.

[0107] Among them, the server 11 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content distribution networks (Content Delivery Network, CDN), and big data and artificial intelligence platforms, etc. Basic cloud computing services. The terminal 12 includes but is not limited to mobile phones, computers, smart voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, etc. It should be noted that the embodiments of the present application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, intelligent transportation, and assisted driving. For example, for the intelligent transportation or assisted driving scenario, the terminal 12 includes but is not limited to mobile phones, computers, smart voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, etc.

[0108] In the embodiments of the present application, the server 11 and the terminal 12 can be used for model training. For example, the server 11 can deploy the trained model on the terminal 12, or the terminal 12 can obtain the data set from the server 11 and train the model deployed on the terminal 12 based on the obtained data set. The present application does not make specific limitations on this.

[0109] It should be understood that Figure 1 This is only an example of an application scenario and should not be understood as a limitation on the present application. For example, Figure 1 Only one server 11 and one terminal 12 are exemplarily included, but in other alternative embodiments, the system framework 100 can include multiple terminals or multiple servers, or even not include terminals or not include servers.

[0110] Figure 2 A schematic flowchart of a model training method 200 based on pseudo labels according to an embodiment of the present application is shown, which can be executed by an electronic device. The electronic device can be implemented as a terminal device or a server. The terminal device can also be referred to as a user terminal, which includes but is not limited to a smartphone, a game console, a desktop computer, a tablet computer, an e-book reader, an MP4 player, an MP4 player, a laptop computer, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, etc. The server can be a standalone physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud services.

[0111] For ease of description, the model training method 200 will be exemplarily described below taking a model training apparatus as an example.

[0112] As Figure 2 shown, the model training method 200 can include:

[0113] S210, the model training apparatus obtains a first sample image set in a target scene and a second sample image set in a non-target scene.

[0114] Exemplarily, the first sample image set can include artificially annotated sample images.

[0115] Exemplarily, the second sample image set can include open source sample images irrelevant to the target scene or sample images irrelevant to the target scene accumulated by the business history.

[0116] Exemplarily, the number of first sample images in the first sample image set can be less than the number of second sample images in the second sample image set, so as to reduce the artificial annotation cost.

[0117] Exemplarily, the target scene can be a scene with high complexity or a scene with more targets. For example, the target scene can be a skiing scene, and the target in the first sample image in the first set of sample images can be a skiing subject (e.g., including a person and skiing equipment), a ski board, or a ski goggle. The skiing equipment includes a ski board, a ski goggle, a ski pole, and the like handheld or worn equipment. For example, the target scene can be a skydiving scene, and the target in the first sample image in the first set of sample images can be a skydiving subject (e.g., including a person and skydiving equipment), a skydiving board, or a skydiving goggle.

[0118] S220, the model training apparatus trains a classification model based on the first set of sample images, labels the target scene, and based on the second set of sample images, labels the non-target scene.

[0119] Exemplarily, the second set of sample images can include open source sample images irrelevant to the target scene or sample images accumulated by business history that can be irrelevant to the target scene. For example, the model training apparatus can train the classification model using two parts of sample images together, one part being open source sample images or sample images accumulated by business history that can be irrelevant to the target scene, and the other part being a small amount of manually labeled sample images under the target scene, and the form of labeling is scene, that is, the model training apparatus can train the classification model using these sample images.

[0120] For example, the second set of sample images can include open source sample images irrelevant to the skiing scene or sample images accumulated by business history that can be irrelevant to the skiing scene. For example, the model training apparatus can train the classification model using two parts of sample images together, one part being open source sample images or sample images accumulated by business history that can be irrelevant to the skiing scene, and the other part being a small amount of manually labeled sample images under the skiing scene, and the form of labeling is scene, that is, the model training apparatus can train the classification model using these sample images.

[0121] S230, the model training apparatus trains a first target detection model based on each first sample image in the first set of sample images, labels the category of the target in each first sample image and the detection frame where the target is located, and based on each second sample image in the second set of sample images, labels the category of the target in each second sample image and the detection frame where the target is located.

[0122] Exemplarily, the second sample image set can include open source sample images irrelevant to the target scene or sample images irrelevant to the target scene accumulated by business history. For example, the model training apparatus can train the first target detection model using two parts of sample images, one part being open source sample images or sample images irrelevant to the target scene accumulated by business history, and the other part being a small amount of sample images in the target scene manually annotated, the annotation being in the form of the category of the target and the detection box in which the target is located, i.e., the model training apparatus can train the first target detection model using these sample images.

[0123] Exemplarily, the second sample image set can include open source sample images irrelevant to the target scene or sample images irrelevant to the target scene accumulated by business history. For example, the model training apparatus can train the first target detection model using two parts of sample images, one part being open source sample images or sample images irrelevant to the target scene accumulated by business history, and the other part being a small amount of sample images in the target scene manually annotated, the annotation being in the form of the category of the target and the detection box in which the target is located, i.e., the model training apparatus can train the first target detection model using these sample images.

[0124] S240, the model training apparatus predicts, based on the unlabeled sample image set, a classification result of each unlabeled sample image in the unlabeled sample image set by using the classification model.

[0125] Exemplarily, the classification result can include the scene of each unlabeled sample image.

[0126] S250, the model training apparatus selects, based on the classification result of each unlabeled sample image, a third sample image set with the scene being the target scene from the unlabeled sample image set, and predicts, based on the third sample image set, the category of the target and the detection box in which the target is located in each third sample image in the third sample image set by using the first target detection model.

[0127] Exemplarily, the detection box predicted by the first target detection model can be represented by the coordinates of the rectangular boundary. For example, the detection box predicted by the first target detection model can be represented by the coordinates of the diagonals of the rectangle. Alternatively, the detection box predicted by the first target detection model can be represented by a mask.

[0128] It should be understood that the mask referred to in the present application can be replaced by a mask, a mask image or a mask image, which is not specifically limited in the present application.

[0129] Exemplarily, the model training apparatus selects, based on the scene of each unlabeled sample image in the set of unlabeled sample images, a third set of sample images with the scene being the target scene, and predicts, based on the third set of sample images, the category of the target in each third sample image in the third set of sample images and the detection box in which the target is located by using the first target detection model.

[0130] It should be noted that the target in each third sample image in the third set of sample images can be one or more targets. Correspondingly, when the target in each third sample image in the third set of sample images is multiple targets, the detection box in which the target in each third sample image in the third set of sample images is located is multiple detection boxes. Taking a skiing scene as an example, the target in each third sample image in the third set of sample images includes a skiing subject (for example, including a person and skiing equipment), a ski, and a ski goggle, and the detection box in which the target in each third sample image in the third set of sample images is located can include a detection box in which the skiing subject (for example, including a person and skiing equipment) is located, a detection box in which the ski is located, and a detection box in which the ski goggle is located.

[0131] S260, the model training apparatus trains a second target detection model based on the third set of sample images, the type of the target in each third sample image, and the detection box in which the target is located.

[0132] Exemplarily, after the model training apparatus predicts the type of the target in each third sample image and the detection box in which the target is located by using the first target detection model, the model training apparatus can construct a label of each third sample image based on the type of the target in each third sample image and the detection box in which the target is located, and train the second target detection model based on each third sample image and the label of each third sample image.

[0133] Taking a skiing scene as an example, the model training apparatus can obtain two parts of sample images, one part being a sample image manually labeled under a skiing scene, and the other part being an open-source sample image irrelevant to the skiing scene or a sample image irrelevant to the skiing scene accumulated in a business history. On the one hand, the form of the label of the two parts of sample images is a scene, that is, the model training apparatus can train the classification model by using these sample images. On the other hand, the form of the label of the two parts of sample images is the category of the target and the detection box in which the target is located, that is, the model training apparatus can train the first target detection model by using these sample images. Then, the model training apparatus can select, based on the classification model, a sample image under the skiing scene from the set of unlabeled sample images, and predict, based on the first target detection model, the label of the selected sample image, that is, the type of the target and the detection box in which the target is located, and train the second target detection model based on the selected sample image and the label of the selected sample image.

[0134] For example, the model training apparatus can obtain two parts of sample images, one part is 1000 sample images of ski scenes manually labeled, and the other part is open source sample images irrelevant to ski scenes or sample images irrelevant to ski scenes accumulated in business history. On the one hand, the labels of the two parts of sample images are in the form of scenes, that is, the model training apparatus can train the classification model with these sample images. On the other hand, the labels of the two parts of sample images are in the form of categories of targets and detection boxes where the targets are located, that is, the model training apparatus can train the first target detection model with these sample images. Then, the model training apparatus can select 100,000 sample images of ski scenes from 1 million unlabeled sample images based on the classification model, predict the labels of the 100,000 sample images, that is, the types of the targets and the detection boxes where the targets are located, based on the first target detection model, and then train the second target detection model based on the selected 100,000 sample images and the labels of the selected 100,000 sample images.

[0135] In this embodiment, the classification model and the first target detection model are first trained using the first sample image set in the target scene, the labels of the first sample image set (i.e., the types of the targets in the first sample images and the detection boxes where the targets are located), the second sample image set in the non-target scene, and the labels of the second sample image set (i.e., the types of the targets in the second sample images and the detection boxes where the targets are located), then the third sample image set is selected from the unlabeled sample image set based on the classification model, and then the labels of the third sample image set (i.e., the types of the targets in the third sample images and the detection boxes where the targets are located) are output by the first target detection model, and then the second target detection model is trained directly based on the third sample image set and the labels of the third sample image set. That is, for the target scene, when training the second target detection model, since the scene requirement of the second sample image set is low, a sample image set that is easy to obtain and has labels can be obtained as the second sample image set, that is, only the first sample image set needs to be labeled, after the classification model and the first target detection model are trained, a large number of unlabeled sample image sets can be converted into labeled third sample image sets by the classification model and the first target detection model, and then the second target detection model is trained based on the labeled third sample image set. Therefore, the scheme of the present application not only ensures sufficient sample images to train the second target detection model, but also reduces the number of manually labeled sample images, and thus reduces the training cost of the model while ensuring the performance of the model.

[0136] In short, the scheme of the present application trains a classification model and a first target detection model using a small set of first sample images and a large set of second sample images, selects a third set of sample images in a target scene from the set of unlabeled sample images using the classification model, obtains labels of the third set of sample images using the first target detection model, and then trains the second target detection model based on the third set of sample images with labels. Therefore, the scheme of the present application not only ensures sufficient sample images to train the second target detection model, but also reduces the number of manually labeled sample images, thereby reducing the training cost of the model while ensuring the performance of the model.

[0137] In addition, since the second target detection model can be based on single image recognition, it can reduce the computational complexity of the second target detection model while ensuring accuracy, thereby improving detection efficiency.

[0138] It should be noted that in other alternative embodiments, the first target detection model can be any large model capable of target detection. For example, the first target detection model can be a large model based on a prompt word. The prompt word is a description of the target. Taking General Large-scale Entity Extraction (GLEE) as an example, GLEE can simultaneously detect and segment the target. If a snowboard is identified, the prompt word can be as simple as: snowboard.

[0139] In some embodiments, the backbone network of the first target detection model includes at least one first feature extraction layer, and the backbone network of the second target detection model includes at least one second feature extraction layer; wherein the number of the at least one first feature extraction layer is greater than the number of the at least one second feature extraction layer, or the number of network parameters of the first feature extraction layer is greater than the number of network parameters of the second feature extraction layer.

[0140] Exemplarily, the first target detection model includes a first backbone network, a first connection layer connected to the first backbone network, and a first activation layer connected to the first connection layer, and the second target detection model includes a second backbone network, a second connection layer connected to the second backbone network, and a second activation layer connected to the second connection layer; wherein the first backbone network includes at least one first feature extraction layer, and the second backbone network includes at least one second feature extraction layer, the number of the at least one first feature extraction layer is greater than the number of the at least one second feature extraction layer, or the number of network parameters of the first feature extraction layer is greater than the number of network parameters of the second feature extraction layer.

[0141] In this embodiment, the number of the at least one first feature extraction layer is greater than the number of the at least one second feature extraction layer, or the number of network parameters of the first feature extraction layer is greater than the number of network parameters of the second feature extraction layer, which means that the second target detection model is a lightweight model relative to the first target detection model, and thus the second target detection model meets the running requirement of deployment on a mobile terminal. In addition, even if the number of layers or the number of network parameters of the feature extraction layer of the lightweight model deployed on a mobile device is small, a large number of unlabeled sample image sets can be converted into labeled third sample image sets through the classification model and the first target detection model, and then the second target detection model is trained based on the labeled third sample image set. Therefore, the scheme of the present application can not only ensure sufficient sample images to train the first target detection model, but also reduce the number of manually labeled sample images, thereby reducing the training cost of the model while ensuring the performance of the model, and thus the second target detection model meets the performance requirement of deployment on a mobile terminal. That is, the model training method provided by the present application can make the first target detection model meet the running requirement and performance requirement of deployment on a mobile terminal.

[0142] In addition, generally, training a lightweight model with good effect requires hundreds of thousands or even millions of training data, while a large model has strong generalization ability and can achieve ideal effect with thousands of training data. Based on this, in the embodiment of the present application, when the first target detection model is used to train the second target detection model, even if the number of first sample images in the first sample image data set is small, the model performance of the second target detection model can be ensured.

[0143] It should be understood that the first feature extraction layer and the second feature extraction layer can be any module or unit capable of feature extraction, as long as the number of the at least one first feature extraction layer is greater than the number of the at least one second feature extraction layer, or the number of network parameters of the first feature extraction layer is greater than the number of network parameters of the second feature extraction layer. The specific implementation manner is not limited in the present application.

[0144] In some embodiments, the first feature extraction layer is a feature extraction layer for attention processing of input features, and the second feature extraction layer is a depth separable convolution layer.

[0145] Exemplarily, when the first feature extraction layer is a feature extraction layer for attention processing of input features, it can be implemented as: a feature extraction layer for self-attention processing of input features, a feature extraction layer for cross-attention processing of input features, or a feature extraction layer for self-attention and cross-attention processing of input features. For example, the first feature extraction layer can be any type of transformer, for example, the first feature extraction layer can be a shifted windows transformer (SwinTransformer).

[0146] It is worth noting that the classification model and the first target detection model in the embodiment are both large models, which can adopt the same backbone network or different backbone networks, and the present application does not make specific limitation thereon.

[0147] In some embodiments, the S250 comprises:

[0148] The model training apparatus obtains, based on the third sample image set, a plurality of class channel corresponding feature maps and a plurality of position channel corresponding feature maps by using the first target detection model; determines, based on the plurality of class channel corresponding feature maps, a class of a target in each third sample image in the third sample image set, and determines, based on the plurality of position channel corresponding feature maps, a detection box in which the target in the each third sample image is located.

[0149] Exemplarily, the model training apparatus obtains, based on the third sample image set, a plurality of class channel corresponding feature maps and a plurality of position channel corresponding feature maps by using the first target detection model; determines, based on the feature map corresponding to each class channel in the plurality of class channel, whether a target is detected, and in the case that a target is detected, determines, based on the plurality of position channel corresponding feature maps, a detection box in which the detected target is located.

[0150] Exemplarily, the plurality of class channels can include a channel corresponding to a detected target in the target scene.

[0151] For example, assuming that the target scene is a skiing scene, the plurality of class channels can include a skiing subject (e.g., including a person and skiing equipment) channel, a ski board channel, or a ski goggles channel.

[0152] Correspondingly, for any third sample image in the third sample image set, the model training apparatus obtains, by using the first target detection model, a feature map corresponding to the skiing subject channel, a feature map corresponding to the snowboard channel, and a feature map corresponding to the ski goggles channel, and a plurality of feature maps corresponding to a plurality of position channels of the any third sample image. Then, the model training apparatus determines, based on the feature map corresponding to the skiing subject channel of the any third sample image, whether the any third sample image contains a skiing subject, and in the case that the any third sample image contains a skiing subject, determines a detection frame of the skiing subject in the any third sample image based on the plurality of feature maps corresponding to the plurality of position channels of the any third sample image. Similarly, the model training apparatus determines, based on the feature map corresponding to the snowboard channel of the any third sample image, whether the any third sample image contains a snowboard, and in the case that the any third sample image contains a snowboard, determines a detection frame of the snowboard in the any third sample image based on the plurality of feature maps corresponding to the plurality of position channels of the any third sample image. The model training apparatus determines, based on the feature map corresponding to the ski goggles channel of the any third sample image, whether the any third sample image contains ski goggles, and in the case that the any third sample image contains ski goggles, determines a detection frame of the ski goggles in the any third sample image based on the plurality of feature maps corresponding to the plurality of position channels of the any third sample image. It should be noted that the plurality of position channels can include diagonal coordinate channels of a rectangle, for example, an x-axis diagonal coordinate channel and a y-axis diagonal coordinate channel of a rectangle, i.e., four position channels.

[0153] Suppose that the feature map corresponding to the skiing subject channel, the feature map corresponding to the snowboard channel, and the feature map corresponding to the ski goggles channel, and the plurality of feature maps corresponding to the plurality of position channels of the any third sample image are all K×L feature maps, then:

[0154] For any third sample image in the third sample image set, the model training apparatus obtains a KxLx7 feature map of the any third sample image by using the first target detection model, where 7 represents a skiing subject channel, a ski board channel, a ski goggles channel, two x-axis coordinate channels of a diagonal of a rectangle, and two y-axis coordinate channels of a diagonal of the rectangle. Based on this, after the model training apparatus obtains the KxLx7 feature map of the any third sample image, the model training apparatus can determine whether the any third sample image contains a skiing subject based on a KxL feature map corresponding to the skiing subject channel of the any third sample image, and in the case where the any third sample image contains a skiing subject, determine a detection box of the skiing subject in the any third sample image based on a KxLx4 feature map corresponding to the four position channels of the any third sample image. Similarly, the model training apparatus determines whether the any third sample image contains a ski board based on a KxL feature map corresponding to the ski board channel of the any third sample image, and in the case where the any third sample image contains a ski board, determines a detection box of the ski board in the any third sample image based on a KxLx4 feature map corresponding to the four position channels of the any third sample image. The model training apparatus determines whether the any third sample image contains a ski goggles based on a KxL feature map corresponding to the ski goggles channel of the any third sample image, and in the case where the any third sample image contains a ski goggles, determines a detection box of the ski goggles in the any third sample image based on a KxLx4 feature map corresponding to the four position channels of the any third sample image.

[0155] In this embodiment, the model training apparatus obtains feature maps corresponding to a plurality of category channels and feature maps corresponding to a plurality of position channels by using the first target detection model based on the third sample image set; determines a category of a target in each third sample image in the third sample image set based on the feature maps corresponding to the plurality of category channels, and determines a detection box in which the target in the each third sample image is located based on the feature maps corresponding to the plurality of position channels. Equivalently, the first target detection model can be used to detect targets corresponding to a plurality of category channels, that is, the labels of the third sample images are enriched, and thus the training effect of the second target detection model and the model performance can be improved.

[0156] It should be noted that for the skiing scene, considering the particularity of the motion scene, when only the character is segmented, due to the particularity of the scene and equipment of the skiing motion, there are problems such as incomplete recognition and extremely unstable effect at the junction of the character and the ski (or ski goggles). In this embodiment, based on the feature map corresponding to the skiing subject channel, the feature map corresponding to the ski channel, and the feature map corresponding to the ski goggles channel, the category of the target in each third sample image in the third sample image set is determined, the skiing subject, the ski, and the ski goggles can be recognized separately, and when the second target detection model is trained as a label, the second target detection model can also detect more flexible targets, and the detection effect is improved.

[0157] In some embodiments, the method 200 further includes:

[0158] The model training apparatus crops the detection box where the target in each first sample image is located to obtain a cropped image of each first sample image; trains a first image segmentation model based on the cropped image of each first sample image and the mask corresponding to the cropped image of each first sample image; crops the detection box where the target in each third sample image is located to obtain a cropped image of each third sample image, and predicts the mask corresponding to the cropped image of each third sample image using the first image segmentation model based on the cropped image of each first sample image; and trains a second image segmentation model based on the cropped image of each third sample image and the mask corresponding to the cropped image of each third sample image.

[0159] It should be noted that object segmentation is an important task in the field of computer vision, which refers to dividing pixels in an image into different regions or objects and associating each region with a specific class or instance. The purpose of object segmentation is to identify and separate the target of interest from a complex image scene while preserving the shape and structure information of the target. Similar to object detection, object segmentation can also be achieved based on deep learning methods, that is, by training a deep learning-based segmentation model such as U-net or MaskR-CNN, and then segmenting the image to be segmented based on the segmentation model. However, training a segmentation model also requires a large amount of manually annotated sample data, which can result in excessive training costs.

[0160] In this embodiment, the detection box in which the target in each first sample image is cropped to obtain a cropped image of each first sample image; based on the cropped image of each first sample image and the mask corresponding to the cropped image of each first sample image, the first image segmentation model is trained; the detection box in which the target in each third sample image is cropped to obtain a cropped image of each third sample image, and based on the cropped image of each first sample image, the first image segmentation model is used to predict the mask corresponding to the cropped image of each third sample image; based on the cropped image of each third sample image and the mask corresponding to the cropped image of each third sample image, the second image segmentation model is trained.

[0161] That is, the first image segmentation model can be trained using the first sample image set under the target scene and the labels of the first sample image set (i.e., the mask corresponding to the cropped image of each first sample image). Thus, after the third sample image set with the target scene is selected from the unlabeled sample image set using the classification model, the labels of the third sample image set (i.e., the mask corresponding to the cropped image of each third sample image) can be obtained using the first image segmentation model. Then, the second image segmentation model can be trained based on the third sample image set and the labels of the third sample image set. That is, for the target scene, only the first sample image set needs to be labeled when training the second image segmentation model. A large number of unlabeled sample image sets can be converted into labeled third sample image sets through the classification model and the first image segmentation model. Then, the second image segmentation model is trained based on the labeled third sample image set. Since the labels of the second sample image set under the non-target scene can be obtained from open source data, the number of sample images that need to be manually labeled can be reduced, and thus the training cost of the second image segmentation model can be reduced.

[0162] In short, the scheme of the present application trains a classification model and a first image segmentation model using a small number of first sample image sets and a large number of general second sample image sets, selects a third sample image set under the target scene from the unlabeled sample image set using the classification model, obtains the labels of the third sample image set using the first image segmentation model, and then trains the second image segmentation model based on the labeled third sample image set. Therefore, the scheme of the present application not only ensures sufficient sample images to train the second image segmentation model, but also reduces the number of sample images that need to be manually labeled, thereby reducing the training cost of the model while ensuring the performance of the model.

[0163] In addition, since the second image segmentation model can be based on a single image recognition, it can reduce the computational complexity of the second image segmentation model while ensuring accuracy, thereby improving the segmentation efficiency.

[0164] Of course, in other alternative embodiments, the cropping of the detection box in which the target in each of the first sample images is located can also be used as a post-processing process of the first target detection model or a pre-processing process of the first image segmentation model, and similarly, the cropping of the detection box in which the target in each of the third sample images is located can also be used as a post-processing process of the first target detection model or a pre-processing process of the first image segmentation model, which is not specifically limited by the present application.

[0165] In some embodiments, the training of the first image segmentation model based on the cropped image of each of the first sample images and the mask corresponding to the cropped image of each of the first sample images comprises:

[0166] The model training device performs down-sampling on the cropped image of each of the first sample images to obtain a down-sampled image of the cropped image of each of the first sample images, and performs down-sampling on the mask corresponding to the cropped image of each of the first sample images to obtain a down-sampled image of the mask corresponding to the cropped image of each of the first sample images; and trains the first image segmentation model based on the down-sampled image of the cropped image of each of the first sample images, with the down-sampled image of the mask corresponding to the cropped image of each of the first sample images as a label; wherein the training of the second image segmentation model based on the cropped image of each of the third sample images and the mask corresponding to the cropped image of each of the third sample images comprises: performing down-sampling on the cropped image of each of the third sample images to obtain a down-sampled image of the cropped image of each of the third sample images, and performing down-sampling on the mask corresponding to the cropped image of each of the third sample images to obtain a down-sampled image of the mask corresponding to the cropped image of each of the third sample images; and training the second image segmentation model based on the down-sampled image of the cropped image of each of the third sample images, with the down-sampled image of the mask corresponding to the cropped image of each of the third sample images as a label.

[0167] Exemplarily, the model training apparatus can down-sample the cropped image of each first sample image according to a preset down-sampling ratio to obtain a down-sampled image of the cropped image of each first sample image, and down-sample the mask corresponding to the cropped image of each first sample image according to the preset down-sampling ratio to obtain a down-sampled image of the mask corresponding to the cropped image of each first sample image; and train the first image segmentation model based on the down-sampled image of the cropped image of each first sample image, with the down-sampled image of the mask corresponding to the cropped image of each first sample image as the label. Correspondingly, the model training apparatus can down-sample the cropped image of each third sample image according to a preset down-sampling ratio to obtain a down-sampled image of the cropped image of each third sample image, and down-sample the mask corresponding to the cropped image of each third sample image according to the preset down-sampling ratio to obtain a down-sampled image of the mask corresponding to the cropped image of each third sample image; and train the second image segmentation model based on the down-sampled image of the cropped image of each third sample image, with the down-sampled image of the mask corresponding to the cropped image of each third sample image as the label.

[0168] In this embodiment, the image segmentation is performed in the manner of first cropping, then down-sampling and then segmentation. The cropping makes the image to be segmented free of redundant background and irrelevant subjects, and thus the information of the target can be retained as much as possible during the down-sampling, for example, small targets and details at the junctions of a person and equipment can be retained as much as possible, and thus the accuracy of the segmentation can be improved.

[0169] In some embodiments, the backbone network of the first image segmentation model comprises at least one third feature extraction layer, and the backbone network of the second image segmentation model comprises at least one fourth feature extraction layer; wherein the number of the at least one third feature extraction layer is greater than the number of the at least one fourth feature extraction layer, or the number of network parameters of the third feature extraction layer is greater than the number of network parameters of the fourth feature extraction layer.

[0170] Exemplarily, the first image segmentation model comprises a third backbone network, a third connection layer connected to the third backbone network, and a third activation layer connected to the third connection layer, and the second image segmentation model comprises a fourth backbone network, a fourth connection layer connected to the fourth backbone network, and a fourth activation layer connected to the fourth connection layer; wherein the third backbone network comprises at least one third feature extraction layer, and the fourth backbone network comprises at least one fourth feature extraction layer, the number of the at least one third feature extraction layer is greater than the number of the at least one fourth feature extraction layer, or the number of network parameters of the third feature extraction layer is greater than the number of network parameters of the fourth feature extraction layer.

[0171] In this embodiment, the number of the at least one third feature extraction layer is greater than the number of the at least one fourth feature extraction layer, or the number of network parameters of the third feature extraction layer is greater than the number of network parameters of the fourth feature extraction layer, which means that the second image segmentation model is a lightweight model relative to the first image segmentation model, and thus the second image segmentation model meets the running requirements for deployment on a mobile terminal. In addition, even if the number of layers or the number of network parameters of the feature extraction layer of the lightweight model deployed on a mobile device is small, a large number of unlabeled sample image sets can be converted into labeled third sample image sets through the classification model and the first image segmentation model, and then the second image segmentation model is trained based on the labeled third sample image set. Therefore, the scheme of the present application can not only ensure sufficient sample images to train the first image segmentation model, but also reduce the number of manually labeled sample images, thereby reducing the training cost of the model while ensuring the performance of the model, and thus the second image segmentation model meets the performance requirements for deployment on a mobile terminal. That is, the model training method provided by the present application can make the first image segmentation model meet the running requirements and performance requirements for deployment on a mobile terminal.

[0172] In addition, generally, training a lightweight model with good effect requires hundreds of thousands or even millions of training data, while a large model can achieve ideal results with thousands of training data due to its strong generalization ability. Based on this, in the embodiments of the present application, when the first image segmentation model is used to train the second image segmentation model, even if the number of first sample images in the first sample image data set is small, the model performance of the second image segmentation model can be ensured.

[0173] It should be understood that the third feature extraction layer and the fourth feature extraction layer can be any module or unit capable of feature extraction, as long as the number of the at least one third feature extraction layer is greater than the number of the at least one fourth feature extraction layer, or the number of network parameters of the third feature extraction layer is greater than the number of network parameters of the fourth feature extraction layer, and the specific implementation manner is not limited in the present application.

[0174] In some embodiments, the third feature extraction layer is a feature extraction layer that performs attention processing on the input features, and the fourth feature extraction layer is a depth separable convolution layer.

[0175] Exemplarily, when the third feature extraction layer is a feature extraction layer for attention processing of input features, it can be implemented as: a feature extraction layer for self-attention processing of input features, a feature extraction layer for cross-attention processing of input features, or a feature extraction layer for self-attention and cross-attention processing of input features. For example, the third feature extraction layer can be any type of transformer, for example, the third feature extraction layer can be a shifted windows transformer (SwinTransformer).

[0176] It is worth noting that the first image segmentation model in the embodiment is a large model like the classification model and the first target detection model described above, which can use the same backbone network or different backbone networks, and the present application does not make specific limitations thereto.

[0177] Figure 3 is a schematic structural diagram of the classification model, the first target detection model or the first image segmentation model provided by the embodiment of the present application.

[0178] As shown in Figure 3 , the backbone network structures of the classification model, the first target detection model and the first image segmentation model are consistent, and are all backbone networks formed by N transformers. On the basis of the backbone network, a layer of conventional 1x1 convolution layer (Conv) or full connection layer (Fc) is added, and finally an activation layer (Softmax) is passed to obtain the corresponding output. It is worth noting that although the backbone networks of the classification model, the first target detection model and the first image segmentation model are the same, their outputs are not the same. For example, the output of the classification model is a binary classification vector, the output of the first target detection model is a feature map corresponding to multiple class channels and a feature map corresponding to multiple position channels, and the output of the first image segmentation model is a mask.

[0179] Figure 4 is a schematic structural diagram of the second target detection model or the second image segmentation model provided by the embodiment of the present application.

[0180] As shown in Figure 4As shown, the backbone networks of the second target detection model and the second image segmentation model are consistent, and are both backbone networks formed by M depth separable convolution layers, each of which can include multiple convolution layers (Conv), each of which is connected with an activation layer (ReLU6). For example, each depth separable convolution layer can include, in sequence, a 1x1 convolution layer, an activation layer, a 3x3 convolution layer, an activation layer, a 1x1 convolution layer, and an activation layer. It is worth noting that although the backbone networks of the second target detection model and the second image segmentation model are the same, their outputs are different. For example, the output of the second target detection model is a feature map corresponding to multiple class channels and a feature map corresponding to multiple position channels, and the output of the second image segmentation model is a mask.

[0181] Figure 5 is an example of the training principle of the second target detection model or the second image segmentation model provided by the embodiments of the present application.

[0182] As shown in Figure 5 , the model training apparatus can train a classification model, a first target detection model, and a first image segmentation model using a small number of first sample image sets in a target scene and a large number of second sample image sets in a non-target scene, and then use the classification model to select a third sample image set in the target scene from the sample image set without labels. After the model training apparatus obtains the third sample image set, on the one hand, it obtains the labels of the third sample image set (i.e., the type of the target in the third sample image and the detection box where the target is located) using the first target detection model, and then trains the second target detection model based on the third sample image set with labels. On the other hand, it obtains the labels of the third sample image set (i.e., the mask corresponding to the cropped image of each third sample image) using the first image segmentation model, and then trains the second image segmentation model based on the third sample image set with labels. Since the labels of the second sample image set in the non-target scene can be obtained from open source data, the number of sample images that need to be manually labeled can be reduced by the scheme of the present application, thereby reducing the training cost of the second target detection model and the second image segmentation model.

[0183] Figure 6 is an example of the training principle of the second target detection model or the second image segmentation model provided by the embodiments of the present application.

[0184] As shown in Figure 6 , after the model training apparatus obtains the first sample image set in the target scene and the second sample image set in the non-target scene, it can train the classification model, the first target detection model, and the first image segmentation model in the following manner:

[0185] training a classification model based on the first sample image set and the target scene as a label, and based on the second sample image set and the non-target scene as a label;

[0186] training a first target detection model based on each first sample image in the first sample image set and the category of the target in each first sample image and the detection frame in which the target is located as a label, and based on each second sample image in the second sample image set and the category of the target in each second sample image and the detection frame in which the target is located as a label;

[0187] cropping the detection frame in which the target in each first sample image is located to obtain a cropped image of each first sample image, and training a first image segmentation model based on the cropped image of each first sample image and the mask corresponding to the cropped image of each first sample image.

[0188] After the model training device trains the classification model, the first target detection model, and the first image segmentation model, it can obtain a set of unlabeled sample images, and then use the classification model to predict the classification result of each unlabeled sample image in the set of unlabeled sample images; based on the classification result of each unlabeled sample image, select a third sample image set in which the scene is the target scene from the set of unlabeled sample images.

[0189] After the model training device obtains the third sample image set, on the one hand, it uses the first target detection model to predict the category of the target in each third sample image in the third sample image set and the detection frame in which the target is located based on the third sample image set; and on the other hand, it crops the detection frame in which the target in each third sample image is located to obtain a cropped image of each third sample image, and uses the first image segmentation model to predict the mask corresponding to the cropped image of each third sample image based on the cropped image of each first sample image; and trains a second image segmentation model based on the cropped image of each third sample image and the mask corresponding to the cropped image of each third sample image.

[0190] In the above process, for the first sample image set, the labels of the first sample images in the first sample image set can be manually annotated, which include scene types, types of targets, and detection boxes in which the targets are located. The detection boxes predicted by the first target detection model can be represented by the coordinates of the diagonals of the rectangles, or can be represented by masks. For the second sample image set, a large number of non-target scene standard sample images can be collected at low cost, which are open source or accumulated in business history. The labels of the second sample images in the second sample image set are similar to those of the first sample images in the first sample image set, and include scene types, types of targets, and detection boxes in which the targets are located. Based on this, the first sample images in the first sample image set can be used as positive samples of the classification model, the first target detection model, and the first image segmentation model, and the second sample images in the second sample image set can be used as negative samples of the classification model, the first target detection model, and the first image segmentation model. The positive sample refers to a sample belonging to a target category or target object. The negative sample refers to a sample not belonging to a target category or target object.

[0191] Specifically, the first sample images in the first sample image set in the target scene are used as positive samples, and the second sample images in the second sample image set in the non-target scene are used as negative samples to train the classification model. The classification model is used to screen the unlabeled sample images in the unlabeled sample image set to obtain a third sample image set, for example, to screen out samples with high confidence in positive examples as third sample images in the third sample image set. On the other hand, the detection target corresponding to the target scene is used as a positive sample, and other detection targets are used as negative samples to train the first target detection model. The first target detection model is used to predict the type of target and the detection box in which the target is located in each third sample image in the third sample image set, and is used as the label of each third sample image to train the second target detection model. On the other hand, only based on the first sample images in the first sample image set in the target scene, the first image segmentation model is trained. After the first target detection model predicts the detection box in which the target is located in each third sample image in the third sample image set, the cropped image of each third sample image is cropped according to the predicted detection box, and the first image segmentation model is used to predict the mask corresponding to the cropped image of each third sample image as the label of the cropped image of each third sample image to train the second image segmentation model.

[0192] Taking a skiing scene as an example, the detection targets corresponding to the skiing scene can include: skiing subjects (such as people and skiing equipment), ski boards, and ski goggles. The skiing equipment includes ski boards, ski goggles, ski poles, and other handheld or wearable equipment. The detailed training process is as follows:

[0193] For the three targets, first, 1000 images of ski scenes are collected, and the types of the targets and the detection boxes in which the targets are located are manually labeled. Then, a large number of labeled images of non-ski scenes are collected at low cost, and the labels need to include the scene type, the target type, and the detection box in which the target is located.

[0194] During the training process, the images of ski scenes are used as positive samples, and the images of non-ski scenes are used as negative samples to train a classification model as a preliminary screening model of unlabeled data. The three targets of ski bodies, ski boards, and ski goggles are used as positive samples, and the remaining targets are used as negative samples to train a first target detection model. In addition, the images of ski scenes are used as samples to train a first image segmentation model.

[0195] After the classification model, the first target detection model, and the first image segmentation model are trained by the model training device, a large number of unlabeled sample images mixed with ski scenes and other scenes are collected, and the prediction results are calculated by the binary classification model. The samples with high confidence in the predicted positive examples are screened out. Then, the prediction results of the ski targets are calculated by the trained first target detection model, which are used as the training labels of a second target detection model. The corresponding cropped images are cropped according to the predicted detection boxes, and the masks are calculated on the cropped images by the first image segmentation model, which are used as the training labels of a second image segmentation model. The trained second target detection model and the second image segmentation model can be used as the detection model and the segmentation model that can be deployed on the mobile terminal.

[0196] It should be noted that, in the present embodiment, the cropping of the detection box predicted by the first target detection model (i.e., the cropping of the detection box in which the target is located in each third sample image) can also be used as the post-processing process of the first target detection model or the pre-processing process of the first image segmentation model, which is not limited in the present application.

[0197] In some embodiments, the S260 includes:

[0198] The model training apparatus divides the third sample image set into at least one first batch and at least one second batch; for the first batch, based on each third sample image in the first batch, the detection box in which the target is located in each third sample image in the first batch is taken as a label, and the second target detection model is trained; for the second batch, based on each third sample image in the second batch, a difference value of any third sample image in the second batch is determined, the difference value of the any third sample image being a difference value of the any third sample image relative to the third sample images in the second batch except the any third sample image; based on the difference value of the any third sample image, the third sample images in the second batch are sorted in ascending order of the difference value, to obtain an arrangement order of the third sample images in the second batch; based on the arrangement order, at least one third sample image with a high ranking in the second batch is determined; and based on the at least one third sample image, the detection box in which the target is located in each third sample image of the at least one third sample image is taken as a label, and the second target detection model is trained.

[0199] Exemplarily, the model training apparatus first divides the third sample image set into at least one first batch and at least one second batch, and the division can be based on the characteristics of the images (such as the complexity of the target, the clarity of the image, etc.) or randomly performed to better manage the training process.

[0200] Exemplarily, for each third sample image in the first batch, the model training apparatus directly uses the detection box in which the target is located in the image as a label to train the second target detection model, and this direct supervision helps the model quickly learn the basic features of the target. For each third sample image in the second batch, the model training apparatus first calculates the difference value of the image and other images in the batch, which can be based on image content, features or other metrics, for evaluating the similarity or difference degree between images; then the third sample images in the second batch are sorted according to the difference values, and the purpose of sorting is to identify those sample images that have smaller differences in content or features from other images, and after sorting, the model training apparatus selects at least one third sample image with a high ranking for training, and these images usually contain more representative or challenging targets, thus helping to improve the generalization ability of the model.

[0201] In this embodiment, the training strategy of dividing batches and combining difference value sorting helps the model to pay more attention to those sample images that are more critical to improving performance in the training process. At the same time, by using different training methods for different batches, the model training apparatus can more flexibly adapt to different training needs and data characteristics.

[0202] Of course, in other alternative embodiments, the model training apparatus can divide the third sample images in the third sample image set into multiple batches; for different batches in the multiple batches, different training strategies can be adopted to train the second target detection model. The above-mentioned adoption of all third sample images in the current batch or third sample images with smaller difference values is only exemplary description and should not be understood as a limitation of the present application.

[0203] In some embodiments, the dividing of the third sample images in the third sample image set into at least one first batch and at least one second batch comprises:

[0204] The model training apparatus divides the third sample images in the third sample image set into multiple second batches; based on the number of the multiple batches and a preset ratio, the multiple batches are divided into the at least one first batch and the at least one second batch; wherein the preset ratio is a ratio of the number of the at least one first batch to the number of the at least one second batch, or the preset ratio is a ratio of the number of the at least one second batch to the number of the at least one first batch.

[0205] Exemplarily, when the model training apparatus divides the third sample image set into multiple second batches, it can be based on the number, characteristics or randomness of the images, aiming to disperse a large number of sample images into multiple batches for management and processing.

[0206] Exemplarily, when the model training apparatus divides the multiple batches into the at least one first batch and the at least one second batch based on the number of the multiple batches and a preset ratio, if the preset ratio is a ratio of the number of the at least one first batch to the number of the at least one second batch, the model training apparatus can take the product of the number of the multiple batches and the preset ratio as the number of the at least one first batch, and then randomly select a corresponding number of batches in the multiple batches as the at least one first batch, and take the remaining batches as the at least one second batch. Similarly, if the preset ratio is a ratio of the number of the at least one second batch to the number of the at least one first batch, the model training apparatus can take the product of the number of the multiple batches and the preset ratio as the number of the at least one second batch, and then randomly select a corresponding number of batches in the multiple batches as the at least one second batch, and take the remaining batches as the at least one first batch.

[0207] Of course, in other alternative embodiments, when the model training apparatus divides the plurality of batches into the at least one first batch and the at least one second batch based on the number of the plurality of batches and the preset ratio, it can also divide them into two parts based on batch characteristics (such as the image features focused on by the batch or the number of images in the batch), and determine the part with fewer batches as the at least one first batch, that is, the number of the at least one first batch can be less than the number of the at least one second batch, which means that the sample images in the first batch can be more representative or important and need to be trained first. This application does not make specific limitations.

[0208] In this embodiment, through this batch division and subdivision method based on the preset ratio, the model training apparatus can more flexibly and efficiently utilize the sample image set for model training, which not only improves the efficiency of the training process, but also helps to improve the performance and generalization ability of the model.

[0209] Figure 7 is a schematic flowchart of a target detection method 300 provided by an embodiment of the present application. The target detection method 300 can be executed by an electronic device. The electronic device can be implemented as a terminal device or a server. The terminal device can also be referred to as a user terminal, which includes but is not limited to a smartphone, a game console, a desktop computer, a tablet computer, an e-book reader, an MP4 player, an MP4 player, a laptop computer, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, etc. The server can be a standalone physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud services.

[0210] For ease of description, the target detection method 300 will be exemplarily described below taking a target detection apparatus as an example.

[0211] As shown in Figure 7 , the target detection method 300 can include:

[0212] S310, the target detection apparatus acquires a to-be-detected image.

[0213] Exemplarily, the to-be-detected image can be an image from various channels, such as a camera, a digital camera, a mobile phone, etc. According to specific application scenarios and requirements, the to-be-detected image can need to be processed and analyzed differently. For example, in a skiing scenario, the to-be-detected image can be an image including one or more targets, which can include a skiing subject (such as including a person and skiing equipment), a ski board, or a ski goggles. The skiing equipment includes a ski board, a ski goggles, a ski pole, and other handheld or wearable equipment.

[0214] S320, the target detection apparatus predicts the type of the target in the to-be-detected image and the detection box in which the target is located by using a second target detection model deployed on the user terminal.

[0215] The second target detection model is trained in the following manner:

[0216] The first sample image set in the target scene and the second sample image set in the non-target scene are obtained; a classification model is trained based on the first sample image set, taking the target scene as the label, and based on the second sample image set, taking the non-target scene as the label; a first target detection model is trained based on each first sample image in the first sample image set, taking the category of the target in the each first sample image and the detection box in which the target is located as the label, and based on each second sample image in the second sample image set, taking the category of the target in the each second sample image and the detection box in which the target is located as the label; the classification results of each unlabeled sample image in the unlabeled sample image set are predicted by using the classification model based on the unlabeled sample image set; based on the classification result of the each unlabeled sample image, a third sample image set in which the scene is the target scene is selected from the unlabeled sample image set, and based on the third sample image set, the category of the target in each third sample image in the third sample image set and the detection box in which the target is located are predicted by using the first target detection model; the second target detection model is trained based on the third sample image set, taking the type of the target in the each third sample image and the detection box in which the target is located as the label.

[0217] In some embodiments, the S320 includes:

[0218] When the target detection apparatus predicts the type of the target in the to-be-detected image and the detection box in which the target is located by using the second target detection model deployed on the user terminal, the second target detection model can be used to obtain feature maps corresponding to a plurality of category channels and feature maps corresponding to a plurality of position channels; the category of the target in the to-be-detected image is determined based on the feature maps corresponding to the plurality of category channels, and the detection box in which the target is located is determined based on the feature maps corresponding to the plurality of position channels.

[0219] For example, the target detection apparatus obtains, based on the to-be-detected image, feature maps corresponding to a plurality of category channels and feature maps corresponding to a plurality of position channels by using the second target detection model; whether a target is detected is determined based on the feature map corresponding to each category channel of the plurality of category channels, and in the case where a target is detected, the detection box in which the detected target is located is determined based on the feature maps corresponding to the plurality of position channels.

[0220] For example, the plurality of category channels can include a channel corresponding to a detected target in the target scene.

[0221] For example, assuming that the target scene is a skiing scene, the plurality of category channels can include a skiing subject (e.g., including a person and skiing equipment) channel, a ski board channel, or a ski goggles channel.

[0222] Correspondingly, for a to-be-detected image in the to-be-detected images, the target detection apparatus obtains, by using the second target detection model, a feature map corresponding to the skiing subject channel, a feature map corresponding to the ski board channel, and a feature map corresponding to the ski goggles channel and a plurality of position channel feature maps of the to-be-detected image. Then, the target detection apparatus determines, based on the feature map corresponding to the skiing subject channel of the to-be-detected image, whether there is a skiing subject in the to-be-detected image, and in the case where there is a skiing subject, determines a detection box of the skiing subject in the to-be-detected image based on the plurality of position channel feature maps of the to-be-detected image. Similarly, the target detection apparatus determines, based on the feature map corresponding to the ski board channel of the to-be-detected image, whether there is a ski board in the to-be-detected image, and in the case where there is a ski board, determines a detection box of the ski board in the to-be-detected image based on the plurality of position channel feature maps of the to-be-detected image. The target detection apparatus determines, based on the feature map corresponding to the ski goggles channel of the to-be-detected image, whether there is a ski goggles in the to-be-detected image, and in the case where there is a ski goggles, determines a detection box of the ski goggles in the to-be-detected image based on the plurality of position channel feature maps of the to-be-detected image. It should be noted that the plurality of position channels can include diagonal coordinate channels of a rectangle, for example, diagonal x-axis coordinate channels and diagonal y-axis coordinate channels of a rectangle, i.e., four position channels.

[0223] Assuming that the feature map corresponding to the skiing subject channel, the feature map corresponding to the ski board channel, and the feature map corresponding to the ski goggles channel and the plurality of position channel feature maps of the to-be-detected image are all KxL feature maps, then:

[0224] For the to-be-detected image in the to-be-detected image, the target detection device utilizes the second target detection model to obtain a KxLx7 feature map of the to-be-detected image, wherein 7 represents a skiing subject channel, a ski board channel, a ski goggles channel, two x-axis coordinate channels of opposite corners of a rectangle, and two y-axis coordinate channels of opposite corners of the rectangle. Based on this, after the target detection device obtains the KxLx7 feature map of the to-be-detected image, the target detection device can determine whether there is a skiing subject in the to-be-detected image based on the KxL feature map corresponding to the skiing subject channel of the to-be-detected image, and in the case where there is a skiing subject, determine the detection frame of the skiing subject in the to-be-detected image based on the KxLx4 feature map corresponding to the four position channels of the to-be-detected image. Similarly, the target detection device determines whether there is a ski board in the to-be-detected image based on the KxL feature map corresponding to the ski board channel of the to-be-detected image, and in the case where there is a ski board, determines the detection frame of the ski board in the to-be-detected image based on the KxLx4 feature map corresponding to the four position channels of the to-be-detected image. The target detection device determines whether there is a ski goggles in the to-be-detected image based on the KxL feature map corresponding to the ski goggles channel of the to-be-detected image, and in the case where there is a ski goggles, determines the detection frame of the ski goggles in the to-be-detected image based on the KxLx4 feature map corresponding to the four position channels of the to-be-detected image.

[0225] In some embodiments, the method can further include:

[0226] The target detection device crops the detection frame in which the target is located in the to-be-detected image to obtain a cropped image of the to-be-detected image.

[0227] The target detection device utilizes the second image segmentation model deployed on the user equipment to predict a mask corresponding to the cropped image of the to-be-detected image based on the cropped image of the to-be-detected image.

[0228] The second image segmentation model is trained in the following training manner:

[0229] The detection frame in which the target is located in each first sample image is cropped to obtain a cropped image of each first sample image. The first image segmentation model is trained based on the cropped image of each first sample image and the mask corresponding to the cropped image of each first sample image. The detection frame in which the target is located in each third sample image is cropped to obtain a cropped image of each third sample image, and the mask corresponding to the cropped image of each third sample image is predicted based on the cropped image of each first sample image and the first image segmentation model. The second image segmentation model is trained based on the cropped image of each third sample image and the mask corresponding to the cropped image of each third sample image.

[0230] It should be understood that the specific training process of the second target detection model and the second image segmentation model in the target detection method 300 can refer to the corresponding description in the model training method 200, which is not limited herein.

[0231] Figure 8 is a schematic flowchart of a target detection method 400 provided by an embodiment of the present application. The target detection method 400 can be executed by an electronic device. The electronic device can be implemented as a terminal device or a server. The terminal device can also be referred to as a user terminal, which includes but is not limited to a smart phone, a game console, a desktop computer, a tablet computer, an e-book reader, an MP4 player, an MP4 player, a laptop computer, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, etc. The server can be a standalone physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud services.

[0232] For ease of description, the target detection method 400 will be exemplarily described below by taking a target detection apparatus as an example.

[0233] As shown in Figure 8 , the target detection method 400 can include:

[0234] S410, obtaining a to-be-detected image.

[0235] Exemplarily, the to-be-detected image can be an image from various channels, such as a camera, a digital camera, a mobile phone, etc. According to specific application scenarios and requirements, the to-be-detected image can need to be processed and analyzed differently. For example, in a skiing scenario, the to-be-detected image can be an image including one or more targets, which can include a skiing subject (e.g., including a person and skiing equipment), a ski board, or a ski goggles. The skiing equipment includes a ski board, a ski goggles, a ski pole, and the like.

[0236] S420, predicting the type of the target in the to-be-detected image and the detection box where the target is located by using a second target detection model.

[0237] Exemplarily, the second target detection model is obtained by training in the following manner:

[0238] obtain a first sample image set in a target scene and a second sample image set in a non-target scene; train a classification model based on the first sample image set with the target scene as a label and based on the second sample image set with the non-target scene as a label; train a first target detection model based on each first sample image in the first sample image set with a category of a target in the each first sample image and a detection box where the target is located as a label and based on each second sample image in the second sample image set with a category of a target in the each second sample image and a detection box where the target is located as a label; predict a classification result of each unlabeled sample image in an unlabeled sample image set based on the classification model and the unlabeled sample image set; select a third sample image set in the unlabeled sample image set with a scene of each third sample image in the third sample image set being the target scene based on the classification result of the each third sample image, and predict a category of a target in each third sample image in the third sample image set and a detection box where the target is located based on the first target detection model and the third sample image set; and train a second target detection model based on the third sample image set with the category of the target in the each third sample image and the detection box where the target is located as a label.

[0239] Exemplarily, when predicting the category of the target in the to-be-detected image and the detection box where the target is located by using the second target detection model deployed on the user terminal, the second target detection model can be used to obtain feature maps corresponding to a plurality of category channels and feature maps corresponding to a plurality of position channels; the category of the target in the to-be-detected image is determined based on the feature maps corresponding to the plurality of category channels, and the detection box where the target is located in the to-be-detected image is determined based on the feature maps corresponding to the plurality of position channels.

[0240] Exemplarily, the plurality of category channels can include channels corresponding to detection targets in the target scene.

[0241] For example, assuming that the target scene is a skiing scene, the plurality of category channels can include a skiing subject (for example, including a person and skiing equipment) channel, a ski board channel, or a ski goggles channel.

[0242] Assuming that the feature map corresponding to the skiing subject channel, the feature map corresponding to the snowboard channel, and the feature map corresponding to the ski goggles channel and the feature map corresponding to the plurality of position channels of the to-be-detected image are KxL feature maps, then: for the to-be-detected image in the to-be-detected image, the target detection device can obtain a KxLx7 feature map of the to-be-detected image by using the second target detection model, where 7 represents the skiing subject channel, the snowboard channel, the ski goggles channel, the two x-axis coordinate channels of the opposite corners of the rectangle, and the two y-axis coordinate channels of the opposite corners of the rectangle. Based on this, the target detection device can determine the category of the target in the to-be-detected image and the detection box where the target in the to-be-detected image is located based on the KxL feature map.

[0243] S430, detecting the target?

[0244] Exemplarily, after the target detection device predicts the KxLx7 feature map of the to-be-detected image by using the second target detection model, the target detection device can determine whether there is a skiing subject in the to-be-detected image based on the KxL feature map corresponding to the skiing subject channel of the to-be-detected image. Similarly, the target detection device determines whether there is a snowboard in the to-be-detected image based on the KxL feature map corresponding to the snowboard channel of the to-be-detected image. The target detection device determines whether there is a ski goggles in the to-be-detected image based on the KxL feature map corresponding to the ski goggles channel of the to-be-detected image.

[0245] S440, cropping the detection box to obtain a cropped image of the to-be-detected image.

[0246] Exemplarily, in the case where there is a skiing subject, the target detection device determines the detection box of the skiing subject in the to-be-detected image based on the KxLx4 feature map corresponding to the four position channels of the to-be-detected image. Similarly, in the case where there is a snowboard, the target detection device determines the detection box of the snowboard in the to-be-detected image based on the KxLx4 feature map corresponding to the four position channels of the to-be-detected image. In the case where there is a ski goggles, the target detection device determines the detection box of the ski goggles in the to-be-detected image based on the KxLx4 feature map corresponding to the four position channels of the to-be-detected image.

[0247] S450, predicting a mask corresponding to the cropped image of the to-be-detected image by using a second image segmentation model.

[0248] The target detection device predicts a mask corresponding to the cropped image of the to-be-detected image by using a second image segmentation model.

[0249] S460, outputting the mask.

[0250] The target detection device outputs the mask corresponding to the cropped image of the to-be-detected image.

[0251] S470, end.

[0252] In this embodiment, the input is an image to be detected. First, it is determined whether there is a target by using the second target detection model. If there is a target, the mask of the target is calculated by using the second image segmentation model. If there is no target, the mask is empty. Then, the segmentation of the target is realized.

[0253] The preferred embodiments of the present application are described in detail above with reference to the drawings, but the present application is not limited to the specific details of the embodiments described above. Within the technical concept of the present application, various simple modifications can be made to the technical solutions of the present application, and these simple modifications all belong to the protection scope of the present application. For example, in the specific technical features described in the above embodiments, any suitable combination can be made without contradiction. In order to avoid unnecessary repetition, various possible combinations are not described again in the present application. For another example, various different embodiments of the present application can also be combined arbitrarily, as long as it does not deviate from the idea of the present application, and it should also be considered as disclosed in the present application.

[0254] It should also be understood that in various method embodiments of the present application, the size of the serial number of the processes described above does not mean the order of execution. The execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0255] The above describes the method provided by the embodiments of the present application. The following describes the device provided by the embodiments of the present application.

[0256] Figure 9 is a schematic block diagram of the model training device 500 provided by the embodiments of the present application.

[0257] As shown in Figure 9 , the model training device 500 can include:

[0258] The acquisition unit 510 is configured to acquire a first sample image set in a target scene and a second sample image set in a non-target scene.

[0259] The first training unit 520 is configured to train a classification model based on the first sample image set, with the target scene as a label, and based on the second sample image set, with the non-target scene as a label.

[0260] The second training unit 530 is configured to train a first target detection model based on each first sample image in the first sample image set, with the category of the target in the each first sample image and the detection frame where the target is located as a label, and based on each second sample image in the second sample image set, with the category of the target in the each second sample image and the detection frame where the target is located as a label.

[0261] The first prediction unit 540 is configured to predict, based on the set of unlabeled sample images, a classification result of each of the set of unlabeled sample images by using the classification model.

[0262] The second prediction unit 550 is configured to select, based on the classification result of each of the set of unlabeled sample images, a third set of sample images with a scene being the target scene from the set of unlabeled sample images, and predict, based on the third set of sample images, a category of an object in each of the third set of sample images and a detection box in which the object is located by using the first object detection model.

[0263] The third training unit 560 is configured to train a second object detection model based on the third set of sample images, and take the category of the object in each of the third set of sample images and the detection box in which the object is located as a label.

[0264] In some embodiments, the backbone network of the first object detection model comprises at least one first feature extraction layer, and the backbone network of the second object detection model comprises at least one second feature extraction layer.

[0265] In some embodiments, the number of the at least one first feature extraction layer is greater than the number of the at least one second feature extraction layer, or the number of network parameters of the first feature extraction layer is greater than the number of network parameters of the second feature extraction layer.

[0266] In some embodiments, the first feature extraction layer is a feature extraction layer for attention processing of input features, and the second feature extraction layer is a depth separable convolution layer.

[0267] In some embodiments, the second prediction unit 550 is specifically configured to:

[0268] obtain, based on the third set of sample images and by using the first object detection model, feature maps corresponding to a plurality of category channels and feature maps corresponding to a plurality of position channels;

[0269] determine, based on the feature maps corresponding to the plurality of category channels, the category of the object in each of the third set of sample images, and determine, based on the feature maps corresponding to the plurality of position channels, the detection box in which the object is located in each of the third set of sample images.

[0270] In some embodiments, the second training unit 530 is further configured to:

[0271] crop the detection box in which the object is located in each of the first set of sample images to obtain a cropped image of each of the first set of sample images.

[0272] The third training unit 560 is further configured to train the first image segmentation model based on the cropped image of each first sample image and the mask corresponding to the cropped image of each first sample image.

[0273] The third training unit 560 is further configured to crop the bounding box in which the target is located in each third sample image to obtain a cropped image of each third sample image, and predict a mask corresponding to the cropped image of each third sample image by using the first image segmentation model based on the cropped image of each first sample image.

[0274] The third training unit 560 is further configured to train the second image segmentation model based on the cropped image of each third sample image and the mask corresponding to the cropped image of each third sample image.

[0275] In some embodiments, the second training unit 530 is specifically configured to:

[0276] The second training unit 530 is specifically configured to down-sample the cropped image of each first sample image to obtain a down-sampled image of the cropped image of each first sample image, and down-sample the mask corresponding to the cropped image of each first sample image to obtain a down-sampled image of the mask corresponding to the cropped image of each first sample image.

[0277] The second training unit 530 is specifically configured to train the first image segmentation model based on the down-sampled image of the cropped image of each first sample image, and take the down-sampled image of the mask corresponding to the cropped image of each first sample image as a label.

[0278] In some embodiments, the third training unit 560 is specifically configured to:

[0279] The third training unit 560 is specifically configured to down-sample the cropped image of each third sample image to obtain a down-sampled image of the cropped image of each third sample image, and down-sample the mask corresponding to the cropped image of each third sample image to obtain a down-sampled image of the mask corresponding to the cropped image of each third sample image.

[0280] The third training unit 560 is specifically configured to train the second image segmentation model based on the down-sampled image of the cropped image of each third sample image, and take the down-sampled image of the mask corresponding to the cropped image of each third sample image as a label.

[0281] In some embodiments, the backbone network of the first image segmentation model comprises at least one third feature extraction layer, and the backbone network of the second image segmentation model comprises at least one fourth feature extraction layer.

[0282] In some embodiments, the number of the at least one third feature extraction layer is greater than the number of the at least one fourth feature extraction layer, or the number of network parameters of the third feature extraction layer is greater than the number of network parameters of the fourth feature extraction layer.

[0283] In some embodiments, the third feature extraction layer is a feature extraction layer for attention processing of input features, and the fourth feature extraction layer is a deep separable convolution layer.

[0284] In some embodiments, the third training unit 560 is specifically configured to:

[0285] divide the third sample images in the third sample image set into at least one first batch and at least one second batch;

[0286] for the first batch, based on each third sample image in the first batch, train the second target detection model by taking the bounding box in which the target is located in each third sample image in the first batch as a label;

[0287] for the second batch, based on each third sample image in the second batch, determine a difference value of any third sample image in the second batch, the difference value of the any third sample image being a difference value of the any third sample image relative to the third sample images in the second batch except the any third sample image;

[0288] based on the difference value of the any third sample image, sort the third sample images in the second batch in ascending order of the difference value to obtain an arrangement order of the third sample images in the second batch;

[0289] based on the arrangement order, determine at least one third sample image in the second batch that is arranged at a front position;

[0290] based on the at least one third sample image, train the second target detection model by taking the bounding box in which the target is located in each third sample image in the at least one third sample image as a label.

[0291] In some embodiments, the third training unit 560 is specifically configured to:

[0292] divide the third sample images in the third sample image set into a plurality of second batches;

[0293] based on the number of the plurality of batches and a preset ratio, divide the plurality of batches into the at least one first batch and the at least one second batch;

[0294] wherein the preset ratio is a ratio of the number of the at least one first batch to the number of the at least one second batch, or the preset ratio is a ratio of the number of the at least one second batch to the number of the at least one first batch.

[0295] It should be understood that the device embodiments and the method embodiments can correspond to each other, and similar descriptions can be referred to the method embodiments. To avoid repetition, no longer described here. Specifically, the model training device 500 can correspond to the corresponding subject in the model training method 200 of performing the embodiments of the present application, and each unit in the model training device 500 is respectively for realizing the corresponding process in the model training method 200. For the sake of brevity, no longer described here.

[0296] Figure 10 is a schematic block diagram of the target detection device 600 provided by the embodiments of the present application.

[0297] As shown in Figure 10 , the target detection device 600 can include:

[0298] The acquisition unit 610 is configured to acquire a to-be-detected image.

[0299] The prediction unit 620 is configured to predict the type of the target in the to-be-detected image and the detection box where the target is located by using a second target detection model deployed on the user terminal.

[0300] The second target detection model is obtained by training in the following manner:

[0301] Obtain a first sample image set in a target scene and a second sample image set in a non-target scene.

[0302] Train a classification model based on the first sample image set with the target scene as the label and based on the second sample image set with the non-target scene as the label.

[0303] Train a first target detection model based on each first sample image in the first sample image set with the category of the target in the each first sample image and the detection box where the target is located as the label and based on each second sample image in the second sample image set with the category of the target in the each second sample image and the detection box where the target is located as the label.

[0304] Based on the unlabeled sample image set, predict the classification result of each unlabeled sample image in the unlabeled sample image set by using the classification model.

[0305] Based on the classification result of the each unlabeled sample image, select a third sample image set in the unlabeled sample image set with the scene being the target scene, and based on the third sample image set, predict the category of the target in each third sample image in the third sample image set and the detection box where the target is located by using the first target detection model.

[0306] Based on the third sample image set, the type of the target in each third sample image and the detection frame in which the target is located are taken as labels to train the second target detection model.

[0307] It should be understood that the device embodiments and the method embodiments can correspond to each other, and similar descriptions can be referred to the method embodiments. To avoid repetition, no longer described here. Specifically, the target detection device 600 can correspond to the corresponding subject in the target detection method 300 or the target detection method 400 for executing the embodiments of the present application, and each unit in the target detection device 600 is respectively for realizing the corresponding process in the target detection method 300 or the target detection method 400. In order to be brief, no longer described here.

[0308] It should also be understood that each unit in the model training device 500 or the target detection device 600 involved in the embodiments of the present application is based on logical function division. In actual application, the function of one unit can also be realized by multiple units, or the functions of multiple units can be realized by one unit. Even, these functions can also be assisted by one or more other units. For example, part or all of the model training device 500 or the target detection device 600 are combined into one or several other units. For another example, some unit(s) in the model training device 500 or the target detection device 600 can also be further split into multiple units with smaller functions to constitute, which can realize the same operation without affecting the realization of the technical effects of the embodiments of the present application. For another example, the model training device 500 or the target detection device 600 can also include other units. In actual application, these functions can also be assisted by other units, and can be realized by multiple units.

[0309] It should also be understood that the term "module" or "unit" involved in the embodiments of the present application refers to a computer program or a part of a computer program with a predetermined function, and works with other related parts to achieve a predetermined target, and can be realized by using software, hardware (such as processing circuit or memory) or combination thereof. Similarly, one processor (or multiple processors or memory) can be used to realize one or more modules or units. In addition, each module or unit can be a part of the whole module or unit containing the function of the module or unit.

[0310] According to another embodiment of the present application, the model training apparatus 500 or the target detection apparatus 600 involved in the embodiments of the present application and the method of the embodiments of the present application can be constructed and implemented by running a computer program (including program codes) capable of performing steps involved in the corresponding method on a general computing device including processing elements and storage elements such as a central processing unit (CPU), random access memory (RAM), read-only memory (ROM), etc. The computer program can be recorded on a computer readable storage medium, for example, and loaded into an electronic device through the computer readable storage medium, and the computer program is used to implement the corresponding method of the embodiments of the present application. In other words, the units involved above can be implemented in the form of hardware, in the form of instructions of software, or in the form of a combination of software and hardware. Specifically, the steps of the method embodiments in the embodiments of the present application can be completed by the integrated logic circuit of hardware in the processor and / or the instructions of software, and the steps of the method disclosed in the embodiments of the present application can be directly embodied as execution completed by a hardware decoding processor, or execution completed by a combination of hardware and software in the decoding processor. Alternatively, the software can be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, register, etc. The software in the storage can be run by the processor to complete the steps in the method embodiments involved above.

[0311] Figure 11 is a schematic structural diagram of the electronic device 700 provided by the embodiments of the present application.

[0312] As shown in Figure 11 , the electronic device 700 at least includes a processor 710 and a computer readable storage medium 720. The processor 710 and the computer readable storage medium 720 can be connected by a bus or other means. The computer readable storage medium 720 is used to store a computer program 721, the computer program 721 including computer instructions, and the processor 710 is used to execute the computer instructions stored in the computer readable storage medium 720. The processor 710 is the computing core and control core of the electronic device 700, which is suitable for implementing one or more computer instructions, and is particularly suitable for loading and executing one or more computer instructions to implement a corresponding method flow or a corresponding function.

[0313] As an example, the processor 710 can also be referred to as a central processing unit (CPU). The processor 710 can include, without limitation, a general-purpose processor, a

[0314] As an example, the computer-readable storage medium 720 can be a high-speed RAM memory, or a non-volatile memory such as at least one disk memory; optionally, it can also be at least one computer-readable storage medium located remotely from the aforementioned processor 710. In particular, the computer-readable storage medium 720 includes, without limitation, a volatile memory and / or a non-volatile memory. The non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically EPROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM) used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synch link DRAM (SLDRAM), and direct Rambus RAM (DR RAM).

[0315] As shown in Figure 11 The electronic device 700 can also include a transceiver 730.

[0316] The transceiver 730 can include a transmitter and a receiver. The transceiver 730 can further include an antenna, and the number of the antenna can be one or more.

[0317] It should be understood that the various components in the electronic device 700 are connected through a bus system, and the bus system includes, in addition to a data bus, a power supply bus, a control bus, and a status signal bus. It is worth noting that the electronic device 700 can be any electronic device with data processing capability; the first computer instructions are stored in the computer readable storage medium 720; the first computer instructions stored in the computer readable storage medium 720 are loaded and executed by the processor 710 to implement the corresponding steps in the method embodiments shown in the method embodiments of the present application. Figure 2 、 7 In a specific implementation, the first computer instructions in the computer readable storage medium 720 are loaded and executed by the processor 710 to implement the corresponding steps, and to avoid repetition, details are not described here.

[0318] According to another aspect of the present application, the embodiments of the present application provide a chip. The chip can be an integrated circuit chip, which has the processing capability of signals, and can implement or execute the disclosed methods, steps and logic block diagrams in the embodiments of the present application. The chip can also be called a system chip, a system chip, a chip system or a system on chip, etc. The chip can be applied to various electronic devices capable of installing the chip, so that the device installed with the chip can execute the corresponding steps in the disclosed methods or logic block diagrams in the embodiments of the present application. For example, the chip can be adapted to implement one or more computer instructions, and is particularly adapted to load and execute one or more computer instructions to implement the corresponding method flow or corresponding function.

[0319] According to another aspect of the present application, the embodiments of the present application provide a computer readable storage medium (Memory). The computer readable storage medium is a memory device of a computer, used to store programs and data. It can be understood that the computer readable storage medium here can include the built-in storage medium in the computer, and of course, it can also include the extended storage medium supported by the computer. The computer readable storage medium provides a storage space, and the storage space stores the operating system of the electronic device. The storage space stores computer instructions adapted to be loaded and executed by the processor, and the computer instructions are read and executed by the processor of the computer device, so that the computer device executes the corresponding steps in the disclosed methods or logic block diagrams in the embodiments of the present application.

[0320] According to another aspect of the present application, the embodiments of the present application provide a computer program product or computer program. The computer program product or computer program includes computer instructions stored in a computer readable storage medium. The processor of the computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions, so that the computer device executes the corresponding steps in the disclosed methods or logic block diagrams in the embodiments of the present application. In other words, when the schemes provided by the present application are implemented by using software, the schemes can be implemented in the form of a computer program product or computer program in whole or in part. The computer program product or computer program includes one or more computer instructions. When the computer program instructions are loaded and executed on the computer, the processes of the embodiments of the present application or the functions of the embodiments of the present application are run in whole or in part.

[0321] It is worth noting that the computer involved in the present application can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions involved in the present application can be stored in a computer readable storage medium or transmitted from one computer readable storage medium to another computer readable storage medium, for example, the computer instructions can be transmitted from one website site, computer, server or data center to another website site, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.).

[0322] Those skilled in the art can realize that the units and flow steps of the examples described in combination with the embodiments disclosed herein can be realized in electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solutions. In other words, those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0323] Finally, it should be pointed out that the above is only a specific embodiment of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical range disclosed by the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims. For example, in the above specific embodiments, various specific technical features described in the above specific embodiments can be combined in any appropriate manner without contradiction. For example, various embodiments of the present application can also be combined in any manner without departing from the basic idea of the present application, which should also be considered as disclosed by the present application.

Claims

1. A method for model training based on pseudo labels, characterized in that, include: Acquire a first sample image set under a target scene and a second sample image set under a non-target scene; Training a classification model based on the first sample image set with the target scene as a label, and based on the second sample image set with the non-target scene as a label; Based on each first sample image in the first sample image set, using the category of the target and the detection box where the target is located in each first sample image as a label, and based on each second sample image in the second sample image set, using the category of the target and the detection box where the target is located in each second sample image as a label, training a first object detection model; Based on the unlabeled sample image set, using the classification model, predicting the classification result of each unlabeled sample image in the unlabeled sample image set; Based on the classification result of each of the unlabeled sample images, selecting a third sample image set whose scene is the target scene from the unlabeled sample image set; and based on the third sample image set, using the first object detection model, predicting the category of the object and the detection box where the object is located in each third sample image in the third sample image set; Based on the third sample image set, a second object detection model is trained using the type of the object and the detection box where the object is located in each of the third sample images as labels.

2. The method of claim 1, wherein, The backbone network of the first target detection model includes at least one first feature extraction layer, and the backbone network of the second target detection model includes at least one second feature extraction layer; The number of the at least one first feature extraction layer is greater than the number of the at least one second feature extraction layer, or the number of network parameters of the first feature extraction layer is greater than the number of network parameters of the second feature extraction layer.

3. The method of claim 2, wherein, The first feature extraction layer is a feature extraction layer that performs attention processing on input features, and the second feature extraction layer is a depth-separable convolutional layer.

4. The method of claim 1, wherein, The method of predicting the category of the target and the detection frame where the target is located in each third sample image in the third sample image set by using the first target detection model based on the third sample image set includes: Based on the third sample image set, using the first object detection model, obtaining feature maps corresponding to multiple category channels and feature maps corresponding to multiple position channels; Based on the feature maps corresponding to the multiple category channels, determine the category of the target in each third sample image in the third sample image set, and based on the feature maps corresponding to the multiple position channels, determine the detection box where the target is located in each third sample image.

5. The method of claim 1, wherein, The method further comprises: Cropping the detection frame where the target is located in each of the first sample images to obtain a cropped image of each of the first sample images; training a first image segmentation model based on the cropped image of each first sample image and the mask corresponding to the cropped image of each first sample image; cropping a detection box in which the target in each third sample image is located to obtain a cropped image of each third sample image, and predicting a mask corresponding to the cropped image of each third sample image by using the first image segmentation model based on the cropped image of each first sample image; training a second image segmentation model based on the cropped image of each third sample image and the mask corresponding to the cropped image of each third sample image.

6. The method of claim 5, wherein, The training of the first image segmentation model based on the cropped image of each first sample image and the mask corresponding to the cropped image of each first sample image comprises: down-sampling the cropped image of each first sample image to obtain a down-sampled image of the cropped image of each first sample image, and down-sampling the mask corresponding to the cropped image of each first sample image to obtain a down-sampled image of the mask corresponding to the cropped image of each first sample image; training the first image segmentation model based on the down-sampled image of the cropped image of each first sample image and taking the down-sampled image of the mask corresponding to the cropped image of each first sample image as a label; The training of the second image segmentation model based on the cropped image of each third sample image and the mask corresponding to the cropped image of each third sample image comprises: down-sampling the cropped image of each third sample image to obtain a down-sampled image of the cropped image of each third sample image, and down-sampling the mask corresponding to the cropped image of each third sample image to obtain a down-sampled image of the mask corresponding to the cropped image of each third sample image; training the second image segmentation model based on the down-sampled image of the cropped image of each third sample image and taking the down-sampled image of the mask corresponding to the cropped image of each third sample image as a label.

7. The method of claim 5, wherein, The backbone network of the first image segmentation model comprises at least one third feature extraction layer, and the backbone network of the second image segmentation model comprises at least one fourth feature extraction layer; The number of the at least one third feature extraction layer is greater than the number of the at least one fourth feature extraction layer, or the number of network parameters of the third feature extraction layer is greater than the number of network parameters of the fourth feature extraction layer.

8. The method of claim 7, wherein, The third feature extraction layer is a feature extraction layer for attention processing of input features, and the fourth feature extraction layer is a depth separable convolution layer.

9. The method according to any one of claims 1 to 8, characterized in that, The training of the second target detection model based on the third sample image set and taking the detection box in which the target in each third sample image is located as a label comprises: dividing the third sample image set into at least one first batch and at least one second batch; for the first batch, training the second target detection model based on each third sample image in the first batch and taking the detection box in which the target in each third sample image in the first batch is located as a label; and for the second batch, training the second target detection model based on each third sample image in the second batch and taking the detection box in which the target in each third sample image in the second batch is located as a label. For the second batch, based on each third sample image in the second batch, a difference value of any third sample image in the second batch is determined, and the difference value of the any third sample image is a difference value of the any third sample image relative to third sample images in the second batch except the any third sample image; Based on the difference value of the any third sample image, the third sample images in the second batch are sorted in ascending order of the difference value, to obtain an arrangement order of the third sample images in the second batch; Based on the arrangement order, at least one third sample image with a high ranking in the second batch is determined; Based on the at least one third sample image, the second target detection model is trained by taking a detection frame in which a target is located in each of the at least one third sample image as a label.

10. The method of claim 9, wherein, The third sample images in the third sample image set are divided into at least one first batch and at least one second batch, including: The third sample images in the third sample image set are divided into multiple second batches; Based on the number of the multiple batches and a preset ratio, the multiple batches are divided into the at least one first batch and the at least one second batch; The preset ratio is a ratio of the number of the at least one first batch to the number of the at least one second batch, or the preset ratio is a ratio of the number of the at least one second batch to the number of the at least one first batch.

11. A target detection method characterized by, Including: An image to be detected is obtained; A type of a target in the image to be detected and a detection frame in which the target is located are predicted by using a second target detection model deployed on a user terminal; The second target detection model is trained in the following manner: A first sample image set under a target scene and a second sample image set under a non-target scene are obtained; A classification model is trained based on the first sample image set by taking the target scene as a label and based on the second sample image set by taking the non-target scene as a label; A first target detection model is trained based on each first sample image in the first sample image set by taking a category of a target in the each first sample image and a detection frame in which the target is located as a label, and based on each second sample image in the second sample image set by taking a category of a target in the each second sample image and a detection frame in which the target is located as a label; A classification result of each unlabeled sample image in an unlabeled sample image set is predicted by using the classification model based on the unlabeled sample image set; Based on the classification result of the each unlabeled sample image, a third sample image set in which a scene is the target scene is selected from the unlabeled sample image set, and a category of a target in each third sample image in the third sample image set and a detection frame in which the target is located are predicted by using the first target detection model based on the third sample image set; The second target detection model is trained based on the third sample image set by taking a type of a target in the each third sample image and a detection frame in which the target is located as a label. 12.A pseudo-label based model training apparatus, comprising: Including: An image to be detected is obtained; A type of a target in the image to be detected and a detection frame in which the target is located are predicted by using a second target detection model deployed on a user terminal; The second target detection model is trained in the following manner: A first sample image set under a target scene and a second sample image set under a non-target scene are obtained; A classification model is trained based on the first sample image set by taking the target scene as a label and based on the second sample image set by taking the non-target scene as a label; A first target detection model is trained based on each first sample image in the first sample image set by taking a category of a target in the each first sample image and a detection frame in which the target is located as a label, and based on each second sample image in the second sample image set by taking a category of a target in the each second sample image and a detection frame in which the target is located as a label; A classification result of each unlabeled sample image in an unlabeled sample image set is predicted by using the classification model based on the unlabeled sample image set; Based on the classification result of the each unlabeled sample image, a third sample image set in which a scene is the target scene is selected from the unlabeled sample image set, and a category of a target in each third sample image in the third sample image set and a detection frame in which the target is located are predicted by using the first target detection model based on the third sample image set; The second target detection model is trained based on the third sample image set by taking a type of a target in the each third sample image and a detection frame in which the target is located as a label. The acquisition unit is configured to acquire a first sample image set in a target scene and a second sample image set in a non-target scene; The first training unit is configured to train a classification model based on the first sample image set, with the target scene as a label, and based on the second sample image set, with the non-target scene as a label; The second training unit is configured to train a first target detection model based on each first sample image in the first sample image set, with a category of a target in the each first sample image and a detection frame in which the target is located as a label, and based on each second sample image in the second sample image set, with a category of a target in the each second sample image and a detection frame in which the target is located as a label; The first prediction unit is configured to predict a classification result of each unlabeled sample image in an unlabeled sample image set based on the unlabeled sample image set by using the classification model; The second prediction unit is configured to select a third sample image set in which a scene is the target scene from the unlabeled sample image set based on the classification result of the each unlabeled sample image, and predict a category of a target in each third sample image in the third sample image set and a detection frame in which the target is located based on the third sample image set by using the first target detection model; The third training unit is configured to train a second target detection model based on the third sample image set, with the category of the target in the each third sample image and the detection frame in which the target is located as a label.

13. A target detection apparatus characterized by comprising: The acquisition unit is configured to acquire a to-be-detected image; The prediction unit is configured to predict a category of a target in the to-be-detected image and a detection frame in which the target is located by using a second target detection model deployed on a user terminal; The second target detection model is trained in the following manner: The acquisition unit is configured to acquire a first sample image set in a target scene and a second sample image set in a non-target scene; The first training unit is configured to train a classification model based on the first sample image set, with the target scene as a label, and based on the second sample image set, with the non-target scene as a label; The second training unit is configured to train a first target detection model based on each first sample image in the first sample image set, with a category of a target in the each first sample image and a detection frame in which the target is located as a label, and based on each second sample image in the second sample image set, with a category of a target in the each second sample image and a detection frame in which the target is located as a label; The first prediction unit is configured to predict a classification result of each unlabeled sample image in an unlabeled sample image set based on the unlabeled sample image set by using the classification model; The second prediction unit is configured to select a third sample image set in which a scene is the target scene from the unlabeled sample image set based on the classification result of the each unlabeled sample image, and predict a category of a target in each third sample image in the third sample image set and a detection frame in which the target is located based on the third sample image set by using the first target detection model; ​ Based on the third sample image set, a type of the target in each third sample image and a detection frame in which the target is located are taken as labels to train the second target detection model.

14. An electronic device, comprising: The method comprises: a processor adapted to execute a computer program; a computer readable storage medium having stored therein a computer program, which, when executed by the processor, implements the method of any one of claims 1 to 12.

15. A computer-readable storage medium, characterized in that, A computer program product for storing a computer program which, when executed on a computer, causes the computer to perform the method of any one of claims 1 to 12.