Training Method, Device and Electronic Device of a Model
By using the encoding and decoding modules of the teacher model to determine the loss value in the object detection scenario, the initial student model is corrected, and the problem of low accuracy among heterogeneous models is solved, and the detection accuracy of the student model is improved.
Patent Information
- Application Number
- CN202310369995.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-07
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2043-04-07
AI Technical Summary
In the object detection scenario, when the teacher model and the student model are heterogeneous, the object detection model obtained by knowledge distillation has low accuracy.
By using the initial student model to object detection of sample images, combined with the encoding and decoding modules of the teacher model, the loss value is determined, and the initial student model is corrected based on the loss value until the accurate student model is obtained.
The detection accuracy of student models is improved and the consistency of the detection results of teacher models and student models is ensured.
Smart Images

Figure CN116597269B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technologies, in particular to the fields of object detection, image processing, etc., and specifically relates to a training method, device, and electronic device for a model. Background Art
[0002] In an object detection scenario, usually through a knowledge distillation training method, the knowledge of a teacher model with more parameters and a larger computational amount is transferred to a student model with fewer parameters and a smaller computational amount. However, when the teacher model and the student model are heterogeneous, through the knowledge distillation method, the accuracy of the obtained object detection model is relatively low. Summary of the Invention
[0003] The present disclosure provides a training method, device, and electronic device for a model.
[0004] According to one aspect of the present disclosure, there is provided a training method for a model, including:
[0005] Performing object detection on a sample image by using an initial student model to obtain a first object detection box and / or a first type corresponding to an object of interest in the sample image;
[0006] Encoding the sample image by using an encoding module in a teacher model to obtain a feature vector group corresponding to the sample image;
[0007] Decoding, by using a decoding module of the teacher model, the feature vector corresponding to the first object detection box in the feature vector group to obtain a second object detection box and / or a second type;
[0008] Determining a loss value according to a first difference between the first object detection box and the second object detection box and / or a second difference between the first type and the second type;
[0009] Correcting the initial student model based on the loss value until a student model is obtained.
[0010] According to another aspect of the present disclosure, there is provided a training device for a model, including:
[0011] A first detection module, configured to perform object detection on a sample image by using an initial student model to obtain a first object detection box and / or a first type corresponding to an object of interest in the sample image;
[0012] A second detection module, configured to encode the sample image by using an encoding module in a teacher model to obtain a feature vector group corresponding to the sample image;
[0013] The above-mentioned second detection module is configured to decode, by using a decoding module of the teacher model, the feature vector corresponding to the first object detection box in the feature vector group to obtain a second object detection box and / or a second type;
[0014] A determination module, configured to determine a loss value according to a first difference between a first target detection box and a second target detection box, and / or a second difference between a first type and a second type;
[0015] A correction module, configured to correct an initial student model based on the loss value until a student model is obtained.
[0016] According to another aspect of the present disclosure, there is provided an electronic device, including:
[0017] At least one processor; and
[0018] A memory communicatively connected to the at least one processor; wherein,
[0019] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method of the above embodiments.
[0020] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method according to the above embodiments.
[0021] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0023] Figure 1 is a schematic flowchart of a method for training a model provided by an embodiment of the present disclosure;
[0024] Figure 2 is a schematic flowchart of another method for training a model provided by an embodiment of the present disclosure;
[0025] Figure 3 is a schematic flowchart of another method for training a model provided by an embodiment of the present disclosure;
[0026] Figure 4 is a schematic structural diagram of another model training device provided by an embodiment of the present disclosure;
[0027] Figure 5 is a block diagram of an electronic device for implementing the training of the model according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0028] Exemplary embodiments of the present disclosure will be described below with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted in the following description for clarity and conciseness.
[0029] Artificial intelligence is a discipline that studies the use of computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.). It has technical fields at both the hardware level and the software level. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, as well as deep learning, big data processing technology, and knowledge graph technology.
[0030] When the teacher model and the student model are heterogeneous, the detection results, the structure and physical meaning of the feature maps of the teacher model and the student model are different. Based on the differences between the detection results with different meanings, the determined loss value is inaccurate, and thus the student model is corrected based on the inaccurate loss value, resulting in a relatively low accuracy of the student model.
[0031] In the present disclosure, the initial student model is used to detect a sample image to output a first target detection box, and the teacher model is used to detect the same position of the first target detection box in the sample image to output a second target detection box, so as to ensure the consistency of the candidate boxes in the teacher model and the student model. Thus, the detection results of the student model and the teacher model correspond to each other, which is beneficial to improving the accuracy of the student model.
[0032] Next, with reference to the accompanying drawings, the training method, device, electronic device, and storage medium of the model according to the embodiments of the present disclosure will be described in detail.
[0033] It should be noted that the training method of the model implemented in the present disclosure is configured in a model training device (hereinafter simply referred to as the training device) for illustration. The training device can be applied to any electronic device so that the electronic device can perform the function of model training.
[0034] Among them, the electronic device can be any device with computing capabilities. For example, it can be a personal computer (PC for short), a mobile terminal, etc. The mobile terminal can be a hardware device such as a tablet computer, a personal digital assistant, a wearable device, etc. with various operating systems, touch screens, and / or display screens.
[0035] Figure 1Schematic flowchart of a method for training a model provided by an embodiment of the present disclosure.
[0036] As Figure 1 shown, the method includes:
[0037] Step 101, using an initial student model to perform object detection on a sample image, and obtaining a first object detection box and / or a first type corresponding to an object of interest in the sample image.
[0038] Among them, the student model is a model for object detection. The network structure complexity and scale of the teacher model for object detection are greater than those of the student model for object detection, and the network structure of the student model is simpler.
[0039] In the present disclosure, the sample image can be input into the initial student model, and the first object detection box and / or the first type of the object of interest in the sample image output by the student model can be obtained. Among them, the first object detection box includes the position information of the object of interest in the sample image. The object of interest can be one or more. The type of the object of interest can be one or more, such as people, vehicles, animals, etc.
[0040] Optionally, the initial student model can be a one-stage detection model or a two-stage detection model, and the present disclosure does not limit this.
[0041] Optionally, in the case where the initial student model is a two-stage detection model, the object detection box (proposal) output by the region proposal network (RPN) in the initial student model can be determined as the first object detection box.
[0042] Step 102, using the encoding module in the teacher model to encode the sample image, and obtaining a feature vector group corresponding to the sample image.
[0043] In the present disclosure, the sample image can be input into the teacher model, and the teacher model can perform feature extraction on the sample image to determine the feature map corresponding to the sample image. Then, the teacher model encodes the feature map to obtain the feature vector group corresponding to the sample image. In addition, during the process of encoding the feature map by the teacher model, the position information corresponding to each feature vector in the sample image can be generated.
[0044] Step 103, using the decoding module of the teacher model to decode the feature vector corresponding to the first object detection box in the feature vector group, so as to obtain a second object detection box and / or a second type.
[0045] In the present disclosure, the position information corresponding to each feature vector can be compared with the first target detection box to obtain the feature vector corresponding to the first target detection box in the feature vector group. The decoding module of the teacher model specifically decodes the feature vector corresponding to the first target detection box and performs classification regression processing, so as to determine the second target detection box and / or the second type corresponding to the feature vector.
[0046] Optionally, the object query of the teacher model can be set to the first target detection box. Thus, the teacher model can obtain the corresponding feature vector according to the first target detection box for decoding and classification regression, so as to obtain the second target detection box and / or the second type.
[0047] Step 104: Determine the loss value according to the first difference between the first target detection box and the second target detection box and / or the second difference between the first type and the second type.
[0048] In the present disclosure, based on a preset algorithm, such as Intersection over Union (IoU), etc., the first difference between the first target detection box and the second target detection box can be determined. Based on a preset difference algorithm, such as Kullback-Leibler divergence (KL divergence), Mean Squared Error (MSE), etc., the second difference between the first type and the second type can be determined. Then, based on the first difference and / or the second difference, the loss value is determined.
[0049] Alternatively, the first difference and the second difference can also be weighted, and the weighted sum of the first difference and the second difference is determined as the loss value.
[0050] Step 105: Modify the initial student model based on the loss value until the student model is obtained.
[0051] In the present disclosure, when the loss value is greater than a preset threshold, the parameters of the initial student model can be adjusted based on the loss value, and the adjusted initial student model is continuously trained using the sample images until the number of sample images for training the initial student model reaches a preset number and the training stops, and the student model for target detection is obtained.
[0052] In the present disclosure, after using the initial student model to perform object detection on a sample image to obtain the first target detection box and / or the first type corresponding to the target object in the sample image, the encoding module in the teacher model is used to encode the sample image to obtain a feature vector group corresponding to the sample image, and the decoding module of the teacher model is used to decode the feature vector corresponding to the first target detection box in the feature vector group to obtain the second target detection box and / or the second type. Then, according to the first difference between the first target detection box and the second target detection box and / or the second difference between the first type and the second type, a loss value is determined, and the initial student model is corrected based on the loss value until the student model is obtained. Thus, the initial student model is used to detect the sample image to output the first target detection box, and the teacher model is used to detect the same position of the first target detection box in the sample image to output the second target detection box, so as to ensure the consistency of the candidate boxes in the teacher model and the student model. Thereby, the detection results of the student model and the teacher model correspond to each other, which is beneficial to improving the accuracy of the student model.
[0053] Figure 2 It is a schematic flowchart of a method for training a model provided by an embodiment of the present disclosure.
[0054] As Figure 2 shown, the method includes:
[0055] Step 201, using the initial student model to perform object detection on a sample image to obtain the first target detection box and / or the first type corresponding to the target object in the sample image.
[0056] Step 202, using the encoding module in the teacher model to encode the sample image to obtain a feature vector group corresponding to the sample image.
[0057] Step 203, using the decoding module of the teacher model to decode the feature vector corresponding to the first target detection box in the feature vector group to obtain the second target detection box and / or the second type.
[0058] In the present application, for the specific implementation processes of steps 201 - 203, reference can be made to the detailed description of any embodiment of the present disclosure, which will not be elaborated herein.
[0059] Step 204, in the case where there is an overlapping part between the border of the second target detection box and the border of the first target detection box, using a preset weight value to weight the first difference between the first target detection box and the second target detection box and the second difference between the first type and the second type, and determining the weighted sum of the first difference and the second difference as the loss value.
[0060] In the present disclosure, the first target detection box and the second target detection box may be different. When there is an overlapping part between the border of the second target detection box and the border of the first target detection box, it indicates that the first target detection box only contains a part of the target object, and the first target detection box is inaccurate. At this time, a relatively large weight value can be assigned to the first one, and a relatively small weight value can be assigned to the second one, so as to increase the influence of the first difference on the training of the student model, thereby improving the accuracy of the student model obtained by training.
[0061] Step 205: Modify the initial student model based on the loss value until the student model is obtained.
[0062] In the present application, for the specific implementation process of step 205, reference can be made to the detailed description of any embodiment of the present disclosure, which will not be elaborated herein.
[0063] In the present disclosure, the initial student model is used to detect the sample image to output the first target detection box, and the teacher model is used to detect the same position of the first target detection box in the sample image to output the second target detection box, so as to ensure the consistency of the candidate boxes in the teacher model and the student model. Thus, the detection results of the student model and the teacher model correspond to each other, which is beneficial to improving the accuracy of the student model.
[0064] Figure 3 It is a schematic flowchart of a method for training a model provided by an embodiment of the present disclosure.
[0065] As Figure 3 shown, the method includes:
[0066] Step 301: Use the initial student model to perform target detection on the sample image to obtain the first target detection box and / or the first type corresponding to the target object in the sample image.
[0067] In the present application, for the specific implementation process of step 301, reference can be made to the detailed description of any embodiment of the present disclosure, which will not be elaborated herein.
[0068] Step 302: Screen the multiple first target detection boxes when there are multiple first target detection boxes.
[0069] In the present disclosure, for the target object at the same position in the sample image, the initial student model may output multiple first target detection boxes. To improve the training efficiency, a preset algorithm, such as the Non-maximum suppression (NMS) algorithm, can be used to screen the multiple first target detection boxes, and only the screened first target detection boxes are further processed. Improve the training efficiency.
[0070] Optionally, the initial student model may output the probability that each first target detection box contains a target object (i.e., foreground probability). If the probability corresponding to a certain first target detection box is less than the first threshold, then delete this first target detection box. While reducing redundant first target detection boxes, the first target detection boxes corresponding to the target objects in the sample image are retained. This is conducive to improving the training efficiency of the student model.
[0071] Optionally, it is also possible to calculate the distance between the center position of each first target detection box and the center positions of other first target detection boxes. If a certain distance is less than the threshold, then determine the two first target detection boxes corresponding to this distance as associated target detection boxes. Or, if a certain distance is less than the threshold and the area difference between the two first target detection boxes corresponding to this distance is less than the threshold, then determine the two first target detection boxes corresponding to this distance as associated target detection boxes. Thus, it is possible to determine the third target detection box associated with each first target detection box among multiple first target detection boxes.
[0072] After that, calculate the overlapping area between each first target detection box and the associated third target detection box, and calculate the ratio of the area of each first target detection box to the corresponding overlapping area (i.e., overlapping rate). If the overlapping rate corresponding to a certain first target detection box is greater than the second threshold, it indicates that this first target detection box and the associated third target detection box are similar. At this time, this redundant first target detection box can be deleted. Or, after determining the overlapping rate of the area of each first target detection box and the corresponding overlapping area, if the overlapping rate corresponding to a certain first target detection box is greater than the second threshold and the maximum overlapping rate of the third target detection box associated with this first target detection box is greater than the second threshold, then delete this first target detection box and the target detection box with the largest overlapping rate among the third target detection boxes associated with this first target detection box. Thus, redundant first target detection boxes are deleted, which is conducive to improving the training efficiency of the student model.
[0073] Optionally, if the overlapping rate corresponding to a certain first target detection box is greater than the second threshold and the overlapping rate corresponding to the third target detection box overlapping with this first target detection box is also greater than the second threshold, then delete the target detection box with the largest overlapping rate among this first target detection box and the third target detection box. This is conducive to improving the training efficiency of the student model.
[0074] Optionally, when deleting each first target detection box, it is possible to recalculate the overlapping rate corresponding to each remaining first target detection box. After that, based on the new overlapping rate, screen the remaining first target detection boxes until first target detection boxes with an overlapping rate less than the second threshold are obtained.
[0075] Optionally, use a preset algorithm (e.g., IoU) to determine the difference between each first target detection box and the labeled detection box corresponding to the sample image, and delete the first target detection boxes with corresponding differences greater than the third threshold. This can delete redundant first target detection boxes and is beneficial to improving the training efficiency of the student model.
[0076] Step 303: Use the encoding module in the teacher model to encode the sample image to obtain a feature vector group corresponding to the sample image.
[0077] Step 304: Use the decoding module in the teacher model to decode the feature vector corresponding to the first target detection box in the feature vector group to obtain a second target detection box and / or a second type.
[0078] Step 305: Determine a loss value according to the first difference between the first target detection box and the second target detection box, and / or the second difference between the first type and the second type.
[0079] Step 306: Modify the initial student model based on the loss value until the student model is obtained.
[0080] In this application, for the specific implementation process of steps 303 - 306, reference can be made to the detailed description of any embodiment of the present disclosure, which will not be elaborated here.
[0081] In the present disclosure, after using the initial student model to perform object detection on the sample image to obtain the first target detection box and / or the first type corresponding to the target object in the sample image, in the case where there are multiple first target detection boxes, screen the multiple first target detection boxes, and then use the encoding module in the teacher model to encode the sample image to obtain a feature vector group corresponding to the sample image, and use the decoding module in the teacher model to decode the feature vector corresponding to the first target detection box in the feature vector group to obtain a second target detection box and / or a second type. Then, determine a loss value according to the first difference between the first target detection box and the second target detection box, and / or the second difference between the first type and the second type, and modify the initial student model based on the loss value until the student model is obtained. This can improve the training efficiency and is beneficial to improving the accuracy of the student model.
[0082] To implement the above embodiments, the embodiments of the present disclosure also propose a training device for a model.
[0083] Figure 4 It is a schematic structural diagram of a training device for a model provided by an embodiment of the present disclosure.
[0084] As Figure 4 shown, the training device 400 for the model includes: a first detection module 410, a second detection module 420, a determination module 430, and a modification module 440.
[0085] The first detection module 410 is configured to perform object detection on a sample image by using an initial student model, and obtain a first object detection box and / or a first type corresponding to an object of interest in the sample image.
[0086] The second detection module 420 is configured to encode the sample image by using an encoding module in a teacher model, and obtain a feature vector group corresponding to the sample image.
[0087] The above-mentioned second detection module 420 is configured to decode the feature vector corresponding to the first object detection box in the feature vector group by using a decoding module of the teacher model, so as to obtain a second object detection box and / or a second type.
[0088] The determination module 430 is configured to determine a loss value according to a first difference between the first object detection box and the second object detection box, and / or a second difference between the first type and the second type.
[0089] The correction module 440 is configured to correct the initial student model based on the loss value until a student model is obtained.
[0090] In a possible implementation manner of the embodiment of the present disclosure, the above-mentioned determination module 430 is configured to:
[0091] When there is an overlapping part between the border of the second object detection box and the border of the first object detection box, use a preset weight value to weight the first difference and the second difference, and determine the weighted sum of the first difference and the second difference as the loss value.
[0092] In a possible implementation manner of the embodiment of the present disclosure, further included is:
[0093] A screening module is configured to screen multiple first object detection boxes when there are multiple first object detection boxes.
[0094] In a possible implementation manner of the embodiment of the present disclosure, the above-mentioned screening module is configured to:
[0095] Obtain the probability that each first object detection box contains an object of interest.
[0096] When the probability corresponding to any one of the first object detection boxes is less than a first threshold, delete any one of the first object detection boxes.
[0097] In a possible implementation manner of the embodiment of the present disclosure, the above-mentioned screening module is configured to:
[0098] Determine a third object detection box associated with each first object detection box among the multiple first object detection boxes according to the distance between the center position of each first object detection box and the center positions of other first object detection boxes.
[0099] Determine the overlapping area between each first target detection box and the associated third target detection box;
[0100] Determine the overlapping rate between the area of each first target detection box and the corresponding overlapping area;
[0101] Filter the multiple first target detection boxes according to the overlapping rate corresponding to each first target detection box.
[0102] In a possible implementation manner of this embodiment of the present disclosure, the above filtering module is further configured to:
[0103] In the case where the overlapping rate corresponding to any first target detection box is greater than the second threshold, delete any first target detection box; or,
[0104] In the case where the overlapping rate corresponding to any first target detection box is greater than the second threshold and the maximum overlapping rate of the third target detection box associated with any first target detection box is greater than the second threshold, delete any first target detection box and the target detection box with the largest overlapping rate among the third target detection boxes associated with any first target detection box.
[0105] In a possible implementation manner of this embodiment of the present disclosure, the above filtering module is configured to:
[0106] Determine the difference between each first target detection box and the labeled detection box corresponding to the sample image;
[0107] Delete the first target detection box corresponding to the difference greater than the third threshold.
[0108] It should be noted that the explanatory description of the foregoing method embodiment for training the model is also applicable to the device in this embodiment, so it will not be repeated here.
[0109] In the present disclosure, after using an initial student model to perform object detection on a sample image to obtain a first target detection box and / or a first type corresponding to a target object in the sample image, the encoding module in the teacher model is then used to encode the sample image to obtain a feature vector group corresponding to the sample image, and the decoding module of the teacher model is used to decode the feature vector corresponding to the first target detection box in the feature vector group to obtain a second target detection box and / or a second type. After that, based on a first difference between the first target detection box and the second target detection box, and / or a second difference between the first type and the second type, a loss value is determined, and the initial student model is corrected based on the loss value until a student model is obtained. Thus, the initial student model is used to detect the sample image to output a first target detection box, and the teacher model is used to detect the same position of the first target detection box in the sample image to output a second target detection box, so as to ensure the consistency of the candidate boxes in the teacher model and the student model. Thereby, the detection results of the student model and the teacher model correspond to each other, which is beneficial to improving the accuracy of the student model.
[0110] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device and a readable storage medium.
[0111] Figure 5 A schematic block diagram of an exemplary electronic device 500 that can be used to implement the embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0112] As Figure 5 shown, the device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 502 or a computer program loaded from a storage unit 508 into a RAM (Random Access Memory) 503. In the RAM 503, various programs and data required for the operation of the device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An I / O (Input / Output) interface 505 is also connected to the bus 504.
[0113] Multiple components in device 500 are connected to I / O interface 505, including: input unit 506, such as a keyboard, a mouse, etc.; output unit 507, such as various types of displays, speakers, etc.; storage unit 508, such as a disk, an optical disc, etc.; and communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. Communication unit 509 allows device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0114] Computing unit 501 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of computing unit 501 include but are not limited to CPU (Central Processing Unit), GPU (Graphic Processing Units), various dedicated AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc. Computing unit 501 executes the various methods and processes described above, such as the training method of the model. For example, in some embodiments, the training method of the model can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by computing unit 501, one or more steps of the training method of the model described above can be executed. Alternatively, in other embodiments, computing unit 501 can be configured to execute the training method of the model by any other suitable means (e.g., by means of firmware).
[0115] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application Specific Standard Products), SOCs (System On Chip), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that can receive data and instructions from, and transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0116] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.
[0117] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a RAM, a ROM, an EPROM (Electrically Programmable Read-Only Memory), or a flash memory, an optical fiber, a CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0118] For providing interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (Cathode-Ray Tube) or an LCD (Liquid Crystal Display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0119] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by any form or medium of digital data communication (e.g., a communication network). Examples of the communication network include: a LAN (Local Area Network), a WAN (Wide Area Network), the Internet, and a blockchain network.
[0120] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services (Virtual Private Server). The server may also be a server of a distributed system or a server combined with a blockchain.
[0121] It should be understood that various forms of processes shown above can be used, with steps reordered, added or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.
[0122] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A training method for a model, comprising: Performing object detection on a sample image using an initial student model to obtain a first target detection box corresponding to a target object in the sample image; Encoding the sample image using an encoding module in a teacher model to obtain a feature vector group corresponding to the sample image; Decoding the feature vector corresponding to the first target detection box in the feature vector group using a decoding module of the teacher model to obtain a second target detection box; Determining a loss value according to a first difference between the first target detection box and the second target detection box; Correcting the initial student model based on the loss value until a student model is obtained; The method further comprises: In the case where there are multiple first target detection boxes, determining a third target detection box associated with each first target detection box among the multiple first target detection boxes according to the distance between the center position of each first target detection box and the center positions of other first target detection boxes; Determining the overlapping area between each first target detection box and the associated third target detection box; Determining the overlapping rate of the area of each first target detection box and the corresponding overlapping area; Screening the multiple first target detection boxes according to the overlapping rate corresponding to each first target detection box.
2. The method according to claim 1, wherein The method further comprises: Performing object detection on a sample image using an initial student model to obtain a first type corresponding to a target object in the sample image; Decoding the feature vector corresponding to the first target detection box in the feature vector group using a decoding module of the teacher model to obtain a second type; The determining the loss value includes: Determining a loss value according to a first difference between the first target detection box and the second target detection box, and a second difference between the first type and the second type.
3. The method according to claim 2, wherein The determining the loss value according to a first difference between the first target detection box and the second target detection box, and a second difference between the first type and the second type includes: In the case where there is an overlapping part between the border of the second target detection box and the border of the first target detection box, weighting the first difference and the second difference using a preset weight value, and determining the weighted sum of the first difference and the second difference as the loss value.
4. The method according to claim 1, wherein, The screening the multiple first target detection boxes includes: Obtaining the probability that each first target detection box contains a target object; Deleting any first target detection box in the case where the probability corresponding to any first target detection box is less than a first threshold.
5. The method according to claim 1, wherein The screening the multiple first target detection boxes according to the overlapping rate corresponding to each first target detection box includes: Deleting any first target detection box in the case where the overlapping rate corresponding to any first target detection box is greater than a second threshold; or, In the case where the overlap rate corresponding to any one of the first target detection frames is greater than a second threshold, and the maximum overlap rate of the third target detection frame associated with any one of the first target detection frames is greater than the second threshold, delete the target detection frame with the largest overlap rate among any one of the first target detection frames and the third target detection frame associated with any one of the first target detection frames.
6. The method according to claim 1, wherein, The screening of the multiple first target detection frames includes: Determine the difference between each first target detection frame and the labeled detection frame corresponding to the sample image; Delete the first target detection frame corresponding to the difference greater than a third threshold.
7. A training device for a model, including: A first detection module, configured to perform target detection on a sample image by using an initial student model, and obtain a first target detection frame corresponding to a target object in the sample image; A second detection module, configured to encode the sample image by using an encoding module in a teacher model, and obtain a feature vector group corresponding to the sample image; The second detection module is configured to decode the feature vector corresponding to the first target detection frame in the feature vector group by using a decoding module of the teacher model, so as to obtain a second target detection frame; A determination module, configured to determine a loss value according to a first difference between the first target detection frame and the second target detection frame; A correction module, configured to correct the initial student model based on the loss value until a student model is obtained; The device further includes: A screening module, configured to, when there are multiple first target detection frames, determine a third target detection frame associated with each first target detection frame among the multiple first target detection frames according to the distance between the center position of each first target detection frame and the center positions of other first target detection frames; Determine the overlapping area between each first target detection frame and the associated third target detection frame; Determine the overlap rate of the area of each first target detection frame and the corresponding overlapping area; Screen the multiple first target detection frames according to the overlap rate corresponding to each first target detection frame.
8. The device according to claim 7, characterized in that, The first detection module is further configured to perform target detection on a sample image by using an initial student model, and obtain a first type corresponding to the target object; The second detection module is further configured to decode the feature vector corresponding to the first target detection frame in the feature vector group by using a decoding module of the teacher model, so as to obtain a second type; The determination module is further configured to determine a loss value according to a first difference between the first target detection frame and the second target detection frame, and a second difference between the first type and the second type.
9. The device according to claim 8, wherein, The determination module is configured to: In the case where there is an overlapping part between the border of the second target detection frame and the border of the first target detection frame, use a preset weight value to weight the first difference and the second difference, and determine the weighted sum of the first difference and the second difference as the loss value.
10. The device according to claim 7, wherein, The screening module is configured to: Obtain the probability that each first target detection frame contains a target object; In the case where the probability corresponding to any one of the first target detection frames is less than the first threshold, delete the any one of the first target detection frames.
11. The device according to claim 7, wherein, The screening module is further configured to: In the case where the overlap rate corresponding to any one of the first target detection frames is greater than the second threshold, delete the any one of the first target detection frames; Or, In the case where the overlap rate corresponding to any one of the first target detection frames is greater than the second threshold, and the maximum overlap rate of the third target detection frame associated with the any one of the first target detection frames is greater than the second threshold, delete the target detection frame with the largest overlap rate among the any one of the first target detection frames and the third target detection frame associated with the any one of the first target detection frames.
12. The device according to claim 7, wherein The screening module is configured to: Determine the difference between each of the first target detection frames and the labeled detection frame corresponding to the sample image; Delete the first target detection frame corresponding to the difference greater than the third threshold.
13. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-6.
14. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-6.
15. A computer program product, comprising a computer program, where the computer program implements the method according to any one of claims 1-6 when executed by a processor.
Citation Information
Patent Citations
Target detection result optimization method and device
CN108960174A
Model distillation method, target detection method and related equipment
CN115273136A