Object Identification Device and Object Identification Method

The artificial intelligence object recognition model is trained through the generative adversarial network to generate noise samples, which solves the problem of time-consuming and labor-consuming artificial marking, realizes efficient automatic marking and model simplification, and reduces the demand for artificial resources.

CN114708645BActive Publication Date: 2025-07-08WISTRON CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202110087912.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-12-17
Filing Date
2021-01-22
Publication Date
2025-07-08
Estimated Expiration
2041-01-25

AI Technical Summary

Technical Problem

In the prior art, a large amount of manual labeling data is required before training of an artificial intelligence object recognition model, resulting in problems that consume human resources and time.

Method used

The object identification device and method are used to generate tracking samples and adversarial samples, and noise samples are generated using a generative adversarial network, the teacher model is trained and the student model is initialized. The student model adjusts parameters according to the teacher model until the difference in the output result is less than the learning threshold value, reducing the number of manually marked samples.

Benefits of technology

It realizes the simplicity and robustness of the object recognition model, reduces the time and resource requirements of manual tagging, and can automatically and efficiently mark a large number of training images, solving the problem of time-consuming manual tagging.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114708645B_ABST
    Figure CN114708645B_ABST
Patent Text Reader

Abstract

The present disclosure provides an object recognition device and an object recognition method. The object recognition method includes that a student model adjusts a plurality of parameters according to a teacher model. When the vector difference between the output result of the adjusted student model and the output result of the teacher model is less than a learning threshold value, it is regarded that the student model has completed training, and the student model is extracted as an object recognition model. The space required by the student model is smaller than that of the teacher model. Thereby, a large number of training pictures and labels can be efficiently generated, achieving the effect of not requiring a large amount of manual labeling time for manual labeling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to an identification device and an identification method, and more particularly to an object identification device and an object identification method suitable for determining specific objects in an image. Background Art

[0002] Currently, the labeling work of artificial intelligence (AI) models is mostly contracted independently by specialized enterprises and carried out manually. Especially in countries such as China, India, and Southeast Asia, there are more and more companies that specifically commission manual labeling. Before training all AI object identification models on the market, a large amount of data must be accumulated and a large amount of manual labeling is required. Therefore, it is very labor-intensive and requires a lot of time for manual labeling.

[0003] Therefore, how to use an automatic labeling generation tool to generate a large number of pictures and automatically label them has become one of the problems to be solved in this field. Summary of the Invention

[0004] An embodiment of the present disclosure provides an object identification device including a processor and a storage device. The processor is used to access a program stored in the storage device to implement a preprocessing module, a teacher model training module, and a student model training module. The preprocessing module is used to generate a tracking sample and an adversarial sample. The teacher model training module is used to generate a teacher model. The student model training module initializes a student model based on the teacher model. Among them, the student model adjusts a plurality of parameters according to the teacher model and the adversarial sample. In response to the vector difference between the output result of the adjusted student model and the output result of the teacher model being less than a learning threshold value, it is considered that the student model has completed training, and the student model is extracted as an object identification model.

[0005] An embodiment of the present disclosure provides an object identification method, including: generating a tracking sample and an adversarial sample; generating a teacher model based on the tracking sample; and initializing a student model based on the teacher model; among them, the student model adjusts a plurality of parameters according to the teacher model and the adversarial sample. In response to the vector difference between the output result of the adjusted student model and the output result of the teacher model being less than a learning threshold value, it is considered that the student model has completed training, and the student model is extracted as an object identification model.

[0006] As described above, in some embodiments, the object recognition device and the object recognition method make the number of convolutional layers and neurons of the student model, which is the object recognition model, less than those of the teacher model. Therefore, the object recognition model has model simplicity, and adversarial samples are used in the process of establishing the student model, which can make the object recognition model have model robustness. Furthermore, during the entire process of the student model, the required manually labeled samples are much less than the number of adversarial samples. Therefore, it has the dilution of the number of artificial samples, achieving a reduction in the time and resources for manual labeling. Thereby, the object recognition device and the object recognition method only need to input the video or multiple images of the target object in any environment, and can automatically track and label a large number of objects. This solves the most time-consuming labeling link in the field of artificial intelligence object recognition. Therefore, a large number of training pictures and labels can be efficiently generated, achieving the effect of not requiring a large amount of manual labeling time for manual labeling. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Figure 1 FIG. is a block diagram of an object recognition device according to an embodiment of the present invention.

[0008] Figure 2 FIG. is a schematic diagram of an object recognition method according to an embodiment of the present invention.

[0009] Figure 3 FIG. is a flowchart of an object recognition method according to an embodiment of the present invention.

[0010] Figure 4A FIG. is a schematic diagram of generating a teacher model and adversarial samples according to an embodiment of the present invention.

[0011] Figure 4B FIG. is a schematic diagram of generating an object recognition model according to an embodiment of the present invention.

[0012] DESCRIPTION OF THE REFERENCE NUMERALS:

[0013] 100: Object recognition device;

[0014] PR: Processor;

[0015] ST: Storage device;

[0016] 10: Preprocessing module;

[0017] 20: Teacher model training module;

[0018] 30: Student model training module;

[0019] PRD: Data preprocessing;

[0020] MT: Model training;

[0021] DB: Adversarial samples

[0022] DA: Tracking sample;

[0023] DC: Manually labeled sample;

[0024] ORI: Data collection;

[0025] STM: Student model;

[0026] TTM: Teacher model;

[0027] STV: Student model verification;

[0028] 200, 300: Object identification method;

[0029] 310 - 340, 410 - 474, 510 - 590: Steps. Detailed implementation manner

[0030] The following description is the preferred implementation manner for implementing the invention, which aims to describe the basic spirit of the present invention but is not used to limit the present invention. The actual content of the invention must refer to the scope of the patent application hereafter.

[0031] It must be understood that the words "comprising", "including", etc. used in this specification are used to indicate the existence of specific technical features, numerical values, method steps, operation processes, elements, and / or components, but do not exclude the addition of more technical features, numerical values, method steps, operation processes, elements, components, or any combination of the above.

[0032] In the patent application, words such as "first", "second", "third", etc. are used to modify the elements in the patent application, and are not used to indicate the priority order, precedence relationship, or that one element precedes another element, or the chronological order when performing method steps, but are only used to distinguish elements with the same name.

[0033] Please refer to Figure 1 , Figure 1 is a block diagram showing an object identification device according to an embodiment of the present invention. The object identification device 100 includes a processor PR and a storage device ST. In one embodiment, the processor PR accesses and executes the program stored in the storage device ST to implement a preprocessing module 10, a teacher model training module 20, and a student model training module 30. In one embodiment, the preprocessing module 10, the teacher model training module 20, and the student model training module 30 can be implemented by software or firmware individually or together.

[0034] In one embodiment, the storage device ST can be implemented as a read-only memory, a flash memory, a floppy disk, a hard disk, an optical disc, a USB flash drive, a magnetic tape, a database accessible via a network, or a storage medium having the same function that can be easily conceived by those skilled in the art.

[0035] In one embodiment, the preprocessing module 10, the teacher model training module 20, and the student model training module 30 can be implemented by one or more processors individually or together. The processor can be implemented by an integrated circuit such as a microcontroller, a microprocessor, a Digital Signal Processor (DSP), a Field Programmable Gate Array (FPGA), an Application Specific Integrated Circuit (ASIC), or a logic circuit. In one embodiment, the preprocessing module 10, the teacher model training module 20, and the student model training module 30 can be implemented by a hardware circuit individually or together. For example, the preprocessing module 10, the teacher model training module 20, and the student model training module 30 can be composed of active devices (such as switches, transistors) and passive devices (such as resistors, capacitors, inductors). In one embodiment, the processor PR is used to access the respective operation results of the preprocessing module 10, the teacher model training module 20, and the student model training module 30 in the storage device ST.

[0036] Please refer to Figure 2 and Figure 3 , Figure 2 which is a schematic diagram showing an object recognition method according to an embodiment of the present invention. Figure 3 which is a flowchart showing an object recognition method according to an embodiment of the present invention. The object recognition method 200 and the object recognition method 300 can be implemented by Figure 1 devices. It can be seen from Figure 2 that the object recognition method 200 can be divided into processes such as data collection ORI, data preprocessing PRD, model training MT, and student model verification STV. The following will be described with the steps in Figure 3 in cooperation with Figure 2 for illustration.

[0037] In one embodiment, the processor PR is used to access the preprocessing module 10, the teacher model training module 20, and the student model training module 30, or to access and execute the programs / algorithms in the storage device ST, so as to implement the preprocessing module 10, the teacher model training module 20, and the student model training module 30. In one embodiment, when the preprocessing module 10, the teacher model training module 20, and the student model training module 30 are implemented by hardware (such as a chip or a circuit), they can perform self-computation after receiving data or signals, and send the processing results back to the processor PR. In one embodiment, when the preprocessing module 10, the teacher model training module 20, and the student model training module 30 are implemented by software (such as an algorithm), the processor PR executes the algorithms in the preprocessing module 10, the teacher model training module 20, and the student model training module 30 to obtain the processing results.

[0038] In step 310, the preprocessing module 10 is used to generate a tracking sample and an adversarial sample.

[0039] In one embodiment, after the preprocessing module 10 receives the data collection ORI (the data collection ORI can be obtained, for example, by shooting through a lens, or accessing multiple images or a video in a database), the user first selects the objects (such as a person) in several frames of the video or several images through an object recognition input interface (such as a touch screen, a keyboard, a mouse, etc.). Then, the preprocessing module 10 uses an optical flow algorithm to track an object selected in each frame to generate a tracking sample DA. Among them, the optical flow algorithm is about tracking objects in the visual field, describing the movement of the observed object, surface, or edge caused by the movement relative to the observer. After the user selects several frames or several images to define the selected object, the optical flow algorithm can track the selected object in each frame or other images to generate a tracking sample DA (that is, after tracking the images with the selected object in multiple images, object frame data corresponding to the selected object is generated. For example, a person is selected in the first image. Even if the person moves about 1 cm relative to the first image in the second image, the optical flow algorithm can still select this person in the second image, and so on. In this way, the movement trajectory of this person can be tracked by using multiple images to generate a tracking sample DA). In another embodiment, tracking samples DA can be generated for different types of objects respectively, and multiple tracking samples DA corresponding to multiple types of objects can be obtained.

[0040] In one embodiment, the tracking sample DA includes an Extensible Markup Language (XML) file regarding the selected position relative to the selected object in the original image. The content is the center point coordinates (x, y) of the selected position marked in the screen, the width W of the selected position, the height H of the selected position, and the object category (such as a person). The tracking sample DA is used as the data set for training the teacher model TTM.

[0041] In one embodiment, the preprocessing module 10 adds a noise to the tracking sample DA to generate an adversarial sample DB.

[0042] In one embodiment, the preprocessing module 10 inputs the tracking sample DA into a generative adversarial network (GAN) or an adversarial generative adversarial network (adv-GAN). The generative adversarial network or the adversarial generative adversarial network outputs the adversarial sample DB.

[0043] In one embodiment, the preprocessing module 10 selects the images corresponding to the selected positions in the tracking sample DA, inputs the images corresponding to the selected positions into the adv-GAN to increase the amount of images. The adv-GAN is used to add meaningful noise (i.e., effective information that misleads the teacher model TTM) to generate noise images, and then paste one or more noise images back to the tracking sample DA to generate multiple different adversarial samples DB. For example, after adding noise to the images corresponding to the selected positions and pasting them back to the original tracking sample DA to generate the adversarial sample DB, to the naked eye of the user, it still appears to be a human being selected, but the probability that the selected position in the adversarial sample DB is judged as a cat by the teacher model TTM is 90%, and the probability of being judged as a human is 15%. Then, in the subsequent training step, the parameters of the teacher model TTM are adjusted (such as strengthening the weight regarding human features) until a teacher model TTM is trained to identify the selected position in the adversarial sample DB as a human (for example, the probability of being judged as a cat by the teacher model TTM is 10%, and the probability of being judged as a human is 95%). In another embodiment, the noise images pasted back to the tracking sample DA include noise images of images of different object categories (such as bottles, boxes, signs, etc.). By pasting the noise images including different object categories back to the tracking sample DA to generate the adversarial sample DB, and then using the adversarial sample DB to train the teacher model TTM, the teacher model TTM will be able to identify different categories of objects that coexist in the image.

[0044] By generating an adversarial sample database DB, the number of samples for training the teacher model TTM can be increased, and by adding noise to the adversarial sample database DB and the corresponding correct answers, the teacher model TTM can automatically adjust its parameters to improve the accuracy of object judgment. In another embodiment, the adversarial sample database DB can be increased with object images of different categories. After adding noise to the object images of different categories and pasting them into the adversarial sample database DB, the trained teacher model TTM can recognize multiple objects.

[0045] In step 320, the teacher model training module 20 is used to initially train a preliminary version of a teacher model TTM with the tracking sample DA. At this time, the teacher model TTM's perception of data tends to the form of the tracking sample DA.

[0046] In one embodiment, the teacher model TTM is retrained with the data form of the adversarial sample database DB. At this time, the neural parameters in the teacher model TTM are updated towards the adversarial sample database DB, and the teacher model TTM can be applied to some object detection models, such as the YOLO series, which is a one-stage prediction form, to increase the model accuracy. The reason is that when training the teacher model TTM, the data form has been retrained in two dimensions. Therefore, this training method of domain adaptation reduces the possibility of gradient dispersion. Since the YOLO model requires a large amount of training data, by inputting the tracking sample DA and a large number of adversarial sample databases DB, the robustness of the neural network of the teacher model TTM can be increased. In one embodiment, YOLO is an object detection method that can determine the position and category of objects in a graph by performing a convolutional neural network (CNN) architecture on the picture only once, thus improving the recognition speed.

[0047] In step 330, the student model training module 30 initializes a student model STM according to the teacher model TTM.

[0048] In one embodiment, the student model STM uses the same model framework as the teacher model TTM, but the size of the student model STM can be designed to have a smaller storage capacity (such as fewer weight parameters). In one embodiment, the neural framework construction methods of the student model and the teacher model are similar, only reducing some of the framework construction layers. For example, if the teacher model TTM is built using the YOLO model, then when the student model STM is initialized, it also uses the architecture of the YOLO model. In subsequent steps, the student model STM will learn following the teacher model TTM. However, the input data for both the teacher TTM and the student model STM are adversarial samples DB. During the training process, the student model STM is trained to approximate the teacher model TTM to improve the accuracy of the student model STM in object recognition. Among them, the student model STM learning following the teacher model TTM means that the student model STM continuously adjusts multiple parameters, such as the bias and the weights corresponding to multiple input samples. The student model STM makes the multiple output results (such as a student tensor containing multiple probabilities) approach the output result of the teacher model TTM (such as a teacher tensor containing multiple probabilities).

[0049] In one embodiment, the number of parameters set in the student model STM is less than the number of parameters in the teacher model TTM. In one embodiment, the number of convolutional layers and the number of neurons in the deep learning model set in the student model STM are less than the number of convolutional layers and the number of neurons in the teacher model TTM. Therefore, the number of weight parameters corresponding to the number of convolutional layers and the number of neurons in the student model STM is also less than the number of weight parameters corresponding to the number of convolutional layers and the number of neurons in the teacher model TTM. Therefore, the storage space required by the student model STM is less than that of the teacher model TTM. Furthermore, since the number of convolutional layers and the number of neurons in the student model STM are also less than those in the teacher model TTM, the operation speed of the student model STM will be faster than that of the teacher model TTM.

[0050] In step 340, the student model STM adjusts a plurality of parameters according to the teacher model TTM and the adversarial samples DB (the adversarial samples DB are the data for training the student model STM and the teacher model TTM, and the student model STM learns from the teacher model TTM). In response to the vector difference between a student tensor of the adjusted student model STM and a teacher tensor of the teacher model TTM being less than a learning threshold value (such as 0.1), it is considered that the student model STM has completed training, and the student model STM is extracted as an object recognition model. In one embodiment, before extracting the student model STM as an object recognition model, it may include more detailed methods to improve the recognition accuracy of the student model STM, which will be described in detail later.

[0051] In one embodiment, by Figure 2It can be seen that during the operation of the student model STM, the manually labeled sample DC is not directly input into the student model STM. Therefore, the student model verification STV can be performed through the manually labeled sample DC. For example, the student model training module 30 inputs the manually labeled sample DC (such as a person being boxed) into the student model STM. If the correct output probability of the student model STM for the boxed position being a person is 99% and the probability of being a cat is 0.1%, it can be considered that the accuracy of the student model STM is sufficient to identify objects.

[0052] In one embodiment, the parameters adjusted by the student model STM based on the teacher model TTM can be the bias weight and multiple weights corresponding to multiple inputs. The student tensor refers to multiple probabilities output by the student model STM. For example, the probability that the boxed position in the input image is a person is 70%, the probability of being a cat is 10%, and the probability of being a dog is 10%. Similarly, the teacher tensor refers to multiple probabilities output by the teacher model TTM. For example, the probability that the boxed position in the input image is a person is 90%, the probability of being a cat is 5%, and the probability of being a dog is 5%. The probabilities here refer to the respective probabilities that the boxed positions are a person, a cat, and a dog, so they are all independent and unrelated probabilities.

[0053] In one embodiment, the vector difference can be a practical operation method of a loss function. The vector difference can be calculated using methods such as the Mean Square Error (MSE) and the Mean Absolute Error (MAE). The predicted value (usually denoted as y) in these methods is, for example, the student tensor, and the true value (usually denoted as ) is, for example, the teacher tensor, and the vector difference between the two is calculated. Since these methods are existing methods, they will not be elaborated here. In one embodiment, the range of the vector difference is between 0 and 1.

[0054] In one embodiment, when the vector difference between a student tensor of the adjusted student model STM and a teacher tensor of the teacher model TTM is less than the learning threshold value (such as 0.1), it is considered that the student model STM has completed training, and then the student model STM is extracted as an object recognition model. Since the student model STM has the characteristics of fast operation speed and small storage space, and the student model STM approaches the vector difference to the teacher model TTM, the recognition accuracy of the student model STM is also almost as high as that of the teacher model TTM trained with a large amount of data.

[0055] In one embodiment, in step 340, the preprocessing module 10 is further configured to receive the manually labeled sample DC. When the vector difference between the student tensor of the student model STM and the teacher tensor of the teacher model TTM is less than a learning threshold value (for example, 0.2, which is only an example here and the value can be adjusted according to the implementation), the manually labeled sample DC is input into the teacher model TTM to generate a further trained teacher model.

[0056] In some embodiments, when the vector difference between the student tensor of the student model STM and the teacher tensor of the teacher model TTM is less than the learning threshold value, it means that the execution results of the student model STM and the teacher model TTM are similar. Therefore, the teacher model TTM needs to be trained (this is called further training) with the manually labeled sample DC; and when the vector difference between the student tensor of the student model STM and the post-training tensor of the further trained teacher model is less than the learning threshold value, it means that the execution results of the student model STM and the further trained teacher model are similar, and at this time the student model STM is regarded as having been trained.

[0057] In one embodiment, the teacher model training module 20 inputs the manually labeled sample DC into the further trained teacher model. When the vector difference (or loss function) between the post-training tensor output by the further trained teacher model and the manually labeled sample is less than a further training threshold value, it is regarded that the further trained teacher model has completed training.

[0058] In one embodiment, the number of the manually labeled samples DC is less than the number of the adversarial samples DB.

[0059] By inputting the manually labeled sample DC into the teacher model TTM, the teacher model TTM can learn the dependence of the labeled object (such as a person) on the background (such as a street scene). When the vector difference (or loss function) between the post-training tensor output by the further trained teacher model and the manually labeled sample is less than a further training threshold value, it is regarded that the further trained teacher model has completed training. Then, the teacher model training module 20 uses the further trained teacher model to lead the student model TTM to approximate the further trained teacher model, and the student model TTM will learn from the further trained teacher model again.

[0060] In one embodiment, the student model training module 30 adjusts the parameters of the student model TTM according to the further trained teacher model (the student model TTM learns from the further trained teacher model), such as adjusting the bias value and / or adjusting the weights corresponding to multiple inputs. When the vector difference between the student tensor of the adjusted student model and a post-training teacher tensor of the further trained teacher model is less than the learning threshold value, it is regarded that the student model TTM has completed training, and the student model TTM is extracted as the object recognition model.

[0061] This can improve the recognition rate of the object (such as a person) in the image when the student model TTM analyzes the images of the actual environment.

[0062] In one embodiment, the stop condition for step 340 is that the iterative operation of step 340 reaches a specific number of times (for example, preset to 70 times), indicating that the student model STM has been adjusted 70 times before it is accurate enough. The student model training module 30 extracts the student model STM as an object recognition model.

[0063] Please refer to Figure 4A and Figure 4B , Figure 4A FIG. is a schematic diagram showing the generation of a teacher model TTM and adversarial samples DB according to an embodiment of the present invention. Figure 4B FIG. is a flowchart showing the generation of an object recognition model according to an embodiment of the present invention.

[0064] In step 410, pictures of the target object are recorded. For example, one or more pedestrians walking on the road are photographed by a camera.

[0065] In step 420, the range of the target object is selected by using a mouse, and the selected range of the target object is regarded as the selection position. However, it is not limited to using a mouse. If the object recognition device 100 includes a touch screen, the range of the selected target can be received by the touch screen. For example, the user uses a finger to select the target (such as a person) on the touch screen. At this time, the preprocessing module 10 can know the length and width of the selection position of the selected target in the entire frame and the center point coordinates of the selection position, thereby generating an artificial labeled sample DC. Herein, the selection position may refer to a selected range.

[0066] In one embodiment, in step 420, the user can select the selected target in multiple frames or images, so that the preprocessing module 10 can know the length and width of the selection position of the selected target in each frame or image and the center point coordinates of the selection position.

[0067] In one embodiment, the user can select multiple types of selected targets (target objects), such as selecting multiple people or cats.

[0068] In step 430, the preprocessing module 10 uses an optical flow algorithm in combination with a feature pyramid to perform optical flow tracking on the pixel area of the selection position.

[0069] In one embodiment, since the user has selected the selection position of at least one frame (such as a person), the preprocessing module 10 can use the optical flow algorithm to continue tracking the selection position in subsequent frames (for example, if a person walks to the right in the next frame, the preprocessing module 10 can use the optical flow algorithm to track the selection position of the person in this frame). The feature pyramid network is a feature extractor designed according to the concept of the feature pyramid, aiming to improve the accuracy and speed of finding the selection position.

[0070] In step 440, the preprocessing module 10 uses an image processing algorithm to find the target edge contour for the selected target in motion.

[0071] In step 450, the preprocessing module 10 optimizes the most suitable tracking selection position.

[0072] In one embodiment, due to the large displacement of the object movement, the range of the aforementioned selected position may be too large or there may be noise. Therefore, by using, for example, a binarization algorithm, an edge detection algorithm (Edge detection), etc., the continuous edges of the object are found to find the target edge contour (such as the contour of a person). For example, the processor PR uses the Open Source Computer Vision Library (OpenCV) to perform motion detection. Since the selected position has been processed by binarization, the preprocessing module 10 can calculate the minimized selected position (minimized rectangle). Thereby, the processor PR can converge the selected position to an appropriate size according to the target edge contour as the most suitable tracking selection position, and then perform the tracking selection position to improve the tracking accuracy.

[0073] In step 460, the preprocessing module 10 generates a tracking sample DA.

[0074] For example, the preprocessing module 10 then uses an optical flow algorithm to track a selected object in each frame to generate a large number of automatically generated tracking samples DA, and a large number of tracking samples DA can be automatically generated without manual selection.

[0075] In step 462, the preprocessing module 10 inputs the tracking sample DA into the initial teacher model to train the initial teacher model.

[0076] In one embodiment, the initial teacher model is just a framework (such as the framework of the YOLO model).

[0077] In step 464, the teacher model training module 20 generates a teacher model TTM. This teacher model TTM has learned the tracking sample DA.

[0078] In step 470, the preprocessing module 10 selects (crops) the image of the selected position. In other words, the preprocessing module 10 will select the selected position from the entire frame or image.

[0079] In step 472, the preprocessing module 10 generates a false sample of the selected position.

[0080] In one embodiment, noise can be added to the image at the original box selection position. Preferably, the adv-GAN algorithm can be used to add meaningful noise (i.e., effective information that misleads the teacher model TTM) to generate a noise map (i.e., a false sample at the box selection position).

[0081] In step 474, the preprocessing module 10 pastes one or more false samples at the box selection position back to the tracking sample DA to generate multiple different adversarial samples DB.

[0082] In this example, a large number of adversarial samples DB can be generated for training the teacher model TTM, enabling the teacher model TTM to adjust its parameters (let the teacher model TTM learn the adversarial samples DB) until the teacher model TTM can also correctly identify the box selection positions in the adversarial samples DB.

[0083] As can be seen from the above, through Figure 4A the above process, the tracking sample DA, the teacher model TTM, and the adversarial samples DB are generated. Next, please refer to Figure 4B . In one embodiment, Figure 4B each step in

[0084] can also be executed by the processor PR.

[0085] In step 510, the preprocessing module 10 reads the adversarial samples DB and inputs them into the student model and the teacher model.

[0086] In one embodiment, the adversarial samples DB account for approximately 70% of the overall sample volume, and the manually labeled samples DC account for approximately 30% of the overall sample volume.

[0087] In step 530, the preprocessing module 10 reads the teacher model TTM.

[0088] In step 540, the student model training module 30 constructs an initial student model. At this time, the initial student model adopts the same framework as the teacher model TTM.

[0089] In step 550, the student model training module 30 uses the adversarial samples DB to train the initial student model to generate a student model STM.

[0090] In one embodiment, the student model training module 30 initializes a student model STM based on the teacher model TTM. The student model STM takes the teacher model TTM as a standard and adjusts its parameters (the student model STM learns from the teacher model TTM) so that the output tensor of the student model STM is close to the output tensor of the teacher model TTM.

[0091] In one embodiment, the error between the current student model STM and the previous version of the student model STM is less than an error threshold value (e.g., 5%), indicating that the training of the current student model STM has tended to converge and enters step S560.

[0092] In step 560, the student model STM outputs a student tensor.

[0093] In step 570, the teacher model TTM outputs a teacher tensor.

[0094] In step 580, the processor PR determines whether the vector difference between the student tensor of the adjusted student model STM and the teacher tensor of the teacher model TTM is less than a learning threshold value. If so, it means that the loss function between the student model STM and the teacher model TTM is small and the gap is similar, and it enters step 590. If not, the training process A is performed to continue to let the student model STM learn from the teacher model TTM and continue to adjust the parameters of the student model STM.

[0095] In step 590, the processor PR extracts the latest trained student model STM.

[0096] After step 590 is executed, the training process B is performed. The teacher model TTM is trained by the manually labeled sample DC to improve the accuracy of the teacher model TTM, and then the student model STM continues to learn from the teacher model TTM and continues to adjust the parameters of the student model STM.

[0097] In step 572, the teacher model training module 20 inputs the manually labeled sample DC into the teacher model TTM to generate an advanced teacher model.

[0098] In step 574, the teacher model training module 20 determines whether the vector difference between an advanced tensor output by the advanced teacher model and the manually labeled sample DC is less than an advanced threshold value. If not, it returns to step 572, and continues to input the manually labeled sample DC or the newly added manually labeled sample DC into the advanced teacher model and further trains the advanced teacher model. If so, it means that the advanced teacher model has completed training. The teacher model TTM in step 530 is replaced with the advanced teacher model, so that the student model STM continues to learn from the advanced teacher model, and the student model STM adjusts its parameters to approximate the tensor of the advanced teacher model. When the vector difference between the student tensor of the adjusted student model STM and an advanced teacher tensor of the advanced teacher model is less than the learning threshold value, it is regarded that the student model has completed training, and the student model training module 30 extracts the student model as an object recognition model.

[0099] In one embodiment, when an unknown image is input into the object recognition model, the object recognition model can recognize or frame the position and / or quantity of specific objects in this unknown image. In another embodiment, the object recognition model can recognize or frame the position and / or quantity of objects of different categories in this unknown image.

[0100] As can be seen from the above, the object recognition device and the object recognition method make the number of convolutional layers and neurons of the student model as the object recognition model less than those of the teacher model. Therefore, the object recognition model has model simplicity. Furthermore, the object recognition device and the object recognition method in this case use adversarial samples in the process of establishing the student model, which can make the object recognition model have model robustness. During the entire process of the student model, the required manually labeled samples are much less than the number of adversarial samples. Therefore, it has the dilution of the number of artificial samples, achieving the reduction of the time and resources for manual labeling.

[0101] Thereby, the object recognition device and the object recognition method only need to input pictures or multiple images of the target object in any environment, and can automatically track and label a large number of objects, solving the most time-consuming labeling link in the field of artificial intelligence object recognition. Therefore, it can efficiently generate a large number of training pictures and labels, achieving the effect of not requiring a large amount of manual labeling time for manual labeling.

Claims

1. An object identification device, characterized in that, Comprising: a processor; and a storage device, the processor is used to access a program stored in the storage device to implement a preprocessing module, a teacher model training module, and a student model training module; wherein the preprocessing module is used to generate a tracking sample and an adversarial sample; it includes: tracking a selected object in each frame according to an optical flow algorithm to generate a tracking sample DA; inputting the tracking sample DA into a generative adversarial network or an adversarial sample generation method to generate an adversarial sample DB by the generative adversarial network or the adversarial sample generation method; the teacher model training module trains a teacher model with the tracking sample; and the student model training module initializes a student model according to the teacher model; wherein the student model adjusts a plurality of parameters of the student model according to the teacher model and the adversarial sample, and if the vector difference between the output result of the student model and the output result of the teacher model is less than a learning threshold value, it is considered that the student model has completed training, and the student model is extracted as an object recognition model; the preprocessing module is used to receive an artificially labeled sample. When the vector difference between the output result of the student model and the output result of the teacher model is less than the learning threshold value, the teacher model training module inputs the artificially labeled sample into the teacher model for training to generate a further teacher model; the teacher model training module inputs the artificially labeled sample into the further teacher model; the student model training module adjusts a plurality of parameters of the student model according to the further teacher model. If the vector difference between the output result of the student model and the output result of the further teacher model is less than the learning threshold value, it is considered that the student model has completed training, and the student model is extracted as the object recognition model.

2. The object identification device according to claim 1, characterized in that, The number of convolutional layers and neurons of the deep learning model set by the student model is less than the number of convolutional layers and neurons of the teacher model, and the number of weight parameters corresponding to the number of convolutional layers and neurons of the student model is also less than the number of weight parameters corresponding to the number of convolutional layers and neurons of the teacher model.

3. The object identification device according to claim 2, characterized in that, If the vector difference between a further tensor output by the further teacher model and the artificially labeled sample is less than a further threshold value, it is considered that the further teacher model has completed training.

4. The object identification device according to claim 3, characterized in that, The student model training module adjusts the plurality of parameters of the student model according to the further teacher model. If the vector difference between a student tensor of the student model and a teacher tensor of the further teacher model is less than the learning threshold value, it is considered that the student model has completed training, and the student model is extracted as the object recognition model.

5. The object recognition device according to claim 1, characterized in that, The preprocessing module tracks a selected object in each frame according to an optical flow algorithm to generate the tracking sample; wherein the preprocessing module adds noise to the tracking sample or inputs the tracking sample into a generative adversarial network to generate the adversarial sample.

6. The object identification device according to claim 5, characterized in that, The preprocessing module adds the tracking sample to a noise map, and the noise map includes images of different object categories.

7. The object recognition device according to claim 1, characterized in that, The student model adjusts the partial weights and a plurality of weights such that the output result output by the student model approaches the output result of the teacher model.

8. The object identification device according to claim 1, characterized in that The output result output by the student model is a student tensor, and the output result output by the teacher model is a teacher tensor.

9. An object identification method, characterized in that, It includes: Generating a tracking sample and an adversarial sample; It includes: tracking a selected object in each frame according to an optical flow algorithm to generate a tracking sample DA; inputting the tracking sample DA into a generative adversarial network or an adversarial sample generation method, and the generative adversarial network or the adversarial sample generation method outputs an adversarial sample DB; Training a teacher model according to the tracking sample; and Initializing a student model according to the teacher model; Wherein the student model adjusts a plurality of parameters of the student model according to the teacher model and the adversarial sample. In response to the vector difference between the output result of the student model and the output result of the teacher model being less than a learning threshold value, it is considered that the training of the student model is completed, and the student model is extracted as an object recognition model; The pre-processing module is used to receive an artificially marked sample. When the vector difference between the output result of the student model and the output result of the teacher model is less than the learning threshold value, the teacher model training module inputs the artificially marked sample into the teacher model for training to generate a further teacher model; The teacher model training module inputs the artificially marked sample into the further teacher model; The student model training module adjusts a plurality of parameters of the student model according to the further teacher model. In response to the vector difference between the output result of the student model and the output result of the further teacher model being less than the learning threshold value, it is considered that the training of the student model is completed, and the student model is extracted as the object recognition model.

10. The object identification method according to claim 9, wherein The number of convolutional layers and neurons of the deep learning model set by the student model is less than the number of convolutional layers and neurons of the teacher model, and the number of weight parameters corresponding to the number of convolutional layers and neurons of the student model is also less than the number of weight parameters corresponding to the number of convolutional layers and neurons of the teacher model.

11. The object recognition method according to claim 10, wherein, It further includes: In response to the vector difference between a further tensor output by the further teacher model and the artificially marked sample being less than a further threshold value, it is considered that the training of the further teacher model is completed.

12. The object identification method according to claim 11, wherein It further includes: Adjusting the plurality of parameters of the student model according to the further teacher model; In response to the vector difference between a student tensor of the student model and a teacher tensor of the further teacher model being less than the learning threshold value, it is considered that the training of the student model is completed, and the student model is extracted as the object recognition model.

13. The object identification method according to claim 9, wherein It further includes: Tracking a selected object in each frame according to an optical flow algorithm to generate the tracking sample; and Adding noise to the tracking sample or inputting the tracking sample into a generative adversarial network to generate the adversarial sample.

14. The object identification method according to claim 13, characterized in that, It further includes: Adding the tracking sample to a noise map, and the noise map includes images of different object categories.

15. The object identification method according to claim 9, wherein The student model adjusts the partial weights and a plurality of weights so that the output result output by the student model approaches the output result of the teacher model.

16. The object identification method according to claim 9, wherein The output result output by the student model is a student tensor, and the output result output by the teacher model is a teacher tensor.

Citation Information

Patent Citations

  • Image classification method and device, readable storage medium and terminal equipment

    CN110147456A

  • Image recognition model construction method and image recognition method and device

    CN110837846A

  • Falling detection method and electronic system using the same

    CN110895671A