Training methods for target detection models, classroom behavior detection methods, and related equipment
By improving the regression process and loss function of the CenterNet object detection model, the problems of insufficient target detection accuracy and long training in the existing technology are solved, and more efficient and accurate classroom behavior detection is achieved, and teaching quality is improved.
Patent Information
- Application Number
- CN202111325070.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-10
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2041-11-10
AI Technical Summary
The existing object detection algorithms have problems such as insufficient accuracy, long training time and difficulty in adapting to complex and changing teaching environments in classroom teaching scenarios.
The CenterNet-based object detection model is adopted to improve the training efficiency and accuracy of the object detection model by improving the regression process and loss function, including single-dimensional regression and multi-task loss function.
It effectively improves the training efficiency and detection accuracy of the object detection model, can more accurately identify classroom behavior, reflect students' classroom participation, and improve teaching quality.
Smart Images

Figure CN114005013B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a target detection model training method, a classroom behavior detection method and related equipment. Background Art
[0002] Behavior detection based on target detection is an important application of artificial intelligence technology. With the popularization of information-based education, using artificial intelligence technology to detect classroom behavior and analyze students' classroom participation is conducive to teaching evaluation and improving teaching quality.
[0003] Existing target detection algorithms are divided into two categories: those based on traditional machine learning and those based on deep learning. Taking the detection of hand-raising behavior in class as an example, the target detection algorithm based on traditional machine learning needs to use a designed feature template to extract the hand-raising feature. Due to the unlearnability of the feature template, it cannot adapt to the complex and changeable situations of classroom teaching.
[0004] The target detection algorithm based on deep learning includes the algorithm for detecting key points of the human body and the two-stage target detection algorithm. The algorithm for detecting key points of the human body first uses the key point detection algorithm to extract the human skeleton features, and then uses the classification model to determine whether it is a hand-raising behavior; because the key point detection algorithm is easily affected by occlusion, and occlusion is more frequent in classroom teaching scenes, the performance of the algorithm for detecting key points of the human body is seriously insufficient.
[0005] The two-stage target detection algorithm can solve the challenge of different hand-raising postures in classroom teaching scenarios, but the steps are cumbersome, the training is time-consuming, and the accuracy is insufficient, so it cannot be promoted and used on a large scale.
[0006] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present invention, and therefore may include information that does not constitute the prior art known to ordinary technicians in the field. Summary of the invention
[0007] In view of this, the present invention provides a target detection model training method, a classroom behavior detection method and related equipment, which are based on CenterNet and can effectively improve the training efficiency of the target detection model and improve the target detection accuracy by improving the regression process and loss function.
[0008] One aspect of the present invention provides a method for training a target detection model, comprising: constructing an initial network model based on a center network CenterNet; processing a sample image through the initial network model to obtain the category and center point position of each target object in the sample image, and performing a single-dimensional regression based on the center point position to obtain the detection box information of each target object; based on a sample image set including a plurality of the sample images, controlling the training of the initial network model with a multi-task loss function including classification loss, detection box offset loss and center point offset loss as a constraint condition to obtain a target detection model.
[0009] In some embodiments, performing a one-dimensional regression based on the center point position to obtain the detection frame information of each of the target objects includes: performing a one-dimensional regression of the size based on the center point position to obtain the width or height of the detection frame of the corresponding target object; obtaining the width and height of the detection frame of the corresponding target object based on a preset aspect ratio and the width or height of the detection frame of the corresponding target object, wherein the preset aspect ratio is determined based on the actual category of the corresponding target object.
[0010] In some embodiments, the loss function L(δ) of the center point offset loss is:
[0011]
[0012] Among them, δ is the offset of the center point position of the kth target object, is the width and height of the detection box of the kth target object, s k are the width and height of the real box of the kth target object, k ranges from 1 to N, and N is the number of target objects in the sample image;
[0013] The gradient function L′(δ) of the loss function L(δ) is:
[0014]
[0015] In some embodiments, the loss function L of the detection box offset loss is size for:
[0016]
[0017] The loss function L of the classification loss k for:
[0018]
[0019] in, is the detection value of the target object of category c at the coordinate point (x, y), Y xycis the true value of the target object of category c at the coordinate point (x, y), and α and β are hyperparameters;
[0020] The multi-task loss function L det for:
[0021] L det =L k +λ size L size +λ δ L(δ);
[0022] Among them, λ size and λ δ They are the loss functions L size and the weight of L(δ).
[0023] Another aspect of the present invention provides a training device for a target detection model, comprising: a model construction module, used to construct an initial network model based on a center network CenterNet; an image prediction module, used to process sample images through the initial network model, obtain the category and center point position of each target object in the sample image, and perform single-dimensional regression based on the center point position to obtain the detection box information of each target object; a loss constraint module, used to control the training of the initial network model based on a sample image set including multiple sample images, using a multi-task loss function including classification loss, detection box offset loss and center point offset loss as a constraint condition to obtain a target detection model.
[0024] Another aspect of the present invention provides a classroom behavior detection method, comprising: obtaining an image to be detected; inputting the image to be detected into the target detection model generated by training using the training method described in any of the above embodiments, and obtaining the target category and target detection frame information output by the target detection model; and obtaining a classroom behavior detection result based on the target category and the target detection frame information.
[0025] In some embodiments, the classroom behavior detection result is a retrieval result of hand-raising behavior.
[0026] Another aspect of the present invention provides a classroom behavior detection device, comprising: an image acquisition module, used to obtain an image to be detected; an image processing module, used to input the image to be detected into the target detection model generated by the training method described in any of the above embodiments, and obtain the target category and target detection frame information output by the target detection model; a behavior analysis module, used to obtain the classroom behavior detection result based on the target category and the target detection frame information.
[0027] Another aspect of the present invention provides an electronic device, comprising: a processor; a memory, wherein the memory stores executable instructions; wherein, when the executable instructions are executed by the processor, the training method of the target detection model as described in any of the above embodiments is implemented, and / or the classroom behavior detection method as described in any of the above embodiments is implemented.
[0028] Yet another aspect of the present invention provides a computer-readable storage medium for storing a program, which, when executed by a processor, implements the target detection model training method as described in any of the above embodiments, and / or implements the classroom behavior detection method as described in any of the above embodiments.
[0029] The beneficial effects of the present invention compared with the prior art include at least:
[0030] The present invention is based on the anchor-free target detection algorithm CenterNet, which can handle the occlusion phenomenon in the image well; by improving the regression process, single-dimensional regression is performed based on the preset aspect ratio of the specific behavior box, which effectively improves the training and detection efficiency and accuracy; by improving the loss function, the performance occupation of too difficult and too easy samples in machine training is suppressed, so that the training process can learn more features of normal samples, further improve the training efficiency of the target detection model, and improve the detection accuracy;
[0031] When the target detection model is applied to classroom behavior detection, it can accurately and efficiently identify target behaviors based on classroom teaching videos, effectively reflect students' classroom participation, and improve teaching quality.
[0032] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] The accompanying drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present invention, and together with the specification are used to explain the principles of the present invention. Obviously, the accompanying drawings described below are only some embodiments of the present invention, and for those of ordinary skill in the art, other accompanying drawings can be obtained based on these accompanying drawings without creative work.
[0034] Figure 1 A schematic diagram showing the steps of a method for training a target detection model in one embodiment of the present invention;
[0035] Figure 2 A schematic diagram showing a process of processing a sample image by an initial network model in one embodiment of the present invention;
[0036] Figure 3A curve relationship diagram showing the loss function value and the actual deviation value of the center point offset loss function L(δ) of the present invention and the traditional smooth L1 loss function;
[0037] Figure 4 A curve relationship diagram showing the gradient function value and the actual deviation value of the center point offset loss function L(δ) of the present invention and the traditional smooth L1 loss function;
[0038] Figure 5 A schematic diagram showing a module of a training device for a target detection model in one embodiment of the present invention;
[0039] Figure 6 A schematic diagram showing the steps of a classroom behavior detection method in one embodiment of the present invention;
[0040] Figure 7 A module schematic diagram of a classroom behavior detection device in one embodiment of the present invention is shown;
[0041] Figure 8 A schematic structural diagram of an electronic device in an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0042] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as limited to the embodiments set forth herein. Instead, these embodiments are provided to make the present invention comprehensive and complete and to fully convey the concepts of the example embodiments to those skilled in the art.
[0043] The accompanying drawings are only schematic illustrations of the present invention and are not necessarily drawn to scale. The same reference numerals in the figures represent the same or similar parts, and their repeated description will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0044] In addition, the processes shown in the accompanying drawings are only exemplary and do not necessarily include all the steps. For example, some steps can be decomposed, some steps can be combined or partially combined, and the actual execution order may change according to the actual situation. It should be noted that the embodiments of the present invention and the features in different embodiments can be combined with each other without conflict.
[0045] Figure 1 The main steps of the training method of the target detection model in one embodiment are shown, referring to Figure 1As shown, the training method of the target detection model includes: step S110, building an initial network model based on the center network CenterNet; step S120, processing the sample image through the initial network model to obtain the category and center point position of each target object in the sample image, and performing single-dimensional regression according to the center point position to obtain the detection box information of each target object; step S130, according to a sample image set including multiple sample images, using a multi-task loss function including classification loss, detection box offset loss and center point offset loss as a constraint condition to control the training of the initial network model to obtain the target detection model.
[0046] The center network CenterNet is an anchor-free target detection algorithm with fast detection speed and high accuracy. The backbone network of CenterNet is a deep layer aggregation network (Deep Layer Aggregation, referred to as DLA), which can be used to extract semantic and spatial features, and has a good effect on small target detection. In classroom behavior detection, such as hand-raising behavior detection, the resolution of the target behavior is small, and the DLA network can effectively detect small targets. The DLA network contains two basic structures: iterative deep aggregation (Iterative Deep Aggregation, referred to as IDA) and hierarchical deep aggregation (Hierarchical Deep Aggregation, referred to as HAD). Iterative deep aggregation IDA can extract semantic features of targets of different scales, and hierarchical deep aggregation HAD can fuse channel and depth information to extract semantic information at each stage. The structure and principle of the center network CenterNet, the deep aggregation network DLA, the hierarchical deep aggregation HAD and the iterative deep aggregation IDA are existing technologies, so they will not be described in detail. The present invention mainly improves the regression process and loss function of the algorithm model based on CenterNet, which will be described in detail below.
[0047] Figure 2 The processing process of the sample image by the initial network model in one embodiment is shown, referring to Figure 2As shown, the sample image 210 of this embodiment includes two target objects of the same category (hand-raising behavior). Before the sample image 210 is input into the initial network model, its size is adjusted to facilitate model processing. For example, the sample image 210' is adjusted to a pixel size of [512×512] and input into the initial network model. The initial network model performs hierarchical deep aggregation and iterative deep aggregation on the sample image 210' through the deep aggregation network 220 to obtain the category and center point position of each target object. C1, C2, C3 and C4 are different convolution blocks. The initial network model will output the center point positions of target objects of the same category as a heat map 230. The two points in the heat map 230 correspond to the center point positions of the two detected target objects. The pixel size of the heat map 230 is [128×128].
[0048] Figure 2 The sample image 210 shown includes two target objects, but the present invention is not limited thereto. The sample image may not include the target object, and the initial network model identifies the sample image including the target object and performs subsequent processing.
[0049] The detection frame information of each target object is obtained by performing a single-dimensional regression according to the center point position, including: performing a single-dimensional regression of the size according to the center point position to obtain the width or height of the detection frame of the corresponding target object; according to a preset aspect ratio and the width or height of the detection frame of the corresponding target object, the preset aspect ratio is determined according to the actual category of the corresponding target object, to obtain the width and height of the detection frame of the corresponding target object.
[0050] Reference Figure 2 As shown, in this embodiment, after obtaining the heat map 230, a single-dimensional regression of the width size is performed according to the positions of the two center points to obtain the width information of the two target objects presented in the width map 240 (the actual model output may be size or coordinates). Combined with the preset aspect ratio of the hand-raising behavior (the preset aspect ratios of target objects of different categories are different and can be set as needed), for example, detection frame width: detection frame height = 1:1.5, the height information of the two target objects presented in the height map 250 is inferred. Thus, the heat map 230 is fused with the sample image 210' to obtain a fusion image 260, and then combined with the width information in the width map 240 and the height information in the height map 250, the detection result image 270 is output, in which the two detected target objects are indicated by the detection frame 280.
[0051] Conventional target detection tasks require regression of two dimensions, height and width, which is time-consuming and laborious. The present invention takes into account the aspect ratio of target objects of a specific category, that is, the aspect ratio is basically maintained at a specific constant. Therefore, based on the center point position predicted by CenterNet, only the width is regressed, and then multiplied by the aspect ratio to obtain the final detection frame. Compared with the traditional two-dimensional regression algorithm, the present invention performs single-dimensional size regression and infers the other dimensional sizes of the detection frame through the aspect ratio, which significantly improves the regression accuracy and efficiency.
[0052] Figure 2 What is shown is the process of the initial network model processing a sample image. According to the sample image set including multiple sample images, under the constraints of a multi-task loss function including classification loss, detection box offset loss and center point offset loss, the initial network model is trained and optimized, and its model parameters are continuously adjusted to obtain the target detection model.
[0053] Multi-task loss function L det For: L det =L k +λ size L size +λ δ L(δ), where L k is the loss function of classification loss, L size is the loss function of the detection box offset loss, λ size YesL size The weight is 0.1, L(δ) is the loss function of the center point offset loss, λ δ is the weight of L(δ), which takes the value of 1.
[0054] Classification loss function L k The pixel-based classification loss Logistic regression function is used, which is the existing classification loss function of CenterNet. Classification loss function L k for:
[0055]
[0056] Among them, the initial network model outputs a corresponding number of heat maps according to the number of detected categories. Each heat map contains the coordinate point of at least one target object in the same category. The subscript xyc of the summation symbol is all the coordinate points of all heat maps. is the detection value of the target object of category c at the coordinate point (x, y), Y xyc is the true value of the target object of category c at the coordinate point (x, y), α and β are hyperparameters; α can be set to 2 and β can be set to 4.
[0057] and is the downsampled coordinate value of the coordinate point (x, y), σ p is the standard deviation. Indicates that for category c, a target object of this category is detected in the current (x, y) coordinate, and Indicates that there is no target object of category c at the current coordinate point. Classification loss function L k The specific principle is already known, so it will not be described in detail.
[0058] Detection box offset loss function L size The center point offset loss function L(δ) is the loss function of the center point position regression, where the detection box offset loss function L size The center point offset loss function L(δ) measures the position deviation of the detection box.
[0059] Detection box offset loss function L size The existing CenterNet loss function is used, specifically:
[0060]
[0061] Where N is the number of target objects in the sample image, is the width and height of the detection box of the kth target object. The width and height of the detection box are obtained through regression. k is the width and height of the real box of the kth target object. The width and height of the real box are obtained by fact-marking the sample image. The detection box offset loss function L size The specific principle is already known, so it will not be described in detail.
[0062] The present invention improves the center point offset loss function L(δ). In the existing CenterNet loss function, the loss function that measures the position deviation of the center point usually adopts the smooth L1 loss function, which cannot suppress the learning of samples that are too difficult or too easy. In machine training, there will be samples that are too difficult or too easy. The reason for the difficulty may be that subjective errors and opinions in the fact labeling process lead to controversial fact labeling results, or it may be because the scene is too complicated. Samples that are too easy can be judged without machine learning. Samples that are too difficult and too easy will consume a lot of unnecessary machine training time. Therefore, the present invention adopts a suppression loss function Suppression Loss (i.e., the center point offset loss function L(δ)), which can suppress the loss of difficult and easy samples, so that the training process can be more focused on normal samples, thereby improving the universality of the model.
[0063] The loss function L(δ) of the center point offset loss is specifically:
[0064]
[0065] Among them, δ is the offset of the center point position of the kth target object,
[0066] The derivative of the center point offset loss function L(δ), that is, the corresponding gradient function L′(δ), is:
[0067]
[0068] Figure 3 The curve relationship between the loss function value and the actual deviation value of the center point offset loss function L(δ) of the present invention and the traditional smooth L1 loss function is shown. Figure 4 The curve relationship between the gradient function value and the actual deviation value of the center point offset loss function L(δ) of the present invention and the traditional smooth L1 loss function is shown. Based on the same sample image set (hand-raising image set), the comparison effect of using the center point offset loss function L(δ) of the present invention and using the traditional smooth L1 loss function is shown in Figure 3 and Figure 4 , where the horizontal axis is the fact deviation value, Figure 3 The vertical axis is the loss function value, Figure 4 The vertical axis is the gradient function value. When the fact deviation value reaches 1, it means that the sample is too difficult; at this time, the gradient function value of the smoothed L1 loss function is also 1 (see Figure 4 The loss function value continues to rise (see curve 410). Figure 3 The convolutional neural network will continue to train; however, when the center point offset loss function L(δ) of the present invention is used, the gradient function value is 0 when the fact deviation value reaches 1 (see Figure 4 The loss function value no longer increases (see curve 420). Figure 3 The convolutional neural network will no longer learn the overly difficult sample. Therefore, the suppression loss function of the present invention can suppress the overly difficult and overly easy training losses, suppress the excessive occupation of computing resources by overly difficult or overly easy samples, and enable the model to learn more features of normal samples, that is, learn more essential features of hand-raising behavior in this embodiment.
[0069] Furthermore, after the target detection model is trained and generated using the above training method, the performance values of the target detection model and the conventional detection model in the prior art are tested for the hand-raising image set, and the comparison results are shown in the following table:
[0070]
[0071] The algorithm innovations of the present invention are tested using ablation experiments. mAP (mean Average Precision, i.e., averaging the AP values of all classes) calculates the precision and recall under different confidence levels, with values between 0 and 1. The closer to 1, the better the performance. As shown in the table above, compared with the conventional detection model CenterNet (ResNet), each innovation used in the present invention contributes to the mAP score, indicating that each algorithm innovation of the present invention has value.
[0072] The embodiment of the present invention also provides a training device for a target detection model, which can be used to implement the training method for the target detection model described in any of the above embodiments. The features and principles of the training method described in any of the above embodiments can be applied to the following training device embodiments. In the following training device embodiments, the features and principles of the training process of the target detection model that have been explained will not be repeated.
[0073] Figure 5 The main modules of the training device of the target detection model in one embodiment are shown, referring to Figure 5 As shown, the training device 500 for the target detection model includes: a model construction module 510, which is used to construct an initial network model based on the center network CenterNet; an image prediction module 520, which is used to process the sample image through the initial network model, obtain the category and center point position of each target object in the sample image, and perform single-dimensional regression according to the center point position to obtain the detection box information of each target object; a loss constraint module 530, which is used to control the training of the initial network model according to a sample image set including multiple sample images, using a multi-task loss function including classification loss, detection box offset loss and center point offset loss as a constraint condition to obtain the target detection model.
[0074] Furthermore, the target detection model training device 500 may also include modules for implementing other process steps of the above-mentioned target detection model training method embodiments. The specific principles of each module can refer to the description of the above-mentioned target detection model training method embodiments, and will not be repeated here.
[0075] The training device of the target detection model of the present invention can handle the occlusion phenomenon in the image well based on the Anchor-free target detection algorithm CenterNet; by improving the regression process, single-dimensional regression is performed based on the preset aspect ratio of the specific behavior box, which effectively improves the training and detection efficiency and accuracy; by improving the loss function, the performance occupation of too difficult and too easy samples in machine training is suppressed, so that the training process can learn more features of normal samples, further improve the training efficiency of the target detection model, and improve the detection accuracy.
[0076] The embodiment of the present invention also provides a classroom behavior detection method, which uses a target detection model generated by training the target detection model training method described in any of the above embodiments to detect classroom behavior. The characteristics and principles of the target detection model described in any of the above embodiments can be applied to the following classroom behavior detection method embodiment. In the following classroom behavior detection method embodiment, the characteristics and principles of the target detection model that have been explained will not be repeated.
[0077] Figure 6 The main steps of the classroom behavior detection method in one embodiment are shown, referring to Figure 6 As shown, the classroom behavior detection method includes: step S610, obtaining an image to be detected; step S620, inputting the image to be detected into a trained target detection model to obtain the target category and target detection frame information output by the target detection model; step S630, obtaining a classroom behavior detection result according to the target category and target detection frame information. The classroom behavior detection result is specifically a retrieval result of hand-raising behavior.
[0078] The target category and target detection box information output by the target detection model can be found in Figure 2 As shown in the detection result image 270, the target detection model will output the detection boxes of target objects of the same category in one result image. For example, the detection boxes of all hand-raising behaviors will be output in one result image, so that the detection results of hand-raising behaviors can be easily obtained according to the output of the target detection model.
[0079] In the above-mentioned target detection model training method embodiment, the hand-raising image set, that is, the target object is the hand-raising behavior, is mainly used as an example for explanation; during actual model training, training can be carried out based on multiple image sets such as the hand-raising image set and the table-lying image set, so that the target detection model can learn the characteristics of target objects of different categories, and thus be applied to the detection of various types of classroom behaviors.
[0080] The classroom behavior detection method of the present invention can accurately and efficiently identify target behaviors based on classroom teaching videos, effectively reflect students' classroom participation, and improve teaching quality.
[0081] In specific applications, a camera that can capture a panoramic view of students can be installed in the classroom to obtain student videos. The program developed based on the training method of the target detection model of the present invention and the classroom behavior detection method is deployed on a server / workstation with a graphics processor GPU (supporting CUDA, Compute Unified Device Architecture, unified computing device architecture), which can be deployed locally in the school or on the cloud. According to the actual classroom scene detection effect, such as the hand-raising behavior detection effect, the present invention can quickly and accurately detect the hand-raising behavior in the image and effectively reduce the false detection of the background. According to the identified hand-raising behavior, it can be combined with other classroom teaching behaviors for analysis to evaluate indicators such as classroom teaching activity and teacher-student behavior conversion rate to form an objective evaluation of classroom teaching.
[0082] Furthermore, the detection method of the present invention can also be extended to other industry scenarios, such as determining the initiative and enthusiasm of participants in a meeting, determining whether the operations of employees in a factory comply with safety regulations, etc. As long as the model is trained according to different target behavior image sets, detection and analysis of different target behaviors can be achieved.
[0083] The embodiment of the present invention further provides a classroom behavior detection device, which can be used to implement the classroom behavior detection method described in any of the above embodiments. The characteristics and principles of classroom behavior detection described in any of the above embodiments can be applied to the following classroom behavior detection device embodiment. In the following classroom behavior detection device embodiment, the characteristics and principles of classroom behavior detection that have been explained will not be repeated.
[0084] Figure 7 The main modules of the classroom behavior detection device in one embodiment are shown, referring to Figure 7 As shown, the classroom behavior detection device 700 includes: an image acquisition module 710, used to obtain the image to be detected; an image processing module 720, used to input the image to be detected into a trained target detection model to obtain the target category and target detection frame information output by the target detection model; a behavior analysis module 730, used to obtain the classroom behavior detection result based on the target category and target detection frame information.
[0085] The classroom behavior detection device of the present invention can accurately and efficiently identify target behaviors based on classroom teaching videos, effectively reflect students' classroom participation, and improve teaching quality.
[0086] In summary, the target detection model training method and classroom behavior detection method of the present invention adopt the Anchor-free one-stage target detection algorithm CenterNet, which can reduce the application of candidate box design to the results and speed up the reasoning speed; improve the regression accuracy and training efficiency through a single-dimensional regression strategy; suppress the waste of computing resources by too difficult and too easy samples through a loss suppression function, improve machine training efficiency, realize concentrated learning of the essential characteristics of the target object, improve the universal detection effect of behavior detection, and improve recognition accuracy.
[0087] An embodiment of the present invention also provides an electronic device, including a processor and a memory, wherein executable instructions are stored in the memory, and when the executable instructions are executed by the processor, the target detection model training method and / or classroom behavior detection method described in any of the above embodiments are implemented.
[0088] The electronic device of the present invention is based on the anchor-free target detection algorithm CenterNet, which can well handle the occlusion phenomenon in the image; by improving the regression process, single-dimensional regression is performed based on the preset aspect ratio of the specific behavior box, which effectively improves the training and detection efficiency and accuracy; by improving the loss function, the performance occupation of too difficult and too easy samples in machine training is suppressed, so that the training process can learn more features of normal samples, further improve the training efficiency of the target detection model, and improve the detection accuracy; when the target detection model is applied to classroom behavior detection, it can accurately and efficiently identify the target behavior based on the classroom teaching video, effectively reflect the students' classroom participation, and improve the teaching quality.
[0089] Figure 8 is a schematic diagram of the structure of an electronic device in an embodiment of the present invention. It should be understood that: Figure 8 The modules are only schematically shown, and these modules may be virtual software modules or actual hardware modules. The combination and splitting of these modules and the addition of other modules are all within the protection scope of the present invention.
[0090] like Figure 8 As shown, the electronic device 800 is in the form of a general computing device. The components of the electronic device 800 include but are not limited to: at least one processing unit 810, at least one storage unit 820, a bus 830 connecting different platform components (including the storage unit 8620 and the processing unit 810), a display unit 840, etc.
[0091] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 810, so that the processing unit 810 executes the steps of the target detection model training method and / or classroom behavior detection method described in any of the above embodiments.
[0092] The storage unit 820 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 8201 and / or a cache storage unit 8202 , and may further include a read-only storage unit (ROM) 8203 .
[0093] The storage unit 820 may also include a program / utility 8204 having one or more program modules 8205, such program modules 8205 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
[0094] Bus 830 may represent one or more of several types of bus structures, including a memory unit bus or memory unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.
[0095] The electronic device 800 can also communicate with one or more external devices 900, which can be one or more of a keyboard, a pointing device, a Bluetooth device, etc. These external devices 900 enable the user to communicate interactively with the electronic device 800. The electronic device 800 can also communicate with one or more other computing devices, and the computer device shown includes a router and a modem. This communication can be performed through an input / output (I / O) interface 850. In addition, the electronic device 800 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN) and / or a public network, such as the Internet) through a network adapter 860. The network adapter 860 can communicate with other modules of the electronic device 800 through a bus 830. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in conjunction with the electronic device 800, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage platforms.
[0096] The embodiment of the present invention also provides a computer-readable storage medium for storing a program, which, when executed, implements the target detection model training method and / or classroom behavior detection method described in any of the above embodiments. In some possible implementations, various aspects of the present invention can also be implemented in the form of a program product, which includes a program code. When the program product is run on a terminal device, the program code is used to enable the terminal device to execute the target detection model training method and / or classroom behavior detection method described in any of the above embodiments.
[0097] The computer-readable storage medium of the present invention is based on the anchor-free target detection algorithm CenterNet, which can well handle the occlusion phenomenon in the image; by improving the regression process, single-dimensional regression is performed based on the preset aspect ratio of the specific behavior box, which effectively improves the training and detection efficiency and accuracy; by improving the loss function, the performance occupation of too difficult and too easy samples in machine training is suppressed, so that the training process can learn more features of normal samples, further improve the training efficiency of the target detection model, and improve the detection accuracy; when the target detection model is applied to classroom behavior detection, it can accurately and efficiently identify the target behavior based on the classroom teaching video, effectively reflect the students' classroom participation, and improve the teaching quality.
[0098] The program product may be a portable compact disk read-only memory (CD-ROM) and include program code, and may be run on a terminal device, such as a personal computer. However, the program product of the present invention is not limited thereto, and may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system, apparatus, or device.
[0099] The program product may adopt any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of readable storage media include, but are not limited to: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0100] The readable storage medium may include a data signal propagated in a baseband or as part of a carrier wave, wherein the readable program code is carried. The data signal of this propagation may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable storage medium may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, device, or device. The program code contained on the readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the above.
[0101] Program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device, such as through the Internet using an Internet service provider.
[0102] The above contents are further detailed descriptions of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as falling within the protection scope of the present invention.
Claims
1. A method for training a target detection model, characterized in that: include: Build the initial network model based on the center network CenterNet; Processing the sample image through the initial network model to obtain the category and center point position of each target object in the sample image, and performing single-dimensional regression according to the center point position to obtain detection box information of each target object; The step of performing a one-dimensional regression according to the center point position to obtain the detection frame information of each target object includes: performing a one-dimensional regression of the size according to the center point position to obtain the width or height of the detection frame of the corresponding target object; performing a one-dimensional regression of the size according to the center point position to obtain the width or height of the detection frame of the corresponding target object according to a preset aspect ratio and the width or height of the detection frame of the corresponding target object, wherein the preset aspect ratio is determined according to the real category of the corresponding target object, to obtain the width and height of the detection frame of the corresponding target object; According to a sample image set including a plurality of the sample images, a multi-task loss function including classification loss, detection frame offset loss and center point offset loss is used as a constraint condition to control the training of the initial network model to obtain a target detection model.
2. The training method according to claim 1, characterized in that: The loss function L(δ) of the center point offset loss is: Among them, δ is the offset of the center point position of the kth target object, is the width and height of the detection box of the kth target object, s k are the width and height of the real box of the kth target object, k ranges from 1 to N, and N is the number of target objects in the sample image; The gradient function L'(δ) of the loss function L(δ) is:
3. The training method according to claim 2, characterized in that: The loss function L of the detection box offset loss size for: The loss function L of the classification loss k for: in, is the detection value of the target object of category c at the coordinate point (x, y), Y xyc is the true value of the target object of category c at the coordinate point (x, y), and α and β are hyperparameters; The multi-task loss function L det for: L det =L k +λ size L size +λ δ L(δ); Among them, λ size and λ δ They are the loss functions L size and the weight of L(δ).
4. A training device for a target detection model, characterized in that: include: Model building module, used to build the initial network model based on the center network CenterNet; An image prediction module, used to process the sample image through the initial network model, obtain the category and center point position of each target object in the sample image, and perform single-dimensional regression according to the center point position to obtain the detection box information of each target object; The image prediction module performs a one-dimensional regression according to the center point position to obtain the detection frame information of each target object, including: performing a one-dimensional regression of the size according to the center point position to obtain the width or height of the detection frame of the corresponding target object; according to a preset aspect ratio and the width or height of the detection frame of the corresponding target object, the preset aspect ratio is determined according to the real category of the corresponding target object, to obtain the width and height of the detection frame of the corresponding target object; The loss constraint module is used to control the training of the initial network model according to a sample image set including multiple sample images, using a multi-task loss function including classification loss, detection frame offset loss and center point offset loss as a constraint condition to obtain a target detection model.
5. A classroom behavior detection method, characterized in that: include: Obtaining an image to be detected; Inputting the image to be detected into the target detection model generated by training using the training method according to any one of claims 1 to 3, and obtaining the target category and target detection frame information output by the target detection model; A classroom behavior detection result is obtained according to the target category and the target detection frame information.
6. The classroom behavior detection method according to claim 5, characterized in that: The classroom behavior detection result is a retrieval result of hand-raising behavior.
7. A classroom behavior detection device, characterized in that: include: An image acquisition module, used for acquiring an image to be detected; An image processing module, used for inputting the image to be detected into the target detection model generated by the training method according to any one of claims 1 to 3, and obtaining the target category and target detection frame information output by the target detection model; The behavior analysis module is used to obtain classroom behavior detection results based on the target category and the target detection frame information.
8. An electronic device, characterized in that: include: A processor; a memory storing executable instructions; When the executable instructions are executed by the processor, the training method of the target detection model as described in any one of claims 1 to 3 is implemented, and / or the classroom behavior detection method as described in claim 5 or 6 is implemented.
9. A computer-readable storage medium for storing a program, characterized in that: When the program is executed by the processor, it implements the training method of the target detection model as described in any one of claims 1-3, and / or implements the classroom behavior detection method as described in claim 5 or 6.
Citation Information
Patent Citations
Video action detection method based on central point trajectory prediction
CN111259779A
Pluggable aerial image target positioning detection method, system and device
CN112364843A