Method and device for training a reinforcement detection model
By replacing the convolutional layers of the RetinaNet model with dilated convolutional layers, and combining the balance formula and multiple loss functions to train the model, the problem of low accuracy in machine learning for detecting rebar was solved, and higher detection accuracy was achieved.
Patent Information
- Application Number
- CN202211642914.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-20
- Publication Date
- 2026-01-20
- Estimated Expiration
- 2042-12-20
AI Technical Summary
Existing machine learning methods for detecting rebar have low accuracy.
The ordinary convolutional layers in the RetinaNet model are replaced with dilated convolutional layers. The balance of the center points is calculated using the balance formula for classification. During the classification process, the first loss function is used to train the rebar detection model. The coordinates of the detection box are regressed based on the center point corresponding to each label in the training samples, and the model is trained using the second loss function. The third loss function is used to train the model based on the intersection-union ratio of adjacent detection boxes.
This improved the accuracy of rebar inspection.
Smart Images

Figure CN115965850B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of machine learning, and particularly relates to a training method and device of a steel bar detection model. BACKGROUND
[0002] Applying machine learning to steel bar detection can reduce the workload of workers. The mainstream multi-target detection is generally used for steel bar detection, and there is no optimization and improvement based on the characteristics of steel bar detection. Steel bar detection is different from general target detection, and has characteristics such as dense, non-overlapping and small target. If only the method of multi-target detection in the prior art is used without improvement, the accuracy of steel bar detection will be low.
[0003] In the process of implementing the present disclosure, the inventors have found that the related art has at least the following technical problem: the accuracy of the method of commonly used machine learning for detecting steel bars is low. SUMMARY
[0004] Therefore, the embodiments of the present disclosure provide a training method and device of a steel bar detection model, an electronic device and a computer readable storage medium, to solve the problem of low accuracy of the method of commonly used machine learning for detecting steel bars in the prior art.
[0005] In a first aspect, the embodiments of the present disclosure provide a training method of a steel bar detection model, comprising: replacing an ordinary convolution layer in a RetinaNet model with a dilated convolution layer to obtain a steel bar detection model; obtaining a steel bar training data set, inputting a training sample in the steel bar training data set into the steel bar detection model, and outputting a feature map of the training sample, wherein the feature map contains a plurality of center points; calculating a balance degree of each center point by using a balance degree formula, classifying the center point based on the balance degree of each center point, and training the steel bar detection model by using a first loss function in the process of classification, wherein the class of each center point includes foreground and background; based on each label in the training sample corresponding to one or more center points classified as foreground, the coordinates of the detection box corresponding to the label are regressed, wherein the training sample contains a plurality of labels, and the steel bar detection model is trained by using a second loss function in the process of regression; based on the intersection over union of each two adjacent detection boxes in all detection boxes in the training sample, the steel bar detection model is trained by using a third loss function.
[0006] In a second aspect, the present disclosure provides a device for training a steel bar detection model, comprising: a construction module configured to replace a normal convolution layer in a RetinaNet model with a dilated convolution layer to obtain a steel bar detection model; an acquisition module configured to acquire a steel bar training dataset, input a training sample in the steel bar training dataset into the steel bar detection model, and output a feature map of the training sample, wherein the feature map contains a plurality of center points; a classification module configured to calculate a balance degree of each center point using a balance degree formula, classify each center point based on the balance degree of the center point, and train the steel bar detection model using a first loss function in the process of classification, wherein the class of each center point includes foreground and background; a regression module configured to regress coordinates of a detection box corresponding to each label in the training sample based on one or more center points classified as foreground corresponding to the label, wherein the training sample contains a plurality of labels, and train the steel bar detection model using a second loss function in the process of regression; and a training module configured to train the steel bar detection model using a third loss function based on an intersection over union of each two adjacent detection boxes in all detection boxes in the training sample.
[0007] In a third aspect, the present disclosure provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the computer program.
[0008] In a fourth aspect, the present disclosure provides a computer-readable storage medium, which stores a computer program, wherein the computer program is executed by a processor to implement the steps of the above method.
[0009] Compared with the prior art, the beneficial effects of the embodiments of the present disclosure are that: because the embodiments of the present disclosure replace the ordinary convolution layer in the RetinaNet model with the empty convolution layer to obtain a steel bar detection model; a steel bar training data set is obtained, a training sample in the steel bar training data set is input into the steel bar detection model, and a feature map of the training sample is output, wherein the feature map contains a plurality of center points; the balance degree of each center point is calculated by using a balance degree formula, each center point is classified based on the balance degree of the center point, and the steel bar detection model is trained by using a first loss function in the classification process, wherein the class of each center point includes foreground and background; the coordinates of the detection box corresponding to each label in the training sample are regressed based on one or more center points classified as foreground corresponding to the label, wherein the training sample contains a plurality of labels, and the steel bar detection model is trained by using a second loss function in the regression process; the third loss function is used to train the steel bar detection model based on the intersection over union of each two adjacent detection boxes in all detection boxes in the training sample, so that the above technical means can solve the problem of low accuracy of the method for detecting steel bars by commonly used machine learning in the prior art, and the accuracy of detecting steel bars is improved. BRIEF DESCRIPTION OF DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.
[0011] Figure 1 is a scene schematic diagram of an application scenario of the embodiments of the present disclosure;
[0012] Figure 2 is a flowchart of a training method of a steel bar detection model provided by the embodiments of the present disclosure;
[0013] Figure 3 is a structural schematic diagram of a steel bar detection model training device provided by the embodiments of the present disclosure;
[0014] Figure 4 is a structural schematic diagram of an electronic device provided by the embodiments of the present disclosure. DETAILED DESCRIPTION
[0015] In the following description, for purposes of explanation and not limitation, specific details are set forth such as particular architectures, techniques, etc. in order to provide a thorough understanding of the embodiments of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure can be practiced in other embodiments that depart from these specific details. In other instances, detailed descriptions of well-known systems, devices, circuits, and methods are omitted so as not to obscure the description of the present disclosure with unnecessary detail.
[0016] A method and device for training a steel bar detection model according to an embodiment of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0017] Figure 1 FIG. 1 is a schematic diagram of an application scenario of an embodiment of the present disclosure. The application scenario can include terminal devices 101, 102, and 103, a server 104, and a network 105.
[0018] The terminal devices 101, 102, and 103 can be hardware or software. When the terminal devices 101, 102, and 103 are hardware, they can be various electronic devices with a display screen and supporting communication with the server 104, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like; when the terminal devices 101, 102, and 103 are software, they can be installed in the electronic devices as above. The terminal devices 101, 102, and 103 can be implemented as multiple software or software modules, or as a single software or software module, and the embodiments of the present disclosure do not limit this. Further, various applications can be installed on the terminal devices 101, 102, and 103, such as data processing applications, instant messaging tools, social platform software, search applications, shopping applications, and the like.
[0019] The server 104 can be a server that provides various services, for example, a background server that receives a request sent by a terminal device that establishes a communication connection therewith. The background server can receive and analyze the request sent by the terminal device, and generate a processing result. The server 104 can be a single server, a server cluster composed of several servers, or a cloud computing service center, and the embodiments of the present disclosure do not limit this.
[0020] It should be noted that the server 104 can be hardware or software. When the server 104 is hardware, it can be various electronic devices that provide various services for the terminal devices 101, 102, and 103. When the server 104 is software, it can be multiple software or software modules that provide various services for the terminal devices 101, 102, and 103, or a single software or software module that provides various services for the terminal devices 101, 102, and 103, and the embodiments of the present disclosure do not limit this.
[0021] The network 105 can be a wired network using coaxial cables, twisted-pair wires, and optical fibers, or a wireless network that can realize interconnection of various communication devices without wiring, for example, Bluetooth, Near Field Communication (NFC), Infrared, etc., and the embodiments of the present disclosure do not make any limitation in this regard.
[0022] The user can establish a communication connection with the server 104 via the network 105 through the terminal devices 101, 102, and 103 to receive or send information, etc. It should be noted that the specific types, quantities, and combinations of the terminal devices 101, 102, and 103, the server 104, and the network 105 can be adjusted according to the actual needs of the application scenarios, and the embodiments of the present disclosure do not make any limitation in this regard.
[0023] Figure 2 is a flowchart of a training method of a steel bar detection model provided by the embodiments of the present disclosure. Figure 2 The training method of the steel bar detection model can be executed by Figure 1 a computer or a server, or software on the computer or the server. As Figure 2 shown, the training method of the steel bar detection model includes:
[0024] S201, replacing a normal convolution layer in a RetinaNet model with a dilated convolution layer to obtain a steel bar detection model;
[0025] S202, obtaining a steel bar training data set, inputting a training sample in the steel bar training data set into the steel bar detection model, and outputting a feature map of the training sample, wherein the feature map contains a plurality of center points;
[0026] S203, calculating a balance degree of each center point by using a balance degree formula, classifying the center point based on the balance degree of each center point, and training the steel bar detection model by using a first loss function in the process of classification, wherein the class of each center point includes foreground and background;
[0027] S204, based on each label in the training sample corresponding to one or more center points classified as foreground, regressing coordinates of a detection box corresponding to the label, wherein the training sample contains a plurality of labels, and the steel bar detection model is trained by using a second loss function in the process of regression;
[0028] S205, based on an intersection-over-union of each two adjacent detection boxes in all detection boxes in the training sample, training the steel bar detection model by using a third loss function.
[0029] The convolutional layer used in the RetinaNet model is a normal convolutional layer. In order to increase the attention to adjacent information in the steel bar detection, the normal convolutional layer is replaced by a dilated convolutional layer in the embodiments of the present disclosure, and the modified RetinaNet model is used as the steel bar detection model. The training sample in the steel bar training data set is input into the steel bar detection model, and the feature map of the training sample is output by the feature extraction network of the steel bar detection model. The steel bar detection model has other network layers after the feature extraction network, and these knowledge is well known in the machine model and will not be described here.
[0030] It should be noted that the steel bar training data set includes a large number of training samples, each training sample has a feature map, and each feature map contains a plurality of center points. For the sake of convenience, only one training sample will be considered in the following description. The center point is a point on the feature map, and each center point needs to be predicted whether it is on the steel bar. If a center point is on the steel bar, the class of the center point is foreground, otherwise the class of the center point is background. The embodiments of the present disclosure judge that the center point with a balance degree greater than a preset threshold is on the steel bar, and the center point with a balance degree less than or equal to the preset threshold is not on the steel bar. The first loss function is used to calculate the probability of correct classification of the center point, and the loss value calculated by the first loss function is used to train the network of the classification part of the steel bar detection model. The first loss function can use the focalloss function.
[0031] The label includes a marking part and a frame part. The center points classified as foreground in the label are determined as the center points corresponding to the label, and then the coordinates of the detection frame corresponding to the label are regressed according to the center points corresponding to the label. The second loss function is used to calculate the difference between the label and its corresponding detection frame, and the loss value calculated by the second loss function is used to train the network of the regression part of the steel bar detection model. The second loss function can use the smooth_l1 function.
[0032] The third loss function is used to train the whole steel bar detection model.
[0033] According to the technical scheme provided by the embodiment of the present disclosure, the ordinary convolution layer in the RetinaNet model is replaced by the atrous convolution layer to obtain a steel bar detection model; a steel bar training data set is obtained, and a training sample in the steel bar training data set is input into the steel bar detection model to output a feature map of the training sample, wherein the feature map contains a plurality of center points; the balance degree of each center point is calculated by using a balance degree formula, each center point is classified based on the balance degree of the center point, and the steel bar detection model is trained by using a first loss function in the classification process, wherein the category of each center point includes foreground and background; the coordinates of the detection box corresponding to each label in the training sample are regressed based on one or more center points classified as foreground corresponding to the label, wherein the training sample contains a plurality of labels, and the steel bar detection model is trained by using a second loss function in the regression process; the steel bar detection model is trained by using a third loss function based on the intersection over union of each two adjacent detection boxes in all detection boxes in the training sample. Therefore, by using the above technical means, the problem of low accuracy of the method for detecting steel bars by using machine learning in the prior art can be solved, and the accuracy of detecting steel bars is improved.
[0034] The balance degree of each center point is calculated by using a balance degree formula, including: the distance of each center point to the upper edge, lower edge, left edge and right edge of the label corresponding to the center point is calculated respectively; the balance degree of each center point is calculated by using the balance degree formula based on the distance of each center point to the upper edge, lower edge, left edge and right edge of the label corresponding to the center point.
[0035] The balance degree p of each center point is calculated by using a balance degree formula:
[0036]
[0037] Wherein, the distance of each center point to the left edge of the label corresponding to the center point is l, the distance of each center point to the right edge of the label corresponding to the center point is r, the distance of each center point to the upper edge of the label corresponding to the center point is t, and the distance of each center point to the lower edge of the label corresponding to the center point is b.
[0038] 0.5 in the formula is the index of the whole part in parentheses.
[0039] Each center point is classified based on the balance degree of the center point, including: the center point with a balance degree greater than a preset threshold is classified as foreground; and the center point with a balance degree less than or equal to the preset threshold is classified as background.
[0040] The method comprises the following steps: determining the center point classified as foreground within each label as the center point corresponding to the label; and regressing the coordinates of the detection box corresponding to each label based on the box height and box width of each label and the coordinates of the one or more center points corresponding to the label.
[0041] The four values (the box height and box width of the label and the coordinates of the center point, and the coordinates of the center point are two values) are calculated for each training sample in the embodiments of the present disclosure.
[0042] The method comprises the following steps: determining the center point classified as foreground within each label as the center point corresponding to the label; and regressing the coordinates of the detection box corresponding to each label based on the box height and box width of each label and the coordinates of the one or more center points corresponding to the label.
[0043] The center points adjacent to the center point classified as foreground within each label in the up, down, left and right four directions are also determined as the center points corresponding to the label, which is technically corresponding to replacing the ordinary convolution layer in the RetinaNet model with the atrous convolution layer. Replacing the ordinary convolution layer in the RetinaNet model with the atrous convolution layer is to calculate the information of the adjacent center points. The embodiments of the present disclosure calculate 12 values for each training sample.
[0044] The third loss function:
[0045]
[0046]
[0047] Wherein, n is the total number of detection boxes in the training sample, b i is the i-th detection box, b j is the j-th detection box, gt i is the label corresponding to the i-th detection box, m is the total number of labels in the training sample, and IoU is the intersection over union function.
[0048] The Max() function is the maximum of IoU(b i ,b j )-α and 0.
[0049] All the optional technical solutions described above can be combined to form optional embodiments of the present application, which will not be described one by one.
[0050] The following is an embodiment of the device of the present disclosure, which can be used to execute the method embodiment of the present disclosure. For details not disclosed in the device embodiment of the present disclosure, please refer to the method embodiment of the present disclosure.
[0051] Figure 3 FIG. 1 is a schematic diagram of a steel bar detection model training device provided by an embodiment of the present disclosure. As shown in the figure, the steel bar detection model training device comprises: Figure 3
[0052] The construction module 301 is configured to replace the ordinary convolution layer in the RetinaNet model with a dilated convolution layer to obtain a steel bar detection model.
[0053] The acquisition module 302 is configured to acquire a steel bar training data set, input a training sample in the steel bar training data set into the steel bar detection model, and output a feature map of the training sample, wherein the feature map contains a plurality of center points.
[0054] The classification module 303 is configured to calculate the balance degree of each center point using a balance degree formula, classify each center point based on the balance degree of the center point, and train the steel bar detection model using a first loss function in the process of classification, wherein the class of each center point includes foreground and background.
[0055] The regression module 304 is configured to regress the coordinates of the detection box corresponding to each label in the training sample based on one or more center points classified as foreground corresponding to the label, wherein the training sample contains a plurality of labels, and the steel bar detection model is trained using a second loss function in the process of regression.
[0056] The training module 305 is configured to train the steel bar detection model using a third loss function based on the intersection over union of each two adjacent detection boxes in all detection boxes in the training sample.
[0057] The convolution layer used in the RetinaNet model is an ordinary convolution layer. In order to increase the attention to adjacent information in steel bar detection, the ordinary convolution layer is replaced with a dilated convolution layer in the present embodiment, and the modified RetinaNet model is used as a steel bar detection model. The training sample in the steel bar training data set is input into the steel bar detection model, which outputs the feature map of the training sample using the feature extraction network of the steel bar detection model. The steel bar detection model has other network layers after the feature extraction network, which are all common knowledge in machine models and will not be described here.
[0058] It should be noted that the steel bar training data set includes a large number of training samples, each training sample has a feature map, and each feature map includes a plurality of center points. In order to facilitate understanding, the following can be regarded as only one training sample. The center point is a point on the feature map, and each center point needs to be predicted whether it is on the steel bar. If a center point is on the steel bar, the category of the center point is foreground, otherwise the category of the center point is background. The embodiment of the present disclosure judges that the center point with a balance degree greater than a preset threshold is on the steel bar, and judges that the center point with a balance degree less than or equal to the preset threshold is not on the steel bar. The first loss function is used to calculate the probability of correctly classifying the center point. The loss value calculated by the first loss function is used to train the network of the classification part of the steel bar detection model. The first loss function can use the focalloss function.
[0059] The label includes a marking part and a frame part. The center points classified as foreground in the label are determined as the center points corresponding to the label, and then the coordinates of the detection frame corresponding to the label are regressed according to the center points corresponding to the label. The second loss function is used to calculate the difference between the label and the detection frame corresponding to the label. The difference is the loss value calculated by the second loss function, and the network of the regression part of the steel bar detection model is trained according to the loss value. The second loss function can use the smooth_l1 function.
[0060] The third loss function is used to train the whole steel bar detection model.
[0061] According to the technical scheme provided by the embodiment of the present disclosure, the ordinary convolution layer in the RetinaNet model is replaced by the empty convolution layer to obtain the steel bar detection model. The steel bar training data set is obtained, and the training sample in the steel bar training data set is input into the steel bar detection model to output the feature map of the training sample, wherein the feature map includes a plurality of center points. The balance degree of each center point is calculated by using the balance degree formula, and each center point is classified based on the balance degree of the center point, and the first loss function is used to train the steel bar detection model in the classification process. The category of each center point includes foreground and background. The coordinates of the detection frame corresponding to each label in the training sample are regressed based on one or more center points classified as foreground corresponding to the label, wherein the training sample includes a plurality of labels, and the second loss function is used to train the steel bar detection model in the regression process. The third loss function is used to train the steel bar detection model based on the intersection over union of each two adjacent detection frames in all detection frames in the training sample. Therefore, by using the above technical means, the problem of low accuracy of the method for detecting steel bars by using machine learning in the prior art can be solved, and the accuracy of detecting steel bars is improved.
[0062] Optionally, the classification module 303 is further configured to calculate distances of each center point to an upper edge, a lower edge, a left edge and a right edge of a label corresponding to the center point respectively; and calculate a balance degree of each center point based on the distances of each center point to the upper edge, the lower edge, the left edge and the right edge of the label corresponding to the center point by using a balance degree formula.
[0063] Optionally, the classification module 303 is further configured to calculate a balance degree p of each center point by using a balance degree formula:
[0064]
[0065] wherein the distance of each center point to a left edge of a label corresponding to the center point is l, the distance of each center point to a right edge of the label corresponding to the center point is r, the distance of each center point to an upper edge of the label corresponding to the center point is t, and the distance of each center point to a lower edge of the label corresponding to the center point is b.
[0066] 0.5 in the formula is an index of the whole part in the parentheses.
[0067] Optionally, the classification module 303 is further configured to classify a center point with a balance degree greater than a preset threshold as foreground, and classify a center point with a balance degree less than or equal to the preset threshold as background.
[0068] Optionally, the regression module 304 is further configured to determine a center point classified as foreground and located within each label as a center point corresponding to the label; and regress coordinates of a detection box corresponding to each label based on a box height and a box width of the label and coordinates of one or more center points corresponding to the label.
[0069] Each training sample in the embodiments of the present disclosure calculates four values (a box height and a box width of a label, and coordinates of a center point, which are two values).
[0070] Optionally, the regression module 304 is further configured to determine a center point classified as foreground and located within each label, and a center point adjacent to the center point in four directions of up, down, left and right as center points corresponding to the label; and regress coordinates of a detection box corresponding to each label based on a box height and a box width of the label and coordinates of one or more center points corresponding to the label.
[0071] The center point adjacent to the center point classified as foreground and located within each label in four directions of up, down, left and right is also determined as a center point corresponding to the label, which is technically corresponding to replacing a normal convolution layer in a RetinaNet model with a dilated convolution layer. Replacing the normal convolution layer in the RetinaNet model with the dilated convolution layer is to calculate information of adjacent center points. Each training sample in the embodiments of the present disclosure calculates 12 values.
[0072] The third loss function is:
[0073]
[0074]
[0075] wherein n is the total number of bounding boxes in the training sample, b i is the i-th bounding box, b j is the j-th bounding box, gt i is the label corresponding to the i-th bounding box, m is the total number of labels in the training sample, and IoU is an intersection over union function.
[0076] The Max() function is the maximum of IoU(b i ,b j )-a and 0.
[0077] It should be understood that the size of the serial number of each step in the above embodiments does not mean the order of execution, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present disclosure.
[0078] Figure 4 is a schematic diagram of an electronic device 4 provided by the embodiments of the present disclosure. As shown in Figure 4 , the electronic device 4 of this embodiment includes a processor 401, a memory 402, and a computer program 403 stored in the memory 402 and executable on the processor 401. The processor 401 implements the steps in each of the above method embodiments when executing the computer program 403. Alternatively, the processor 401 implements the functions of each module / unit in each of the above device embodiments when executing the computer program 403.
[0079] The electronic device 4 can be a desktop computer, a notebook computer, a palm computer, a cloud server, and the like. The electronic device 4 can include but is not limited to the processor 401 and the memory 402. Those skilled in the art can understand that Figure 4 , the electronic device 4 is only an example and does not constitute a limitation on the electronic device 4, and can include more or fewer components or different components than those shown.
[0080] The processor 401 can be a central processing unit (CPU), or other general purpose processors, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, etc.
[0081] The memory 402 can be an internal storage unit of the electronic device 4, for example, a hard disk or a memory of the electronic device 4. The memory 402 can also be an external storage device of the electronic device 4, for example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 4. The memory 402 can also include both the internal storage unit and the external storage device of the electronic device 4. The memory 402 is used to store computer programs and other programs and data required by the electronic device.
[0082] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of functional units and modules is taken as an example, and in actual application, the above functions can be completed by different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the above described functions. Each functional unit and module in the embodiment can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The above integrated unit can be realized in the form of hardware or in the form of software functional unit.
[0083] The integrated modules / units, if implemented in the form of software functional units and sold or used as independent products, can be stored in a computer readable storage medium. Based on such understanding, all or part of the processes in the above-mentioned embodiment methods can also be completed by instructing related hardware through a computer program, and the computer program can be stored in a computer readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. The computer program can include computer program code, which can be in the form of source code, object code, executable files or some intermediate forms, etc. The computer readable medium can include any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (Read-Only Memory, ROM), random access memory (Random Access Memory, RAM), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in the computer readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction, for example, in some jurisdictions, according to legislation and patent practice, the computer readable medium does not include electric carrier signal and telecommunication signal.
[0084] The above embodiments are only used to illustrate the technical solutions of the present disclosure, rather than limit them; although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should be included in the protection scope of the present disclosure.
Claims
1. A method for training a reinforcement detection model, characterized in that, The application comprises the following steps: replacing the ordinary convolution layer in the RetinaNet model with a dilated convolution layer to obtain a steel bar detection model; obtaining a steel bar training data set, inputting a training sample in the steel bar training data set into the steel bar detection model, and outputting a feature map of the training sample, wherein the feature map contains a plurality of center points; calculating the balance degree p of each center point by using a balance degree formula, wherein the distance of each center point to the left edge box of the corresponding label of the center point is l, the distance of each center point to the right edge box of the corresponding label of the center point is r, the distance of each center point to the upper edge box of the corresponding label of the center point is t, and the distance of each center point to the lower edge box of the corresponding label of the center point is b, classifying the center point based on the balance degree of each center point, and training the steel bar detection model by using a first loss function in the classification process, wherein the class of each center point includes foreground and background; the first loss function uses a focal loss function; regressing the coordinates of the detection box corresponding to each label in the training sample based on one or more center points classified as foreground corresponding to the label, wherein the training sample contains a plurality of labels, and training the steel bar detection model by using a second loss function in the regression process; the second loss function uses a smooth_l1 function; training the steel bar detection model by using a third loss function based on the intersection over union of each two adjacent detection boxes in all detection boxes in the training sample; the third loss function: calculating the balance degree of each center point by using a balance degree formula, comprising: wherein n is the total number of bounding boxes in the training sample, b i is the i-th bounding box, b j is the j-th bounding box, gt i is the label corresponding to the i-th bounding box, m is the total number of labels in the training sample, and IoU is the intersection over union function.
2. The method of claim 1, wherein, respectively calculating the distance of each center point to the upper edge box, the lower edge box, the left edge box and the right edge box of the corresponding label of the center point; calculating the balance degree of each center point by using the balance degree formula based on the distance of each center point to the upper edge box, the lower edge box, the left edge box and the right edge box of the corresponding label of the center point. classifying the center point based on the balance degree of each center point, comprising:
3. The method of claim 1, wherein, classifying the center point with the balance degree greater than a preset threshold as foreground; classifying the center point with the balance degree less than or equal to the preset threshold as background. regressing the coordinates of the detection box corresponding to each label in the training sample based on one or more center points classified as foreground corresponding to the label, comprising:
4. The method of claim 1, wherein, determining the center points classified as foreground within each label as the center points corresponding to the label; regressing the coordinates of the detection box corresponding to each label based on the box height and box width of the label and the coordinates of one or more center points corresponding to the label. regressing the coordinates of the detection box corresponding to each label in the training sample based on one or more center points classified as foreground corresponding to the label, comprising:
5. The method of claim 1, wherein, determining the center points classified as foreground within each label and the center points adjacent to the center points classified as foreground in the upper, lower, left and right four directions within each label as the center points corresponding to the label; regressing the coordinates of the detection box corresponding to each label based on the box height and box width of the label and the coordinates of one or more center points corresponding to the label. The application comprises the following steps:
6. A training device for a rebar detection model, characterized in that, The construction module is configured to replace a common convolution layer in a RetinaNet model with a hole convolution layer to obtain a steel bar detection model. The acquisition module is configured to acquire a steel bar training data set, input training samples in the steel bar training data set into the steel bar detection model, and output feature maps of the training samples, wherein the feature maps contain a plurality of center points. The classification module is configured to calculate a balance degree p of each center point by using a balance degree formula: wherein a distance of each center point to a left edge box of a corresponding label of the center point is l, a distance of each center point to a right edge box of the corresponding label of the center point is r, a distance of each center point to an upper edge box of the corresponding label of the center point is t, and a distance of each center point to a lower edge box of the corresponding label of the center point is b, classify the center point based on the balance degree of the center point, and train the steel bar detection model by using a first loss function in the process of classification, wherein a class of each center point includes foreground and background; the first loss function uses a focal loss function; The regression module is configured to regress coordinates of a detection box corresponding to each label in the training sample based on one or more center points classified as foreground corresponding to the label, wherein the training sample contains a plurality of labels, and train the steel bar detection model by using a second loss function in the process of regression; the second loss function uses a smooth_l1 function. The training module is configured to train the steel bar detection model by using a third loss function based on an intersection over union of each two adjacent detection boxes in all detection boxes in the training sample; and the third loss function is: wherein n is the total number of bounding boxes in the training sample, b i is the i-th bounding box, b j is the j-th bounding box, gt i is the label corresponding to the i-th bounding box, m is the total number of labels in the training sample, and IoU is an intersection over union function.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the steps of the method in any one of claims 1 to 5.
8. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 7. The computer program is executed by the processor to implement the steps of the method in any one of claims 1 to 5.
Citation Information
Patent Citations
Detection model training method and device, target detection method and device and electronic system
CN113239982A
Steel bar detection method based on target detection
CN115035372A