Method and System for Constructing Object Detection Model Based on Data Closed Loop

Through the construction method of object detection model based on data closed loop, using feature extraction rules and YOLOv5 network architecture, automated road defects and garbage detection are realized, time-consuming and labor-consuming problems of manual inspection, data collection efficiency and quality are improved, and cost is reduced.

CN115565156BActive Publication Date: 2025-07-04COWA TECHNOLOGY CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211233693.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-10
Publication Date
2025-07-04
Estimated Expiration
2042-10-10

AI Technical Summary

Technical Problem

In the prior art, road defects and garbage detection rely on manual inspection, resulting in time-consuming and labor-consuming, high detection rate, low data collection quality, long project iteration cycle and high cost.

Method used

The object detection model construction method based on data closed-loop is adopted, and feature extraction rules and YOLOv5 network architecture are used, combined with data acquisition equipment and operation equipment to realize real-time data acquisition and automated screening, forming a closed-loop data acquisition-model training-data acquisition process.

Benefits of technology

It reduces manual intervention, improves data collection efficiency and quality, reduces missed detection rates, shortens project cycles, reduces costs, and realizes automated road defects and garbage detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115565156B_ABST
    Figure CN115565156B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for constructing an object detection model based on a data closed-loop, including: formulating rules and determining an architecture in advance, and initializing dataset labels; training an object detection network to obtain an intermediate detection model, converting and accelerating the intermediate detection model to obtain an object detection model, collecting initial data, which is N frames of data to be detected, and initializing the data to be detected j as 1; using the object detection model to identify the j-th frame of data to be detected; using a feature extraction rule to identify the j-th frame of data to be detected, and judging the identification target. If it exists, store the j-th frame of data to be detected and the M1 frames of data to be detected before it and the M2 frames of data to be detected after it as the screened data; obtain the final object detection model and deploy it on a device for operation. The present invention automatically collects data, liberates manual labor, reduces costs, and improves driving safety; monitors the road surface in real time and saves pictures of interest, reduces the missed detection rate, and increases the data quantity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and specifically, to a method and system for constructing an object detection model based on a data closed-loop. Background Art

[0002] The urban road environment is an important part of building a beautiful city. The road environment mainly includes two indicators: one is the integrity of the road itself, and the other is the sanitation condition on the road. With the development of China's economy, the average household car ownership has been continuously increasing, which not only causes heavy traffic on the road, but also brings great pressure to the road environment. This pressure inevitably causes some pavement defects, such as cracks, damages, missing guiding lines, etc. In addition, the frequent use of the road also poses more stringent requirements for road sanitation management.

[0003] Currently, the process for road defects still follows the procedure of first manual inspection and then repair. However, manual inspection is not only time-consuming and laborious, but also results in a large number of missed detections. For the road sanitation condition, it depends on the real-time cleaning by sanitation workers and sanitation vehicles. Due to the decline in the number of sanitation human resources and the increase in costs, an automated solution is urgently needed. Therefore, it is imperative to use automated technologies represented by artificial intelligence to solve the detection and treatment of road defects and garbage.

[0004] As is well known, the realization of artificial intelligence depends on models built on data, which means that a large amount of road data collection becomes a necessary prerequisite. Currently, the collection of road defect and garbage data basically relies on manual methods. For example, cameras are set on vehicles, and once the driver discovers a target to be collected during driving, the driver presses a button to save a video segment of about 20 seconds before and after the current moment. After obtaining the video data, the data is then screened manually, and finally, the screened data is used for model training.

[0005] However, this method has the following disadvantages: When the driver collects data, due to the driving task, real-time collection cannot be achieved, and a large number of targets may be missed; the driver's willingness to cooperate with the collection task is low, which may lead to the absence of targets on the entire route; the driver does not have relevant business knowledge and is unclear about the targets to be collected, resulting in low-quality collected data; the collected data is many 20-second video segments, which will increase the workload of data screening. The above disadvantages will lead to a long cycle and high cost in the entire project iteration process.

[0006] Therefore, it is very necessary to apply automated and intelligent solutions to the data collection process. How to combine data collection and model construction more closely and intelligently is a task with high value and high challenges. Summary of the Invention

[0007] Aiming at the defects in the prior art, the purpose of the present invention is to provide a method and system for constructing an object detection model based on a data closed-loop.

[0008] A method for constructing an object detection model based on a data closed-loop provided by the present invention includes:

[0009] Step S1: For the target to be recognized, formulate a feature extraction rule in advance, determine the architecture of the object detection network, construct a first data set using public data, and initialize the data set label i to 1.

[0010] Step S2: Iteratively train the object detection network on the i-th data set. After the training is completed, obtain the i-th intermediate detection model, and perform conversion acceleration on the i-th intermediate detection model to obtain the i-th object detection model. If the mean average precision of the i-th object detection model is greater than the preset precision threshold, the i-th object detection model is the final object detection model and jump to step S9; otherwise, execute step S3.

[0011] Step S3: Deploy the i-th object detection model to the acquisition device, use the acquisition device to collect data to obtain initial data, and then convert the initial data into N frames of data to be detected, and initialize the label j of the data to be detected to 1.

[0012] Step S4: Use the i-th object detection model to identify the j-th frame of data to be detected, and determine whether the target to be recognized exists therein. If it exists, jump to step S6; otherwise, execute step S5.

[0013] Step S5: Use the feature extraction rule to identify the j-th frame of data to be detected, and determine whether the target to be recognized exists therein. If it exists, execute step S6; otherwise, jump to step S7.

[0014] Step S6: Store the j-th frame of data to be detected, the M1 frames of data to be detected before it, and the M2 frames of data to be detected after it as the filtered data.

[0015] Step S7: Increment the value of j by 1, and determine whether the new value of j is less than or equal to N. If so, jump back to step S4 and sequentially execute the relevant steps again; if not, execute step S8.

[0016] Step S8: Label all the filtered data and add it to the i-th data set to obtain the (i + 1)-th data set, then increment i by 1 and jump back to step S2.

[0017] Step S9: Deploy the final object detection model on the operation device to perform object detection operations.

[0018] Preferably, the feature extraction rule is designed according to the features of the recognition target, and the features include color, shape, and texture, which are used to extract the unique properties of the recognition target for differential recognition;

[0019] The feature extraction rule first filters out objects other than the road based on the curb to prevent misrecognition, and then executes corresponding strategies for different types of objects to be recognized to extract features;

[0020] The architecture of the object detection network is YOLOv5.

[0021] Preferably, the object detection network is iteratively trained on the i-th dataset, and the i-th intermediate detection model is obtained after training, including:

[0022] The i-th dataset is divided into the i-th training set and the i-th validation set;

[0023] The object detection network is iteratively trained on the i-th training set. Each complete iteration is regarded as one round, and a total of K rounds of training are performed;

[0024] After each round of training, the object detection network of the corresponding round is verified on the i-th validation set to obtain the mean average precision of the corresponding round;

[0025] The object detection network corresponding to the round with the highest mean average precision among the K rounds is selected as the i-th intermediate detection model.

[0026] Preferably, the conversion and acceleration of the i-th intermediate detection model to obtain the i-th object detection model includes:

[0027] The framework of the i-th intermediate detection model is converted from PyTorch to TensorRT;

[0028] The data precision of the converted model is reduced from FP32 to FP16 for quantization acceleration to obtain the i-th object detection model.

[0029] Preferably, the format of the initial data is video, and the format of the data to be detected is picture;

[0030] The acquisition device includes a manned or unmanned vehicle equipped with a camera, which collects the environmental video during driving in real time as the initial data;

[0031] The conversion of the initial data into N frames of data to be detected means that the video of the initial data is converted into consecutive frames of pictures using video processing tools as the data to be detected;

[0032] The operation device includes a manned or unmanned vehicle equipped with a camera;

[0033] The target detection operation includes: the operation device collects the environmental video during driving in real time through a camera; converts the environmental video into multiple frames of pictures to be detected; the final target detection model performs frame-by-frame recognition on the pictures to be detected, and when the target to be recognized is detected, uploads the corresponding frame picture and the corresponding position information to the cloud.

[0034] A target detection model construction system based on data closed-loop according to the present invention, which executes the method for constructing a target detection model based on data closed-loop, includes:

[0035] Pre-preparation module: For the target to be recognized, a feature extraction rule is formulated in advance, the architecture of the target detection network is determined, and a first data set is constructed using public data;

[0036] Training and verification module: Iteratively train the target detection network on the data set, obtain an intermediate detection model after training is completed, convert and accelerate the intermediate model to obtain a target detection model. If the mean average precision of the target detection model is greater than the preset precision threshold, then the target detection model is the final target detection model;

[0037] Collection and conversion module: Deploy the target detection model to the collection device, use the collection device to collect data to obtain initial data, and then convert the initial data into data to be detected;

[0038] Recognition and screening module: Use the target detection model and the feature extraction rule to perform frame-by-frame recognition on the data to be detected. If the target to be recognized is detected in a certain frame of data, store the frame of data and the preset number of frames of data before and after it as the screened data;

[0039] Data update module: Label all the screened data and add it to the original data set to obtain a new data set;

[0040] Detection operation module: Deploy the final target detection model on the operation device to perform target detection operations.

[0041] Preferably, the feature extraction rule is designed according to the features of the recognized target, and the features include color, shape, and texture, which are used to extract the unique properties of the recognized target for discrimination and recognition;

[0042] The feature extraction rule first filters out objects other than the road based on the road edge to prevent misrecognition, and then executes corresponding strategies for different types of targets to be recognized for feature extraction;

[0043] The architecture of the target detection network is YOLOv5.

[0044] Preferably, the implementation process of the training and verification module includes:

[0045] Divide the said data set into a training set and a validation set;

[0046] Iteratively train the target detection network on the said training set. Each complete iteration is regarded as one round, and a total of K rounds of training are performed;

[0047] After each round of training, validate the target detection network of the corresponding round on the said validation set to obtain the mean average precision (mAP) of the corresponding round;

[0048] Select the target detection network corresponding to the round with the highest mAP among the K rounds as the intermediate detection model;

[0049] Convert the framework of the said intermediate detection model from PyTorch to TensorRT;

[0050] Reduce the data precision of the converted model from FP32 to FP16 for quantization acceleration to obtain the target detection model;

[0051] Judge whether the mAP of the said target detection model is greater than a preset precision threshold. If so, end the data collection and model training work, and the said target detection model is the final target detection model. If not, continue the data collection and model training work.

[0052] Preferably, the format of the initial data is video, and the format of the data to be detected is picture;

[0053] The said acquisition device includes a manned or unmanned vehicle equipped with a camera, which collects the environmental video during driving in real time as the initial data;

[0054] The conversion of the said initial data into N frames of data to be detected means converting the video of the initial data into consecutive frames of pictures using video processing tools as the data to be detected.

[0055] Preferably, the said operation device includes a manned or unmanned vehicle equipped with a camera;

[0056] The said target detection operation includes: the operation device collects the environmental video during driving in real time through the camera; converts the environmental video into multiple frames of pictures to be detected; the final target detection model performs frame-by-frame recognition on the pictures to be detected, and when the target to be recognized is detected, uploads the corresponding frame of picture and the corresponding position information to the cloud.

[0057] Compared with the prior art, the present invention has the following beneficial effects:

[0058] 1. It saves driver labor, reduces costs, and improves driving safety;

[0059] 2. The provided solution monitors the road surface in real time and saves the images of interest, greatly reducing the missed detection rate and increasing the quantity of collected data;

[0060] 3. Directly utilize the learning ability of the model to obtain a large amount of data that the model is prone to misdetect and has low confidence, and can also filter out the data that is easy to learn, greatly reducing the data screening cost, improving the data training efficiency, and enhancing the quality of the collected data; on the other hand, rules for identifying relevant target features can be formulated to assist the model in identification, obtaining the targets that the model is prone to miss, making the data distribution more comprehensive;

[0061] 4. The target detection model obtained by using the present invention is the initial version for the final project deployment (real-time online identification of road defects and garbage through the model), which can effectively verify the feasibility of the project and avoid useless work;

[0062] 5. The data obtained by automatic collection is sent for model training, and then the model in turn improves the efficiency of the collection system. The data collection - model training - data collection process forms a closed loop, greatly accelerating the iteration speed, shortening the project cycle, and reducing the project cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Other features, objectives, and advantages of the present invention will become more apparent by reading the following detailed description of the non-limiting embodiments with reference to the accompanying drawings:

[0064] Figure 1 It is a flowchart of an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0065] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that those of ordinary skill in the art can make several changes and improvements without departing from the concept of the present invention. These all belong to the protection scope of the present invention.

[0066] Embodiment:

[0067] According to a method for constructing a target detection model based on a data closed loop provided by the present invention, as Figure 1 shown, it includes:

[0068] Step S1: For the target to be identified, pre-formulate feature extraction rules, determine the architecture of the target detection network, construct a first data set using public data, and initialize the data set label i as 1;

[0069] Step S2: Iteratively train the target detection network on the i-th dataset. After the training is completed, obtain the i-th intermediate detection model, and perform conversion acceleration on the i-th intermediate detection model to obtain the i-th target detection model. If the mean average precision of the i-th target detection model is greater than the preset precision threshold, then the i-th target detection model is the final target detection model and jump to Step S9; otherwise, execute Step S3.

[0070] Step S3: Deploy the i-th target detection model to the acquisition device, use the acquisition device to collect data to obtain initial data, then convert the initial data into N frames of data to be detected, and initialize the label j of the data to be detected as 1.

[0071] Step S4: Use the i-th target detection model to identify the j-th frame of data to be detected, and determine whether the target to be identified exists therein. If it exists, jump to Step S6; otherwise, execute Step S5.

[0072] Step S5: Use the feature extraction rule to identify the j-th frame of data to be detected, and determine whether the target to be identified exists therein. If it exists, execute Step S6; otherwise, jump to Step S7.

[0073] Step S6: Store the j-th frame of data to be detected, the M1 frames of data to be detected before it, and the M2 frames of data to be detected after it as the screened data.

[0074] Step S7: Increment the value of j by 1, and determine whether the new value of j is less than or equal to N. If so, jump back to Step S4 and sequentially execute the relevant steps again; if not, execute Step S8.

[0075] Step S8: After labeling all the screened data, add it to the i-th dataset to obtain the i+1-th dataset, then increment i by 1 and jump back to Step S2.

[0076] Step S9: Deploy the final target detection model on the operation device to perform target detection operations.

[0077] In this embodiment, the feature extraction rule is designed according to the features of the target to be identified. The features include but are not limited to color, shape, texture, etc., and are used to extract the unique properties of the target to be identified for differential identification.

[0078] The feature extraction rule first filters out objects other than the road based on the road edge to prevent misidentification, and then executes corresponding strategies for feature extraction for different types of targets to be identified.

[0079] The architecture of the target detection network is YOLOv5.

[0080] In one embodiment, M1 and M2 are not equal. When j - M1 < 0, that is, the number of data frames before the j-th frame of data to be detected is less than M1 frames, then all the data before the j-th frame of data to be detected and the M2 frames of data to be detected after it are extracted and stored as the filtered data.

[0081] When j + M2 > N, that is, the number of data frames after the j-th frame of data to be detected is less than M2 frames, then all the data after the j-th frame of data to be detected and the M1 frames of data to be detected before it are extracted and stored as the filtered data.

[0082] Iteratively training the target detection network on the i-th data set and obtaining the i-th intermediate detection model after training completion includes:

[0083] Dividing the i-th data set into the i-th training set and the i-th validation set;

[0084] Iteratively training the target detection network on the i-th training set, with each complete iteration being one round, and a total of K rounds of training;

[0085] After each round of training, validating the target detection network of the corresponding round on the i-th validation set to obtain the mean average precision of the corresponding round;

[0086] Selecting the target detection network corresponding to the round with the highest mean average precision among the K rounds as the i-th intermediate detection model.

[0087] In this embodiment, the above iterative training of the target detection network on the data set and obtaining the intermediate detection model after training completion is specifically that the iterative training is preset to 300 rounds, with one complete iteration of the entire data set being one round. After each round of training, validation is performed on the validation set to obtain the evaluation metric mAP of this iteration round. When the 300-round iterative training is completed, the target detection network corresponding to the round with the best evaluation metric mAP is selected as the intermediate detection model.

[0088] The explanations of relevant terms are as follows:

[0089] mAP: An evaluation metric of a mainstream target detection model, used to measure the model performance; the higher the mAP, the better the model performance.

[0090] Converting and accelerating the i-th intermediate detection model to obtain the i-th target detection model includes:

[0091] Converting the framework of the i-th intermediate detection model from PyTorch to TensorRT; reducing the data precision of the converted model from FP32 to FP16 for quantization acceleration to obtain the i-th target detection model.

[0092] In this embodiment, the framework of the intermediate detection model is converted from PyTorch to TensorRT, and the data precision is reduced from FP32 to FP16. The purpose is to accelerate through quantization, improve the recognition speed of the model, greatly shorten the project cycle, and thus reduce the project cost.

[0093] The explanations of relevant terms are as follows:

[0094] Pytorch: A scientific computing package based on Python, providing two advanced functions: powerful GPU-accelerated tensor computing (such as NumPy); a deep neural network containing an automatic differentiation system. It is mainly used to build deep learning models.

[0095] TensorRT: A high-performance deep learning inference optimizer, a library launched by NVIDIA for accelerating model inference on large-scale data centers, embedded platforms, or autonomous driving platforms.

[0096] FP32 / FP16: FP16 refers to a data type encoded and stored using 2 bytes (16 bits); similarly, FP32 refers to using 4 bytes (32 bits). Compared with FP32, FP16 in TensorRT can achieve nearly twice the speed improvement.

[0097] The format of the initial data is video, and the format of the data to be detected is picture;

[0098] The acquisition device includes a manned or unmanned vehicle equipped with a camera, which collects the environmental video during driving in real time as the initial data;

[0099] Converting the initial data into N frames of data to be detected means using a video processing tool to convert the video of the initial data into consecutive frames of pictures as the data to be detected.

[0100] In this embodiment, the format of the initial data is RTSP stream, which cannot be directly applied to model recognition. Through an existing video processing tool, this embodiment uses FFmpeg to convert it into a format of consecutive frames of pictures, and finally uses it as the data to be detected for model recognition and detection.

[0101] The explanations of relevant terms are as follows:

[0102] RTSP stream: A protocol for real-time stream transmission, often used in video surveillance services;

[0103] FFmpeg: An open-source computer program that can be used to record, convert digital audio and video, and convert them into streams. It has very powerful functions including video acquisition, video format conversion, video capture, etc.

[0104] The operation device includes a manned or unmanned vehicle equipped with a camera;

[0105] The target detection operation includes: the operation device collects the environmental video during driving in real time through the camera; converts the environmental video into multiple frames of pictures to be detected; the final target detection model identifies each frame of the pictures to be detected one by one, and when the target to be identified is detected, uploads the corresponding frame of picture and the corresponding position information to the cloud.

[0106] A target detection model construction system based on data closed-loop includes:

[0107] A pre-preparation module: for the target to be identified, formulate a feature extraction rule in advance, determine the architecture of the target detection network, and construct a first data set using public data;

[0108] A training and verification module: iteratively train the target detection network on the data set, obtain an intermediate detection model after training is completed, convert and accelerate the intermediate model to obtain a target detection model. If the mean average precision of the target detection model is greater than a preset precision threshold, then the target detection model is the final target detection model;

[0109] A collection and conversion module: deploy the target detection model to the collection device, use the collection device to collect data to obtain initial data, and then convert the initial data into data to be detected;

[0110] An identification and screening module: respectively use the target detection model and the feature extraction rule to identify each frame of the data to be detected one by one. As long as one of them detects the target to be identified in a certain frame of data, store the frame of data and several frames of data before and after it as the screened data;

[0111] A data update module: label all the screened data and add it to the original data set to obtain a new data set;

[0112] A detection operation module: deploy the final target detection model on the operation device to perform target detection operations.

[0113] The feature extraction rule is manually designed according to the features of the target to be identified including color, shape, texture, etc., and is used to extract the unique properties of the target to be identified for differential identification;

[0114] The feature extraction rule first filters out objects other than roads based on the road edge to prevent misidentification, and then executes corresponding strategies for different types of targets to be identified for feature extraction.

[0115] The architecture of the target detection network is YOLOv5.

[0116] The implementation process of the training and validation module includes:

[0117] Divide the dataset into a training set and a validation set;

[0118] Iteratively train the object detection network on the training set. Each complete iteration is regarded as one round, and a total of K rounds of training are performed;

[0119] After each round of training, validate the object detection network of the corresponding round on the validation set to obtain the mean average precision (mAP) of the corresponding round;

[0120] Select the object detection network corresponding to the round with the highest mAP among the K rounds as the intermediate detection model;

[0121] Convert the framework of the intermediate detection model from PyTorch to TensorRT;

[0122] Reduce the data precision of the converted model from FP32 to FP16 for quantization acceleration to obtain the object detection model;

[0123] Judge whether the mAP of the object detection model is greater than the preset precision threshold. If so, end the data collection and model training work, and the object detection model is the final object detection model. If not, continue the data collection and model training work.

[0124] The format of the initial data is video, and the format of the data to be detected is picture;

[0125] The acquisition device includes a manned or unmanned vehicle equipped with a camera, which collects the environmental video during driving in real time as the initial data;

[0126] Converting the initial data into N frames of data to be detected means using video processing tools to convert the video of the initial data into consecutive frames of pictures as the data to be detected.

[0127] The operation device includes a manned or unmanned vehicle equipped with a camera;

[0128] The object detection operation includes: the operation device collects the environmental video during driving in real time through the camera; converts the environmental video into multiple frames of pictures to be detected; the final object detection model performs frame-by-frame recognition on the pictures to be detected, and when the target to be recognized is detected, uploads the corresponding frame of picture and the corresponding position information to the cloud.

[0129] Those skilled in the art know that, in addition to implementing the systems, devices, and their respective modules provided by the present invention in the form of pure computer-readable program code, it is entirely possible to make the systems, devices, and their respective modules provided by the present invention be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. by logically programming the method steps. Therefore, the systems, devices, and their respective modules provided by the present invention can be considered as a kind of hardware component, and the modules included therein for implementing various programs can also be regarded as the structures within the hardware component; the modules for implementing various functions can also be regarded as either software programs for implementing the method or the structures within the hardware component.

[0130] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.

Claims

1. A method for constructing an object detection model based on a data closed-loop, characterized in that Including: Step S1: For the target to be recognized, pre - formulate feature extraction rules, determine the architecture of the target detection network, construct a first data set using public data, and initialize the data set label i as 1; Step S2: Iteratively train the target detection network on the i dataset. After the training is completed, obtain the i intermediate detection model. Convert and accelerate the i intermediate detection model to obtain the i target detection model. If the mean average precision of the i target detection model is greater than the preset precision threshold, then the i target detection model is the final target detection model and jump to Step S9; otherwise, execute Step S3. Step S3: Deploy the i target detection model to the acquisition device, use the acquisition device to collect data to obtain initial data, and then convert the initial data into N frame data to be detected, and initialize the label of the data to be detected j as 1; Step S4: Use the i target detection model to identify the j frame of data to be detected, and determine whether the target to be identified exists therein. If it exists, jump to step S6; otherwise, execute step S5. Step S5: Use the feature extraction rule to identify the j frame of data to be detected, and determine whether the target to be identified exists therein. If it exists, execute Step S6; otherwise, jump to Step S7. Step S6: Store the j frame of data to be detected and the M 1 frame of data to be detected before it and the M 2 frames of data to be detected after it as the filtered data; Step S7: Increment the value of j by 1, and check if the new j value is less than or equal to N . If so, jump back to step S4 and re-execute the relevant steps in sequence; if not, execute step S8. Step S8: After annotating all the filtered data, add it to the i dataset to obtain the i +(1) dataset, then i increment the value by 1 and jump back to Step S2; Step S9: Deploy the final target detection model on the operation device to perform target detection operations; The format of the initial data is video, and the format of the data to be detected is picture; The acquisition device includes a manned or unmanned vehicle equipped with a camera, which collects the environmental video during driving in real time as the initial data; Said converting the initial data into N frame data to be detected means converting the video of the initial data into consecutive frame pictures by using a video processing tool as the data to be detected; The operation device includes a manned or unmanned vehicle equipped with a camera; The target detection operation includes: the operation device collects the environmental video during driving in real time through the camera; converts the environmental video into multiple frames of data to be detected; the final target detection model performs frame-by-frame recognition on the data to be detected, and when the target to be recognized is detected, uploads the corresponding frame of picture and the corresponding position information to the cloud.

2. The method for constructing a target detection model based on a data closed loop according to claim 1, wherein The feature extraction rule is designed according to the features of the target to be recognized, and the features include color, shape, and texture, which are used to extract the unique properties of the target to be recognized for discrimination and recognition; The feature extraction rule first filters out objects other than the road based on the road edge to prevent misrecognition, and then executes corresponding strategies for different types of targets to be recognized for feature extraction; The architecture of the target detection network is YOLOv5.

3. The method for constructing a target detection model based on a data closed-loop according to claim 1, wherein, The above-mentioned i iterative training of the target detection network is performed on the dataset, and after the training is completed, the i intermediate detection model is obtained. Including: Divide the i data set into a i training set and a i validation set; On the i training set, iterative training is performed on the object detection network. Each complete iteration is regarded as one round, and a total of K rounds are trained; After each round of training, on the i validation set, the object detection network of the corresponding round is verified to obtain the mean average precision of the corresponding round; Select K The object detection network corresponding to the round with the highest average precision mean in the round is used as the i intermediate detection model.

4. The method for constructing a target detection model based on a data closed-loop according to claim 1, wherein The conversion and acceleration of the i intermediate detection model yields the i target detection model, including: Convert the framework of the i intermediate detection model from PyTorch to TensorRT; Reduce the data precision of the converted model from FP32 to FP16 for quantization acceleration to obtain the i object detection model.

5. A target detection model construction system based on a data closed-loop, characterized in that, Execute the method for constructing a target detection model based on data closed-loop according to claim 1, including: Pre-preparation module: For the target to be recognized, formulate a feature extraction rule in advance, determine the architecture of the target detection network, and construct a first data set using public data; Training and verification module: Iteratively train the target detection network on the data set. After training is completed, obtain an intermediate detection model, convert and accelerate the intermediate detection model to obtain a target detection model. If the mean average precision of the target detection model is greater than the preset precision threshold, then the target detection model is the final target detection model; Acquisition and conversion module: Deploy the target detection model to the acquisition device, use the acquisition device to collect data to obtain initial data, and then convert the initial data into data to be detected; Recognition and screening module: Use the target detection model and the feature extraction rule to perform frame-by-frame recognition on the data to be detected respectively. If the target to be recognized is detected in a certain frame of data, store the frame of data and the data of each preset number of frames before and after it as the screened data; Data update module: Label all the screened data and add it to the original data set to obtain a new data set; Detection operation module: Deploy the final target detection model on the operation device to perform target detection operations.

Citation Information

Patent Citations

  • Picture recognition model acquisition method and device, electronic equipment and storage medium

    CN113392886A

  • Marine organism intelligent detection and identification method

    CN114596584A