A method and terminal for constructing image recognition model in safety supervision scenario
By building a front-layer network model and training an all-round network model on the initial network model, the problem of slow training of image recognition model in the security monitoring scenario in the existing technology is solved, the high migration ability and low computing resource requirements of the model are achieved, and the security monitoring efficiency is improved.
Patent Information
- Application Number
- CN202311242883.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-25
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2043-09-25
AI Technical Summary
In the prior art In the training process of image recognition models in security monitoring scenarios, the input of new samples needs to be retrained in the shared feature network, resulting in slow overall training progress.
The method of building a front-layer network model and training an all-round network model on the initial network model is adopted. By locking the all-round network model and splicing the front-layer network model in front of it, the demand for training data for specific scenarios is reduced and fine-tuned to adapt to new security monitoring scenarios.
It improves the migration capability of the model, reduces the demand for computing resources, and reduces the dependence on training data in specific scenarios, and improves the efficiency of infrastructure violation monitoring and security monitoring during power operations.
Smart Images

Figure CN117173636B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electric power operation, and in particular to a method and a terminal for constructing an image recognition model in a safety supervision scenario. Background Art
[0002] It is very important to monitor infrastructure violations or security during power operations. Currently, the safety images taken on site in real time are usually processed based on common models to output recognition results to achieve real-time safety monitoring of power operations. For example, the Chinese patent "An image management method and device based on a multi-task machine learning model" with publication number CN111813532A discloses a shared feature expression network input into a multi-task machine learning model to obtain task output features; and the task output features are respectively input into multiple sub-task networks to obtain recognition results; and then the image data is processed based on the recognition results.
[0003] That is, by sharing the feature network among multiple sub-task models, multiple tasks can achieve task separation and reduce the model size. However, during the model training process, the input of new samples often requires retraining the shared feature network, resulting in slow overall training progress. Summary of the invention
[0004] The technical problem to be solved by the present invention is to provide a method and a terminal for constructing an image recognition model in a safety and supervision scenario, thereby reducing computing resources and improving migration capabilities.
[0005] In order to solve the above technical problems, the technical solution adopted by the present invention is:
[0006] A method for constructing an image recognition model in a safety supervision scenario comprises the following steps:
[0007] S1, build the front layer network model;
[0008] S2, build the initial network model;
[0009] S3, using multiple types of data sets to train the initial network model to obtain a universal network model;
[0010] S4, locking the universal network model, and splicing the front-layer network model in front of the universal network model;
[0011] S5. Fine-tune the spliced universal network model to obtain a final image recognition model.
[0012] In order to solve the above technical problems, another technical solution adopted by the present invention is:
[0013] A terminal for building an image recognition model in a safety and supervision scenario, comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps in the above-mentioned method for building an image recognition model in a safety and supervision scenario when executing the computer program.
[0014] The beneficial effects of the present invention are as follows: the present invention provides a method and terminal for constructing an image recognition model in a safety and supervision scenario, by first constructing a front-layer network model and using a large number of safety and supervision data sets of various categories to train an all-purpose network model, and then locking the all-purpose network model and splicing the front-layer network model to the front of the all-purpose network model. Since the all-purpose network model has been trained with a large amount of data and has the ability to label various feature objects, it is possible to migrate these excellent labeling capabilities to various safety and supervision scenarios through the front-layer network, adapt to the image processing needs of different security monitoring, power operations and other fields, and have strong migration capabilities; at the same time, when modifying the method, there is no need to retrain the entire model, only the front-layer network model needs to be modified, which greatly reduces the demand for training data for specific scenarios, and the front-layer network model only needs to be fine-tuned and can be directly applied to new safety and supervision scenarios, greatly reducing computing resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 A flowchart of a method for constructing an image recognition model in a safety monitoring scenario according to an embodiment of the present invention;
[0016] Figure 2 The present invention is a schematic diagram of a terminal for constructing an image recognition model in a safety monitoring scenario according to an embodiment of the present invention.
[0017] Description of labels:
[0018] 1. A terminal for building an image recognition model in a safety and supervision scenario; 2. Memory; 3. Processor. DETAILED DESCRIPTION
[0019] In order to explain the technical content, achieved objectives and effects of the present invention in detail, the following is an explanation in combination with the implementation modes and the accompanying drawings.
[0020] Please refer to Figure 1 , a method for constructing an image recognition model in a safety supervision scenario, comprising the steps of:
[0021] S1, build the front layer network model;
[0022] S2, build the initial network model;
[0023] S3, using multiple types of data sets to train the initial network model to obtain a universal network model;
[0024] S4, locking the universal network model, and splicing the front-layer network model in front of the universal network model;
[0025] S5. Fine-tune the spliced universal network model to obtain a final image recognition model.
[0026] From the above description, it can be seen that the beneficial effects of the present invention are: by first constructing a front-layer network model and using a large number of safety and supervision data sets of various categories to train the universal network model, and then locking the universal network model and splicing the front-layer network model to the front of the universal network model, since the universal network model has been trained with a large amount of data and has the ability to label various feature objects, it is possible to migrate these excellent labeling capabilities to various safety and supervision scenarios through the front-layer network, adapt to different image processing needs in security monitoring, power operations and other fields, and have strong migration capabilities; at the same time, when modifying the method, there is no need to retrain the entire model, only the front-layer network model needs to be modified, which greatly reduces the demand for training data for specific scenarios, and the front-layer network model only needs to be fine-tuned and can be directly applied to new safety and supervision scenarios, greatly reducing computing resources.
[0027] Furthermore, the step S1 is specifically as follows:
[0028] A front-layer network model LowTastNet consisting of two convolutional network layers is constructed, and a special operator layer is added between the two convolutional network layers. The operator of the special operator layer is as follows:
[0029] 0.3*sign(x)+0.7*relu(x)(1);
[0030] Where x is a real number, the value of sign(x) is 1, -1 or 0, indicating the positive or negative value of x and whether it is 0, respectively, and relu(x) = max(0, x), indicating the activation function.
[0031] Furthermore, the step S1 further includes:
[0032] The front layer network model is defined to include three layers from 0 to 2, wherein the 0th layer and the 2nd layer are convolutional network layers, and the 1st layer is a special operator layer;
[0033] The operators defining the two convolutional network layers are both Conv2d.
[0034] From the above description, we can see that a special operator layer is added between the two convolutional network layers of the front-layer network model to enable the front-layer network to have over-migration capability and diversity.
[0035] Furthermore, the step S2 is specifically as follows:
[0036] Construct an initial network model GenpNett consisting of five single convolutional network layers, one matrix convolutional network layer, three maximum pooling layers and a general detection head.
[0037] Furthermore, the step S2 further includes:
[0038] The initial network model is defined to include ten layers from 0 to 9, wherein five single convolutional network layers are respectively located at the 0th, 2nd, 4th, 5th and 6th layers of the initial network model, one matrix convolutional network layer is located at the 3rd layer of the initial network model, three maximum pooling layers are respectively located at the 1st, 7th and 8th layers of the initial network model, and one universal detection head is located at the 9th layer of the initial network model;
[0039] The operators defining five single convolutional network layers and one matrix convolutional network layer are all Conv2d;
[0040] The operators defining the three maximum pooling layers are all maxpool;
[0041] Define a general detection head operator as Detect.
[0042] From the above description, we can see that constructing the initial network model GenpNet can detect common things, is compatible with most safety and supervision scenarios, and is also convenient for subsequent use of various categories of safety and supervision data sets to train a universal network model.
[0043] Furthermore, the multiple categories of data sets in step S3 include COCO data set, Pascal VOC data set, ImageNet data set, KITTI data set, Open Images data set, SUN data set, Cityscapes data set, DOTA data set and VisDrone data set.
[0044] Furthermore, the step S3 further includes:
[0045] The data set for training the universal network model merges labels with the same semantics according to different practical classifications, and labels with different semantics represent a classification box.
[0046] From the above description, we can know that the COCO dataset is a large-scale image detection, semantic segmentation and image caption generation dataset, containing more than 330K images, 1.5 million objects, 80 object categories and 91 material categories; the Pascal VOC dataset is one of the foundations of object detection technology, containing 20 categories and 11,530 images for training and verification; the ImageNet dataset is a large-scale visualization database for visual object recognition software research, containing more than 14 million image URLs manually annotated by ImageNet; the KITTI dataset was jointly established by the Karlsruhe Institute of Technology and Toyota America Technical Research Institute, and is currently the world's largest computer vision algorithm evaluation dataset for autonomous driving scenarios; Open The Images dataset is an open source image dataset released by Google. Its latest version contains more than 9 million images, all of which are labeled with categories. The SUN dataset is a dataset used to train and evaluate image super-resolution algorithms. It contains 817 high-resolution images from different types, such as architecture, landscape, nature, etc. The Cityscapes dataset has 5,000 images of driving scenes in urban environments. The DOTA dataset contains 2,806 aerial images with a size of approximately 4k*4k, including 15 categories and a total of 188,282 instances. The VisDrone dataset is a large-scale dataset for drone visual recognition. It includes a wealth of scenes and object categories, with the aim of promoting the development of drone vision and intelligence. These multi-category datasets contain high-resolution images, videos, annotations, and metadata, which can provide a variety of image data samples for training the all-round network model, covering various safety and protection scenarios including power operation scenarios, forming sample-rich training sets, test sets, and verification sets, so as to improve the final trained all-round network model to be suitable for most safety and supervision scenarios, reduce the cost of large amounts of data collection and labeling when the model is subsequently used for target detection and security monitoring operations, and at the same time reduce dependence on high computing resources, effectively save data and computing resources, and improve the efficiency of infrastructure violation monitoring and security monitoring during power operations.
[0047] Furthermore, the step S4 is specifically as follows:
[0048] S41, when training the initial network model, setting the parameter weights of the initial network model not to be modifiable based on gradients, but enabling gradient propagation, and locking the universal network model;
[0049] S42, splicing the front layer network model in front of the locked all-purpose network model, inputting the image feature F0 into the front layer network to obtain feature F1, and then inputting F1 into the all-purpose network model, and the all-purpose network model is responsible for outputting classification and box.
[0050] From the above description, it can be seen that when a large number of key data sets are used to train the universal network model, the constructed initial network model is first locked to ensure that there is no need to correct the weights of the general universal network model in the subsequent fine-tuning stage, which greatly reduces computing resources; at the same time, by directly splicing the front-layer network model before the universal network model, it only needs to modify the front-layer network model in the future and can be directly applied to new safety and supervision scenarios, thereby further improving the efficiency of infrastructure violation monitoring and security monitoring during power operations.
[0051] Furthermore, the fine-tuning process in step S5 is as follows:
[0052] The learning rate of the spliced universal network model is controlled, and the model is tested and verified using real-time safety monitoring data until the test set error and the verification set error no longer decrease.
[0053] From the above description, it can be seen that by finally testing and training the model based on a small amount of safety and supervision data in a specific scenario, ensuring that the errors of the test set and the validation set remain stable, it can be directly applied to new safety and supervision scenarios, further improving the efficiency of infrastructure violation monitoring and security monitoring during power operations.
[0054] Please refer to Figure 2 A terminal for building an image recognition model in a safety and supervision scenario includes a processor, a memory, and a computer program stored in the memory and executable on the processor. The terminal is characterized in that when the processor executes the computer program, the steps in the above-mentioned method for building an image recognition model in a safety and supervision scenario are implemented.
[0055] From the above description, it can be seen that the beneficial effects of the present invention are: based on the same technical concept, in conjunction with the above-mentioned method for constructing an image recognition model in a safety and supervision scenario, a terminal for constructing an image recognition model in a safety and supervision scenario is provided, by first constructing a front-layer network model and using a large number of safety and supervision data sets of various categories to train the all-round network model, and then locking the all-round network model and splicing the front-layer network model to the front of the all-round network model. Since the all-round network model has been trained with a large amount of data and has the ability to label various feature objects, it is possible to migrate these excellent labeling capabilities to various safety and supervision scenarios through the front-layer network, adapt to different image processing needs in security monitoring, power operations and other fields, and have strong migration capabilities; at the same time, when modifying the method, there is no need to retrain the entire model, only the front-layer network model needs to be modified, which greatly reduces the demand for training data for specific scenarios, and the front-layer network model only needs to be fine-tuned and can be directly applied to new safety and supervision scenarios, greatly reducing computing resources.
[0056] The present invention provides a method and terminal for constructing an image recognition model in a safety monitoring scenario, which is applied to infrastructure violation monitoring and security monitoring in power operation scenarios, and is described in detail in the following specific embodiments.
[0057] Please refer to Figure 1 , Embodiment 1 of the present invention is:
[0058] A method for constructing an image recognition model in a safety supervision scenario, such as Figure 1 As shown, the steps include:
[0059] S1. Build the front-layer network model, specifically:
[0060] A front-layer network model LowTastNet consisting of two convolutional network layers is constructed, and a special operator layer is added between the two convolutional network layers. The operator of the special operator layer is as follows (1):
[0061] 0.3*sign(x)+0.7*relu(x)(1).
[0062] In the above formula, x is a real number, the value of sign(x) is 1, -1 or 0, indicating the positive or negative value of x and whether it is 0, respectively, and relu(x) = max(0, x), which represents the activation function.
[0063] That is, a special operator layer is added between the two convolutional network layers of the front-layer network model to enable the front-layer network to have cross-transfer capability and diversity.
[0064] The structure of the front-layer network model constructed in this embodiment and the parameters of each layer are shown in Table 1 below:
[0065] Table 1: LowTastNet
[0066] layer Operator aisle size enter Output 0 Conv2d 20 3*3*1 416*416*3 416*416*20 1 0.3*sign(x)+0.7*relu(x) 416*416*20 416*416*20 2 Conv2d 3 3*3*1 416*416*20 416*416*3
[0067] As shown in Table 1, the front-layer network model is defined to include three layers from 0 to 2, where layers 0 and 2 are convolutional network layers, and layer 1 is a special operator layer.
[0068] At the same time, the operators of the two convolutional network layers are defined as Conv2d, where the channel of the convolutional network layer of layer 0 is 20, the convolution kernel size is 3*3*1, the input is 416*416*3, and the output is 416*416*20; the input and output of the special operator of layer 1 are both 416*416*20; the channel of the convolutional network layer of layer 2 is defined as 3, the convolution kernel size is 3*3*1, the input is 416*416*20, and the output is 416*416*3.
[0069] In this embodiment, (416*416) in the input and output represents the feature map resolution, and 3 and 20 are channels.
[0070] S2. Build the initial network model, specifically:
[0071] Construct the initial network model GenpNett consisting of five single convolutional network layers, a matrix convolutional network layer containing (3*5), three maximum pooling layers and a universal detection head.
[0072] That is, to build the initial network model GenpNet, which can detect common things, is compatible with most safety and supervision scenarios, and is also convenient for subsequent use of various categories of safety and supervision data sets to train the all-round network model.
[0073] The structure of the initial network model constructed in this embodiment and the parameters of each layer are shown in Table 2 below:
[0074] Table 2: GenpNett
[0075]
[0076]
[0077] As shown in Table 2, the initial network model is defined to include ten layers from 0 to 9, of which five single convolutional network layers are located at the 0th, 2nd, 4th, 5th and 6th layers of the initial network model, one matrix convolutional network layer is located at the 3rd layer of the initial network model, three maximum pooling layers are located at the 1st, 7th and 8th layers of the initial network model, and one universal detection head is located at the 9th layer of the initial network model.
[0078] At the same time, the operators of five single convolutional network layers and one matrix convolutional network layer are all Conv2d, the channels are all 32, the convolution kernel sizes are all 3*3*1, the input of the 0th convolutional network layer is 416*416*3, the output is 416*416*32, and the input and output of other convolutional network layers are all 104*104*32; the operators of three maximum pooling layers are all maxpool, the convolution kernel sizes of the 1st and 7th maximum pooling layers are both 4*4 / 2, and the convolution kernel size of the 8th maximum pooling layer is 2*2 / 2, the input of the 1st maximum pooling layer is 416*416*32, and the output is 104*104*32, the input of the 7th maximum pooling layer is 104*104*32, and the output is 26*26*32, and the input of the 8th maximum pooling layer is 26*26*32, and the output is 13*13*32.
[0079] In this embodiment, (416*416), (104*104), (26*26) and (13*13) in the input and output represent feature map resolutions, and 3 and 32 represent channels.
[0080] Then define a universal detection head operator as Detect, and the output is (5018+5)*13*13, where 5018 represents the number of categories, 5 represents the four dimensions and a confidence level of the target box, and (13*13) represents the feature map resolution.
[0081] S3. Use multiple categories of data sets to train the initial network model to obtain a universal network model.
[0082] S4. Lock the universal network model and splice the front-layer network model in front of the universal network model.
[0083] S5. Fine-tune the spliced universal network model to obtain the final image recognition model.
[0084] That is, in this embodiment, by first constructing a front-layer network model and using a large number of safety and supervision data sets of various categories to train the universal network model, and then locking the universal network model and splicing the front-layer network model to the front of the universal network model, since the universal network model has been trained with a large amount of data and has the ability to label various feature objects, it is possible to migrate these excellent labeling capabilities to various safety and supervision scenarios through the front-layer network, adapt to different image processing needs in security monitoring, power operations and other fields, and have strong migration capabilities; at the same time, there is no need to retrain the entire model when modifying the method, only the front-layer network model needs to be modified, which greatly reduces the demand for training data for specific scenarios, and the front-layer network model only needs to be fine-tuned and can be directly applied to new safety and supervision scenarios, greatly reducing computing resources.
[0085] Embodiment 2 of the present invention is:
[0086] A terminal for building an image recognition model in a safety and supervision scenario. Based on the above-mentioned embodiment 1, in this embodiment, the multiple categories of data sets in step S3 include a COCO data set, a Pascal VOC data set, an ImageNet data set, a KITTI data set, an Open Images data set, a SUN data set, a Cityscapes data set, a DOTA data set, and a VisDrone data set.
[0087] Among them, the COCO dataset is a large-scale image detection, semantic segmentation, and image caption generation dataset, containing more than 330K images, 1.5 million objects, 80 object categories, and 91 material categories.
[0088] The Pascal VOC dataset is one of the foundations of target detection technology. It contains 20 categories and 11,530 images for training and verification.
[0089] The ImageNet dataset is a large-scale visual database used for visual object recognition software research, containing more than 14 million image URLs manually annotated by ImageNet.
[0090] The KITTI dataset was jointly created by Karlsruhe Institute of Technology and Toyota Research Institute of America. It is currently the world's largest computer vision algorithm evaluation dataset for autonomous driving scenarios.
[0091] The Open Images dataset is an open source image dataset released by Google. Its latest version contains more than 9 million images, all of which are labeled with categories.
[0092] The SUN dataset is a dataset used to train and evaluate image super-resolution algorithms. It contains 817 high-resolution images from different types, such as architecture, landscape, nature, etc.
[0093] The Cityscapes dataset has 5,000 images of driving scenes in urban environments.
[0094] The DOTA dataset contains 2806 aerial images with a size of approximately 4k*4k, including 15 categories and a total of 188,282 instances.
[0095] The VisDrone dataset is a large-scale dataset for drone vision recognition. It includes a rich range of scenes and object categories. Its purpose is to promote the development of drone vision and intelligence. The VisDrone dataset contains 14 object categories, including pedestrians, vehicles, bicycles, pedestrian occlusion, vehicle occlusion, bicycle occlusion, pedestrian occlusion head, vehicle occlusion head, bicycle occlusion head, traffic cones, other moving objects, buildings, trees, and benches.
[0096] Meanwhile, in this embodiment, step S3 also includes:
[0097] The dataset for training the universal network model merges labels with the same semantics according to different practical classifications, with a total of 5018 categories. Labels with different semantics represent a classification box.
[0098] That is, in this embodiment, the above-mentioned multiple categories of data sets include high-resolution images, videos, annotations and metadata, etc., which can provide a variety of image data samples covering various safety and protection scenarios including power operation scenarios for training the universal network model, forming sample-rich training sets, test sets and verification sets, so as to improve the applicability of the ultimately trained universal network model to most safety and supervision scenarios, reduce the cost of large amounts of data collection and labeling required when the model is subsequently used for target detection and security monitoring operations, and at the same time reduce dependence on high computing resources, effectively save data and computing resources, and improve the efficiency of infrastructure violation monitoring and security monitoring during power operations.
[0099] Embodiment 3 of the present invention is:
[0100] A terminal for building an image recognition model in a safety supervision scenario, based on the above-mentioned embodiment 1 or embodiment 2, in this embodiment, step S4 is specifically:
[0101] S41. When training the initial network model, the parameter weights of the initial network model cannot be modified based on the gradient, but gradient propagation can be performed to lock the universal network model.
[0102] That is, when a large number of key data sets are used to train a universal network model, the constructed initial network model is first locked to ensure that there is no need to correct the weights of the general universal network model in the subsequent fine-tuning stage, thereby greatly reducing computing resources.
[0103] S42. Splice the front-layer network model in front of the locked all-purpose network model, input the image feature F0 into the front-layer network, obtain the feature F1, and then input F1 into the all-purpose network model. The all-purpose network model is responsible for outputting the classification and the frame.
[0104] That is, by directly splicing the front-layer network model in front of the all-purpose network model, it can be directly applied to new safety and supervision scenarios by only modifying the front-layer network model, thereby further improving the efficiency of infrastructure violation monitoring and security monitoring during power operations.
[0105] Finally, in this embodiment, the fine-tuning process in step S5 is defined as:
[0106] The learning rate of the spliced universal network model is controlled, where the learning rate is limited to 0.00001, and the model is tested and verified using real-time safety monitoring data until the test set error and the verification set error no longer decrease.
[0107] That is, by finally testing and training the assembled model based on a small amount of safety and supervision data in a specific scenario, and ensuring that the errors of the test set and the validation set remain stable, the model can be directly applied to new safety and supervision scenarios, thereby further improving the efficiency of infrastructure violation monitoring and security monitoring during power operations.
[0108] Please refer to Figure 2 , Embodiment 2 of the present invention is:
[0109] A terminal 1 for building an image recognition model in a safety and supervision scenario includes a memory 2, a processor 3, and a computer program stored in the memory 2 and executable on the processor 3. When executing the computer program, the processor 3 implements the steps in a method for building an image recognition model in a safety and supervision scenario in any one of the above-mentioned embodiments 1 to 3.
[0110] In summary, the present invention provides a method and terminal for constructing an image recognition model in a safety supervision scenario, which has the following beneficial effects:
[0111] 1. Use various types of data sets to train the universal network model. Subsequently, the universal network model spliced with the previous network model is directly used to perform infrastructure violation monitoring and security monitoring in safety supervision scenarios, avoiding the need for large amounts of data, reducing the cost of data collection and annotation, and at the same time reducing dependence on high computing resources, saving time and cost.
[0112] 2. The all-round model has been trained with a large amount of data and has the ability to label various feature objects. These excellent labeling capabilities can be transferred to safety and supervision scenarios through the front-layer network to adapt to the image processing needs of different security monitoring, power operations and other fields, and has strong portability.
[0113] 3. This method does not require retraining the entire model, but only the front-layer network needs to be trained, which greatly reduces the need for scene-specific training data. After fine-tuning the front-layer network with a small number of samples, it can be directly applied to new safety and supervision scenarios.
[0114] 4. In the fine-tuning stage, the general target detection network does not need to perform weight correction during the back-propagation process, which further significantly reduces computing resources.
[0115] The above descriptions are merely embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent transformations made using the contents of the present invention's specification and drawings, or directly or indirectly applied in related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A method for constructing an image recognition model in a safety supervision scenario. It is characterized in that Includes steps: S1, build the front layer network model; S2, build the initial network model; S3, using multiple types of data sets to train the initial network model to obtain a universal network model; The data set for training the universal network model merges labels with the same semantics according to different practical classifications, and labels with different semantics represent a classification box; S4, locking the universal network model, and splicing the front-layer network model in front of the universal network model; S5, fine-tuning the spliced universal network model to obtain a final image recognition model; The step S1 is specifically as follows: A front-layer network model LowTastNet consisting of two convolutional network layers is constructed, and a special operator layer is added between the two convolutional network layers. The operator of the special operator layer is as follows: (1); Where x is a real number, the value of sign(x) is 1, -1 or 0, indicating the positive or negative value of x and whether it is 0, respectively, and relu(x)=max(0,x), indicating the activation function; The step S2 is specifically as follows: Construct an initial network model GenpNett consisting of five single convolutional network layers, one matrix convolutional network layer, three maximum pooling layers and a general detection head; The initial network model is defined to include ten layers from 0 to 9, wherein five single convolutional network layers are respectively located at the 0th, 2nd, 4th, 5th and 6th layers of the initial network model, one matrix convolutional network layer is located at the 3rd layer of the initial network model, three maximum pooling layers are respectively located at the 1st, 7th and 8th layers of the initial network model, and one universal detection head is located at the 9th layer of the initial network model; The operators defining five single convolutional network layers and one matrix convolutional network layer are all Conv2d; The operators defining the three maximum pooling layers are all maxpool; Define a general detection head operator as Detect; The multiple categories of data sets in step S3 include COCO data set, Pascal VOC data set, ImageNet data set, KITTI data set, Open Images data set, SUN data set, Cityscapes data set, DOTA data set and VisDrone data set; The step S4 is specifically as follows: S41, when training the initial network model, setting the parameter weights of the initial network model to be non-modifiable based on gradients, but enabling gradient propagation, and locking the universal network model; S42, splicing the front layer network model in front of the locked all-purpose network model, inputting the image feature F0 into the front layer network to obtain feature F1, and then inputting F1 into the all-purpose network model, and the all-purpose network model is responsible for outputting classification and box.
2. According to the method for constructing an image recognition model in a safety supervision scenario according to claim 1, It is characterized in that The step S1 further comprises: The front layer network model is defined to include three layers from 0 to 2, wherein the 0th and 2nd layers are convolutional network layers, and the 1st layer is a special operator layer; The operators defining the two convolutional network layers are both Conv2d.
3. According to the method for constructing an image recognition model in a safety supervision scenario according to claim 1, It is characterized in that The fine-tuning process in step S5 is as follows: The learning rate of the spliced universal network model is controlled, and the model is tested and verified using real-time safety monitoring data until the test set error and the verification set error no longer decrease.
4. A terminal for building an image recognition model in a safety and supervision scenario. It is characterized in that The method comprises a processor, a memory and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method implements the steps in a method for constructing an image recognition model in a safety and supervision scenario as described in any one of claims 1 to 3 above.
Citation Information
Patent Citations
Image management method and device based on multi-task machine learning model
CN111813532A
Training method and recognition method of social relation recognition model and related equipment
CN112668509A
Multi-modal fusion classification optimization method considering inter-modal semantic distance measurement
CN113343974A