A Method, System, Medium and Device for Deploying YOLOv5 on a CPU Platform
By modifying the YOLOv5 source code and output structure, the problem of OpenCV version adaptation during YOLOv5 deployment on the CPU platform is solved, and effective deployment on lower versions of OpenCV is achieved, reducing cost and complexity.
Patent Information
- Application Number
- CN202210262944.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-17
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-03-17
AI Technical Summary
When deploying YOLOv5 target detection network on the CPU platform, it is difficult for the existing technology to effectively adapt to the OpenCV version, resulting in deployment difficulties.
By modifying the YOLOv5 source code, replacing the original slice operation with the slice operation supported by OpenCV customized by the contract module, and modifying the output structure, the three output multi-dimensional arrays are dimensionally reduced and then merged, expanding the adaptation range of the OpenCV version.
It realizes that YOLOv5 can be deployed on the CPU platform only by supporting OpenCV3.4.13 and above, reducing deployment costs and complexity.
Smart Images

Figure CN114637530B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and in particular, to a method, system, medium and device for deploying YOLOv5 on a CPU platform. Background Art
[0002] With the advent of the intelligent era, there will be more and more fields and scenarios that use object detection, such as smoking detection in public places and safety guarantee detection in industrial production. In the process of building a smart city, due to the limited hardware conditions of old cities, using overly complex deployment methods is likely to cause problems such as high costs and difficult maintenance. Based on comprehensive considerations of cost, performance, etc., in the prior art, a solution of deploying the YOLOv5 object detection network is adopted to minimize the deployment cost as much as possible under the condition of sacrificing a little performance.
[0003] However, one of the common solutions in the prior art for deploying the YOLOv5 object detection network is the openvino deployment solution. However, the openvino deployment solution requires additional authorization and installation from Intel, with high costs and cumbersome operations. Using a GPU to deploy the YOLOv5 object detection network is also a deployable solution. However, the high price of graphics cards results in relatively high deployment costs. Therefore, in terms of cost savings, simply using a CPU to deploy the YOLOv5 object detection network has become a better choice.
[0004] However, when using a CPU to deploy an object detection network, the OpenCV version recommended by the YOLOv5 source code generally needs to be OpenCV 4.5 or higher. Therefore, when faced with a more basic version below OpenCV 4.5, it is difficult to achieve good deployment of YOLOv5. The dependency conditions for YOLOv5 deployment need to be further reduced, and the adaptation range of the OpenCV version during YOLOv5 deployment needs to be further expanded. Summary of the Invention
[0005] In view of at least one defect or improvement requirement of the prior art mentioned in the background art, the present invention provides a method, system, medium and device for deploying YOLOv5 on a CPU platform to solve the technical problem of how to expand the adaptation range of the OpenCV version when deploying YOLOv5 on a CPU platform.
[0006] To solve the above technical problems, in a first aspect, the present invention provides a method for deploying YOLOv5 on a CPU platform, including:
[0007] Modifying the original slicing operation in the YOLOv5 source code into a slicing operation supported by OpenCV customized by the contract module;
[0008] Modify the output structure of YOLOv5. After performing dimensionality reduction on the multi-dimensional arrays corresponding to the three outputs of YOLOv5 respectively, merge them into one output;
[0009] Traverse the one-dimensional arrays of each row of the dimensionality-reduced array to obtain the detection results that meet the conditions.
[0010] According to the method for deploying YOLOv5 on the CPU platform provided by the present invention, the modification of the original slicing operation in the YOLOv5 source code to the slicing operation supported by OpenCV customized by the contract module is specifically as follows:
[0011] Modify the Focus module in the common.py file of the YOLOv5 source code. Replace the original slicing operation with the method of the contract module in the code of the Focus module.
[0012] According to the method for deploying YOLOv5 on the CPU platform provided by the present invention, the modification of the output structure of YOLOv5, after performing dimensionality reduction on the multi-dimensional arrays corresponding to the three outputs of YOLOv5 respectively and then merging them into one output, is specifically as follows:
[0013] Modify the detect module in yolo.py. In the forward function, perform dimensionality reduction on the multi-dimensional arrays corresponding to the three outputs of YOLOv5 respectively, and then use torch.cat() to merge the three dimensionality-reduced arrays and output them.
[0014] According to the method for deploying YOLOv5 on the CPU platform provided by the present invention, the output objects of the three outputs of YOLOv5 are respectively:
[0015] Detection results of small targets, detection results of medium targets, and detection results of large targets.
[0016] According to the method for deploying YOLOv5 on the CPU platform provided by the present invention, after traversing the one-dimensional arrays of each row of the dimensionality-reduced array to obtain the detection results that meet the conditions, it further includes:
[0017] Use the non-maximum suppression algorithm to filter the overlapping boxes of the detection results, and draw boxes and output after obtaining the optimal detection results.
[0018] In a second aspect, the present invention provides a system for deploying YOLOv5 on the CPU platform, including:
[0019] Slicing module: used to modify the original slicing operation in the YOLOv5 source code to the slicing operation supported by OpenCV customized by the contract module;
[0020] Merging module: used to modify the output structure of YOLOv5, and after performing dimensionality reduction on the multi-dimensional arrays corresponding to the three outputs of YOLOv5 respectively, merge them into one output;
[0021] Traversal module: used to traverse the one-dimensional arrays of each row of the array after dimensionality reduction to obtain the detection results that meet the conditions.
[0022] According to the system for deploying YOLOv5 on the CPU platform provided by the present invention, the system further includes:
[0023] Filtering module: used to filter overlapping boxes of the detection results using the non-maximum suppression algorithm, and draw a frame for output after obtaining the optimal detection results.
[0024] In a third aspect, the present invention provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it can implement the steps of any one of the above methods.
[0025] In a fourth aspect, the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it can implement the steps of any one of the above methods.
[0026] Compared with the prior art, the beneficial effects of the present invention:
[0027] The method of the present invention mainly modifies the original slicing operation in the YOLOv5 source code into the slicing operation supported by OpenCV customized by the contract module, modifies the output structure of YOLOv5, and after performing dimensionality reduction on the multi-dimensional arrays corresponding to the three outputs of YOLOv5 respectively, merge them into one output. Therefore, only the CPU platform needs to support OpenCV version 3.4.13 or above, which expands the adaptation range of the OpenCV version when deploying YOLOv5 on the CPU platform, and only the lower version of the most basic open-source OpenCV can be used to implement the deployment of YOLOv5 on the CPU platform. Description of the Drawings
[0028] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0029] Figure 1 It is a flowchart of the method for deploying YOLOv5 on the CPU platform provided by the embodiment of the present invention;
[0030] Figure 2 It is a block diagram of an electronic device provided by an embodiment of the present invention, which is suitable for implementing the method described above. Detailed implementation manners
[0031] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings. Apparently, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0032] YOLO is a target detection network, and YOLOv5 is a derivative version based on it, and there are dozens of derivative versions. Deploying the YOLOv5 target detection network can achieve reducing the deployment cost as much as possible under the condition of sacrificing a little performance, and it is applicable to demand scenarios with not too high performance requirements and extremely sensitive to costs.
[0033] OpenCV is a cross-platform computer vision and machine learning software library distributed under the Apache 2.0 license (open source), and it can run on operating systems such as Linux, Windows, Android, and Mac OS. It is lightweight and efficient - composed of a series of C functions and a small number of C++ classes, and at the same time provides interfaces in languages such as Python, Ruby, and MATLAB, implementing many general algorithms in image processing and computer vision. There are many versions of OpenCV and it is often updated. The deployment method of OpenCV recommended by the YOLOv5 source code requires at least the version of OpenCV 4.5. However, not all servers in any scenario support the version of OpenCV 4.5 and above. Based on the limitation of the adaptation range of this OpenCV version in the actual application scenario, the present invention proposes a method for deploying YOLOv5 on the CPU platform to expand the adaptation range of the OpenCV version.
[0034] As Figure 1 shown, in one embodiment, a method for deploying YOLOv5 on the CPU platform includes the following steps S1 - S3:
[0035] S1. Modify the original slicing operation in the YOLOv5 source code into a slicing operation supported by OpenCV customized by the contract module.
[0036] To utilize the ONNX format model exported by YOLOv5, it is necessary to first modify the YOLOv5 source code so that the exported model can be adapted to the dnn module of OpenCV. The specific operation is to use the slicing operation of the contract module to replace the slicing operation not supported by OpenCV. Since OpenCV reports an error when reading the ONNX model exported by the export in the YOLOv5 source code, it is necessary to modify the Focus module in the common.py file of the YOLOv5 source code and replace the original slicing operation with the method of the contract module in the code of the Focus module.
[0037] #return self.conv(torch.cat([x[...,::2,::2],x[...,1::2,::2],x[...,::2,1::2],x[...,1::2,1::2]],1))
[0038] Modify the above default original slicing operation to the custom OpenCV-supported slicing operation of the contract module
[0039] #return self.conv(self.contract(x))
[0040] By modifying the slicing operation not supported by OpenCV, the model exported by YOLOv5 can be recognized by the dnn module of OpenCV. The next step is to modify the output structure of the model network to facilitate subsequent coding for parsing.
[0041] S2. Modify the output structure of YOLOv5. After reducing the dimensions of the multi-dimensional arrays corresponding to the three outputs of YOLOv5 respectively, merge them into one output.
[0042] The original output structure of the YOLOv5 network model is three multi-dimensional vectors. This structure contains the sequential recognition and detection results of large, medium, and small targets. Although it is clear, the complexity of parsing the final detection results is high. Because in order to facilitate the final unified analysis of the bounding box results, these three outputs need to be integrated into a reduced-dimensional vector. The YOLOv5 network has three exits. In one embodiment, OUTPUTS: [1,3,80,80,85], [1,3,40,40,85], [1,3,20,20,85].
[0043] According to the model structure file ( / models / yolov5s.yaml), these three exits are respectively used to output the detection results of small, medium, and large targets. Therefore, in order to improve efficiency and facilitate deployment, a single output can be traversed.
[0044] Modify the detect module of yolo.py to transform the output format in the forward function. In one embodiment, at
[0045] #x[i] = x[i].view(bs,self.na,self.no,ny,nx).permute(0,1,3,4,2).contiguous()
[0046] add a line after
[0047] #x[i] = x[i].view(bs*self.na*ny*nx,self.no).contiguous()
[0048] each output can be merged into a two-dimensional array after dimensionality reduction, and then the three outputs can be concatenated using torch.cat() to merge the two-dimensional array with the format [25200.85].
[0049] Then, use the export.py script to export the network model to the onnx format. After modifying the YOLOv5 model, an onnx model that can be loaded by the dnn module of OpenCV is obtained, and then the output results obtained from the model need to be decoded.
[0050] S3. Traverse the one-dimensional arrays of each row of the dimensionality-reduced array to obtain the detection results that meet the conditions.
[0051] The dnn module of OpenCV provides a function interface readNetFromONNX() to read the neural network model. The following statement can be used to obtain the model:
[0052] net = readNetFromONNX(netPath)
[0053] Define a flag to set the inference engine to use the CPU or GPU. The entire inference process needs to first set the network input, then send it into the network for inference, and finally parse the network output. Among them, traversing the network output and parsing to obtain the final detection results is the most difficult step because the network output is changed to a two-dimensional array format of 25200*85. This step requires traversing each one-dimensional array with a length of 85 in each row and obtaining the detection results that meet the conditions.
[0054] Finally, there will be many output boxes overlapping in the obtained detection results. Therefore, preferably, the non-maximum suppression algorithm (NMS) can be used to filter the overlapping boxes, and finally draw the boxes and output the optimal detection results.
[0055] In one embodiment, the present invention further provides a system for deploying YOLOv5 on a CPU platform, including:
[0056] Slicing module: used to modify the original slicing operation in the YOLOv5 source code into a slicing operation supported by OpenCV customized by the contract module;
[0057] Merging module: used to modify the output structure of YOLOv5, and perform dimensionality reduction processing on the multi-dimensional arrays corresponding to the three outputs of YOLOv5 respectively and then merge them into one output;
[0058] Traversal module: used to traverse each one-dimensional array in each row of the dimensionality-reduced array to obtain detection results that meet the conditions.
[0059] Preferably, the system further includes:
[0060] Filtering module: used to filter overlapping boxes of the detection results using the non-maximum suppression algorithm, and draw a frame for output after obtaining the optimal detection results.
[0061] Figure 2 A block diagram of an electronic device suitable for implementing the method described above according to an embodiment of the present disclosure is schematically shown. Figure 2 The shown electronic device is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present disclosure.
[0062] As Figure 2 shown, the electronic device 1000 described in this embodiment includes: a processor 1001, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1002 or the program loaded from the storage part 1008 into the random access memory (RAM) 1003. The processor 1001 can include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), and so on. The processor 1001 can also include on-board memory for caching purposes. The processor 1001 can include a single processing unit or multiple processing units for performing different actions of the method flow according to the embodiments of the present disclosure.
[0063] In the RAM 1003, various programs and data required for the operation of the system 1000 are stored. The processor 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. The processor 1001 performs various operations of the method flow according to the embodiments of the present disclosure by executing the programs in the ROM 1002 and / or the RAM 1003. It should be noted that the programs may also be stored in one or more memories other than the ROM 1002 and the RAM 1003. The processor 1001 may also perform various operations of the method flow according to the embodiments of the present disclosure by executing the programs stored in the one or more memories.
[0064] According to an embodiment of the present disclosure, the electronic device 1000 may further include an input / output (I / O) interface 1005, and the input / output (I / O) interface 1005 is also connected to the bus 1004. The system 1000 may further include one or more of the following components connected to the I / O interface 1005: an input portion 1006 including a keyboard, a mouse, etc.; an output portion 1007 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage portion 1008 including a hard disk, etc.; and a communication portion 1009 including a network interface card such as a LAN card, a modem, etc. The communication portion 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the I / O interface 1005 as needed. A removable medium 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1010 as needed so that a computer program read from it can be installed into the storage portion 1008 as needed.
[0065] The method flow according to the embodiments of the present disclosure may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program includes program codes for performing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from a network through the communication portion 1009, and / or installed from the removable medium 1011. When the computer program is executed by the processor 1001, the above-described functions defined in the system according to the embodiments of the present disclosure are executed. According to an embodiment of the present disclosure, the above-described systems, devices, modules, etc. may be implemented by computer program modules.
[0066] An embodiment of the present invention further provides a computer-readable storage medium, which may be included in the device / system described in the above embodiment; or may exist alone without being assembled into the device / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiments of the present disclosure is implemented.
[0067] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memories (RAMs), read-only memories (ROMs), erasable programmable read-only memories (EPROMs or flash memories), portable compact disk read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the above. In an embodiment of the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present disclosure, the computer-readable storage medium may include one or more memories other than the ROM 1002 and / or the RAM 1003 described above.
[0068] It should be noted that in each embodiment of the present invention, each functional module may be integrated into a processing module, or each module may exist physically alone, or two or more modules may be integrated into one module. The above integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product.
[0069] The flowcharts or block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0070] Those skilled in the art will understand that the features recited in the various embodiments and / or claims of the present disclosure can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly recited in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features recited in the various embodiments and / or claims of the present disclosure can be combined and / or combined in various ways, and all such combinations and / or combinations fall within the scope of the present disclosure.
[0071] Although the present disclosure has been shown and described with reference to specific exemplary embodiments of the present disclosure, those skilled in the art should understand that various changes in form and detail can be made to the present disclosure without departing from the spirit and scope of the present disclosure as defined by the appended claims and their equivalents. Therefore, the scope of the present disclosure should not be limited to the above embodiments, but should be determined not only by the appended claims, but also by the equivalents of the appended claims.
Claims
1. A method for deploying YOLOv5 on a CPU platform, characterized in that, Including: Modify the original slicing operation in the YOLOv5 source code to the slicing operation supported by OpenCV customized by the contract module; Modify the output structure of YOLOv5. After performing dimensionality reduction on the multi-dimensional arrays corresponding to the three outputs of YOLOv5 respectively, merge them into one output; Traverse the one-dimensional arrays of each row of the dimensionality-reduced array to obtain the detection results that meet the conditions; The specific method of modifying the original slicing operation in the YOLOv5 source code to the slicing operation supported by OpenCV customized by the contract module is as follows: Modify the Focus module in the common.py file of the YOLOv5 source code. Replace the original slicing operation with the method of the contract module in the code of the Focus module; The specific method of modifying the output structure of YOLOv5. After performing dimensionality reduction on the multi-dimensional arrays corresponding to the three outputs of YOLOv5 respectively, merge them into one output is as follows: Modify the detect module of yolo.py. In the forward function, perform dimensionality reduction on the multi-dimensional arrays corresponding to the three outputs of YOLOv5 respectively, and then use torch.cat() to merge the three dimensionality-reduced arrays and output.
2. The method for deploying YOLOv5 on a CPU platform according to claim 1, characterized in that, The output objects of the three outputs of YOLOv5 are respectively: Detection results of small targets, detection results of medium targets, and detection results of large targets.
3. The method for deploying YOLOv5 on a CPU platform according to claim 1, characterized in that, After traversing the one-dimensional arrays of each row of the dimensionality-reduced array to obtain the detection results that meet the conditions, it further includes: Use the non-maximum suppression algorithm to filter the overlapping boxes of the detection results, and draw a frame and output after obtaining the optimal detection results.
4. A system for deploying YOLOv5 on a CPU platform, which executes the method for deploying YOLOv5 on a CPU platform according to claim 1, characterized in that, Including: Slicing module: used to modify the original slicing operation in the YOLOv5 source code to the slicing operation supported by OpenCV customized by the contract module; Merging module: used to modify the output structure of YOLOv5. After performing dimensionality reduction on the multi-dimensional arrays corresponding to the three outputs of YOLOv5 respectively, merge them into one output; Traversal module: used to traverse the one-dimensional arrays of each row of the dimensionality-reduced array to obtain the detection results that meet the conditions.
5. The system for deploying YOLOv5 on a CPU platform according to claim 4, characterized in that, The system further includes: Filtering module: used to use the non-maximum suppression algorithm to filter the overlapping boxes of the detection results, and draw a frame and output after obtaining the optimal detection results.
6. A non-transitory computer-readable storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by a processor, it can implement the steps of the method according to any one of claims 1 to 3.
7. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it can implement the steps of the method according to any one of claims 1 to 3.
Citation Information
Patent Citations
Method and system for monitoring people falling into water in real time based on video stream
CN114120174A
Build Deployment Automation for Information Technology Management
US20150248280A1