Ball picking machine control method and device based on yolov5 algorithm and storage medium
By using a ball recognition network based on the YOLOv5 algorithm and wind-powered pickup technology, the problems of low accuracy in badminton shuttlecock recognition and low pickup efficiency have been solved, achieving efficient and low-damage badminton shuttlecock pickup.
Patent Information
- Application Number
- CN202310700031.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-13
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2043-06-13
AI Technical Summary
In existing technologies, the accuracy of badminton shuttlecock recognition and the efficiency of shuttlecock retrieval are low. The shuttlecock retrieval machine has low detection accuracy and low retrieval efficiency, and cannot effectively adapt to the irregular shape and fragile characteristics of badminton shuttlecock.
A ball recognition network based on the YOLOv5 algorithm is adopted. By fusing shallow and deep features, the accuracy and efficiency of badminton shuttlecock detection are improved. The shuttlecock is picked up by wind suction. Combined with the preset movement driving mode and recognition position planning, it quickly moves to the target shuttlecock for pickup.
It improves the effectiveness and efficiency of badminton shuttlecock recognition, reduces damage to shuttlecocks during retrieval, and achieves high-precision shuttlecock recognition and efficient shuttlecock retrieval.
Smart Images

Figure CN116630864B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, in particular to a ball picking machine control method and device based on YOLOV5 algorithm and a storage medium. BACKGROUND
[0002] Badminton is a popular sport in the world, and more and more people are joining the badminton training team. However, the repeated picking of badminton increases the invalid physical consumption of the trainees and wastes the effective training time.
[0003] In the field of ball picking machines, the known ball picking machines are mainly for regular balls such as basketballs, footballs, and table tennis balls. In the field of ball picking machines, there is still a lack of effective ball picking machines for picking irregular and easily damaged balls such as badminton. The ball picking machines in the related art plan routes through remote control / fixed route cruising, and have low picking efficiency, long picking time, and low site adaptability. At the same time, in the related art, the scheme of identifying irregular and easily damaged balls such as badminton by using artificial intelligence or neural networks has low ball detection accuracy, low ball picking efficiency, and poor use effect due to the small size of the balls and the large variation in the damage degree of the balls during use.
[0004] Currently, there is no effective solution to the problem of low detection accuracy and low ball picking efficiency in the related art for identifying badminton. SUMMARY
[0005] The embodiments of the present application provide a ball picking machine control method and device based on YOLOV5 algorithm, a storage medium, and an electronic device to at least solve the problem of low detection accuracy and low ball picking efficiency in the related art for identifying badminton.
[0006] In a first aspect, the embodiments of the present application provide a control method of a ball picking machine based on a YOLOV5 algorithm, comprising: acquiring a video in a badminton training process, wherein the video comprises a plurality of video image frames; detecting candidate balls in each of the video image frames by using a pre-trained ball recognition network, and determining a bounding box corresponding to the candidate balls, wherein the ball recognition network is based on the YOLOV5 algorithm, and a neural network trained by using a preset fused shallow feature as a shallow feature of the YOLOV5 algorithm, the fused shallow feature is generated by fusing a shallow feature extracted from a sample image by using the YOLOV5 algorithm and a preset deep feature, and the deep feature comprises preset coarse-grained semantic information for representing a badminton composition feature; screening a target ball from the candidate balls according to distance information of the bounding box and a center line corresponding to the video image frame; and controlling the ball picking machine to move to a real scene field position corresponding to a position of the target ball according to a preset movement driving mode, and sucking a badminton at the real scene field position by using wind power.
[0007] In a second aspect, the embodiments of the present application provide a control device of a ball picking machine based on YOLOV5, comprising:
[0008] an acquisition module configured to acquire a video in a badminton training process, wherein the video comprises a plurality of video image frames;
[0009] an identification module configured to detect candidate balls in each of the video image frames by using a pre-trained ball recognition network, and determine a bounding box corresponding to the candidate balls, wherein the ball recognition network is based on the YOLOV5 algorithm, and a neural network trained by using a preset fused shallow feature as a shallow feature of the YOLOV5 algorithm, the fused shallow feature is generated by fusing a shallow feature extracted from a sample image frame by using the YOLOV5 algorithm and a preset deep feature, and the deep feature comprises preset coarse-grained semantic information for representing a badminton composition feature;
[0010] a screening module configured to screen a target ball from the candidate balls according to distance information of the bounding box and a center line corresponding to the video image frame;
[0011] a processing module configured to control the ball picking machine to move to a real scene field position corresponding to a position of the target ball according to a preset movement driving mode, and suck a badminton at the real scene field position by using wind power.
[0012] In a third aspect, an electronic device is provided, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the steps of the ball-picking robot control method based on the YOLOV5 algorithm when executing the computer program.
[0013] In a fourth aspect, a storage medium is provided, which stores a computer program, and the program, when executed by a processor, implements the steps of the ball-picking robot control method based on the YOLOV5 algorithm according to the first aspect.
[0014] Compared with the related art, the ball-picking robot control method, device, storage medium, and electronic device based on the YOLOV5 algorithm provided by the embodiments of the present application have the beneficial effects of enhancing the ball recognition ability and detection accuracy, and improving the ball-picking efficiency. Specifically, by acquiring a video in a badminton training process, the video includes multiple video image frames; a pre-trained ball recognition network is used to detect candidate balls in each video image frame and determine the bounding box corresponding to the candidate ball, the ball recognition network is based on the YOLOV5 algorithm, and is a neural network trained by using a preset fused shallow feature as a shallow feature of the YOLOV5 algorithm, the fused shallow feature is generated by fusing a shallow feature extracted from a sample image by the YOLOV5 algorithm and a preset deep feature, and the deep feature includes preset coarse-grained semantic information for representing the composition features of a badminton; target balls are screened out from the candidate balls according to distance information of the bounding box and a center line corresponding to the video image frame; a ball-picking robot is controlled to move to a real scene location corresponding to a location where the target ball is located according to a preset movement driving mode, and a badminton located at the real scene location is sucked by wind power; the deep feature with the preset coarse-grained semantic information for representing the composition features of the badminton is fused with the shallow feature, so that the fused shallow feature has high pixels and multiple semantics, and the ball recognition network is trained by using the fused shallow feature, thereby enhancing the robustness and detection ability of the ball recognition network for small objects, improving the detection accuracy, and improving the effectiveness and efficiency of badminton recognition; at the same time, a driving path is planned based on the position of the recognized ball, and the ball-picking robot is quickly moved to the badminton to be picked up according to the corresponding movement driving mode, and then the badminton is picked up by wind power, thereby improving the ball-picking efficiency and reducing the damage to the ball during the picking process. The embodiments of the present application solve the problem of low detection accuracy and ball-picking efficiency in the related art.
[0015] The details of one or more embodiments of the present application are presented in the following drawings and description to make other features, objects, and advantages of the present application more apparent. BRIEF DESCRIPTION OF DRAWINGS
[0016] The accompanying drawings, which are included to provide a further understanding of the application and are incorporated in and constitute a part of this application, illustrate embodiments of the present application and together with the description serve to explain the present application. In the drawings:
[0017] Figure 1 is a hardware structure block diagram of a terminal of a ball picking machine control method based on a YOLOV5 algorithm according to an embodiment of the present application;
[0018] Figure 2 is a flowchart of a ball picking machine control method based on a YOLOV5 algorithm according to an embodiment of the present application;
[0019] Figure 3 is a flowchart of a ball picking machine control method according to a preferred embodiment of the present application;
[0020] Figure 4 is a structure block diagram of a control device of a ball picking machine based on YOLOV5 according to an embodiment of the present application. DETAILED DESCRIPTION
[0021] In order to make the objects, technical solutions and advantages of the present application clearer, the present application is described and explained below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and should not be used to limit the present application. Based on the embodiments provided in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of the present application. In addition, it can be understood that although the efforts made in this development process can be complex and lengthy, some design, manufacture or production changes made on the basis of the technical content disclosed in the present application by those of ordinary skill in the art related to the content disclosed in the present application are only routine technical means and should not be understood as insufficient disclosure of the present application.
[0022] In the present application, the phrase "embodiment" means that the specific features, structures or characteristics described in combination with the embodiment can be included in at least one embodiment of the present application. The appearance of this phrase at various places in the specification does not necessarily mean the same embodiment, nor is it an independent or alternative embodiment to other embodiments. Those of ordinary skill in the art explicitly and implicitly understand that the embodiments described in the present application can be combined with other embodiments without conflict.
[0023] Unless otherwise defined, technical terms and scientific terms used in the present application shall have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. The terms "a", "an", "one", "this", and the like, as used in the present application, do not denote number restriction, can represent singular or plural. The terms "include", "contain", "have" and any variations thereof in the present application are intended to cover non-exclusive inclusion; for example, a process, method, system, product or device including a series of steps or modules (units) is not limited to the listed steps or units, but can also include steps or units not listed, or can also include other steps or units inherent to the process, method, product or device. The "multi-link" in the present application refers to more than or equal to two links. "And / or" describes the association between the associated objects, which means that there can be three relationships, for example, "A and / or B" can represent three cases: A exists alone, A and B exist together, and B exists alone. The terms "first", "second", "third" and the like in the present application only distinguish similar objects, and do not represent a specific order for the objects.
[0024] Before the embodiments of the present application are described, the YOLOV5 algorithm involved in the embodiments of the present application is described as follows:
[0025] The YOLOV5 algorithm is an image recognition training and deployment framework based on pytorch, which has the advantages of fast speed, high performance and easy use. Even in a small computer with only CPU, the pre-trained model can be smoothly deployed or run. The network structure of YOLOV5 mainly consists of four parts, which are input end Input, backbone network Backbone, Neck and output end Prediction, wherein,
[0026] The Input uses Mosaic data enhancement, adaptive anchor box calculation and picture size processing. Mosaic data enhancement splices multiple pictures in a random scaling, random cropping and random arrangement manner, greatly enriches the detection data set, and increases many small targets, so that the robustness of the network structure is better.
[0027] The Backbone includes three parts of Focus, CSP and SPP. The Focus is used to perform slicing operation, for example, the size of an image of 608*608*3 will become 304*304*12 after slicing operation in the Focus structure, and then the convolution operation of 32 convolution kernels is performed to generate a feature map of 304*304*32; there are two kinds of CSP in the YOLOV5 model, namely CSP1_x and CSP2_x, wherein x in CSP1_x represents that there are x residual components Resunit in CSP1, and x in CSP2_x represents that there are 2x CBLs in CSP2, and x affects the depth of the network structure; in the Backbone, spatial vector pyramid pooling SPP is used to perform maximum pooling on the feature map, and different scale feature maps are spliced together.
[0028] The Neck is between the Backbone and the Prediction, and is used to connect and extract fused features. The Neck uses a feature pyramid network FPN+path aggregation network PAN structure, the FPN is from top to bottom, and uses up-sampling to transmit and fuse data to obtain a predicted feature map, and the PAN uses a feature pyramid to convey location information from bottom to top.
[0029] The Prediction includes a bounding box loss function and non-maximum suppression NMS, YOLOV5 uses GIOU_Loss as a loss function, increases the intersection scale measurement method to solve the problem that the predicted frame and the target frame are not intersected, and distinguishes the different intersection conditions of the overlapping predicted frame and the target frame when the IOU is the same. In selecting the optimal predicted frame, the NMS operation is needed, and the weighted NMS method is used in YOLOV5 to remove the redundant predicted frame, and the GIOU_Loss loss function formula used by the YOLOV5 of the embodiment of the application is as follows:
[0030]
[0031] The embodiments of the application are specifically described below based on the drawings.
[0032] The method embodiments provided in the embodiment can be executed in a terminal, a computer or a similar computing device. Taking the running on the terminal as an example, Figure 1 is a hardware structure block diagram of the terminal of the YOLOV5 algorithm-based ball picking machine control method of the embodiment of the application. As shown in Figure 1 , the terminal can include one or more Figure 1The terminal shown in FIG. 1 includes only one processor 102 (the processor 102 can include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Optionally, the terminal can further include a transmission device 106 for communication function and an input / output device 108. Those skilled in the art can understand that, Figure 1 The structure shown in FIG. 1 is only schematic and does not limit the structure of the terminal. For example, the terminal can include more or fewer components than those shown in FIG. 1, or have a different configuration from that shown in FIG. 1. Figure 1 Figure 1 The structure shown in FIG. 1 is only schematic and does not limit the structure of the terminal. For example, the terminal can include more or fewer components than those shown in FIG. 1, or have a different configuration from that shown in FIG. 1.
[0033] The memory 104 can be used to store computer programs, such as software programs of application software and modules, such as the computer program corresponding to the ball-picking machine control method based on the YOLOV5 algorithm in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the computer programs stored in the memory 104, that is, implements the method described above. The memory 104 can include a high-speed random access memory, and can further include a non-volatile memory, such as one or more magnetic storage devices, a flash memory, or other non-volatile solid-state memories. In some examples, the memory 104 can further include a memory remotely arranged with respect to the processor 102, and these remote memories can be connected to the terminal 10 through a network. Examples of the network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0034] The transmission device 106 is used to receive or send data via a network. Specific examples of the network can include a wireless network provided by a communication provider of the terminal 10. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, NIC) which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (Radio Frequency, RF) module which is used to communicate with the Internet in a wireless manner.
[0035] The embodiments of the present application provide a ball-picking machine control method based on the YOLOV5 algorithm running on the terminal, Figure 2 is a flowchart of the ball-picking machine control method based on the YOLOV5 algorithm according to the embodiments of the present application, as shown in FIG. 2, the flow includes the following steps: Figure 2
[0036] In step S201, a video in a badminton training process is acquired, wherein the video includes multiple video image frames.
[0037] In the embodiment, the video is collected by the monitoring device (for example, a camera) installed on the ball picking robot. The ball picking robot moves in the corresponding sports field and collects multiple frames of video image frames of the badminton training process in the sports field in real time through the installed monitoring device. In the embodiment, in order to capture the scattered badminton in a larger range, the ball picking robot is provided with a long-focus camera and a wide-angle camera. Therefore, the collected video includes the videos collected by the two cameras, and when the ball is identified, the badminton is detected in the visual angle region corresponding to the two cameras in turn.
[0038] In step S202, a pre-trained ball recognition network is used to detect candidate balls in each frame of video image and determine the bounding box corresponding to the candidate ball. The ball recognition network is a neural network trained based on the YOLOV5 algorithm and using a preset fused shallow feature as a shallow feature of the YOLOV5 algorithm. The fused shallow feature is generated by fusing a shallow feature extracted from a sample image by the YOLOV5 algorithm and a preset deep feature. The deep feature includes preset coarse-grained semantic information for representing the composition features of the badminton.
[0039] In the embodiment, the ball recognition network is used to identify the target in the video image frame of the video, that is, to detect whether there is a badminton in the visual angle region corresponding to the monitoring device in the video image frame. In the embodiment, the target identified by the ball recognition network is represented by a rectangular bounding box, and the bounding box labels the target class information and the target confidence corresponding to the identified target. After the bounding box corresponding to the target in the video image frame is identified, according to the target class information and the target confidence of the target, the target whose target class information is a set class (for example, a badminton) and whose target confidence meets a set threshold is selected as a candidate ball, and the bounding box corresponding to the target is the bounding box corresponding to the candidate ball described in the embodiment.
[0040] In the embodiment, the collected video image frames corresponding to different cameras are different, and the detection of the target and the object ball is preferentially performed in the visual angle region of the video image frame collected by the corresponding camera in a set detection order. For example, the target detection is preferentially performed in the visual angle region of the video image frame collected by the wide-angle camera, and then the target detection is performed in the visual angle region of the video image frame collected by the long-focus camera.
[0041] It should be understood that since the targets (badminton) identified and picked up by the ball picking machine in the embodiment of the present application are small objects, although a sufficient number of small object data sets and a larger image input model can be used to train and test the obtained network model for detection, the network model trained in this way will sacrifice detection speed when performing target recognition. For example, the detection speed of the 640 input model is more than twice that of the 1280 input model. In this embodiment, in order to improve the detection speed, the network structure of YOLOV5 is improved to improve the algorithm's detection ability for small objects; it is understandable that high-resolution shallow features are very important for detecting small objects; the feature map of the shallow features retains relatively complete feature information of small objects due to its high resolution, which can facilitate the detection of small objects, while the deep features are closer to the output and are some coarse-grained information, containing more abstract information, that is, semantic information, but semantic information can better characterize the characteristics of small objects, such as the components of badminton such as feathers, heads and rubber strips. It is understandable that the pixels of shallow features are high but the semantic information is small, and the pixels of deep features are low but have more semantic information; in this embodiment, the shallow features extracted from the CSP1_1 module of YOLOV5 (corresponding to shallow information with high resolution) are compared with the pre-extracted deep features (low resolution, more semantic information) for shallow information (fine-grained information, such as image color information, texture The YOLOV5 algorithm compares the shallow features of the deep features with the shallow features (information, edge information, and angular information). When the similarity of the shallow information of the two is within the preset similarity threshold range, the two are considered to be similar. The semantic information of the corresponding deep features is assigned to the corresponding shallow features, so that the shallow features have high resolution while also containing certain semantic information, that is, the shallow features are fused. Then, the YOLOV5 algorithm connects the fused shallow features to the first CSP1_3 module to improve the accuracy of small object detection. It can be understood that training the corresponding network based on the YOLOV5 algorithm is well-known and feasible, and training the ball recognition network of the embodiment of the present application based on the YOLOV5 algorithm does not constitute an unclear limitation on the scheme of the ball picking machine control method of the embodiment of the present application. In this embodiment, on the basis of the existing network trained and matched based on the YOLOV5 algorithm, by introducing the fusion of shallow features, the trained ball recognition network does not need to perform convolution and pooling on the entire image during target detection, thereby speeding up the reasoning speed.
[0042] Step S203 : Filtering the target sphere from the candidate spheres according to the distance information between the bounding box and the center line of the corresponding video image frame.
[0043] In the embodiment, after the candidate balls are identified, i.e., the bounding boxes are obtained, the target ball that is the first to be picked up currently needs to be selected from the candidate balls. In the embodiment, the target ball is the ball closest to the ball picking machine. In order to determine the ball closest to the ball picking machine, the relative position relationship or distance between the bounding box representing the candidate ball and the baseline, i.e., the middle line of the video image frame in which the candidate ball appears, is calculated, and the ball closest to the baseline is selected according to the relative position relationship or distance, so as to obtain the target ball that is preferentially picked up. In the embodiment, the middle line of the video image frame is a known position parameter. Specifically, the position parameter of the visual angle region of the corresponding camera located on the ball picking machine can be known in advance. The video image frame and the visual angle region of the camera are in one-to-one mapping, i.e., each video image frame represents the picture or image appearing in the visual angle region of the corresponding camera. When the position parameter of the visual angle region of the corresponding camera is known, the position and coordinate data of the middle line of the video image frame are known, and there is no need to detect the middle line of the video image frame in real time by using an external middle line detection algorithm. In the embodiment, the position of the middle line of the video image frame can also be used to specify the position of the ball picking machine. According to the positional relationship between the middle line of the video image frame and the centroid of the bounding box of the candidate ball appearing in the visual angle region of the camera, the positional relationship between the actual feather ball to be picked up and the ball picking machine can be determined.
[0044] In the embodiment, after the candidate balls are identified and the bounding box corresponding to each candidate ball is determined, the relevant coordinate data corresponding to each bounding box is determined and known, for example, the four vertex coordinates of the rectangular bounding box. Then, based on the relevant coordinate data corresponding to the bounding box and the known position and coordinate data of the middle line of the video image frame, the distance information of the bounding box relative to the middle line of the video image frame can be calculated, so as to determine the ball closest to the middle line, and thus obtain the target ball.
[0045] In some optional embodiments, the distance between the centroid of the bounding box and the corresponding middle line is used to calculate the distance between the bounding box and the middle line of the video image frame. In the calculated distances, the bounding box closest to the middle line is selected in a traversal manner. For example, the candidate ball is denoted as X i , the set of the plurality of candidate balls X i is denoted as S, the centroid of the bounding box of each candidate ball is denoted as (x i , y i ), the corresponding middle line is denoted as l, each candidate ball X i in S is traversed, the distance d i between the centroid of each candidate ball X i and the corresponding middle line is calculated, and min{d1, d2, …, d n} is selected as the target ball Xm .
[0046] Step S204, according to the preset moving driving mode, control the ball picking machine to drive to the real scene field position corresponding to the position of the target ball, and take the shuttlecock in the real scene field position by wind.
[0047] In this embodiment, when the target ball and its corresponding bounding box are determined, the appropriate moving driving mode is selected according to the position of the target ball, so that the ball picking machine performs differential steering, and the corresponding middle line and the center of mass of the target ball are on the same straight line, and then the ball picking machine drives in a straight line to the real scene field position of the target ball.
[0048] In this embodiment, after the ball picking machine drives in a straight line to the real scene field position of the target ball, the opened fan works at a set wind speed, thereby generating appropriate suction force to safely recover the shuttlecock.
[0049] It should be noted that the posture, mass and use of the shuttlecock will affect the picking of the shuttlecock; in this embodiment, at least based on the weight and contour of the shuttlecock, the shuttlecock is approximated as a cone with one side close to the ground, according to Bernoulli effect and Bernoulli equation in fluid dynamics, it can be deduced that the place with high air flow rate has low pressure, wherein the Bernoulli equation is:
[0050]
[0051] Wherein, p is the pressure of a point in the fluid, v is the flow rate of the fluid at that point, p is the fluid density, g is the acceleration of gravity, h is the height of the point, and C is a constant.
[0052] The above Bernoulli equation can also be expressed as:
[0053]
[0054] Therefore, under the action of the appropriate fan, the wind speed changes and the wind speed above and below the shuttlecock is not the same, so that the shuttlecock obtains upward lift; in this embodiment, the fan is controlled to work at the wind speed determined based on the Bernoulli equation, so that the shuttlecock is safely taken without damaging the shuttlecock, and multiple shuttlecocks are quickly collected.
[0055] Through the steps S201 to S204, the video in the badminton training process is acquired, the video includes multiple video image frames; a pre-trained ball recognition network is used to detect candidate balls in each video image frame and determine the bounding box corresponding to the candidate ball, the ball recognition network is based on the YOLOV5 algorithm, and a neural network trained by using a preset fusion shallow feature as a shallow feature of the YOLOV5 algorithm, the fusion shallow feature is generated by fusing a shallow feature extracted from a sample image by using the YOLOV5 algorithm and a preset deep feature, and the deep feature includes preset coarse-grained semantic information for representing a badminton composition feature; the target ball is selected from the candidate balls according to distance information of the bounding box and a center line of the corresponding video image frame; the ball picking machine is controlled to move to a real scene location corresponding to a location where the target ball is located according to a preset movement driving mode, and the badminton ball on the real scene location is picked up by using wind power; by fusing the deep feature with the preset coarse-grained semantic information for representing the badminton composition feature and the shallow feature, the fused shallow feature has high pixels and multiple semantics, the ball recognition network is trained by using the fused shallow feature, the robustness and the detection capability of small objects of the ball recognition network are enhanced, the detection accuracy is improved, and the effectiveness and efficiency of badminton ball recognition are improved; meanwhile, the driving path is planned based on the position of the recognized ball, the ball picking machine moves to the badminton ball to be picked up according to the corresponding movement driving mode, the ball is picked up by using wind power, the ball picking efficiency is improved, and the damage to the ball in the process of picking up the badminton ball is reduced, the problem of low detection accuracy and ball picking efficiency in the related art is solved, and the beneficial effects of enhancing the ball recognition capability and the detection accuracy and improving the ball picking efficiency are achieved.
[0056] It should be noted that in the embodiments of the present application, after the ball picking machine starts to work, the fan does not work and searches for the nearby badminton ball by moving around the ground in a circle; after the ball picking machine recognizes the badminton ball, the ball picking machine will move towards the direction where the badminton ball is recognized and start the ball suction fan when it is about to reach the position of the badminton ball, when the badminton ball approaches the suction air duct opening of the ball picking machine, the badminton ball will be pressed into the interior of the ball picking machine by the air pressure difference and fall into the built-in badminton ball barrel of the ball picking machine in a stacked manner.
[0057] In some embodiments, the target ball is selected from the candidate balls according to distance information of the bounding box and a center line of the corresponding video image frame, including the following steps:
[0058] Step 21, determining the straight line distance from the center of mass of the bounding box of each candidate ball to the center line of the corresponding video image frame.
[0059] In the present embodiment, the distance of the bounding box from the center line of the video image frame is calculated based on the distance of the center of mass of the bounding box from the corresponding center line.
[0060] Step 22, in order of distance from small to large, select the candidate sphere whose straight line distance is less than the distance threshold value, to obtain the target sphere.
[0061] In this embodiment, all the identified candidate spheres are traversed, and the centroid of the bounding box of each candidate sphere and the straight line distance of the corresponding midline are calculated respectively. Then, according to the size of the straight line distance, the candidate sphere whose straight line distance meets the set requirement (the set distance threshold value) is selected as the target sphere.
[0062] In some optional embodiments, the candidate sphere corresponding to the smallest straight line distance is selected from the plurality of straight line distances arranged in order of distance from small to large, to obtain the target sphere.
[0063] By determining the straight line distance from the centroid of the bounding box of each candidate sphere to the midline of the corresponding video image frame in steps 21 to 23 above; in order of distance from small to large, the candidate sphere whose straight line distance is less than the distance threshold value is selected, to obtain the target sphere, the determination of the first picked shuttlecock after completing the sphere recognition is realized, and the effectiveness of the shuttlecock recognition and the efficiency of the ball picking are improved.
[0064] In some embodiments, determining the straight line distance from the centroid of the bounding box of each candidate sphere to the midline of the corresponding video image frame comprises the following steps:
[0065] Step 31, obtaining the first coordinate data corresponding to the bounding box, and determining the first centroid coordinate corresponding to the bounding box according to the first coordinate data, wherein the first coordinate data is used to represent the vertex coordinates of the bounding box.
[0066] In this embodiment, after the pre-trained sphere recognition network is used to identify the candidate sphere and the bounding box corresponding to the candidate sphere, the first coordinate data (for example, the coordinates of the four vertices of the rectangular frame) corresponding to the bounding box is also known. Then, the first centroid coordinate of the bounding box is calculated and determined through the known first coordinate data corresponding to the bounding box.
[0067] Step 32, determining the second coordinate data corresponding to the midline of each frame of video image according to the preset window parameter, wherein the window parameter is used to represent the angle size corresponding to the monitoring device for collecting the video.
[0068] In this embodiment, the preset window parameter refers to the position parameter of the angle region of the corresponding camera on the ball picking machine. The position parameter is known in advance. The video image frame and the angle region of the camera are one-to-one mapping, that is, each frame of video image frame corresponds to the picture or image appearing in the angle region of the corresponding camera. In the case that the position parameter of the angle region of the corresponding camera is known, the position and coordinate data of the midline of the video image frame are known, that is, the second coordinate data is known in advance.
[0069] Step 33, according to the first centroid coordinate and the second coordinate data, calculating the distance between the bounding box and the center line of the video image frame where the bounding box is located, obtaining the straight line distance from the centroid of the corresponding bounding box to the center line of the corresponding video image frame.
[0070] In the embodiment, the calculation method or formula for calculating the distance between the bounding box and the center line of the video image frame where the bounding box is located based on the first centroid coordinate and the second coordinate data is not limited in the embodiment, and it needs to be understood that the corresponding calculation method or formula does not constitute an unclear limitation on the calculation of the corresponding distance, that is, the calculation method or formula for calculating the distance between the bounding box and the center line of the video image frame where the bounding box is located based on the first centroid coordinate and the second coordinate data is implementable.
[0071] In the embodiment, the calculation method or formula for calculating the distance between the bounding box and the center line of the video image frame where the bounding box is located based on the first centroid coordinate and the second coordinate data is not limited in the embodiment, and it needs to be understood that the corresponding calculation method or formula does not constitute an unclear limitation on the calculation of the corresponding distance, that is, the calculation method or formula for calculating the distance between the bounding box and the center line of the video image frame where the bounding box is located based on the first centroid coordinate and the second coordinate data is implementable.
[0072] In some embodiments, the movement driving mode includes a differential steering driving function, and the robot is controlled to drive to the real scene field position corresponding to the position of the target ball according to the preset movement driving mode, including the following steps:
[0073] Step 41, determining the target bounding box corresponding to the target ball, and determining the second centroid coordinate corresponding to the target bounding box according to the obtained second coordinate data corresponding to the target bounding box, wherein the second coordinate data is used to represent the vertex coordinates of the target bounding box.
[0074] Step 42, judging the positional relationship between the target bounding box and the center line of the video image frame where the target bounding box is located, and determining the target differential steering driving function from the preset plurality of differential steering driving functions according to the judgment result.
[0075] Step 43, controlling the robot to adjust the pose to align the real scene field position corresponding to the position of the target ball according to the target speed determined according to the target differential steering driving function and the second centroid coordinate, and controlling the robot to drive to the real scene field position.
[0076] In the embodiment, after the target sphere and the corresponding target bounding box are determined, a parameter (set as x) is used to determine the position of the target sphere relative to the corresponding center line, that is, to determine the positional relationship between the target bounding box and the center line of the video image frame in which the target bounding box is located. For example, when the target sphere is located on the left side of the center line, the corresponding value of the parameter is negative, for example, x ∈ (-1, 0). For another example, when the target sphere is located on the right side of the center line, the corresponding value of the parameter is positive, for example, x ∈ (0, 1).
[0077] In the embodiment, after the positional relationship between the target bounding box and the center line of the video image frame in which the target bounding box is located is determined, the target differential steering drive function is obtained according to the determination result (the value of x) and the position of the target sphere (corresponding to the second centroid coordinate) based on the following differential steering drive function formula f(x),
[0078]
[0079] For example, when x ∈ (-1, -0.1), the target differential steering drive function is f(x) = -k (-1 ≤ x ≤ -0.5). Then, the ball-picking robot is controlled to drive the corresponding wheel motor speed difference to change. It can be understood that the shorter the distance from the centroid of the target bounding box of the target sphere to the corresponding center line, the smaller the absolute value of the wheel motor speed difference of the ball-picking robot, and the smaller the distance adjusted by the ball-picking robot. When the ball-picking robot is adjusted to the corresponding center line coinciding with the centroid of the target bounding box of the target sphere, the wheel motor speed difference is 0.
[0080] In the above steps, the target bounding box corresponding to the target sphere is determined, and the second centroid coordinate corresponding to the target bounding box is determined according to the obtained second coordinate data corresponding to the target bounding box, wherein the second coordinate data is used to represent the vertex coordinates of the target bounding box. The positional relationship between the target bounding box and the center line of the video image frame in which the target bounding box is located is determined, and the target differential steering drive function is determined from the plurality of preset differential steering drive functions according to the determination result. The ball-picking robot is controlled to adjust the pose to align the position of the target sphere in the real scene field position according to the target speed determined by the target differential steering drive function and the second centroid coordinate, and the ball-picking robot is controlled to drive to the real scene field position, so as to adjust the ball-picking robot and the target sphere to be in a straight line, thereby quickly driving to the target sphere to complete the picking, and improving the ball-picking efficiency.
[0081] In some embodiments, the video image frame includes a wide-angle video frame and a long-focus video frame. A pre-trained sphere recognition network is used to detect candidate spheres in each video image frame and determine the bounding box corresponding to the candidate sphere, including the following steps:
[0082] Step 51, sequentially selecting a wide-angle video frame and a telephoto video frame from the video image frames.
[0083] Step 52, sequentially performing object sphere detection in the wide-angle video frame and the telephoto video frame by using a sphere recognition network to obtain a detection result, wherein the detection result includes a bounding box of the object sphere and a confidence corresponding to the object sphere.
[0084] Step 53, selecting an object sphere with a confidence greater than a preset confidence threshold according to the confidence corresponding to the object sphere to obtain a candidate sphere, and taking the bounding box of the object sphere with the confidence greater than the preset confidence threshold as a bounding box corresponding to the candidate sphere.
[0085] In some embodiments, the shallow features extracted from the sample image frames by using the YOLOV5 algorithm and the preset deep features are fused to generate fused shallow features, including the following steps:
[0086] Step 61, extracting initial shallow features from the preprocessed sample image by using the YOLOV5 algorithm, wherein the preprocessing includes one of the following: data enhancement, adaptive anchor frame calculation, picture random scaling, image random cropping, image random arrangement and splicing, the initial shallow features include first fine-grained information of at least one of the following: image color information, image texture information, image edge information, image corner information, and the sample image includes video image frames obtained from an existing badminton training video.
[0087] Step 62, obtaining second fine-grained information and a plurality of target semantic information corresponding to the preset deep features, wherein the second fine-grained information includes at least one of the following: image color information, image texture information, image edge information, image corner information, and the target semantic information is used to represent the composition features of the badminton.
[0088] Step 63, judging whether the similarity of the second fine-grained information and the first fine-grained information is within a preset similarity threshold interval.
[0089] Step 64, in the case where it is judged that the similarity is within the preset similarity threshold interval, assigning the plurality of target semantic information possessed by the deep features to the initial shallow features to obtain the fused shallow features.
[0090] In the embodiment, because the high-resolution shallow features are very important for detecting small objects, the feature maps of the shallow features retain relatively complete feature information of small objects due to their high resolution, which facilitates the detection of small objects, and the deep features are some coarse-grained information and contain more abstract information, i.e., semantic information, but the semantic information can better represent the features of small objects, for example, the constituent features of a shuttlecock, i.e., feathers, a ball head, and a rubber strip, which means that the shallow features have high pixels but less semantic information, and the deep features have low pixels but more semantic information. In the embodiment, the shallow features (corresponding to the shallow information with high resolution) extracted from the CSP1_1 module of YOLOV5 are compared with the pre-extracted deep features (low resolution and more semantic information) in terms of shallow information (fine-grained information, such as image color information, texture information, edge information, and corner information). When the similarity of the shallow information of the two is within a preset similarity threshold interval, it is considered that the two are similar. The semantic information possessed by the corresponding deep features is assigned to the corresponding shallow features, so that the shallow features contain certain semantic information while having high resolution, i.e., the shallow features are fused. Then, the YOLOV5 algorithm connects the fused shallow features to the first CSP1_3 module to improve the accuracy of small object detection. It can be understood that training the corresponding network based on the YOLOV5 algorithm is known and can be implemented, and training the ball recognition network of the embodiment based on the YOLOV5 algorithm does not constitute a limitation on the scheme of the ball picking machine control method of the embodiment. In the embodiment, based on the existing network trained and matched based on the YOLOV5 algorithm, the introduction of the fused shallow features enables the trained ball recognition network to not need to perform convolution and pooling on the entire image during target detection, thereby accelerating the inference speed.
[0091] Figure 3 is a flowchart of the ball picking machine control method of the preferred embodiment of the application, which includes the following steps:
[0092] Step S301, start the suction fan.
[0093] Step S302, determine whether there is a candidate ball in the video image frame corresponding to the wide-angle camera, if yes, execute step S303, otherwise, execute step S305.
[0094] Step S303, the ball picking machine turns and steers while advancing, and then executes step S304.
[0095] Step S304, the ball picking machine sucks the shuttlecock, and then executes step S302.
[0096] Step S305: Determine whether there is a candidate sphere in the video image frame corresponding to the telephoto camera. If yes, execute step S303; otherwise, execute step S306.
[0097] Step S306: the ball picking machine rotates clockwise, and then executes step S302.
[0098] The control method of the ball picking machine according to the preferred embodiment of the present application is described below. The control method includes the following steps:
[0099] Step 1. Turn on the suction fan.
[0100] Step 2: Target determination: Assume that there is a badminton target X in the video image frame corresponding to the wide-angle camera and the video image frame corresponding to the telephoto camera. i , X i The set is denoted as S. If there is a badminton in the video image frames corresponding to the two cameras, the badminton in the video image frame corresponding to the wide-angle camera (the badminton close to the body of the ball picker) is given priority; if there is no badminton in the video image frame corresponding to the wide-angle camera, then determine whether there is a badminton in the video image frame corresponding to the telephoto camera (the badminton far from the body of the ball picker), and control the ball picker through the corresponding mobile drive mode; if there is no candidate ball in the video image frames corresponding to the two cameras, the ball picker is controlled to rotate clockwise.
[0101] Step 3: Select the target. The target set S consists of all the badmintons identified by the ball picking machine. In the target set S, let each target X i The center of mass marked by the rectangle is (x i ,y i ), let the midline of the video image frame corresponding to the wide-angle camera and the video image frame corresponding to the telephoto camera be l. Since the badminton balls are scattered at different distances, the sizes of the rectangular marking frames are different. By traversing the set S, calculate each target X in the set S. i The distance d from the center of mass to the corresponding midline i , select min{d1, d2, ..., d n} as the target sphere X m .
[0102] Step S4: drive the ball picker to turn at a differential speed, so that the target X m The center of mass (x m ,y m ) and the midline of the video image frame corresponding to the wide-angle camera or the video image frame corresponding to the telephoto camera are located in a straight line and then move forward, and finally are successfully picked up by the ball picking machine.
[0103] Step 5, when a qualified shuttlecock is successfully picked up, continue to iterate steps 1 to 4, and after judging that there is no shuttlecock in the field, exit and terminate.
[0104] The embodiment also provides a control device of the YOLOV5-based ball picking machine, which is used to implement the above-mentioned embodiments and preferred embodiments and has been described above. As used below, the terms "module", "unit", "sub-unit", etc. can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware, or a combination of software and hardware is also possible and is contemplated.
[0105] Figure 4 is a structural block diagram of the control device of the YOLOV5-based ball picking machine according to the embodiment of the present application, as Figure 4 shown, the device comprises an acquisition module 41, an identification module 42, a screening module 43, and a processing module 44, wherein,
[0106] The acquisition module 41 is configured to acquire a video in a shuttlecock training process, wherein the video comprises a plurality of video image frames.
[0107] The identification module 42 is coupled to the acquisition module 41 and is configured to detect candidate balls in each video image frame by using a pre-trained ball recognition network, and determine a bounding box corresponding to the candidate ball, wherein the ball recognition network is based on a YOLOV5 algorithm, and is a neural network trained by using a preset fusion shallow feature as a shallow feature of the YOLOV5 algorithm, the fusion shallow feature is generated by fusing a shallow feature extracted from a sample image frame by using the YOLOV5 algorithm and a preset deep feature, and the deep feature comprises preset coarse-grained semantic information for representing a composition feature of the shuttlecock.
[0108] The screening module 43 is coupled to the identification module 42 and is configured to screen a target ball from the candidate ball according to distance information between the bounding box and a center line of the corresponding video image frame.
[0109] The processing module 44 is coupled to the screening module 43 and is configured to control the ball picking machine to move to a real scene field position corresponding to a position where the target ball is located according to a preset movement driving mode, and to suck the shuttlecock on the real scene field position by using wind power.
[0110] The control device of the ball picking machine based on YOLOV5 according to the embodiment of the present application adopts the video obtained in the process of badminton training, the video including multiple video image frames; a pre-trained ball recognition network is used to detect candidate balls in each video image frame and determine the bounding box corresponding to the candidate ball, the ball recognition network being based on the YOLOV5 algorithm and using a preset fused shallow feature as a shallow feature of the YOLOV5 algorithm to train a neural network, the fused shallow feature being generated by fusing a shallow feature extracted from a sample image by the YOLOV5 algorithm and a preset deep feature, the deep feature including preset coarse-grained semantic information for representing the composition feature of the badminton; the target ball is selected from the candidate balls according to the distance information of the bounding box and the center line of the corresponding video image frame; the ball picking machine is controlled to move to the real scene location corresponding to the location of the target ball in a preset moving driving mode, and the badminton on the real scene location is sucked by wind power; the deep feature with the preset coarse-grained semantic information for representing the composition feature of the badminton is fused with the shallow feature, so that the fused shallow feature has high pixels and multiple semantics, and the ball recognition network is trained by using the fused shallow feature, thereby enhancing the robustness and detection capability of small objects of the ball recognition network, improving the detection accuracy, and improving the effectiveness and efficiency of badminton recognition; at the same time, the driving path is planned based on the position of the recognized ball, and the ball picking machine is quickly moved to the badminton to be picked up in the corresponding moving driving mode, and the ball is picked up by wind power, thereby improving the ball picking efficiency and reducing the damage to the ball in the process of picking up the badminton, solving the problem of low detection accuracy and ball picking efficiency in the related art, and achieving the beneficial effects of enhancing the ball recognition capability and detection accuracy and improving the ball picking efficiency.
[0111] In some embodiments, the screening module 43 further includes:
[0112] A first determination unit is configured to determine the straight-line distance from the center of the bounding box of each candidate ball to the center line of the corresponding video image frame.
[0113] A selection unit is coupled to the first determination unit and configured to select, in order of increasing distance, the candidate ball with a straight-line distance less than a distance threshold to obtain the target ball.
[0114] In some embodiments, the selection unit is further configured to select, from the multiple straight-line distances arranged in order of increasing distance, the candidate ball corresponding to the smallest straight-line distance to obtain the target ball.
[0115] In some embodiments, the first determining unit is further configured to obtain first coordinate data corresponding to the bounding box, and determine a first centroid coordinate corresponding to the bounding box according to the first coordinate data, wherein the first coordinate data is used to represent vertex coordinates of the bounding box; determine second coordinate data corresponding to a center line of each frame of video image according to a preset window parameter, wherein the window parameter is used to represent a viewing angle size corresponding to a monitoring device for collecting the video; calculate a distance between the bounding box and the center line of the video image frame where the bounding box is located according to the first centroid coordinate and the second coordinate data, and obtain a straight line distance from the centroid of the corresponding bounding box to the center line of the corresponding video image frame.
[0116] In some embodiments, the mobile driving mode includes a differential steering driving function, and the processing module 44 further includes:
[0117] A second determining unit is configured to determine a target bounding box corresponding to a target ball, and determine a second centroid coordinate corresponding to the target bounding box according to second coordinate data corresponding to the target bounding box, wherein the second coordinate data is used to represent vertex coordinates of the target bounding box.
[0118] A first judging unit, coupled with the second determining unit, is configured to judge a positional relationship between the target bounding box and a center line of a video image frame where the target bounding box is located, and determine a target differential steering driving function from a plurality of preset differential steering driving functions according to a judgment result.
[0119] A control unit, coupled with the first judging unit, is configured to control the ball picking robot to adjust a pose to a real scene field position corresponding to a position where the target ball is located according to a target rotating speed determined according to the target differential steering driving function and the second centroid coordinate, and control the ball picking robot to drive to the real scene field position.
[0120] In some embodiments, the video image frame includes a wide-angle video frame and a telephoto video frame, and the identification module 42 further includes:
[0121] A selecting unit is configured to sequentially select the wide-angle video frame and the telephoto video frame from the video image frame.
[0122] A detecting unit, coupled with the selecting unit, is configured to sequentially perform object ball detection in the wide-angle video frame and the telephoto video frame by using a ball identification network to obtain a detection result, wherein the detection result includes a bounding box of the object ball and a confidence corresponding to the object ball.
[0123] A screening unit, coupled with the detecting unit, is configured to select an object ball with a confidence greater than a preset confidence threshold according to the confidence corresponding to the object ball to obtain a candidate ball, and take the bounding box of the object ball with the confidence greater than the preset confidence threshold as a bounding box corresponding to the candidate ball.
[0124] In some embodiments, the identification module 42 is further configured to extract initial shallow features from the pre-processed sample images using a YOLOV5 algorithm, wherein the pre-processing includes one of data augmentation, adaptive anchor box calculation, picture random scaling, image random cropping, image random arrangement and splicing, and the initial shallow features include first fine-grained information of at least one of image color information, image texture information, image edge information and image corner information, and the sample images include video image frames obtained from existing badminton training videos; obtain second fine-grained information and multiple target semantic information corresponding to the preset deep features, wherein the second fine-grained information includes at least one of image color information, image texture information, image edge information and image corner information, and the target semantic information is used to represent the composition features of the badminton; determine whether the similarity of the second fine-grained information and the first fine-grained information is within a preset similarity threshold interval; in a case where it is determined that the similarity is within the preset similarity threshold interval, assign the multiple target semantic information possessed by the deep features to the initial shallow features to obtain the fusion shallow features.
[0125] The embodiment also provides an electronic device including a memory and a processor, the memory storing a computer program, and the processor being configured to execute the computer program to perform the steps in any of the above method embodiments.
[0126] Optionally, the electronic device can further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.
[0127] Optionally, in the embodiment, the processor can be configured to execute the following steps through the computer program:
[0128] S1, obtaining a video in a badminton training process, wherein the video includes multiple video image frames.
[0129] S2, detecting candidate balls in each video image frame using a pre-trained ball recognition network, and determining a bounding box corresponding to the candidate balls, wherein the ball recognition network is a neural network trained based on a YOLOV5 algorithm and using a preset fusion shallow feature as a shallow feature of the YOLOV5 algorithm, the fusion shallow feature is generated by fusing a shallow feature extracted from a sample image using the YOLOV5 algorithm and a preset deep feature, and the deep feature includes preset coarse-grained semantic information used to represent composition features of the badminton.
[0130] S3, screening target balls from the candidate balls according to distance information of the bounding boxes and a center line of the corresponding video image frames.
[0131] S4, according to the preset mobile driving mode, controls the ball picking machine to drive to the real scene field position corresponding to the position of the target ball, and to suck the shuttlecock at the real scene field position by wind power.
[0132] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation manners, and this embodiment will not be repeated here.
[0133] In addition, in combination with the ball picking machine control method based on the YOLOV5 algorithm in the above embodiments, the present embodiment can provide a storage medium for implementation. The storage medium has a computer program stored thereon; the computer program is executed by a processor to implement any one of the ball picking machine control methods based on the YOLOV5 algorithm in the above embodiments.
[0134] Those skilled in the art should understand that the technical features of the above embodiments can be combined in any way, and in order to make the description concise, all possible combinations of the technical features in the above embodiments are not described, however, as long as the combination of the technical features does not exist contradictory, it should be considered as the scope of the present application.
[0135] The above embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the scope of the patent. It should be noted that for those skilled in the art, without departing from the concept of the present application, some modifications and improvements can be made, which are all within the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. A ball picking machine control method based on YOLOV5 algorithm, characterized in that, The method comprises: acquiring a video in a badminton training process, wherein the video comprises a plurality of video image frames; detecting candidate balls in each of the video image frames by using a pre-trained ball recognition network, and determining a bounding box corresponding to the candidate balls, wherein the ball recognition network is a neural network trained based on a YOLOV5 algorithm and using a preset fused shallow feature as a shallow feature of the YOLOV5 algorithm, the fused shallow feature is generated by fusing a shallow feature extracted from a sample image by using the YOLOV5 algorithm and a preset deep feature, and the deep feature comprises preset semantic information of a coarse granularity for representing a component feature of a badminton; screening a target ball from the candidate balls according to distance information of the bounding box and a center line corresponding to the video image frame; controlling a ball picking robot to move to a real scene location corresponding to a location where the target ball is located according to a preset moving driving mode, and using wind power to suck a badminton located at the real scene location, wherein the fused shallow feature is generated by fusing a shallow feature extracted from a sample image frame by using the YOLOV5 algorithm and a preset deep feature, comprising: extracting an initial shallow feature from a preprocessed sample image by using the YOLOV5 algorithm, wherein the preprocessing comprises one of the following: data enhancement, adaptive anchor frame calculation, picture random scaling, image random cropping, image random arrangement and splicing, the initial shallow feature comprises first fine-grained information of at least one of the following: image color information, image texture information, image edge information and image corner information, and the sample image comprises a video image frame acquired from an existing badminton training video; acquiring second fine-grained information corresponding to the deep feature and a plurality of target semantic information, wherein the second fine-grained information comprises at least one of the following: image color information, image texture information, image edge information and image corner information, and the target semantic information is used to represent a component feature of a badminton; judging whether a similarity of the second fine-grained information and the first fine-grained information is within a preset similarity threshold interval; in a case where it is judged that the similarity is within the preset similarity threshold interval, assigning a plurality of the target semantic information possessed by the deep feature to the initial shallow feature to obtain the fused shallow feature.
2. The method of claim 1, wherein, screening a target ball from the candidate balls according to distance information of the bounding box and a center line corresponding to the video image frame, comprising: determining a straight line distance from a center of the bounding box of each of the candidate balls to the center line of the corresponding video image frame; selecting the candidate ball with a straight line distance less than a distance threshold in a descending order of distance to obtain the target ball.
3. The method of claim 2, wherein, selecting the candidate ball with a straight line distance less than a distance threshold in a descending order of distance, comprising: selecting the candidate ball corresponding to the smallest straight line distance from a plurality of straight line distances arranged in a descending order of distance to obtain the target ball.
4. The method of claim 2, wherein, The straight-line distance of the centroid of the bounding box of each candidate sphere to the center line of the corresponding video image frame is determined, including: Obtaining first coordinate data corresponding to the bounding box, and determining first centroid coordinates corresponding to the bounding box according to the first coordinate data, wherein the first coordinate data is used to represent the vertex coordinates of the bounding box; According to the preset window parameter, the second coordinate data corresponding to the center line of each frame of the video image frame is determined, wherein the window parameter is used to represent the angle size corresponding to the monitoring device for collecting the video; According to the first centroid coordinates and the second coordinate data, the distance between the bounding box and the center line of the video image frame where the bounding box is located is calculated, and the straight-line distance of the centroid of the corresponding bounding box to the center line of the corresponding video image frame is obtained.
5. The method of claim 1, wherein, The moving driving mode includes a differential steering driving function, and the ball picking machine is controlled to drive to the real scene field position corresponding to the position of the target sphere according to the preset moving driving mode, including: Determine the target bounding box corresponding to the target sphere, and determine the second centroid coordinates corresponding to the target bounding box according to the obtained second coordinate data corresponding to the target bounding box, wherein the second coordinate data is used to represent the vertex coordinates of the target bounding box; Determine the position relationship between the target bounding box and the center line of the video image frame where the target bounding box is located, and determine the target differential steering driving function from the preset plurality of differential steering driving functions according to the determination result; According to the target rotational speed determined according to the target differential steering driving function and the second centroid coordinates, the ball picking machine is controlled to adjust the pose to align the real scene field position corresponding to the position of the target sphere, and the ball picking machine is controlled to drive to the real scene field position.
6. The method of claim 1, wherein, The video image frame includes a wide-angle video frame and a long-focus video frame, and a pre-trained sphere recognition network is used to detect candidate spheres in each frame of the video image frame and determine the bounding box corresponding to the candidate sphere, including: The wide-angle video frame and the long-focus video frame are sequentially selected from the video image frame; The sphere recognition network is used to sequentially perform object sphere detection in the wide-angle video frame and the long-focus video frame to obtain detection results, wherein the detection results include the bounding box of the object sphere and the confidence corresponding to the object sphere; According to the confidence corresponding to the object sphere, the object sphere with a confidence greater than a preset confidence threshold is selected to obtain the candidate sphere, and the bounding box of the object sphere with a confidence greater than the preset confidence threshold is taken as the bounding box corresponding to the candidate sphere.
7. A ball picking machine control device based on a YOLOV5 algorithm, characterized in that, Including: An acquisition module is configured to acquire a video in a badminton training process, wherein the video includes a plurality of video image frames; The recognition module is used to use a pre-trained sphere recognition network to detect candidate spheres in each frame of the video image and determine the bounding box corresponding to the candidate sphere, wherein the sphere recognition network is based on the YOLOV5 algorithm and is a neural network trained using preset fused shallow features as shallow features of the YOLOV5 algorithm. The fused shallow features are generated by fusing shallow features extracted from sample image frames by the YOLOV5 algorithm with preset deep features. The deep features include preset coarse-grained semantic information for characterizing the composition features of the badminton. The recognition module is also used to extract initial shallow features from the preprocessed sample image using the YOLOV5 algorithm, wherein the preprocessing includes one of the following: data enhancement, adaptive anchor box calculation, random image scaling, random image cropping, random image arrangement and splicing, and The initial shallow features include at least one of the following first fine-grained information: image color information, image texture information, image edge information, and image angular information, and the sample image includes a video image frame obtained from an existing badminton training video; obtaining second fine-grained information corresponding to the preset deep features and multiple target semantic information, wherein the second fine-grained information includes at least one of the following: image color information, image texture information, image edge information, and image angular information, and the target semantic information is used to characterize the component features of badminton; determining whether the similarity between the second fine-grained information and the first fine-grained information is within a preset similarity threshold interval; and if it is determined that the similarity is within the preset similarity threshold interval, assigning the multiple target semantic information possessed by the deep features to the initial shallow features to obtain the fused shallow features; a screening module, configured to screen out a target sphere from the candidate spheres based on distance information between the bounding box and a center line corresponding to the video image frame; The processing module is used to control the ball picking machine to move to the real scene venue position corresponding to the position of the target ball according to a preset mobile driving mode, and to absorb the badminton at the real scene venue position by wind power. 8.An electronic device comprising a memory and a processor, the electronic device comprising: A computer program is stored in the memory, and the processor is configured to run the computer program to perform the steps of the ball picking machine control method based on the YOLOV5 algorithm according to any one of claims 1 to 6.
9. A computer readable storage medium having stored thereon a computer program, characterized in that, When the computer program is executed by a processor, the steps of the ball picking machine control method based on the YOLOV5 algorithm according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Method for detecting moving small target in complex scene
CN112364865A
Positioning and ball picking system special for tennis match and method of positioning and ball picking system
CN115171026A