The application relates to the technical field of intelligent human-computer interaction, in particular to an aerial handwritten text collection method based on an improved yolov5 network and
monocular vision, which comprises the following steps: constructing an initial fingertip dataset and performing pretreatment to obtain a target fingertip dataset, improving the yolov5 network based on an efficient
pyramid segmentation attention module EPSA and a weighted bidirectional feature
pyramid network BiFPN, obtaining an improved yolov5 network, inputting a
training set of the target fingertip dataset into the improved yolov5 network to obtain a
fingertip detection model, obtaining real-time two-dimensional video images of aerial handwritten text based on a
monocular camera and inputting the
fingertip detection model, and forming the aerial handwritten text based on a coordinate
system virtual sliding technology, so that the technical problems that, in the prior art, a 3D sensor is large and expensive, resulting in insufficient universality, the requirement for a WIFI environment is relatively strict, resulting in relatively large limitation, and a fingertip is a
small target, resulting in low detection accuracy and high omission rate when the yolov5 network detects the fingertip are solved.