This invention belongs to the field of video target monitoring systems, specifically relating to a method for efficiently overlaying text labels based on video target classification. The method includes the following steps: acquiring video
stream data; decoding the video
stream and outputting video frame data; placing the video frame data into a display
queue for further
processing; performing target recognition
inference on the video frames, outputting the target's category name, confidence level information, and the target's position information in the image; querying the target's category name and confidence level information in a
label cache; if a corresponding text
label image exists, it is directly reused; otherwise, a corresponding text
label image is rendered and stored in the label cache; after the query is completed, the images are stitched together to form a complete target label image; the complete target label image is overlaid on the video frames for
annotation; and the annotated video frames are encoded and pushed.