This application discloses a
subtitle generation method, device, and storage medium, relating to the field of
speech recognition technology. The method includes: obtaining a cross-process
file descriptor through a first channel between a
server and a
client, wherein the cross-process
file descriptor is used to characterize the reading end of a second channel; responding to a
subtitle startup command, sending an audio capture command through the first channel, and reading an audio
data stream from the reading end of the second channel according to the cross-process
file descriptor, wherein the audio
data stream is captured by the
server and written to the writing end of the second channel; and
parsing and recognizing the audio
data stream to obtain
subtitle text data. This application constructs an architecture that decouples hierarchical permissions from cross-
process communication, with the
server possessing
system-level permissions responsible for audio capture. This eliminates the need for cumbersome pop-up confirmation operations when the
client application starts subtitles, thus eliminating interruptions to the user's viewing continuity and achieving seamless interaction from startup.