This application relates to the field of audio
processing technologies, and provides a voice
processing method and an electronic device. If the electronic device receives a long-press operation after transcribing a voice file into a text, it indicates that a user may intend to select speech content located in different segments simultaneously. The electronic device may concatenate all displayed speech information to obtain a concatenated character string. Next, the electronic device determines a target character range based on a start position and an end position corresponding to the long-press operation in combination with the concatenated character string. The target character range represents selected characters corresponding to the long-press operation in the concatenated character string. Next, the electronic device displays the concatenated character string in
plain text format, and sets a selection state for characters in the target character range to achieve the cross-segment selection of speech information. The electronic device may perform corresponding
processing, for example, sharing, searching,
copying, among other operations, according to a function control selected by the user, to provide the user with diverse services, thereby enhancing user experience.