Video processing methods, apparatus, devices, storage media, and programs

JP2026520765APending Publication Date: 2026-06-24BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2024-10-09
Publication Date
2026-06-24

Smart Images

  • Figure 2026520765000001_ABST
    Figure 2026520765000001_ABST
Patent Text Reader

Abstract

This disclosure provides a video processing method and apparatus, device and storage medium, the method comprising: performing speech recognition on audio in a target video draft; obtaining an initial text segment of the target video draft; linguistically converting the initial text segment of the target video draft to obtain a target text segment based on the initial language type and target language type of the target video draft; generating a target audio segment based on the target text segment; and generating an edited video draft corresponding to the target video draft based on the target audio segment. Embodiments of this disclosure further generate an edited video draft by converting the initial text segment of the target video draft to a target text segment and generating a target audio segment belonging to a target language type based on the target text segment.
Need to check novelty before this filing date? Find Prior Art

Claims

1. A video processing method, Perform speech recognition on the audio in the target video draft to obtain the initial text segment of the target video draft, Based on the initial language type and target language type of the target video draft, a language conversion is performed on the initial text segment to obtain a target text segment, the initial language type is used to represent the language type to which the target video draft belongs, and the target text segment belongs to the target language type. Based on the target text segment, generate a target audio segment belonging to the target language type, A video processing method comprising generating an edited video draft corresponding to the target video draft based on the target audio segment, wherein there is a correspondence between the target audio segment in the edited video draft and the initial text segment of the target video draft, and the edited video draft belongs to the target language type.

2. Before performing language conversion on the initial text segment to obtain the target text segment based on the initial language type and target language type of the target video draft, Based on the initial text segment of the target video draft, the initial language type of the target video draft is determined, The method according to claim 1, further comprising receiving a target language type input to the target video draft.

3. Determining the initial language type of the target video draft based on the initial text segment of the target video draft is: Convert the characters in the initial text segment of the target video draft into a pre-defined type of code, The method of claim 2, further comprising determining the initial language type of the target video draft based on a code range to which a pre-set type of code corresponding to the character belongs.

4. Based on the initial language type and target language type of the target video draft, performing language conversion on the initial text segment to obtain the target text segment is: The method according to claim 1, further comprising determining the initial text segment of the target video draft as the target text segment if it is determined that the initial language type and the target language type belong to different dialect types within the same language type.

5. Before performing language conversion on the initial text segment to obtain the target text segment based on the initial language type and target language type of the target video draft, Perform speech recognition on the audio in the target video draft and obtain the voiceprint information of the target video draft. The method according to claim 1, further comprising generating different text tracks for different voiceprint information, each for which an initial text segment corresponding to the different voiceprint information is placed.

6. Generating a target audio segment based on the aforementioned target text segment is, To determine the text-to-speech audio information corresponding to the first text track, The method according to claim 5, further comprising using the text-to-speech audio information to generate a target audio segment based on a target text segment on the first text track.

7. Before generating an edited video draft corresponding to the target video draft based on the target audio segment, The playback speed of the target audio segment is adjusted based on the time length of the initial audio segment corresponding to the target audio segment to obtain a modified target audio segment, and the initial audio segment belongs to the target video draft, and the difference between the time length of the modified target audio segment and the time length of the initial audio segment is not greater than a preset time length threshold, Accordingly, generating an edited video draft corresponding to the target video draft based on the target audio segment is: The method according to claim 1, comprising generating an edited video draft corresponding to the target video draft based on the target audio segment after the gear shift.

8. After generating an edited video draft corresponding to the target video draft based on the target audio segment, In response to a selection operation for a first target audio segment in the edited video draft, a plurality of candidate text contents corresponding to the first target audio segment are displayed. In response to a selection operation for a target candidate text content among the multiple candidate text contents, an audio segment corresponding to the target candidate text content is generated. The method according to claim 1, further comprising updating the first target audio segment in the edited result video draft using the audio segment corresponding to the target candidate text content.

9. Before generating an edited video draft corresponding to the target video draft based on the target audio segment, The method further includes determining the target text segment as the subtitle segment for the corresponding target audio segment. Accordingly, generating an edited video draft corresponding to the target video draft based on the target audio segment is: The method according to claim 1, comprising generating an edited video draft corresponding to the target video draft based on the target audio segment and the subtitle segment of the target audio segment.

10. Before performing speech recognition on the audio in the target video draft and obtaining the initial text segment corresponding to the target video draft, This further includes determining the target video draft based on the original video, Accordingly, after generating an edited video draft corresponding to the target video draft based on the target audio segment, The method according to claim 1, further comprising generating a target video belonging to the target language type, corresponding to the original video, based on the edited video draft, in response to a video export operation.

11. A video processing device, A speech recognition module for performing speech recognition on the audio in a target video draft and obtaining an initial text segment of the target video draft, Based on the initial language type and target language type of the target video draft, a language conversion is performed on the initial text segment to obtain a target text segment. The initial language type is used to represent the language type to which the target video draft belongs, and the target text segment is a conversion module belonging to the target language type. A first generation module for generating a target audio segment belonging to the target language type based on the target text segment, A video processing device comprising a second generation module used to generate an edited video draft corresponding to the target video draft based on the target audio segment, wherein there is a correspondence between the target audio segment in the edited video draft and the initial text segment of the target video draft, and the edited video draft belongs to the target language type.

12. A computer-readable storage medium wherein a command is stored in the computer-readable storage medium, and when the command is executed by a terminal device, the terminal device is made to implement the method according to any one of claims 1 to 10.

13. A video processing device comprising memory, a processor, and a computer program stored in the memory and executable by the processor, wherein the processor, when executing the computer program, realizes the method according to any one of claims 1 to 10.