Video processing methods, apparatus, devices, storage media, and programs
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- BEIJING ZITIAO NETWORK TECH CO LTD
- Filing Date
- 2024-10-09
- Publication Date
- 2026-06-24
Smart Images

Figure 2026520765000001_ABST
Abstract
Claims
1. A video processing method, Perform speech recognition on the audio in the target video draft to obtain the initial text segment of the target video draft, Based on the initial language type and target language type of the target video draft, a language conversion is performed on the initial text segment to obtain a target text segment, the initial language type is used to represent the language type to which the target video draft belongs, and the target text segment belongs to the target language type. Based on the target text segment, generate a target audio segment belonging to the target language type, A video processing method comprising generating an edited video draft corresponding to the target video draft based on the target audio segment, wherein there is a correspondence between the target audio segment in the edited video draft and the initial text segment of the target video draft, and the edited video draft belongs to the target language type.
2. Before performing language conversion on the initial text segment to obtain the target text segment based on the initial language type and target language type of the target video draft, Based on the initial text segment of the target video draft, the initial language type of the target video draft is determined, The method according to claim 1, further comprising receiving a target language type input to the target video draft.
3. Determining the initial language type of the target video draft based on the initial text segment of the target video draft is: Convert the characters in the initial text segment of the target video draft into a pre-defined type of code, The method of claim 2, further comprising determining the initial language type of the target video draft based on a code range to which a pre-set type of code corresponding to the character belongs.
4. Based on the initial language type and target language type of the target video draft, performing language conversion on the initial text segment to obtain the target text segment is: The method according to claim 1, further comprising determining the initial text segment of the target video draft as the target text segment if it is determined that the initial language type and the target language type belong to different dialect types within the same language type.
5. Before performing language conversion on the initial text segment to obtain the target text segment based on the initial language type and target language type of the target video draft, Perform speech recognition on the audio in the target video draft and obtain the voiceprint information of the target video draft. The method according to claim 1, further comprising generating different text tracks for different voiceprint information, each for which an initial text segment corresponding to the different voiceprint information is placed.
6. Generating a target audio segment based on the aforementioned target text segment is, To determine the text-to-speech audio information corresponding to the first text track, The method according to claim 5, further comprising using the text-to-speech audio information to generate a target audio segment based on a target text segment on the first text track.
7. Before generating an edited video draft corresponding to the target video draft based on the target audio segment, The playback speed of the target audio segment is adjusted based on the time length of the initial audio segment corresponding to the target audio segment to obtain a modified target audio segment, and the initial audio segment belongs to the target video draft, and the difference between the time length of the modified target audio segment and the time length of the initial audio segment is not greater than a preset time length threshold, Accordingly, generating an edited video draft corresponding to the target video draft based on the target audio segment is: The method according to claim 1, comprising generating an edited video draft corresponding to the target video draft based on the target audio segment after the gear shift.
8. After generating an edited video draft corresponding to the target video draft based on the target audio segment, In response to a selection operation for a first target audio segment in the edited video draft, a plurality of candidate text contents corresponding to the first target audio segment are displayed. In response to a selection operation for a target candidate text content among the multiple candidate text contents, an audio segment corresponding to the target candidate text content is generated. The method according to claim 1, further comprising updating the first target audio segment in the edited result video draft using the audio segment corresponding to the target candidate text content.
9. Before generating an edited video draft corresponding to the target video draft based on the target audio segment, The method further includes determining the target text segment as the subtitle segment for the corresponding target audio segment. Accordingly, generating an edited video draft corresponding to the target video draft based on the target audio segment is: The method according to claim 1, comprising generating an edited video draft corresponding to the target video draft based on the target audio segment and the subtitle segment of the target audio segment.
10. Before performing speech recognition on the audio in the target video draft and obtaining the initial text segment corresponding to the target video draft, This further includes determining the target video draft based on the original video, Accordingly, after generating an edited video draft corresponding to the target video draft based on the target audio segment, The method according to claim 1, further comprising generating a target video belonging to the target language type, corresponding to the original video, based on the edited video draft, in response to a video export operation.
11. A video processing device, A speech recognition module for performing speech recognition on the audio in a target video draft and obtaining an initial text segment of the target video draft, Based on the initial language type and target language type of the target video draft, a language conversion is performed on the initial text segment to obtain a target text segment. The initial language type is used to represent the language type to which the target video draft belongs, and the target text segment is a conversion module belonging to the target language type. A first generation module for generating a target audio segment belonging to the target language type based on the target text segment, A video processing device comprising a second generation module used to generate an edited video draft corresponding to the target video draft based on the target audio segment, wherein there is a correspondence between the target audio segment in the edited video draft and the initial text segment of the target video draft, and the edited video draft belongs to the target language type.
12. A computer-readable storage medium wherein a command is stored in the computer-readable storage medium, and when the command is executed by a terminal device, the terminal device is made to implement the method according to any one of claims 1 to 10.
13. A video processing device comprising memory, a processor, and a computer program stored in the memory and executable by the processor, wherein the processor, when executing the computer program, realizes the method according to any one of claims 1 to 10.