Video processing method, electronic device, and program product

CN122534278APending Publication Date: 2026-08-07BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2026-06-22
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0004]一种视频处理方法、电子设备及程序产品,以克服音画不同的问题

Benefits of technology

[0019]一种可能的视频处理方法、电子设备及程序产品,通过获取第一音频,将所述第一音频切分成多个第一音频片段;基于所述多个第一音频片段生成多个第一视频片段,所述多个第一视频片段对应多个第二音频;基于所述第一音频片段与所述第二音频的特征,将所述多个第一视频片段与所述第一音频片段对齐,得到第二视频。通过将第一音频切分成多个第一音频片段,再基于各第一音频片段生成多个第一视频片段,之后基于第一视频片段自带的第二音频作为发音的参考基准,对第一视频片段和第一音频,生成第二视频。在基于音频生成多个视频片段之后,利用了与口型动作强关联的音频特征而非原始波形对第一音频片段和视频自带的音频进行匹配,因此可以有效提高匹配准确性,避免视频中口型、画面不一致的问题,提高所合成的第二视频的声画一致性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122534278A_ABST
    Figure CN122534278A_ABST
Patent Text Reader

Abstract

A video processing method, electronic equipment and program product, by cutting the first audio into multiple first audio segments, and then generating multiple first video segments based on each first audio segment, and then taking the second audio carried by the first video segment as the reference standard for pronunciation, generating the second video based on the first video segment and the first audio. After generating multiple video segments based on the audio, the audio features strongly associated with the mouth movement are used instead of the original waveform to match the first audio segment and the audio carried by the video, so that the matching accuracy can be effectively improved, the problem of inconsistent mouth shape and picture in the video can be avoided, and the sound and picture consistency of the synthesized second video can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to a video processing method, electronic device, and program product. Background Technology

[0002] Currently, in the field of digital content production, audio-visual alignment technology is widely used in scenarios such as video production, AI cover songs, dubbing synthesis, digital human animation, and music video (MV) creation. Taking short video platforms as an example, users often need to replace the original video with AI-generated audio or generate natural lip-sync animation for digital human characters.

[0003] However, the audio-video alignment schemes in related technologies generate videos with audio-video desynchronization, affecting the quality of the generated videos. Summary of the Invention

[0004] A video processing method, electronic device, and program product to overcome the problem of audio-visual discrepancies.

[0005] Firstly, a video processing method is provided, including:

[0006] A first audio file is obtained and divided into multiple first audio segments; multiple first video segments are generated based on the multiple first audio segments, and the multiple first video segments correspond to multiple second audio files; based on the features of the first audio segments and the second audio files, the multiple first video segments are aligned with the first audio segments to obtain a second video.

[0007] Secondly, a video generation apparatus is provided, comprising:

[0008] The acquisition module is used to acquire the first audio and divide the first audio into multiple first audio segments;

[0009] The processing module is configured to generate multiple first video segments based on the multiple first audio segments, wherein the multiple first video segments correspond to multiple second audio segments;

[0010] The generation module is used to align the plurality of first video segments with the first audio segment based on the features of the first audio segment and the second audio segment to obtain the second video.

[0011] Thirdly, a cloud service system is provided, including:

[0012] The cloud server is equipped with a video generation service module, which is used to receive the first video clip and the first audio uploaded by the user, execute the video processing method described in the first aspect and various possible designs of the first aspect, and return the generated second video to the user terminal.

[0013] The user terminal is used to upload a first video clip and a first audio clip to the cloud server, and to receive a second video clip returned by the cloud server.

[0014] Fourthly, an electronic device is provided, comprising: a processor and a memory;

[0015] The memory stores computer-executed instructions;

[0016] The processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the video processing method as described in the first aspect and various possible designs of the first aspect.

[0017] Fifthly, a computer-readable storage medium is provided, wherein computer-executable instructions are stored therein, and when a processor executes the computer-executable instructions, the video processing method described in the first aspect and various possible designs of the first aspect is implemented.

[0018] Sixthly, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the video processing method as described in the first aspect and various possible designs of the first aspect.

[0019] A possible video processing method, electronic device, and program product involves acquiring a first audio file, segmenting it into multiple first audio segments, generating multiple first video segments based on the multiple first audio segments, each first video segment corresponding to multiple second audio files, and aligning the multiple first video segments with the first audio segments based on the features of the first audio segments and the second audio files to obtain a second video. By segmenting the first audio file into multiple first audio segments, generating multiple first video segments based on each first audio segment, and then using the second audio files inherent in the first video segments as a reference for pronunciation, a second video is generated by combining the first video segments and the first audio files. After generating multiple video segments based on the audio, the matching of the first audio segments and the audio files inherent in the videos is performed using audio features strongly correlated with lip movements rather than the original waveform. This effectively improves matching accuracy, avoids inconsistencies between lip movements and visuals in the video, and enhances the audio-visual consistency of the synthesized second video. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a diagram illustrating an application scenario of a possible video processing method.

[0022] Figure 2 A flowchart illustrating a possible video processing method. Figure 1 ;

[0023] Figure 3 A flowchart illustrating one possible implementation of step S1031;

[0024] Figure 4 A flowchart of one possible implementation of step S1031-1;

[0025] Figure 5 A schematic diagram of a possible articulation origin feature;

[0026] Figure 6 A flowchart illustrating one possible implementation of step S1032;

[0027] Figure 7 This is a schematic diagram of one possible method for generating a second video.

[0028] Figure 8 A flowchart illustrating a possible video processing method. Figure 2 ;

[0029] Figure 9 A flowchart illustrating one possible implementation of step S202;

[0030] Figure 10 A structural block diagram of a possible video generation device;

[0031] Figure 11 A schematic diagram of the structure of a possible electronic device;

[0032] Figure 12 This is a schematic diagram of the hardware structure of a possible electronic device. Detailed Implementation

[0033] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0034] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0035] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0036] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0037] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0038] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0039] One possible video processing method can be applied to applications (APPs) with video generation capabilities, such as video editing applications and short video applications. More specifically, it can be applied to audio-visual synthesis applications. The executing entity of this embodiment can be a terminal device running the aforementioned application with video generation capabilities, a server deploying the server-side component of the aforementioned application, or other electronic devices performing similar functions. Specifically, when the executing entity is a terminal device, the terminal device executes the method provided in this embodiment by running the aforementioned application; when the executing entity is a server, the server-side component of the aforementioned application with video generation capabilities can run partially or entirely on the server, executing the method provided in this embodiment on the server side, while the terminal device runs the client-side component of the application. Communication between the server and the terminal device is based on server-client communication, enabling the terminal device to obtain the execution result of the method provided in this embodiment and display it as needed.

[0040] In some embodiments, the terminal device or server can implement a possible video processing method by running various computer-executable instructions or computer programs. For example, computer-executable instructions can be program-level commands, machine instructions, or software instructions. Computer programs can be native programs or software modules within an operating system; they can be local applications, i.e., programs that need to be installed in the operating system to run; or they can be applets embedded in any app, i.e., programs that run in a browser environment. In summary, the aforementioned computer-executable instructions can be of any form, and the aforementioned computer programs can be of any form of application, module, or plugin, with the specific implementation configured as needed. Furthermore, in implementing a possible video processing method, the terminal device can execute the method by running locally configured computer-executable instructions or computer programs, or by calling computer-executable instructions or computer programs configured on an external server. In some embodiments, the server may be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud storage, cloud communication, cloud database, cloud computing, cloud functions, network services, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms. Among these, cloud services may be interactive processing services that can be invoked by terminal devices.

[0041] Figure 1 This is a diagram illustrating an application scenario of a possible video processing method. (Reference) Figure 1As shown, the terminal device is the execution subject, and the target application with video generation function runs on the terminal device. For example, in the application scenario of short video content creation, the user generates a first video (including original audio and video footage) based on the AI ​​cover audio through the target application. Then, the AI-generated cover audio is replaced in the original video to generate a composite video, and it is ensured that the lip movements of the characters in the video footage of the composite video are strictly synchronized with the cover audio.

[0042] The audio-video alignment schemes in related technologies generate videos with audio-video desynchronization, affecting video quality. To address this issue, this paper proposes a video processing method that focuses on features directly related to lip movements, such as vocal rhythm, timing of mouth opening, and emphasis, rather than overall waveform similarity, to solve the aforementioned problem.

[0043] refer to Figure 2 , Figure 2 A flowchart illustrating a possible video processing method. Figure 1 The above method can be applied to terminal devices or servers. In one possible implementation, for a terminal device executing the method, the terminal device can implement the video processing method by executing program code deployed locally and / or externally. In another possible implementation, a server can be used to deploy functional services based on a possible video processing method, and the terminal device can implement the possible video processing method by accessing the server and calling the corresponding functional services. For example, a possible video processing method includes:

[0044] Step S101: Obtain the first audio and divide the first audio into multiple first audio segments.

[0045] Step S102: Generate multiple first video segments based on multiple first audio segments, with each first video segment corresponding to a multiple second audio segment.

[0046] Step S103: Based on the features of the first audio segment and the second audio segment, align multiple first video segments with the first audio segment to obtain the second video.

[0047] refer to Figure 1The illustrated application scenario diagram presents one possible implementation method, using a terminal device as the execution subject to introduce a possible video processing method. For example, the terminal device runs a target application, which provides video generation functionality to the user. Specifically, the user loads a first audio file through the target application's interface. This first audio file could be, for example, an AI-generated cover song, a voice-over recorded by a professional voice actor, or a digital human speech synthesis audio file. Then, the first audio file is divided into multiple segments, i.e., multiple first audio segments. In one possible implementation, the first audio file can be segmented at equal time intervals to generate multiple first audio segments, in which case each first audio segment has the same duration. In another possible implementation, the first audio file can be segmented into sentences based on the content of human voices and dialogues, with each sentence generating a corresponding first audio segment. For example, each lyric in the first audio file corresponds to one first audio segment.

[0048] Furthermore, in one possible implementation, step S101 is specifically implemented as follows:

[0049] Step S1011: In response to the first operation, at least one segmentation position in the first audio is obtained; the first operation is used to trigger the first control or to set the segmentation position of the first audio, and the first control is used to automatically determine the segmentation position.

[0050] Step S1012: Based on the segmentation position, the first audio is segmented into multiple second audio segments, and the second audio segments correspond to a speech sentence in the first audio.

[0051] Step S1013: Based on the content coherence between the speech sentences corresponding to the second audio segment, merge at least two adjacent second audio segments into a first audio segment.

[0052] For example, in one possible implementation, the terminal device first obtains at least one segmentation position in the first audio by responding to the user's interactive operation (first operation) within the interactive interface. This segmentation position can be manually set through the first operation or selected from the first audio based on the selection logic corresponding to the first control after triggering it. Then, based on this segmentation position, the first audio is segmented into multiple second audio segments, each corresponding to a speech sentence. A speech sentence refers to the speech corresponding to a complete sentence, such as a complete dialogue or a complete lyric (a sentence separated by punctuation marks like periods and commas). Next, based on the content corresponding to adjacent second audio segments, the second audio segments with coherent content are merged into a corresponding long audio segment, i.e., the first audio segment. The first audio segment generated in this way can have different durations (first duration).

[0053] Furthermore, in one possible implementation, after obtaining multiple second audio segments, text extraction can be performed on the second audio segments to obtain corresponding text content. Then, the text content is characterized to obtain feature vectors representing semantics. Next, the vector distance between the feature vectors of the speech sentences corresponding to adjacent second audio segments is calculated to determine the content coherence between the speech sentences, thereby completing the merging of the second audio segments. In the above implementation, the content coherence between second audio segments is determined by the text semantics of the speech sentences corresponding to the second audio segments, thereby achieving the merging of semantically similar second audio segments to form a first audio segment. This makes the semantics corresponding to the first audio segments more similar and more focused on similar content, thus better enabling the construction of a matching first video segment in subsequent steps, allowing the first video segment generated based on the first audio segment to better represent the content of the first audio segment.

[0054] In another possible implementation, this embodiment further includes:

[0055] Based on the second audio segment, a corresponding first frame is generated, which represents the semantic content of the speech sentences in the second audio segment.

[0056] Step S1010A: Based on the second audio segment, generate the corresponding first screen, which represents the semantic content of the speech sentence of the second audio segment.

[0057] Step S1010B: Based on the image similarity of adjacent first images, obtain the content coherence between the speech sentences corresponding to the second audio segment.

[0058] For example, in another possible implementation, after feature extraction based on the second audio segment, a semantic content first image representing the speech sentences of the second audio segment is generated based on the content features of the second audio segment. This involves performing an "audio-to-image" multimodal processing step, which can be implemented by calling a multimodal model. Then, based on the first image corresponding to each second audio segment, the similarity between the images is compared to obtain the content coherence between the speech sentences corresponding to the second audio segment. For example, if the image similarity between the first image P1 (corresponding to second audio segment A) and the first image P2 (corresponding to second audio segment B) is greater than a similarity threshold, then their content is considered coherent. In the subsequent aggregation step of the second audio segments, the coherent second audio segments A and B are aggregated into a first audio segment.

[0059] In the above implementation, the second audio segment is converted into a corresponding first image, and a similarity comparison is performed based on the image to obtain the content coherence between the speech sentences. The first audio segment obtained in this way has better clustering in the dimension of image content, thus reducing the scene jumps in the first video segment when generating the corresponding first video segment based on the first audio segment, and improving the smoothness of the video.

[0060] Furthermore, in one possible implementation, step S102 is specifically implemented as follows:

[0061] Step S1021: In response to the second operation on the second control, obtain the semantic information of the first audio segment.

[0062] Step S1022: Input the semantic information and the first duration corresponding to the first audio segment into the multimodal model to generate a first video segment and a corresponding second audio segment. The first video segment and the corresponding second audio segment match the first duration. The second control is used to trigger the task of generating the first video segment. The pronunciation features of the second audio segment match the visual features of the first video segment.

[0063] For example, after obtaining multiple first audio segments through the above steps, a second control is triggered through an interactive operation (second operation). This second control is, for example, a button control for a "change voice" function. After triggering the second control, the terminal device extracts the semantic information of each first audio segment. The semantic information represents the dialogue content of the first audio segment; it can be text information or the feature vector corresponding to the text information. Then, the semantic information and the first duration corresponding to each first audio segment are input into a multimodal model. Based on the first video segment corresponding to each first audio segment and the semantic information, the multimodal model generates a first video segment of the corresponding length (first video segment) and a corresponding second audio. The second audio is, for example, the original audio that matches the first video segment. Here, "matching" means that the pronunciation nodes of the second audio completely match the actions (e.g., lip movements) of the characters in the first video segment. Therefore, in subsequent steps, the second audio can be used as a reference to perform alignment between the first video segment and the first audio segment.

[0064] Furthermore, in one possible implementation, the second operation includes a first triggering operation and a second triggering operation. The specific implementation of steps S1021 and S1022 includes:

[0065] Step S1021A: In response to the trigger operation on the second control, extract the text content from the first audio segment and display the text content on the interactive interface.

[0066] Step S1021B: In response to the confirmation operation of the third control, input the text content and the first duration corresponding to the first audio segment into the multimodal model to generate the first video segment and the corresponding second audio. The third control is used to modify and / or confirm the text content.

[0067] For example, firstly, the second control is a functional control used to parse the first audio segment and extract the speech content from it. After the user triggers the second control through a triggering operation (e.g., a click), the second control, based on the corresponding processing logic, calls the audio processing model to extract the corresponding text content (e.g., lyrics) from the first audio segment and displays it on the interactive interface. Then, the user continues to modify or confirm the text content by operating the third control. Specifically, the third control can be a modification control; by triggering this control, the text content in the interactive interface can be further modified or replaced, thereby controlling the content of the subsequently generated first video segment. Alternatively, the third control can be a confirmation control; by triggering this control, the text content is confirmed, and subsequent steps are performed to generate the corresponding video content based on the text content. Specifically, the text content and the first duration corresponding to the first audio segment are input into a multimodal model to generate a video and audio with a playback duration of the first duration and content matching the text content, i.e., the first video segment and the corresponding second audio.

[0068] Furthermore, after generating multiple first video segments and corresponding second audio segments through the above steps, audio replacement needs to be performed on each first video segment, that is, replacing the second audio segment corresponding to the first video segment with the first audio segment, thereby realizing the audio-video replacement and synthesis function. Specifically, the first video segment can refer to a video file (video clip) containing a person's voice and the corresponding original sound, such as a user-uploaded original song MV, a live speech video, a film clip, etc., and the second audio is the original audio corresponding to the first audio segment. After obtaining the second audio segment corresponding to the first video segment and the corresponding first audio segment through the above steps, the audio features of the first audio segment and the second audio segment are used to align the first video segment and the first audio segment. Audio features refer to audio attribute features directly related to the mouth movements when a person speaks, rather than the original audio waveform, including but not limited to volume time-varying features, pronunciation start features, etc. In a more specific implementation, after the terminal device obtains the second audio and the first audio segment, it first performs audio format unification preprocessing on both, converting the second audio and the first audio segment into a standard format of mono with a 16kHz sampling rate, eliminating the impact of channel differences and sampling rate inconsistencies on subsequent matching. Then, the audio features related to lip movements of the second audio segment and the first audio segment are extracted respectively. By performing dynamic sliding matching within a preset time range, the similarity of the audio features of the two segments at different candidate time points is calculated. Finally, the candidate time point with the highest similarity is selected as the first time point. Then, based on the first time point, the alignment of the first audio segment and the first video segment is completed.

[0069] In a more specific application scenario, for example, the first video clip is a 3-minute music video of user A's song, and the corresponding first audio clip is a cover song audio clip of user B. In step S103, after preprocessing the first audio clip and the second audio clip, feature comparison is performed. The features are, for example, audio features strongly correlated with lip movements. A first timing point (e.g., 1.66 seconds) is output. Then, based on this first timing point, the first video clip and the first audio clip are aligned. The technical purpose of this process is to replace the traditional waveform matching method by extracting and matching audio features strongly correlated with lip movements, thereby solving the problem of unstable matching caused by differences in timbre and encoding in the existing technology, and providing an accurate time reference for subsequent audio-video alignment.

[0070] In one possible implementation, step S103 is specifically implemented as follows:

[0071] Step S1031: Extract the audio features of the first audio segment, including volume time-varying features and pronunciation start features.

[0072] Step S1032: Based on audio features, perform time offset matching on the first video segment and the second audio segment to obtain the first time point.

[0073] Step S1033: Using the first time point as the starting point, align the frame sequence of the first video segment with the first audio to generate the second video.

[0074] For example, in step S1031, the volume time-varying feature refers to the curve feature describing the change of audio volume over time, which can reflect the rhythm and pause pattern of the audio; the pronunciation start point feature refers to the time point feature describing the sudden rise of volume in the audio, which can reflect key time points directly related to mouth movements such as opening speech and the appearance of stress.

[0075] In one possible implementation, the terminal device first extracts the volume envelope of the preprocessed first audio segment to obtain its volume envelope curve. Then, based on this volume envelope curve, it performs first-order difference calculations to extract all abrupt changes in volume, generating the articulation start-point feature curve of the first audio segment. Simultaneously, the terminal device uses the same method to extract the time-varying volume features and articulation start-point features of the second audio segment from the first video segment.

[0076] In step S1032, time offset matching refers to generating multiple candidate time points within a preset candidate time offset range with a fixed step size, calculating the audio feature similarity between the second audio segment corresponding to each candidate time point and the entire first audio segment, and selecting the candidate time point with the highest similarity as the first time point. The terminal device first determines a preset time offset range (e.g., 0 seconds to 5 seconds), and then generates 250 candidate time points within this range with a step size of 20 milliseconds. For each candidate time point, a segment of the second audio starting from that point with the same duration as the first audio segment is extracted. The volume envelope correlation coefficient and the pronunciation start point correlation coefficient between this segment and the first audio segment are calculated, and a comprehensive similarity score is obtained through weighted fusion. Finally, the candidate time point with the highest comprehensive score is selected as the first time point.

[0077] In the above implementation method, by extracting two complementary lip-sync related audio features and using dynamic sliding matching, the matching degree at different time points is comprehensively and accurately evaluated, thereby improving the ability of the first time point to achieve the best lip-sync effect.

[0078] Furthermore, in one possible implementation, such as Figure 3 As shown, the specific implementation of step S1031 includes:

[0079] Step S1031-1: Obtain the volume envelope of the first audio segment.

[0080] For example, after obtaining the first audio segment, the terminal device can obtain the audio envelope corresponding to the first audio segment through real-time processing. The volume envelope is the time-varying volume feature of the first audio segment. For example, the volume envelope of the first audio segment can be extracted using a pre-trained envelope extraction model, or the N maximum volume value points of the first audio segment can be extracted to construct the volume envelope. In one possible implementation, the specific implementation of step S1031-1 is as follows: Figure 4 As shown, it includes:

[0081] Step S1031-11: Based on the duration of the first audio segment, generate at least two time windows of fixed duration;

[0082] Step S1031-12: Obtain the volume intensity corresponding to each time window. The volume intensity represents the root mean square of the volume value of the first audio segment within the corresponding time window.

[0083] Step S1031-13: Obtain the volume envelope of the first audio segment based on the volume intensity corresponding to the time window.

[0084] For example, in steps S1031-11, a time window refers to dividing a continuous audio signal into multiple discrete short time segments for calculating the volume intensity of the audio segment by segment. The size and sliding step of the time window can be adjusted according to actual needs; for example, the sliding step can be set to half the window size. The terminal device first obtains the total duration of the first audio segment, and then generates a continuous time window sequence covering the entire duration of the first audio segment according to the preset window size and sliding step. For example, for a first audio segment with a duration of 10 seconds, if the window size is set to 40 milliseconds and the sliding step is set to 20 milliseconds, then a total of (10000-40) / 20+1=499 time windows are generated.

[0085] In steps S1031-12, volume intensity refers to the energy level of the audio signal within each time window, calculated using the root mean square (RMS) value. The formula for calculating the RMS value is not elaborated here. The RMS value accurately reflects the volume level perceived by the human ear. Next, the terminal device performs RMS calculation on the audio sampling points within each time window to obtain the volume intensity corresponding to each time window. For example, for a 40-millisecond time window, at a 16kHz sampling rate, there are 640 sampling points. Calculating the RMS value of these 640 sampling points gives the volume intensity of that window. Then, in steps S1031-13, the terminal device connects the volume intensity values ​​corresponding to all time windows in chronological order to form a curve that changes over time, which is the volume envelope curve of the first audio segment.

[0086] In a specific application scenario, for example, a 3-minute cover song audio, step S1031-11 generates 100 time windows of 40 milliseconds each; step S1031-12 calculates the volume intensity of each window; finally, step S1031-13 outputs a volume envelope curve composed of 100 volume intensity values, i.e., the volume envelope.

[0087] In the above implementation, the continuous audio waveform is converted into a volume time-varying feature that reflects the rhythm of the voice by short-window RMS calculation, eliminating the influence of factors unrelated to lip shape, such as timbre and phase, on the matching and improving the stability of the matching.

[0088] In another possible implementation, the volume envelope can be pre-generated data. For example, after the server preprocesses the first audio segment, it generates a volume envelope and saves it as reference information for the first audio segment. The terminal device obtains the volume envelope corresponding to the first audio segment at the same time as obtaining the first audio segment. The server can generate the volume envelope using the scheme described in step S1031-1 above, or other schemes. Other schemes include, for example, extracting the volume envelope of the first audio segment using a pre-trained envelope extraction model, or extracting the N maximum volume value points of the first audio segment and constructing the volume envelope, etc., which will not be elaborated here.

[0089] Step S1031-2: Perform first-order difference calculation on the volume envelope to obtain the positive envelope segment, which indicates the audio segment with increased volume in the first audio segment.

[0090] Step S1031-3: Based on the positive envelope segment, obtain the pronunciation origin feature.

[0091] For example, first-order difference refers to calculating the difference in volume intensity between two adjacent time points in the volume envelope curve, used to detect sudden changes in volume. The positive envelope segment refers to the portion of the first-order difference result with a positive difference, corresponding to the time periods when the volume rises in the audio. These time periods typically correspond to the opening movements or stresses during human speech. The terminal device performs first-order difference calculation on the volume envelope curve obtained in step S1031-1. After obtaining the difference curve, it performs threshold filtering on the difference curve, retaining only the positive portion where the difference is greater than a preset threshold, generating a positive envelope segment. The preset threshold can be dynamically adjusted according to the overall volume level of the audio. In one possible implementation, the preset threshold can be set to 1.5 times the absolute value of the average difference to filter out interference caused by small volume fluctuations.

[0092] Subsequently, the terminal device extracts the starting time points of all positive envelope segments, which are the pronunciation starting points of the first audio segment. In one possible implementation, these pronunciation starting points are arranged in chronological order to form a pronunciation starting point feature curve. For example, this curve has a peak only at the pronunciation starting point position, and zero at other positions. For example, in a more specific application scenario, such as for the volume envelope curve of a cover song audio, step S1031-2 calculates the first-order difference curve, and after filtering, obtains all positive envelope segments with significantly increased volume; step S1031-3 extracts the starting points of these segments to obtain the opening time points and accent time points of all lyrics in the cover song audio, forming a pronunciation starting point feature curve. In the above implementation, the volume change points are extracted by first-order difference, accurately capturing the opening timing and accent position directly related to lip movements. These features are key to achieving accurate lip alignment and can effectively avoid the problems of "sound rushing" or "mouth moving first and sound arriving later" (i.e., misalignment of lip shape and vocal timing, and audio-visual asynchrony).

[0093] Figure 5 A schematic diagram of a possible articulation origin feature, such as Figure 5 As shown, exemplarily, firstly, the volume envelope P1 of the first audio segment is obtained. Then, a first-order difference calculation is performed on the volume envelope Env_1, and the volume abrupt change point is selected to obtain multiple positive envelope segments, such as L1, L2, L3, and L4 shown in the figure. Next, the set of the above positive envelope segments [L1, L2, L3, L4] is used as the pronunciation start point feature. The above pronunciation start point feature indicates the pronunciation point in the first audio segment. By aligning the above pronunciation points, the alignment of the first audio segment and the original audio can be achieved.

[0094] Furthermore, in one possible implementation, such as Figure 6 As shown, the specific implementation of step S1032 includes:

[0095] Step S1032-1: Obtain the preset time offset range.

[0096] Step S1032-2: Generate at least two candidate time points within the time offset range, and use the candidate time points as the starting point to obtain the audio features of the corresponding audio segments in the second audio.

[0097] Step S1032-3: Obtain the similarity between the audio features corresponding to the candidate time point and the audio features of the first audio segment.

[0098] Step S1032-4: Based on the similarity between the audio features corresponding to the candidate time points and the audio features of the first audio segment, select the first time point from at least two candidate time points.

[0099] For example, the preset time offset range refers to the time interval within the first video segment that may match the starting point of the first audio segment. The time offset range can be a fixed preset value or can be set according to the content characteristics and application scenario of the first video segment. For instance, in an AI cover song scenario, the first video segment typically has a few seconds of intro, so the time offset range can be set to 0 to 5 seconds; in a dubbing scenario, the time offset range can be set to 0 to 3 seconds. Of course, it is understood that the terminal device obtains the time offset range based on preset configuration parameters, or the user can manually adjust this range in the operation interface.

[0100] Next, the terminal device generates candidate time points by sliding within a time offset range with a fixed step size. Candidate time points are candidate time points generated within the time offset range for matching; the smaller the step size, the higher the matching accuracy. For example, the step size is set to 20 milliseconds, consistent with the sliding step size for volume envelope extraction, to ensure time axis alignment. Then, for each candidate time point, the terminal device extracts a segment of the second audio that starts from that point and has the same duration as the first audio segment. Then, using the same method as in step S1031, it extracts the volume time-varying features and pronunciation start-point features of this second audio segment.

[0101] Next, the terminal device calculates the similarity of volume time-varying features and pronunciation start-point features between the candidate segments of the second audio and the first audio segment, and then obtains a comprehensive similarity score through weighted fusion. Finally, the terminal device compares the comprehensive similarity scores of all candidate time points and selects the candidate time point with the highest score as the first time point. If multiple candidate time points have the same score and are all the highest, the earliest candidate time point is selected.

[0102] In the above implementation, comprehensive dynamic scanning and matching are performed within a reasonable time range to ensure that the best alignment time point is not missed. At the same time, the accuracy and reliability of the matching results are guaranteed by multi-feature similarity calculation.

[0103] Furthermore, in one possible implementation, the specific implementation of step S1032-3 includes:

[0104] Steps S1032-31: Based on the scene of the first video segment, obtain the weight values ​​of the volume time-varying feature and the pronunciation start point feature respectively.

[0105] Steps S1032-32: Obtain the first similarity between the volume time-varying feature corresponding to the candidate time-series point and the volume time-varying feature of the first audio segment, and obtain the second similarity between the pronunciation start point feature corresponding to the candidate time-series point and the pronunciation start point feature of the first audio segment.

[0106] Steps S1032-33: Based on the weight values, perform a weighted sum of the first similarity and the second similarity to obtain the similarity of the audio features.

[0107] For example, the weight values ​​of volume time-varying features and articulation start-point features are used to balance their importance in similarity calculation. Different scene scenarios have different degrees of dependence on the two features. For example, in a singing scene, vocal rhythm is more important, so the weight of the volume time-varying features should be set higher; in a dialogue scene, the timing of speaking is more critical, so the weight of the articulation start-point features should be set higher. In one possible implementation, the terminal device can automatically obtain the weight values ​​through a preset scene-weight mapping table. For example, the volume envelope weight is 0.65 and the articulation start-point weight is 0.35 for a singing scene; and the volume envelope weight is 0.4 and the articulation start-point weight is 0.6 for a dialogue scene. Of course, it is understood that the weight values ​​can also be manually adjusted by the user according to actual needs.

[0108] Further, for example, the first similarity refers to the Pearson correlation coefficient between the two volume envelope curves, with a value range of [-1, 1]. The closer the value is to 1, the stronger the linear correlation between the two curves, that is, the higher the rhythm matching degree. The second similarity refers to the cosine similarity between the two pronunciation start feature curves, with a value range of [0, 1]. The closer the value is to 1, the higher the peak position matching degree of the two curves, that is, the higher the opening timing matching degree.

[0109] The terminal device calculates the volume envelope correlation coefficient (first similarity) and the cosine similarity of the pronunciation start point (second similarity) between the candidate segments of the second audio and the first audio segment, respectively. Specifically, the terminal device performs a weighted sum of the first and second similarities based on the weight values ​​obtained in steps S1032-31 to obtain a comprehensive audio feature similarity. Furthermore, to avoid misjudgments caused by excessive energy differences, an average difference penalty term can be introduced into the comprehensive similarity, specifically as follows:

[0110] Similarity = Volume envelope weight × First similarity + Pronunciation origin weight × Second similarity - 0.05 × Average difference.

[0111] The average difference refers to the mean square error of the overall energy of the two audio segments.

[0112] In the above implementation, the matching algorithm can adapt to different application scenarios by dynamically adjusting the weights of different features. At the same time, an average difference penalty term is introduced to further improve the accuracy of matching and avoid incorrect matching caused by similar local features but large differences in overall energy.

[0113] For example, after obtaining the first time point, the terminal device extracts the frame sequence of the first video segment based on the first time point, and then uses the first audio segment as a new audio track to synchronously synthesize with the extracted frame sequence to generate the second video. Figure 7 An illustration of a possible method for generating a second video, such as... Figure 7 As shown, exemplarily, after obtaining the first time point t1 through the above steps, the video sequence is segmented using the first time point t1 as the dividing point, resulting in video sequence P1 and video sequence P2. Video sequence P2 starts from the first time point t1. Then, the starting point of the first audio segment is aligned with the first time point t1. Based on the aligned video sequence P2 and the first audio segment, a replacement video segment is generated. This process is the alignment of a first video segment with a first audio segment. The alignment of other first video segments with first audio segments is similar, and corresponding replacement video segments are formed after alignment. Finally, all replacement video segments are aggregated to form the second video.

[0114] In one possible implementation, a first audio file is acquired and segmented into multiple first audio segments. Multiple first video segments are generated based on these segments, each corresponding to a second audio file. Based on the features of the first and second audio segments, the video segments are aligned with the first audio segments to obtain a second video. By segmenting the first audio into multiple segments and generating multiple first video segments based on each segment, and then using the second audio embedded in each video segment as a reference for pronunciation, a second video is generated by matching the first video segments with the first audio. After generating multiple video segments based on the audio, audio features strongly correlated with lip movements, rather than the original waveform, are used to match the first audio segments with the audio embedded in the video. This effectively improves matching accuracy, avoids inconsistencies between lip movements and visuals in the video, and enhances the audio-visual consistency of the synthesized second video.

[0115] refer to Figure 8 , Figure 8 A flowchart illustrating a possible video processing method. Figure 2 The above-mentioned possible video processing methods are... Figure 2 Based on the illustrated embodiment, steps S102-S103 are further refined, and the video processing method includes:

[0116] Step S200: Obtain the first audio and divide the first audio into multiple first audio segments.

[0117] Step S201: Generate multiple first video segments based on multiple first audio segments, with each first video segment corresponding to a multiple second audio segment.

[0118] Step S202: Extract the audio features of the first audio segment, including volume time-varying features and pronunciation start features.

[0119] For example, such as Figure 9 As shown, the specific implementation of step S202 includes:

[0120] Step S2021: Based on the duration of the first audio segment, generate at least two time windows of fixed duration;

[0121] Step S2022: Obtain the volume intensity corresponding to each time window. The volume intensity represents the root mean square of the volume value of the first audio segment within the corresponding time window.

[0122] Step S2023: Based on the volume intensity corresponding to the time serial port, obtain the time-varying volume characteristics of the first audio segment.

[0123] Step S2024: Perform first-order difference calculation on the volume time-varying features to obtain the positive envelope segment, which indicates the audio segment with increased volume in the first audio segment;

[0124] Step S2025: Based on the positive envelope segment, obtain the pronunciation origin feature.

[0125] For example, the implementation of this step is similar to that of step S1021 in the first embodiment. First, the volume time-varying features of the first audio segment are obtained, which is the volume envelope in step S1021. Then, a first-order difference calculation is performed based on the volume time-varying features to obtain the positive envelope segment, and then the pronunciation start point feature is obtained based on the positive envelope segment. For specific implementation details, please refer to the description in the previous embodiments, which will not be repeated here.

[0126] Step S203: Obtain the semantic features of the first audio segment.

[0127] Step S204: Weight the audio features and semantic features to obtain combined features, and perform time offset matching on the first video segment and the second audio based on the combined features to obtain the first time point.

[0128] For example, the semantic features of the first audio segment refer to features that can reflect the semantic content of the audio. In one possible implementation, the semantic features include a first semantic feature based on spectral features and / or a second semantic feature based on text alignment. Semantic features can supplement the deficiencies of audio features, ensuring matching accuracy even when there are significant differences in timbre or subtle changes in pronunciation rhythm. The terminal device extracts one or more semantic features of the first audio segment according to actual needs. On the other hand, it simultaneously uses the same method to extract the semantic features corresponding to the second audio of the first video segment, or reads the pre-generated semantic features corresponding to the second audio.

[0129] The terminal device then weights and combines the audio and semantic features to obtain a combined feature. This combined feature is then used as the comparison target to calculate the comprehensive similarity of each candidate time-series point. The weights of the audio and semantic features can be dynamically adjusted according to the application scenario. For example, in AI cover song scenarios with significant timbre differences, the weight of the semantic features can be appropriately increased. This implementation method, by introducing semantic features, further enhances the robustness of audio-video alignment, enabling it to adapt to more complex application scenarios and addressing the problem of audio feature matching failure in extreme cases.

[0130] Furthermore, in one possible implementation, the semantic features include a first semantic feature, and the specific implementation of step S203 includes:

[0131] Step S203A-1: Extract the spectral features of the first audio segment, whereby the spectral features characterize the timbre and / or pitch of the first audio segment;

[0132] Step S203A-2: Obtain the first semantic feature based on the spectral features.

[0133] Accordingly, the implementation of step S204 includes:

[0134] Step S204A: Weight the audio features and the first semantic features to obtain combined features, and perform time offset matching on the first video segment and the second audio based on the combined features to obtain the first time point.

[0135] For example, in one possible implementation, spectral features refer to features describing the energy distribution of an audio signal at different frequencies. These features can represent the timbre and / or pitch of the first audio segment, such as Mel-frequency cepstral coefficients. This feature can simulate the auditory characteristics of the human ear, effectively capturing the timbre and pitch information of the audio, which reflects the overall similarity of the audio. Subsequently, the spectral features are processed using data structures (e.g., mean and variance normalization) to generate first semantic features. These features are then weighted and combined with audio features to generate combined features, and a first time-series point is determined based on the combined features. The step of determining the first time-series point based on the combined features has been described in previous steps and will not be repeated here. In the above implementation, by introducing spectral features, timbre and pitch information are supplemented, significantly improving the matching robustness of the algorithm in scenarios with large timbre differences and expanding the applicability of the algorithm.

[0136] In another possible implementation, the semantic features include a second semantic feature, and the specific implementation of step S203 includes:

[0137] Step S203B-1: Extract the text alignment of the first audio segment. The text alignment represents the degree of matching between the text content corresponding to the first audio segment and the picture content of the picture sequence.

[0138] Step S203B-2: Obtain the second semantic feature based on text alignment.

[0139] Accordingly, the implementation of step S204 includes:

[0140] Step S204B: Based on the weighted combination of audio features and second semantic features, perform time offset matching on the first video segment and the second audio to obtain the first time point.

[0141] For example, firstly, the terminal device uses an Automatic Speech Recognition (ASR) model to convert the second audio segments of the first audio segment and the first video segment into text, and obtains the timestamp corresponding to each word based on the Connectionist Temporal Classification (CTC) algorithm. Then, it performs semantic matching between the text of the first audio segment and the text of the second audio segment, calculates their text similarity, and combines the timestamp information to obtain the text alignment. Next, the terminal device calculates the second semantic feature similarity between candidate segments of the second audio segment and the first audio segment, then weights and combines this with the audio feature similarity to generate a combined feature. Finally, it determines the first time sequence point based on the combined feature. The specific method for determining the first time sequence point is similar to the specific implementation method for determining the first time sequence point described in the previous embodiments, and will not be repeated here. In the above implementation, by introducing text alignment, the consistency of audio and video content is guaranteed at the semantic level, effectively solving the matching problem caused by differences in pronunciation rhythm, and is particularly suitable for scenarios with high requirements for semantic accuracy, such as dubbing and digital human dialogue.

[0142] In another possible implementation, the semantic features include a first semantic feature and a second semantic feature. The specific implementation of step S203 includes:

[0143] Step S203C-1: Extract the spectral features of the first audio segment, and obtain the first semantic features based on the spectral features. The spectral features characterize the timbre and / or pitch of the first audio segment.

[0144] Step S203C-2: Extract the text alignment of the first audio segment, and based on the text alignment, obtain the second semantic feature. The text alignment represents the degree of matching between the text content corresponding to the first audio segment and the video content of the video sequence.

[0145] Accordingly, the implementation of step S204 includes:

[0146] Step S204C: The audio features, the first semantic features, and the second semantic features are weighted and combined to obtain the combined features. Based on the combined features, the first video segment and the second audio segment are time-off matched to obtain the first time point.

[0147] For example, in the above implementation, the first semantic feature and the second semantic feature of the first audio segment are extracted simultaneously. The first semantic feature, the second semantic feature, and the audio feature are then weighted and combined to obtain the corresponding combined feature. Conversely, the combined feature of the original audio is obtained in the same way. Based on the combined feature, time offset matching is performed on the first video segment and the combined audio to determine the first time point. For example, if the audio feature weight is set to 0.75, the first semantic feature weight to 0.15, and the second semantic feature weight to 0.1, then the comprehensive similarity = 0.75 × audio feature similarity + 0.15 × first semantic feature similarity + 0.1 × second semantic feature similarity. In the above implementation, through multi-dimensional feature fusion, various information related to audio-video alignment is comprehensively covered, giving the algorithm strong robustness and adaptability, and enabling it to meet the high-precision alignment requirements of various complex scenarios.

[0148] In one possible implementation, optionally, after step S204, the following method is further included:

[0149] Step S205: Obtain at least two adjacent frame frames corresponding to the first time point.

[0150] Step S206: Based on the matching degree between the content of the picture frame and the first audio segment, obtain the first picture frame, and fine-tune the first timing point with the timing point of the first picture frame.

[0151] Step S207: Using the finely adjusted first time point as the starting point, align the image sequence and the first audio segment to generate the second video.

[0152] For example, firstly, the terminal device acquires the first video segment frame corresponding to the first time point, as well as several adjacent frames before and after that frame. For instance, it acquires the frame corresponding to the first time point, the previous frame, and the next frame, totaling three adjacent frames. Then, based on the matching degree between the content of the frames and the first audio segment, the first frame is obtained, and the first time point is fine-tuned. The first frame refers to the frame with the highest matching degree to the starting point of the first audio segment. The terminal device uses the time point of each adjacent frame as a candidate fine-tuning point, generates a corresponding preview video segment, and then determines the optimal first frame through content analysis or user feedback. One possible implementation is automatic fine-tuning, whereby the terminal device uses a pre-trained lip-sync detection model to identify the lip-sync state of the person in each adjacent frame (e.g., closed, open, half-open, etc.), and then matches it with the pronunciation state at the starting point of the first audio segment, selecting the frame with the best matching lip-sync state as the first frame. Another possible implementation involves manual adjustment, where the terminal device generates a 1-second preview video segment starting from each adjacent frame for the user to listen to and watch. The user selects the preview video with the best lip-sync effect based on their own perception, and the corresponding frame becomes the first frame. For example, the fine-tuning step size can be a frame-level step size, such as 0.0417 seconds in a 24fps video.

[0153] Finally, the terminal device uses the timing point corresponding to the first frame as the fine-tuned first timing point to align the frame sequence and the first audio segment. For details, please refer to [link / reference needed]. Figure 2 The specific implementation of step S103 in the illustrated embodiment generates the final second video.

[0154] The following is a more specific embodiment. First, the terminal device automatically matches the first timing point to 1.66 seconds, which corresponds to the 40th frame (1.66 seconds ≈ 40 frames at 24fps). Then, frames 39, 40, and 41 are acquired (step S205). Next, lip-sync detection reveals that the lip shape in frame 41 best matches the opening of the first word in the cover audio. Therefore, the first timing point is fine-tuned to 1.708 seconds, which corresponds to the time of frame 41 (step S206). Finally, the final synthesized video is generated starting from 1.708 seconds (step S207).

[0155] In the above implementation, multi-dimensional feature fusion significantly improves the matching accuracy and robustness of the algorithm in complex scenarios, enabling it to adapt to various extreme situations such as large differences in timbre and adjustments in pronunciation rhythm. Frame-level fine-tuning improves alignment accuracy to the frame level, ensuring that the final generated video has precise lip-sync effects, meeting the needs of high-precision applications such as film and television production and digital human animation. The above implementation... Figure 2Based on the illustrated embodiment, semantic feature extraction and multi-feature fusion matching steps are added to further improve the robustness of matching; at the same time, a frame-level fine-tuning step is added, which can accurately optimize the automatic matching results.

[0156] The implementation method of the above step S201 is the same as Figure 2 The implementation of step S102 in the illustrated embodiment is the same, and will not be described in detail here.

[0157] Corresponding to the video processing method in the above embodiments, Figure 10 This is a structural block diagram of a possible video generation device. The method described in the above embodiments can be executed by this video generation device, which can be implemented by software and / or hardware, and can be integrated into an electronic device with certain data processing capabilities. The electronic device may include, but is not limited to, mobile terminals with big data processing capabilities, as well as fixed terminals with big data processing capabilities such as desktop computers and supercomputers.

[0158] For ease of explanation, only the parts relevant to the embodiments are shown. (Refer to...) Figure 10 The video generation device 3 includes:

[0159] The acquisition module 31 is used to acquire the first audio and divide the first audio into multiple first audio segments;

[0160] Processing module 32 is used to generate multiple first video segments based on multiple first audio segments, wherein the multiple first video segments correspond to multiple second audio segments;

[0161] The generation module 33 is used to align multiple first video segments with the first audio segments based on the features of the first audio segment and the second audio segment to obtain the second video.

[0162] In one possible implementation, when the acquisition module 31 divides the first audio into multiple first audio segments, it is specifically used to: obtain at least one segmentation position in the first audio in response to a first operation; the first operation is used to trigger a first control or to set the segmentation position of the first audio, and the first control is used to automatically determine the segmentation position; based on the segmentation position, divide the first audio into multiple second audio segments, each second audio segment corresponding to a speech sentence in the first audio; and according to the content coherence between the speech sentences corresponding to the second audio segments, merge at least two adjacent second audio segments into one first audio segment.

[0163] In one possible implementation, the acquisition module 31 is further configured to: generate a corresponding first frame based on the second audio segment, wherein the first frame represents the semantic content of the speech sentences of the second audio segment; and obtain the content coherence between the speech sentences corresponding to the second audio segment based on the similarity of adjacent first frames.

[0164] In one possible implementation, multiple first audio segments correspond to multiple first durations. When the processing module 32 generates multiple first video segments based on the multiple first audio segments, it is specifically used to: in response to a second operation on the second control, obtain the semantic information of the first audio segments, and input the semantic information and the first duration corresponding to the first audio segments into a multimodal model to generate first video segments and corresponding second audio, wherein the first video segments and corresponding second audio match the first duration, the second control is used to trigger the task of generating the first video segments, and the pronunciation features of the second audio match the picture features of the first video segments.

[0165] In one possible implementation, the second operation includes a first trigger operation and a second trigger operation. When the processing module 32, in response to the second operation on the second control, obtains the semantic information of the first audio segment and inputs the semantic information and the first duration corresponding to the first audio segment into the multimodal model to generate the first video segment and the corresponding second audio, it is specifically used to: extract the text content in the first audio segment in response to the trigger operation on the second control; display the text content on the interactive interface; and input the text content and the first duration corresponding to the first audio segment into the multimodal model in response to the confirmation operation on the third control to generate the first video segment and the corresponding second audio. The third control is used to modify and / or confirm the text content.

[0166] In one possible implementation, when the generation module 33 aligns multiple first video segments with the first audio segment based on the features of the first audio segment and the second audio segment to obtain the second video, it is specifically used to: extract the audio features of the first audio segment, including volume time-varying features and pronunciation start-point features; perform time offset matching on the first video segment and the second audio segment based on the audio features to obtain a first time sequence point; and align the frame sequence of the first video segment and the first audio segment with the first time sequence point as the starting point to generate the second video.

[0167] In one possible implementation, the generation module 33 is further configured to: fine-tune the first time point based on at least two adjacent frame frames corresponding to the first time point to obtain a second time point; when the generation module 33 generates a second video by aligning the frame sequence of the first video segment and the first audio with the first time point as the starting point, it is specifically configured to: align the frame sequence of the first video segment and the first audio with the second time point as the starting point to generate a second video.

[0168] In one possible implementation, when the generation module 33 performs time offset matching on the first video segment and the second audio based on audio features to obtain the first time point, it is specifically used to: obtain a preset time offset range; generate at least two candidate time points within the time offset range, and obtain the audio features of the corresponding audio segment in the second audio using the candidate time points as the starting point; and select the first time point from the at least two candidate time points based on the similarity between the audio features corresponding to the candidate time points and the audio features of the first audio segment.

[0169] In one possible implementation, the generation module 33 is further configured to: obtain weight values ​​of volume time-varying features and pronunciation start-point features based on the scene of the first video segment; obtain a first similarity between the volume time-varying features corresponding to the candidate time-series point and the volume time-varying features of the first audio, and obtain a second similarity between the pronunciation start-point features corresponding to the candidate time-series point and the pronunciation start-point features of the first audio; and, based on the weight values, perform a weighted summation of the first similarity and the second similarity to obtain the similarity of the audio features.

[0170] The acquisition module 31, processing module 32, and generation module 33 are connected in sequence. A possible video generation device 3 can execute the technical solution of the above method embodiment, and its implementation principle and technical effects are similar, so they will not be repeated here.

[0171] Figure 11 This is a schematic diagram of the structure of a possible electronic device, such as... Figure 11 As shown, the electronic device 4 includes:

[0172] Processor 41, and memory 42 communicatively connected to processor 41;

[0173] Memory 42 stores instructions executed by the computer;

[0174] The processor 41 executes computer execution instructions stored in the memory 42 to achieve, for example, Figures 2-9 The video processing method in the illustrated embodiment.

[0175] Optionally, the processor 41 and the memory 42 are connected via a bus 43.

[0176] For relevant instructions, please refer to the corresponding text. Figures 2-9 The relevant descriptions and effects of the steps in the corresponding embodiments are understood, and will not be elaborated on here.

[0177] One possible computer-readable storage medium stores computer-executable instructions that, when executed by a processor, are used to implement the above-disclosed... Figures 2-9 The video processing method provided in any of the corresponding embodiments.

[0178] One possible computer program product includes a computer program that, when executed by a processor, implements the above disclosure. Figures 2-9 The video processing method provided in any of the corresponding embodiments.

[0179] To implement the above embodiments, an electronic device is also provided.

[0180] refer to Figure 12 The diagram illustrates a suitable structure for implementing an electronic device 900, which can be a terminal device or a server. The terminal device can include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, personal digital assistants (PDAs), tablet computers, portable media players (PMPs), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 12 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of use of the above embodiments.

[0181] like Figure 12 As shown, the electronic device 900 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage device 908 into a random access memory (RAM) 903. The RAM 903 also stores various programs and data required for the operation of the electronic device 900. The processing unit 901, ROM 902, and RAM 903 are interconnected via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0182] Typically, the following devices can be connected to I / O interface 905: input devices 906 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 907 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 908 including, for example, magnetic tapes, hard disks, etc.; and communication devices 909. Communication device 909 allows electronic device 900 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 12An electronic device 900 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0183] Specifically, according to the above embodiments, the processes described in the flowcharts above can be implemented as computer software programs. For example, the disclosed embodiments include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 909, or installed from a storage device 908, or installed from a ROM 902. When the computer program is executed by the processing device 901, it performs the functions defined in the methods of the disclosed embodiments.

[0184] It should be noted that the computer-readable medium disclosed above can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In the above disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the above disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0185] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0186] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods shown in the above embodiments.

[0187] Computer program code for performing the operations disclosed above can be written in one or more programming languages ​​or a combination thereof. These programming languages ​​include object-oriented programming languages—such as Java, Smalltalk, and C++—as well as conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0188] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0189] The units or modules described in the embodiments can be implemented in software or hardware. The names of the units or modules do not necessarily limit the specific unit itself.

[0190] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0191] In any possible context, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0192] Firstly, in one possible implementation, a video processing method is provided, including:

[0193] Obtain a first audio file and segment it into multiple first audio segments; generate multiple first video segments based on the multiple first audio segments, with each first video segment corresponding to a multiple second audio file; align the multiple first video segments with the first audio segments based on the features of the first audio segments and the second audio files to obtain the second video file.

[0194] In one possible implementation, dividing the first audio into multiple first audio segments includes: responding to a first operation to obtain at least one segmentation position in the first audio; the first operation is used to trigger a first control or to set the segmentation position of the first audio, and the first control is used to automatically determine the segmentation position; based on the segmentation position, dividing the first audio into multiple second audio segments, each second audio segment corresponding to a speech sentence in the first audio; and merging at least two adjacent second audio segments into a first audio segment according to the content coherence between the speech sentences corresponding to the second audio segments.

[0195] One possible implementation also includes: generating a corresponding first frame based on the second audio segment, the first frame representing the semantic content of the speech sentences of the second audio segment; and obtaining the content coherence between the speech sentences corresponding to the second audio segment based on the similarity of adjacent first frames.

[0196] In one possible implementation, multiple first audio segments correspond to multiple first durations. Generating multiple first video segments based on the multiple first audio segments includes: responding to a second operation on a second control, obtaining semantic information of the first audio segments, and inputting the semantic information and the first duration corresponding to the first audio segments into a multimodal model to generate first video segments and corresponding second audio segments. The first video segments and corresponding second audio segments match the first durations. The second control is used to trigger the task of generating the first video segments, and the pronunciation features of the second audio segments match the visual features of the first video segments.

[0197] In one possible implementation, the second operation includes a first trigger operation and a second trigger operation. In response to the second operation on the second control, semantic information of the first audio segment is obtained, and the semantic information and the first duration corresponding to the first audio segment are input into a multimodal model to generate a first video segment and a corresponding second audio. This includes: in response to the trigger operation on the second control, extracting text content from the first audio segment; displaying the text content on the interactive interface; and in response to the confirmation operation on the third control, inputting the text content and the first duration corresponding to the first audio segment into the multimodal model to generate a first video segment and a corresponding second audio. The third control is used to modify and / or confirm the text content.

[0198] In one possible implementation, based on the features of the first audio segment and the second audio, multiple first video segments are aligned with the first audio segment to obtain the second video, including: extracting audio features of the first audio segment, the audio features including volume time-varying features and pronunciation start point features; performing time offset matching on the first video segment and the second audio based on the audio features to obtain a first time sequence point; and using the first time sequence point as the starting point, aligning the frame sequence of the first video segment and the first audio to generate the second video.

[0199] In one possible implementation, the method further includes: fine-tuning the first time point based on at least two adjacent frame frames corresponding to the first time point to obtain a second time point; and generating a second video by aligning the frame sequence of the first video segment and the first audio with the first time point as the starting point, including: aligning the frame sequence of the first video segment and the first audio with the second time point as the starting point to generate a second video.

[0200] In one possible implementation, based on audio features, time offset matching is performed on the first video segment and the second audio segment to obtain a first time point, including: obtaining a preset time offset range; generating at least two candidate time points within the time offset range, and using the candidate time points as starting points to obtain the audio features of the corresponding audio segment in the second audio segment; and selecting the first time point from the at least two candidate time points based on the similarity between the audio features corresponding to the candidate time points and the audio features of the first audio segment.

[0201] In one possible implementation, the method further includes: obtaining weight values ​​for volume time-varying features and pronunciation start-point features based on the scene of the first video segment; obtaining a first similarity between the volume time-varying features corresponding to the candidate time-series point and the volume time-varying features of the first audio, and obtaining a second similarity between the pronunciation start-point features corresponding to the candidate time-series point and the pronunciation start-point features of the first audio; and weighting and summing the first similarity and the second similarity based on the weight values ​​to obtain the similarity of the audio features.

[0202] Secondly, in one possible implementation, a video generation apparatus is provided, comprising:

[0203] The acquisition module is used to acquire the first audio and divide the first audio into multiple first audio segments;

[0204] The processing module is used to generate multiple first video segments based on multiple first audio segments, and the multiple first video segments correspond to multiple second audio segments;

[0205] The generation module is used to align multiple first video segments with the first audio segments based on the features of the first audio segment and the second audio segment to obtain the second video.

[0206] In one possible implementation, when the acquisition module divides the first audio into multiple first audio segments, it is specifically used to: obtain at least one segmentation position in the first audio in response to a first operation; the first operation is used to trigger a first control or to set the segmentation position of the first audio, and the first control is used to automatically determine the segmentation position; based on the segmentation position, divide the first audio into multiple second audio segments, each second audio segment corresponding to a speech sentence in the first audio; and according to the content coherence between the speech sentences corresponding to the second audio segments, merge at least two adjacent second audio segments into one first audio segment.

[0207] In one possible implementation, the acquisition module is further configured to: generate a corresponding first frame based on the second audio segment, wherein the first frame represents the semantic content of the speech sentences of the second audio segment; and obtain the content coherence between the speech sentences corresponding to the second audio segment based on the similarity of adjacent first frames.

[0208] In one possible implementation, multiple first audio segments correspond to multiple first durations. When the processing module generates multiple first video segments based on the multiple first audio segments, it is specifically used to: in response to a second operation on the second control, obtain the semantic information of the first audio segments, and input the semantic information and the first duration corresponding to the first audio segments into a multimodal model to generate first video segments and corresponding second audio, wherein the first video segments and corresponding second audio match the first duration, the second control is used to trigger the task of generating the first video segments, and the pronunciation features of the second audio match the picture features of the first video segments.

[0209] In one possible implementation, the second operation includes a first trigger operation and a second trigger operation. When the processing module, in response to the second operation on the second control, obtains the semantic information of the first audio segment and inputs the semantic information and the first duration corresponding to the first audio segment into the multimodal model to generate the first video segment and the corresponding second audio, it is specifically used to: extract the text content from the first audio segment in response to the trigger operation on the second control; display the text content on the interactive interface; and, in response to the confirmation operation of the third control, input the text content and the first duration corresponding to the first audio segment into the multimodal model to generate the first video segment and the corresponding second audio. The third control is used to modify and / or confirm the text content.

[0210] In one possible implementation, when the generation module aligns multiple first video segments with the first audio segment based on the features of the first audio segment and the second audio segment to obtain the second video, it specifically performs the following: extracting audio features of the first audio segment, including volume time-varying features and pronunciation start-point features; performing time offset matching on the first video segment and the second audio segment based on the audio features to obtain a first time sequence point; and aligning the frame sequence of the first video segment and the first audio segment with the first time sequence point as the starting point to generate the second video.

[0211] In one possible implementation, the generation module is further configured to: fine-tune the first time-series point based on at least two adjacent frame frames corresponding to the first time-series point to obtain a second time-series point; when the generation module generates a second video by aligning the frame sequence of the first video segment and the first audio with the first time-series point as the starting point, it is specifically configured to: align the frame sequence of the first video segment and the first audio with the second time-series point as the starting point to generate a second video.

[0212] In one possible implementation, when the generation module performs time offset matching on the first video segment and the second audio based on audio features to obtain the first time point, it is specifically used to: obtain a preset time offset range; generate at least two candidate time points within the time offset range, and use the candidate time points as starting points to obtain the audio features of the corresponding audio segments in the second audio; and select the first time point from the at least two candidate time points based on the similarity between the audio features corresponding to the candidate time points and the audio features of the first audio segment.

[0213] In one possible implementation, the generation module is further configured to: obtain weight values ​​of volume time-varying features and articulation start-point features based on the scene of the first video segment; obtain a first similarity between the volume time-varying features corresponding to the candidate time-series point and the volume time-varying features of the first audio, and obtain a second similarity between the articulation start-point features corresponding to the candidate time-series point and the articulation start-point features of the first audio; and, based on the weight values, perform a weighted summation of the first and second similarities to obtain the similarity of the audio features.

[0214] Thirdly, in one possible implementation, a cloud service system is provided, comprising:

[0215] The cloud server is equipped with a video generation service module, which is used to receive the first video clip and the first audio uploaded by the user, execute the video processing methods as described in the first aspect and various possible designs of the first aspect, and return the generated second video to the user terminal.

[0216] The user terminal is used to upload a first video clip and a first audio clip to the cloud server, and to receive a second video clip returned by the cloud server.

[0217] Fourthly, in one possible implementation, an electronic device is provided, comprising: at least one processor and a memory;

[0218] The memory stores the instructions that the computer executes;

[0219] At least one processor executes computer execution instructions stored in memory, causing at least one processor to perform the video processing method as described in the first aspect above and various possible designs of the first aspect.

[0220] Fifthly, in one possible implementation, a computer-readable storage medium is provided, which stores computer-executable instructions that, when executed by a processor, implement the video processing method described in the first aspect and various possible designs of the first aspect.

[0221] In a sixth aspect, in one possible implementation, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the video processing method as described in the first aspect and various possible designs of the first aspect.

[0222] The above description is merely a preferred embodiment and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in the disclosure that have similar functions.

[0223] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. Multitasking and parallel processing may be advantageous in certain environments. Similarly, while several specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the disclosure. Certain features described in the context of a single embodiment may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0224] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. A video processing method, comprising: Obtain the first audio file and divide the first audio file into multiple first audio segments; Multiple first video segments are generated based on the multiple first audio segments, and the multiple first video segments correspond to multiple second audio segments; Based on the features of the first audio segment and the second audio segment, the plurality of first video segments are aligned with the first audio segment to obtain the second video.

2. The method according to claim 1, wherein dividing the first audio into multiple first audio segments includes: In response to the first operation, at least one segmentation position in the first audio is obtained; The first operation is used to trigger the first control, or to set the segmentation position of the first audio, and the first control is used to automatically determine the segmentation position; Based on the segmentation position, the first audio is segmented into multiple second audio segments, and the second audio segments correspond to a speech sentence in the first audio. Based on the content coherence between the speech sentences corresponding to the second audio segment, at least two adjacent second audio segments are merged into a first audio segment.

3. The method according to claim 2, further comprising: Based on the second audio segment, a corresponding first screen is generated, wherein the first screen represents the semantic content of the speech sentence of the second audio segment; Based on the image similarity of adjacent first images, the content coherence between the speech sentences corresponding to the second audio segment is obtained.

4. The method according to claim 1, wherein the plurality of first audio segments correspond to a plurality of first durations, and the step of generating a plurality of first video segments based on the plurality of first audio segments includes: In response to a second operation on the second control, semantic information of the first audio segment is obtained, and the semantic information and the first duration corresponding to the first audio segment are input into a multimodal model to generate a first video segment and a corresponding second audio segment. The first video segment and the corresponding second audio segment match the first duration. The second control is used to trigger the task of generating the first video segment. The pronunciation features of the second audio segment match the visual features of the first video segment.

5. The method according to claim 4, wherein the second operation includes a first triggering operation and a second triggering operation, and the step of obtaining semantic information of the first audio segment in response to the second operation on the second control, and inputting the semantic information and the first duration corresponding to the first audio segment into a multimodal model to generate a first video segment and a corresponding second audio segment, includes: In response to a trigger operation on the second control, extract the text content from the first audio segment; Display the text content on the interactive interface; In response to the confirmation operation of the third control, the text content and the first duration corresponding to the first audio segment are input into the multimodal model to generate a first video segment and a corresponding second audio segment. The third control is used to modify and / or confirm the text content.

6. The method according to claim 1, wherein aligning the plurality of first video segments with the first audio segment based on the features of the first audio segment and the second audio segment to obtain the second video comprises: Extract the audio features of the first audio segment, the audio features including volume time-varying features and articulation start features; Based on the audio features, time offset matching is performed on the first video segment and the second audio to obtain the first time point; Starting from the first time point, align the frame sequence of the first video segment with the first audio segment to generate a second video.

7. The method according to claim 6, further comprising: Based on at least two adjacent frame frames corresponding to the first time point, the first time point is fine-tuned to obtain the second time point; The step of generating a second video by aligning the frame sequence of the first video segment and the first audio segment with the first time point as the starting point includes: Starting from the second time point, the frame sequence of the first video segment and the first audio segment are aligned to generate the second video.

8. The method according to claim 7, wherein performing time offset matching on the first video segment and the second audio based on the audio features to obtain the first time point includes: Get the preset time offset range; At least two candidate time points are generated within the time offset range, and the audio features of the corresponding audio segments in the second audio are obtained using the candidate time points as the starting point. Based on the similarity between the audio features corresponding to the candidate time points and the audio features of the first audio segment, the first time point is selected from the at least two candidate time points.

9. The method according to claim 8, further comprising: Based on the scene of the first video segment, the weight values ​​of the volume time-varying feature and the pronunciation start point feature are obtained respectively; Obtain the first similarity between the volume time-varying feature corresponding to the candidate time-series point and the volume time-varying feature of the first audio segment, and obtain the second similarity between the pronunciation start point feature corresponding to the candidate time-series point and the pronunciation start point feature of the first audio segment; Based on the weight values, the first similarity and the second similarity are weighted and summed to obtain the similarity of the audio features.

10. An electronic device, comprising: Processor and memory; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the video processing method as described in any one of claims 1 to 9.

11. A computer program product comprising a computer program that, when executed by a processor, implements the video processing method as described in any one of claims 1 to 9.