Audio and video processing method and device

By double-speed processing of audio materials and cache audio data, and adjusting the video frame screen time stamp with audio timestamp, the synchronization problem of double-speed preview of materials in video editing drafts is solved, and the effect of audio and video synchronization and quick browsing is achieved.

CN120568147APending Publication Date: 2025-08-29BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410232485.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-29
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

In video editing scenarios, it is difficult for the existing technology to quickly and accurately preview the materials in video editing drafts, especially synchronous processing of audio and video clips, resulting in the edited parts that cannot be quickly browsed during the editing process.

Method used

After double-speed processing of audio materials, the audio data is cached in real time and the video frame screen time stamp is adjusted according to the audio timestamp to achieve double-speed playback of audio and video synchronized, ensuring that the video frame is aligned with the audio timestamp, and supporting overall browsing.

Benefits of technology

It realizes convenient and accurate speed preview playback in video editing drafts, keeps the material attributes unchanged, supports quick browsing of edited parts, and improves editing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120568147A_ABST
    Figure CN120568147A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an audio and video processing method and device, and the method comprises the steps: obtaining an audio material in a video editing draft, carrying out the speed multiplication of the audio material, and obtaining the speed-multiplied audio data; adjusting video materials in the video editing draft according to the audio data to obtain adjusted video data; wherein timestamps of video frame pictures in the audio data and the video data are aligned, and then the audio data and the video data are presented. According to the technical scheme, speed-multiplying special effect processing is carried out on the audio data, and the video frames and the audio timestamps played at the speed-multiplying time are subjected to corresponding audio and picture synchronization, so that the speed-multiplying playing effect of the whole editing timeline is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer and network communication technology, and more particularly to an audio and video processing method and device. Background Art

[0002] In existing video editing scenarios, when previewing and playing audio and video materials in a video editing draft at multiple speeds, it is often possible to support speed changes for individual video clips and individual audio clips. For example, after speed changes are performed on each audio clip and individual video clip using speed change information, the clips are superimposed on the entire editing draft and played sequentially according to the timeline during preview playback.

[0003] However, in the above processing, changing the speed of a single audio clip will change the clip properties, and in the editing process, the edited part is often expected to be played quickly for overall browsing.

[0004] Therefore, how to conveniently and accurately achieve double-speed preview and playback of materials in video editing drafts in editing scenarios has become a technical problem that needs to be solved urgently. Summary of the Invention

[0005] The embodiments of the present disclosure provide an audio and video processing method and device to solve the technical problem of conveniently and accurately realizing double-speed preview and playback of materials in a draft in an editing scenario.

[0006] In a first aspect, an embodiment of the present disclosure provides an audio and video processing method, including:

[0007] Get the audio material from the video editing draft;

[0008] Performing speed processing on the audio material to obtain speeded-up audio data;

[0009] Adjusting the video material in the video editing draft according to the audio data to obtain adjusted video data; wherein the audio data and the video frames in the video data are time-stamp aligned;

[0010] The audio data and the video data are presented.

[0011] In a second aspect, an embodiment of the present disclosure provides an audio and video processing device, including:

[0012] An acquisition unit, used for acquiring audio materials in a video editing draft;

[0013] A first processing unit is configured to perform speed processing on the audio material to obtain speeded audio data;

[0014] A second processing unit is configured to adjust the video material in the video editing draft according to the audio data to obtain adjusted video data; wherein the audio data and the video frames in the video data are time-stamp aligned;

[0015] The third processing unit is configured to present the audio data and the video data.

[0016] In a third aspect, an embodiment of the present disclosure provides an electronic device, including: a processor and a memory;

[0017] The memory stores computer-executable instructions;

[0018] The processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the audio and video processing method described in the possible design of the first aspect above.

[0019] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, in which computer execution instructions are stored. When a processor executes the computer execution instructions, the audio and video processing method described in the possible design of the first aspect above is implemented.

[0020] In a fifth aspect, an embodiment of the present disclosure provides a computer program product, including a computer program, which, when executed by a processor, implements the audio and video processing method as described in the possible design of the first aspect above.

[0021] This embodiment provides an audio and video processing method and device. The method obtains audio material from a video editing draft and processes the audio material at a double speed to obtain sped-up audio data. The method then adjusts the video material in the video editing draft according to the audio data to obtain adjusted video data. The audio data and the video frames in the video data are time-stamp aligned, and the audio and video data are then presented. In this technical solution, the audio data is processed at a double speed, and the video frames are synchronized with the timestamps of the audio played at the double speed, thereby achieving the effect of double-speed playback of the entire editing timeline. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0023] Figure 1 Schematic diagram of the audio and video processing method provided in the embodiment of the present disclosure Figure 1 ;

[0024] Figure 2 Schematic diagram of the audio and video processing method provided in the embodiment of the present disclosure Figure 2 ;

[0025] Figure 3 Schematic diagram of the audio and video processing method provided in the embodiment of the present disclosure Figure 3 ;

[0026] Figure 4 Schematic diagram of the audio and video processing method provided in the embodiment of the present disclosure Figure 4 ;

[0027] Figure 5 A schematic diagram of the structure of an audio and video processing device provided in an embodiment of the present disclosure;

[0028] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0029] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present disclosure without making any creative efforts shall fall within the scope of protection of the present disclosure.

[0030] In existing video editing scenarios, when previewing and playing audio and video materials in a video editing draft at multiple speeds, it is often possible to support speed changes for individual video clips and individual audio clips. For example, after speed changes are performed on each audio clip and individual video clip using speed change information, the clips are superimposed on the entire editing draft and played sequentially according to the timeline during preview playback.

[0031] However, in the above processing, changing the speed of a single audio clip will change the clip properties, and in the editing process, the edited part is often expected to be played quickly for overall browsing.

[0032] Therefore, how to conveniently and accurately achieve double-speed preview and playback of materials in video editing drafts in editing scenarios has become a technical problem that needs to be solved urgently.

[0033] In order to solve the above-mentioned technical problems, the inventor's technical conception is as follows: after the audio data is processed with double-speed special effects, it is cached in real time, that is, the cached audio data is previewed and displayed on the timeline. In order to ensure the synchronization of sound and picture, the video frame can be timestamp-aligned with the audio timestamp played at double speed to achieve the effect of double-speed playback of the entire timeline. Under this implementation, the properties of the material clips will not be changed, and during the editing process, the edited part can be played quickly for overall browsing.

[0034] The audio and video processing method according to the embodiments of the present disclosure is performed by an electronic device, which may be a mobile phone, a computer, a tablet, a server, and the like.

[0035] The following is a specific implementation process of the audio and video processing method, audio and video processing device, and electronic device involved in the embodiments of the present disclosure. Some examples are only for illustrative purposes and are not limiting.

[0036] Figure 1 Schematic diagram of the audio and video processing method provided in the embodiment of the present disclosure Figure 1 .like Figure 1 As shown, the audio and video processing method includes:

[0037] Step 11. Get the audio material in the video editing draft;

[0038] In this step, in the video editing scenario, the audio material in the video editing draft is obtained. The audio material can be voice, music or other materials; it can also be the audio part in the video.

[0039] It should be understood that the number and type of audio materials are not limited, the number can be one or more, and the types can include voice, music, etc.

[0040] Step 12: Process the audio material at a double speed to obtain the double-speed audio data;

[0041] In this step, the audio material is processed at a double speed for preview, and the audio data after double speed processing is cached.

[0042] In some possible implementations, the audio material is decoded to obtain multiple sub-audio data, each sub-audio data includes: P sampling point data, and then the multiple sub-audio data are processed at a double speed to obtain audio data after the double speed, where P is a positive integer.

[0043] The audio data at the doubled speed is then stored in a First-In-First-Out (FIFO) cache.

[0044] In this implementation, based on the FIFO principle, after the previewed video frames are continuously obtained, the audio and video data with synchronized audio and video can be continuously previewed and played.

[0045] Step 13: Adjust the video material in the video editing draft according to the audio data to obtain the adjusted video data;

[0046] The audio data and the video frames in the video data are time-stamp aligned.

[0047] In this step, since the audio frames and video frames whose timestamps are aligned after the speed adjustment are also aligned before the speed adjustment, the timestamp positions before and after the speed adjustment are changed.

[0048] Therefore, it is necessary to adjust the video material in the video editing draft according to the audio data obtained after speed doubling. Based on the audio data, the video frame images that need to be decoded from the video editing draft can be determined first, and then the corresponding video frame images can be decoded to obtain the adjusted video data.

[0049] For example, in a double-speed adjustment, the audio frame and video frame at the 1s position of the timeline before the adjustment will be located at the 0.5s position of the editing timeline after the adjustment, that is, they are aligned at the 0.5s position.

[0050] Step 14: Present the audio data and video data.

[0051] In this step, after the above processing, audio data and video data are obtained, and the video data and audio data can be presented on a video editing processing device.

[0052] In a possible implementation, the audio data and the video data are previewed and played according to the editing timeline. For example, the audio data can be played through the audio module and the video data can be played through the display module on the editing timeline.

[0053] The audio and video processing method provided by the disclosed embodiment obtains audio material from a video editing draft and processes the audio material at a double speed to obtain the double-speed audio data; adjusts the video material in the video editing draft according to the audio data to obtain the adjusted video data; wherein the audio data and the video frame images in the video data are time-stamp aligned, and then the audio data and video data are presented. In this technical solution, the audio data is processed at a double speed to synchronize the audio and video frames with the timestamps of the audio played at the double speed, thereby achieving the effect of double-speed playback of the entire editing timeline.

[0054] Based on the above embodiments, Figure 2 Schematic diagram of the audio and video processing method provided in the embodiment of the present disclosure Figure 2 .like Figure 2As shown, the above step 12 may include:

[0055] Step 21: Decode the audio material to obtain a plurality of sub-audio data, each of which includes: P sampling point data;

[0056] Wherein, P is a positive integer.

[0057] In this step, the audio material is decoded by driving P samples as a group, that is, the data required for each decoding corresponds to the P samples driven, that is, each sub-audio data obtained contains P sampling point data.

[0058] The duration corresponding to each driving P samples is the duration corresponding to the audio data.

[0059] For example, P may be 1024.

[0060] It should be understood that the embodiment of the present application does not limit the number of audio materials, which can be one or more, arranged sequentially on the editing timeline.

[0061] Optionally, the number of the multiple sub-audio data is N, where N is an integer greater than 1.

[0062] Then, step 21 may be: decoding the audio material to obtain N-1 sub-audio data and M sampling point data, where M is an integer greater than 0 and less than P; padding the M sampling point data, and using the padded M sampling points as the Nth sub-audio data;

[0063] The multiple sub-audio data include: N-1 sub-audio data and Nth sub-audio data.

[0064] In one possible implementation, when decoding the audio material, the first N-1 sub-audio data are decoded, and after further decoding, 2 sampling point data are obtained. Then, the 1022 data after the 2 sampling point data are padded with 0 according to 1024 samples, which is a kind of padding operation, to obtain the Nth sub-audio data.

[0065] Step 22: Perform double-speed processing on the plurality of sub-audio data to obtain double-speed audio data.

[0066] In this step, the decoded sub-audio data are processed at a double speed, which can be achieved by changing the speed without changing the pitch.

[0067] For example, if the speed information corresponding to the speed processing is 2x speed, 1024*2 sampling point data is required to obtain the audio data after speed increase (input 1024*2, output 1024).

[0068] Among them, the 1024 sampling point data after speed increase cannot be directly converted to the sampling rate. At this time, it is necessary to multiply it by the speed information to obtain the actual playback time.

[0069] The implementation in this embodiment can be to decode and render multiple audio data, and then perform speed-up special effects processing. The speed-up audio data is stored in the cache FIFO. When the cache meets a pipeline (English: pipeline), it is dequeued from the FIFO. Furthermore, if the speed meets 1024 after multiplication, memory samples (English: samples) of the size corresponding to 1024 sample data are applied.

[0070] Then the audio and video are synchronized, and the player previews and plays the video.

[0071] The audio and video processing method provided by the disclosed embodiments decodes audio material to obtain multiple sub-audio data, each sub-audio data including 1024 sampling points. The multiple sub-audio data are then speed-processed to obtain the speeded audio data. In this technical solution, audio decoding is performed based on the 1024 sampling points as one sub-audio data. The speed information is then used to speed-process the decoded multiple sub-audio data, thereby determining the audio data actually consumed after the audio driver.

[0072] Based on the above embodiments, Figure 3 Schematic diagram of the audio and video processing method provided in the embodiment of the present disclosure Figure 3 .like Figure 3 As shown, step 13 may include:

[0073] Step 31: Determine the timestamp of the video frame in the video material on the editing timeline based on the timestamp corresponding to the audio data;

[0074] In this step, the timestamps before the audio data and video data are adjusted at multiple speeds are aligned. Therefore, after the audio data and video data are adjusted at multiple speeds, their timestamps are also aligned. Therefore, the timestamps corresponding to the video frame that needs to be displayed on the editing timeline can be determined based on the timestamps corresponding to the audio data.

[0075] The preview of the video material in the video editing draft is based on the preview position on the editing timeline. The video frame at the preview position is extracted and synthesized according to the editing information of the video editing draft to render the preview image and then display it. When the video editing draft preview is played, the preview position is changed in sequence according to the editing timeline to render the preview image and display it on the screen in sequence.

[0076] For example, for video material, at 1x speed, the frames corresponding to the video frames are: start frame, play point K, at speed S ( Figure 2Taking 2 as an example), the frames corresponding to the video frame are: start frame, play point K+S.

[0077] Step 32: Decode the video frames of the video material in the video editing draft according to the timestamps of the video frames in the video material to obtain video data.

[0078] In this step, the video frames that need to be decoded in the video material can be determined based on the timestamps of the video frames in the video material, and then these video frames are decoded to obtain video data.

[0079] The audio and video processing method provided by the embodiment of the present disclosure determines the timestamp of the video frame in the video material on the editing timeline according to the timestamp corresponding to the audio data, and decodes the video frame of the video material in the video editing draft according to the timestamp of the video frame in the video material to obtain the video data, so that after the audio speed is increased, the corresponding video speed can be adjusted to ensure the synchronization of the audio and video corresponding to the audio data and the video data.

[0080] Based on the above embodiments, Figure 4 Schematic diagram of the audio and video processing method provided in the embodiment of the present disclosure Figure 4 .like Figure 4 As shown, the above step 32 may include:

[0081] Step 41: If the decoding unit where the timestamp of the first video frame is located is different from the decoding unit where the timestamp of the second video frame is located, after the decoding of the first video frame is completed, the second video frame is decoded in the decoding unit corresponding to the second video frame, and the second video frame is used as the new first video frame. This step is repeated until the video frames in the video material are decoded.

[0082] The first video frame is a video frame currently being decoded, and the second video frame is a video frame to be decoded next.

[0083] In this implementation, before decoding the video material, it is necessary to understand that the encoded video material has determined the corresponding multiple decoding units during encoding. During decoding, the decoder decodes the video frames in the decoding units in sequence according to the editing timeline.

[0084] That is, after the timestamp of the second video frame is determined as described above, when decoding the currently decoded first video frame, it is determined whether the decoding unit to which the timestamp of the second video frame belongs is the same decoding unit as the decoding unit to which the timestamp of the currently decoded first video frame belongs. If they are not the same decoding unit, the decoder is controlled to decode the second video frame from the decoding unit to which the timestamp of the second video frame belongs.

[0085] Optionally, this implementation can be achieved by refreshing the decoder.

[0086] Then, the second video frame is decoded, and the next video frame is used as a new second video frame. The above process is repeated until the video frame in the video material is decoded.

[0087] In the prior art, all video frames in the decoding unit must be decoded to determine the video frame to be previewed, which is inefficient. However, this solution can directly skip the subsequent decoding process of the current decoding unit when the decoding unit to which the timestamp of the next frame to be previewed belongs is different from the decoding unit to which the timestamp of the currently decoded frame to be previewed belongs, which is more efficient.

[0088] Step 42: When the decoding unit where the timestamp of the first video frame is located is the same as the decoding unit where the timestamp of the second video frame is located, after the decoding of the first video frame is completed, continue decoding the second video frame in the decoding unit where the timestamp of the first video frame is located, and repeat this step until the video frame in the decoding unit is completely decoded.

[0089] After the timestamp of the second video frame is determined as described above, when decoding the currently decoded first video frame, it is determined whether the decoding unit to which the timestamp of the second video frame belongs is the same decoding unit as the decoding unit to which the timestamp of the currently decoded first video frame belongs. If they are the same decoding unit, decoding continues in the current decoding unit until the second video frame is decoded.

[0090] Furthermore, the second video frame is used as the new first video frame currently being decoded, and the above judgment is continued until the last frame to be previewed in the current decoding unit is decoded, and the last frame to be previewed in the current decoding unit is used as the new first video frame currently being decoded.

[0091] Step 43: When the decoding unit where the timestamp of the new first video frame is located is different from the decoding unit where the timestamp of the second video frame is located, after the decoding of the new first video frame is completed, the second video frame is decoded in the decoding unit corresponding to the second video frame, and the second video frame is used as the new first video frame. This step is repeated until the video frame in the video material is decoded.

[0092] In this step, if it is determined in the previous step that the decoding unit where the new currently decoded first video frame is located is different from the decoding unit where the timestamp of the second video frame is located, then after the decoding of the new currently decoded first video frame is completed, the decoder is refreshed and decoding is performed in the decoding unit where the timestamp of the second video frame is located until the decoding of the second video frame is completed.

[0093] After that, the second video frame is used as the new first video frame currently decoded, and this step is repeated until the decoding process of all frames to be previewed in the video material is completed, and multiple decoded video frames are obtained.

[0094] The audio and video processing method provided by the embodiment of the present disclosure includes: when the decoding unit where the timestamp of the first video frame is located is different from the decoding unit where the timestamp of the second video frame is located, after the decoding of the first video frame is completed, the second video frame is decoded in the decoding unit corresponding to the second video frame, and the second video frame is used as the new first video frame, and this step is repeated until the video frame in the video material is decoded; when the decoding unit where the timestamp of the first video frame is located is the same as the decoding unit where the timestamp of the second video frame is located, after the decoding of the first video frame is completed, the second video frame is continued to be decoded in the decoding unit where the timestamp of the first video frame is located, and this step is repeated until the video frame in the decoding unit is decoded; when the decoding unit where the timestamp of the new first video frame is located is different from the decoding unit where the timestamp of the second video frame is located, after the decoding of the new first video frame is completed, the second video frame is decoded in the decoding unit corresponding to the second video frame, and the second video frame is used as the new first video frame, and this step is repeated until the video frame in the video material is decoded. In this technical solution, when decoding a video frame, if it is determined that the next video frame is not in the current decoding unit, it directly jumps to the next decoding unit for decoding, thereby improving the efficiency of previewing in the video editing scenario during double-speed playback.

[0095] Based on the above method embodiment, Figure 5 A schematic diagram of the structure of the audio and video processing device provided in the embodiment of the present disclosure is shown in FIG. Figure 5As shown, the audio and video processing device includes:

[0096] An acquisition unit 51 is used to acquire audio material in a video editing draft;

[0097] The first processing unit 52 is used to perform speed processing on the audio material to obtain speeded audio data;

[0098] The second processing unit 53 is configured to adjust the video material in the video editing draft according to the audio data to obtain adjusted video data; wherein the audio data and the video frames in the video data are time-stamp aligned;

[0099] The third processing unit 54 is configured to display the audio data and the video data.

[0100] According to one or more embodiments of the present disclosure, the first processing unit 52 is specifically configured to:

[0101] Decoding the audio material to obtain a plurality of sub-audio data, each sub-audio data including: P sampling point data, where P is a positive integer;

[0102] The multiple sub-audio data are processed at a double speed to obtain the double-speed audio data.

[0103] According to one or more embodiments of the present disclosure, the number of the plurality of sub-audio data is N, where N is an integer greater than 1;

[0104] Accordingly, the first processing unit 52 decodes the audio material to obtain a plurality of sub-audio data, specifically:

[0105] Decode the audio material to obtain N-1 sub-audio data and M sampling point data, where M is an integer greater than 0 and less than P;

[0106] Perform padding operation on the M sampling point data, and use the padded M sampling points as the Nth sub-audio data;

[0107] The multiple sub-audio data include: N-1 sub-audio data and Nth sub-audio data.

[0108] According to one or more embodiments of the present disclosure, the second processing unit 53 is configured to:

[0109] Determine the timestamp of the video frame in the video material on the editing timeline according to the timestamp corresponding to the audio data;

[0110] According to the time stamp of the video frame in the video material, the video frame of the video material in the video editing draft is decoded to obtain video data.

[0111] According to one or more embodiments of the present disclosure, the second processing unit 53 decodes the video frames of the video material in the video editing draft according to the timestamps of the video frames in the video material to obtain video data, specifically for:

[0112] When the decoding unit where the timestamp of the first video frame is located is different from the decoding unit where the timestamp of the second video frame is located, after the decoding of the first video frame is completed, the second video frame is decoded in the decoding unit corresponding to the second video frame, and the second video frame is used as the new first video frame. This step is repeated until the video frame in the video material is decoded. The first video frame is the video frame currently being decoded, and the second video frame is the next video frame to be decoded.

[0113] According to one or more embodiments of the present disclosure, the second processing unit 53 is further configured to:

[0114] When the decoding unit where the timestamp of the first video frame is located is the same as the decoding unit where the timestamp of the second video frame is located, after decoding of the first video frame is completed, continuing to decode the second video frame in the decoding unit where the timestamp of the first video frame is located, repeating this step until the video frame in the decoding unit is completely decoded;

[0115] When the decoding unit where the timestamp of the new first video frame is located is different from the decoding unit where the timestamp of the second video frame is located, after the decoding of the new first video frame is completed, the second video frame is decoded in the decoding unit corresponding to the second video frame, and the second video frame is used as the new first video frame. This step is repeated until the video frame in the video material is decoded.

[0116] According to one or more embodiments of the present disclosure, the third processing unit 54 is configured to:

[0117] Preview and play audio and video data according to the editing timeline.

[0118] The technical solutions and technical effects of the audio and video processing device provided in the embodiments of the present disclosure are similar to those in the above embodiments and will not be repeated here.

[0119] In order to implement the above embodiment, the embodiment of the present disclosure further provides an electronic device. Figure 6 This is a schematic diagram of the structure of the electronic device provided in the embodiment of the present disclosure, with reference to Figure 6 , the electronic device may be a terminal device.

[0120] Among them, the terminal equipment may include but is not limited to mobile terminals such as mobile phones, laptops, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (Portable Android Devices, PADs), portable multimedia players (PMPs), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., as well as fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.

[0121] like Figure 6 As shown, the electronic device may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 61, which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 62 or programs loaded from a storage device 68 into a random access memory (RAM) 63. Various programs and data required for the operation of the electronic device are also stored in RAM 63. The processing device 61, ROM 62, and RAM 63 are connected to each other via a bus 64. An input / output (I / O) interface 65 is also connected to the bus 64.

[0122] Typically, the following devices may be connected to the I / O interface 65: an input device 66 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 67 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 68 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 69. The communication device 69 may allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Although Figure 6 The electronic device is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.

[0123] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication device 69, or installed from the storage device 68, or installed from the ROM 62. When the computer program is executed by the processing device 61, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.

[0124] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0125] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0126] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device executes the method shown in the above embodiment.

[0127] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0128] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0129] The units involved in the embodiments described in this disclosure may be implemented in software or hardware. In some cases, the name of a unit does not limit the unit itself. For example, the first acquisition unit may also be described as a "unit for acquiring at least two Internet Protocol addresses."

[0130] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0131] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0132] In a first aspect, according to one or more embodiments of the present disclosure, there is provided an audio and video processing method, comprising:

[0133] Get the audio material from the video editing draft;

[0134] Performing speed processing on the audio material to obtain speeded-up audio data;

[0135] Adjusting the video material in the video editing draft according to the audio data to obtain adjusted video data; wherein the audio data and the video frames in the video data are time-stamp aligned;

[0136] The audio data and the video data are presented.

[0137] According to one or more embodiments of the present disclosure, the step of processing the audio material at a double speed to obtain the double-speeded audio data includes:

[0138] Decoding the audio material to obtain a plurality of sub-audio data, each sub-audio data including: P sampling point data, where P is a positive integer;

[0139] The multiple sub-audio data are processed at a double speed to obtain the audio data at a double speed.

[0140] According to one or more embodiments of the present disclosure, the number of the plurality of sub-audio data is N, where N is an integer greater than 1;

[0141] Correspondingly, the audio material is decoded to obtain a plurality of sub-audio data, including:

[0142] Decoding the audio material to obtain N-1 sub-audio data and M sampling point data, where M is an integer greater than 0 and less than P;

[0143] Perform padding operation on the M sampling point data, and use the padded M sampling points as the Nth sub-audio data;

[0144] The multiple sub-audio data include: N-1 sub-audio data and Nth sub-audio data.

[0145] According to one or more embodiments of the present disclosure, adjusting the video material in the video editing draft according to the audio data to obtain the adjusted video data includes:

[0146] Determining the timestamp of the video frame in the video material on the editing timeline according to the timestamp corresponding to the audio data;

[0147] The video frame images of the video material in the video editing draft are decoded according to the time stamps of the video frame images in the video material to obtain the video data.

[0148] According to one or more embodiments of the present disclosure, decoding the video frame images of the video material in the video editing draft according to the timestamps of the video frame images in the video material to obtain the video data includes:

[0149] When the decoding unit where the timestamp of the first video frame is located is different from the decoding unit where the timestamp of the second video frame is located, after the decoding of the first video frame is completed, the second video frame is decoded in the decoding unit corresponding to the second video frame, and the second video frame is used as the new first video frame. This step is repeated until the video frame in the video material is decoded. The first video frame is the video frame currently being decoded, and the second video frame is the next video frame to be decoded.

[0150] According to one or more embodiments of the present disclosure, the method further includes:

[0151] When the decoding unit where the timestamp of the first video frame is located is the same as the decoding unit where the timestamp of the second video frame is located, after decoding of the first video frame is completed, continuing to decode the second video frame in the decoding unit where the timestamp of the first video frame is located, and repeating this step until the video frame in the decoding unit is completely decoded;

[0152] When the decoding unit where the timestamp of the new first video frame is located is different from the decoding unit where the timestamp of the second video frame is located, after the decoding of the new first video frame is completed, the second video frame is decoded in the decoding unit corresponding to the second video frame, and the second video frame is used as the new first video frame. This step is repeated until the video frame in the video material is decoded.

[0153] According to one or more embodiments of the present disclosure, presenting the audio data and the video data includes:

[0154] The audio data and the video data are previewed and played according to the editing timeline.

[0155] In a second aspect, according to one or more embodiments of the present disclosure, there is provided an audio and video processing device, including:

[0156] An acquisition unit, used for acquiring audio materials in a video editing draft;

[0157] A first processing unit is configured to perform speed processing on the audio material to obtain speeded audio data;

[0158] A second processing unit is configured to adjust the video material in the video editing draft according to the audio data to obtain adjusted video data; wherein the audio data and the video frames in the video data are time-stamp aligned;

[0159] The third processing unit is configured to display the audio data and the video data.

[0160] According to one or more embodiments of the present disclosure, the first processing unit is specifically configured to:

[0161] Decoding the audio material to obtain a plurality of sub-audio data, each sub-audio data including: P sampling point data, where P is a positive integer;

[0162] The multiple sub-audio data are processed at a double speed to obtain the audio data at a double speed.

[0163] According to one or more embodiments of the present disclosure, the number of the plurality of sub-audio data is N, where N is an integer greater than 1;

[0164] Correspondingly, the first processing unit performs a decoding operation on the audio material to obtain a plurality of sub-audio data, specifically:

[0165] Decoding the audio material to obtain N-1 sub-audio data and M sampling point data, where M is an integer greater than 0 and less than P;

[0166] Perform padding operation on the M sampling point data, and use the padded M sampling points as the Nth sub-audio data;

[0167] The multiple sub-audio data include: N-1 sub-audio data and Nth sub-audio data.

[0168] According to one or more embodiments of the present disclosure, the second processing unit is configured to:

[0169] Determining the timestamp of the video frame in the video material on the editing timeline according to the timestamp corresponding to the audio data;

[0170] The video frame images of the video material in the video editing draft are decoded according to the time stamps of the video frame images in the video material to obtain the video data.

[0171] According to one or more embodiments of the present disclosure, the second processing unit decodes the video frame images of the video material in the video editing draft according to the timestamps of the video frame images in the video material to obtain the video data, specifically for:

[0172] When the decoding unit where the timestamp of the first video frame is located is different from the decoding unit where the timestamp of the second video frame is located, after the decoding of the first video frame is completed, the second video frame is decoded in the decoding unit corresponding to the second video frame, and the second video frame is used as the new first video frame. This step is repeated until the video frame in the video material is decoded. The first video frame is the video frame currently being decoded, and the second video frame is the next video frame to be decoded.

[0173] According to one or more embodiments of the present disclosure, the second processing unit is further configured to:

[0174] When the decoding unit where the timestamp of the first video frame is located is the same as the decoding unit where the timestamp of the second video frame is located, after decoding of the first video frame is completed, continuing to decode the second video frame in the decoding unit where the timestamp of the first video frame is located, and repeating this step until the video frame in the decoding unit is completely decoded;

[0175] When the decoding unit where the timestamp of the new first video frame is located is different from the decoding unit where the timestamp of the second video frame is located, after the decoding of the new first video frame is completed, the second video frame is decoded in the decoding unit corresponding to the second video frame, and the second video frame is used as the new first video frame. This step is repeated until the video frame in the video material is decoded.

[0176] According to one or more embodiments of the present disclosure, the third processing unit is configured to:

[0177] The audio data and the video data are previewed and played according to the editing timeline.

[0178] In a third aspect, according to one or more embodiments of the present disclosure, there is provided an electronic device, comprising: at least one processor and a memory;

[0179] The memory stores computer-executable instructions;

[0180] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the audio and video processing method described in the first aspect and various possible designs of the first aspect.

[0181] In a fourth aspect, according to one or more embodiments of the present disclosure, a computer-readable storage medium is provided, in which computer execution instructions are stored. When a processor executes the computer execution instructions, the audio and video processing method described in the first aspect and various possible designs of the first aspect is implemented.

[0182] In a fifth aspect, according to one or more embodiments of the present disclosure, a computer program product is provided, including a computer program, which, when executed by a processor, implements the audio and video processing method described in the first aspect and various possible designs of the first aspect.

[0183] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.

[0184] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

[0185] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. An audio and video processing method, characterized in that: include: Get the audio material from the video editing draft; Performing speed processing on the audio material to obtain speeded-up audio data; Adjusting the video material in the video editing draft according to the audio data to obtain adjusted video data; wherein the audio data and the video frames in the video data are time-stamp aligned; The audio data and the video data are presented.

2. The method according to claim 1, characterized in that The step of processing the audio material at a double speed to obtain the double-speed audio data includes: Decoding the audio material to obtain a plurality of sub-audio data, each sub-audio data including: P sampling point data, where P is a positive integer; The multiple sub-audio data are processed at a double speed to obtain the audio data at a double speed.

3. The method according to claim 2, characterized in that The number of the plurality of sub-audio data is N, where N is an integer greater than 1; Correspondingly, the audio material is decoded to obtain a plurality of sub-audio data, including: Decoding the audio material to obtain N-1 sub-audio data and M sampling point data, where M is an integer greater than 0 and less than P; Perform padding operation on the M sampling point data, and use the padded M sampling points as the Nth sub-audio data; The multiple sub-audio data include: N-1 sub-audio data and Nth sub-audio data.

4. The method according to any one of claims 1 to 3, characterized in that The step of adjusting the video material in the video editing draft according to the audio data to obtain the adjusted video data includes: Determining the timestamp of the video frame in the video material on the editing timeline according to the timestamp corresponding to the audio data; The video frame images of the video material in the video editing draft are decoded according to the time stamps of the video frame images in the video material to obtain the video data.

5. The method according to claim 4, characterized in that The decoding of the video frame images of the video material in the video editing draft according to the timestamps of the video frame images in the video material includes: When the decoding unit where the timestamp of the first video frame is located is different from the decoding unit where the timestamp of the second video frame is located, after the decoding of the first video frame is completed, the second video frame is decoded in the decoding unit corresponding to the second video frame, and the second video frame is used as the new first video frame. This step is repeated until the video frame in the video material is decoded. The first video frame is the video frame currently being decoded, and the second video frame is the next video frame to be decoded.

6. The method according to claim 5, characterized in that The method further comprises: When the decoding unit where the timestamp of the first video frame is located is the same as the decoding unit where the timestamp of the second video frame is located, after decoding of the first video frame is completed, continuing to decode the second video frame in the decoding unit where the timestamp of the first video frame is located, and repeating this step until the video frame in the decoding unit is completely decoded; When the decoding unit where the timestamp of the new first video frame is located is different from the decoding unit where the timestamp of the second video frame is located, after the decoding of the new first video frame is completed, the second video frame is decoded in the decoding unit corresponding to the second video frame, and the second video frame is used as the new first video frame. This step is repeated until the video frame in the video material is decoded.

7. The method according to any one of claims 1 to 3, characterized in that The presenting of the audio data and the video data comprises: The audio data and the video data are previewed and played according to the editing timeline.

8. An audio and video processing device, characterized in that: include: An acquisition unit, used for acquiring audio materials in a video editing draft; A first processing unit is configured to perform speed processing on the audio material to obtain speeded audio data; A second processing unit is configured to adjust the video material in the video editing draft according to the audio data to obtain adjusted video data; wherein the audio data and the video frames in the video data are time-stamp aligned; The third processing unit is configured to display the audio data and the video data.

9. An electronic device, characterized in that: include: processor and memory; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor executes the audio and video processing method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions. When the processor executes the computer-executable instructions, the audio and video processing method according to any one of claims 1 to 7 is implemented.

11. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method for audio and video processing according to any one of claims 1 to 7 is implemented.