Audio and video processing method and system, electronic equipment and storage medium

By dividing live video streams into independent audio streams and video streams, the problem of difficulty in adjusting live video and sound synchronization in the prior art is solved, flexible audio adjustment and reduced cost of video editing are achieved, and the multiplexing of live streams and large-scale expansion of volume is supported.

CN120223940APending Publication Date: 2025-06-27CTRIP TRAVEL NETWORK TECH SHANGHAI0
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510468154.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art is difficult to adjust the synchronous presentation of pictures and sounds in live broadcasts based on audience interaction, and has high requirements for graphics computing capabilities, resulting in the inability to expand on a large scale. The highly customized live broadcast streams are technically difficult and costly when reusing.

Method used

By dividing video playback into audio stream and video stream, the independent operation and synchronous playback of audio stream and video stream are realized, and users can flexibly adjust the audio without adjusting the video, reducing the technical difficulty and cost of video editing.

Benefits of technology

It realizes the flexibility of adjusting audio based on audience interaction, reduces the dependence on high-performance graphics rendering, reduces the cost and technical difficulty of video editing, and supports flexible reuse of live streams.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120223940A_ABST
    Figure CN120223940A_ABST
Patent Text Reader

Abstract

The invention provides an audio and video processing method and system, electronic equipment and a storage medium. The audio and video processing method comprises the following steps: acquiring a video stream and an audio stream; wherein the video stream comprises at least one video clip, and the audio stream comprises at least one audio clip; simultaneously playing the video stream and the audio stream; in response to an operation for the audio stream, performing corresponding processing on the audio stream; wherein the audio stream is used for playing sound, the video stream is used for playing pictures, and synchronous presentation of the sound and the pictures is realized by simultaneously playing the audio stream and the video stream; in addition, the user can perform independent operation on the audio stream without processing video clips in the video stream, and the defect that synchronous presentation of pictures and sound in live broadcast is difficult to adjust according to audience interaction in the prior art is overcome. Therefore, the user can flexibly adjust the audio stream without adjusting the video stream, and the technical difficulty and cost of video editing are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of video processing, and in particular, to an audio - video processing method, system, electronic device, and storage medium. Background Art

[0002] Currently, the implementation of unmanned live streaming mainly involves real - time rendering and mixing of video and audio into a video stream for push - streaming live broadcast according to a pre - set script. This method has three problems:

[0003] First, the mixed video stream is decoded and played with audio - video synchronization on the player, and the picture and sound are presented completely synchronously. During the live broadcast, due to the adjustment of product talk and the need for user interaction, the audio part needs to be adjusted. This is very likely to cause the audio time to become longer or shorter, and then lead to the problem of sound without picture. And how to fill in the picture has become a difficult problem.

[0004] Second, real - time audio - video rendering and mixing of the live stream are required, so the requirement for graphics computing ability is very high, resulting in the problem of inability to scale up massively.

[0005] Third, the mixed live stream is a highly customized video. If it needs to be reused in other live rooms, it needs to be decoded and edited, which not only has a high technical difficulty but also increases the cost of the live broadcast. Summary of the Invention

[0006] The technical problem to be solved by the present disclosure is to overcome the defect in the prior art that it is difficult to adjust the synchronous presentation of the picture and sound in the live broadcast according to the audience interaction, and to provide an audio - video processing method, system, electronic device, and storage medium.

[0007] The present disclosure solves the above - mentioned technical problem through the following technical solutions:

[0008] In a first aspect, an audio - video processing method is provided. The audio - video processing method includes the following steps:

[0009] Obtain a video stream and an audio stream; wherein, the video stream includes at least one video segment, and the audio stream includes at least one audio segment;

[0010] Play the video stream and the audio stream simultaneously;

[0011] In response to an operation on the audio stream, perform corresponding processing on the audio stream.

[0012] Optionally, the step of obtaining the video stream specifically includes:

[0013] In response to the import of at least one video material, erase the sound in the video material to obtain the corresponding video segment;

[0014] And / or,

[0015] The steps of obtaining the audio stream specifically include:

[0016] In response to the import of the text script, splitting the text script to obtain a plurality of text segments;

[0017] Converting the text segments into corresponding audio segments.

[0018] Optionally, the step of simultaneously playing the video stream and the audio stream specifically includes:

[0019] In response to the start time of the live broadcast room, playing the video stream and the audio stream through different players.

[0020] Optionally, the audio-visual processing method further includes:

[0021] In response to identifying that a viewer enters the live broadcast room, detecting the behavior type of the viewer;

[0022] Generating a reply voice according to the behavior type;

[0023] The step of playing the audio stream specifically includes: in response to the completion of playing of a target audio segment, playing the reply voice; wherein, the target audio segment is any one of the at least one audio segment.

[0024] Optionally, the step of erasing the sound in the video material to obtain a corresponding video segment specifically includes:

[0025] Erasing the sound in the video material;

[0026] Performing transcoding processing on the video material with the sound erased according to a first target parameter to obtain a corresponding video segment;

[0027] And / or,

[0028] The step of converting the text segments into corresponding audio segments specifically includes:

[0029] Converting the text segments into initial audio segments;

[0030] Performing transcoding processing on the initial audio segments according to a second target parameter to obtain corresponding audio segments.

[0031] Optionally, the step of simultaneously playing the video stream and the audio stream specifically includes:

[0032] In response to the completion of playing of the video stream and the non-completion of playing of the audio stream, looping to play the video stream;

[0033] In response to the completion of playing of the audio stream, stopping playing the video stream.

[0034] In a second aspect, an audio - video processing system is provided. The audio - video processing system includes: an acquisition module, a playback module, and a processing module;

[0035] The acquisition module is used to acquire a video stream and an audio stream; wherein, the video stream includes at least one video segment, and the audio stream includes at least one audio segment;

[0036] The playback module is used to play the video stream and the audio stream simultaneously;

[0037] The processing module is used to perform corresponding processing on the audio stream in response to an operation on the audio stream.

[0038] Optionally, the acquisition module is specifically configured to: in response to the import of at least one video material, erase the sound in the video material to obtain a corresponding video segment;

[0039] Optionally, the acquisition module is specifically configured to: in response to the import of a text script, split the text script to obtain a plurality of text segments; and convert the text segments into corresponding audio segments.

[0040] Optionally, the playback module is specifically configured to: in response to the start time of the live - broadcast room, play the video stream and the audio stream through different players.

[0041] Optionally, the audio - video processing system further includes:

[0042] a detection module, configured to detect the behavior type of the audience in response to identifying that the audience enters the live - broadcast room;

[0043] a voice generation module, configured to generate a reply voice according to the behavior type;

[0044] The playback module is specifically configured to: in response to the completion of playing a target audio segment, play the reply voice; wherein, the target audio segment is any one of the at least one audio segment.

[0045] Optionally, the acquisition module is further specifically configured to: erase the sound in the video material; and perform transcoding processing on the video material with the sound erased according to a first target parameter to obtain a corresponding video segment;

[0046] Optionally, the acquisition module is further specifically configured to: convert the text segments into initial audio segments; and perform transcoding processing on the initial audio segments according to a second target parameter to obtain corresponding audio segments.

[0047] Optionally, the playback module body is configured to: in response to the completion of the video stream playback and the incomplete playback of the audio stream, loop-play the video stream; and in response to the completion of the audio stream playback, stop playing the video stream.

[0048] In a third aspect, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the audio-visual processing method described in the first aspect is implemented.

[0049] In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the audio-visual processing method described in the first aspect is implemented.

[0050] Based on common general knowledge in the art, the above preferred conditions can be combined arbitrarily to obtain various preferred examples of the present disclosure.

[0051] The positive and progressive effects of the present disclosure are as follows: To solve the defect in the prior art that it is difficult to adjust the synchronous presentation of the picture and sound in a live broadcast according to the interaction of the audience, in this application, the video playback is divided into an audio stream and a video stream, where the audio stream is used to play sound and the video stream is used to play the picture. By playing the audio stream and the video stream simultaneously, the synchronous presentation of sound and picture is achieved; the separate playback of the video stream and the audio stream also enables the user to perform separate operations on the audio stream without having to process the video stream. Therefore, the user can flexibly adjust the audio without having to adjust the video, reducing the technical difficulty and cost of video editing. Description of the Drawings

[0052] Figure 1 It is a flowchart of an audio-visual processing method provided in Embodiment 1 of the present disclosure;

[0053] Figure 2 It is a specific flowchart of obtaining an audio stream in step S11 provided in Embodiment 1 of the present disclosure;

[0054] Figure 3 It is a partial flowchart of an audio-visual processing method provided in Embodiment 1 of the present disclosure;

[0055] Figure 4 It is a flowchart of step S111 provided in Embodiment 1 of the present disclosure;

[0056] Figure 5 It is a flowchart of step S113 provided in Embodiment 1 of the present disclosure;

[0057] Figure 6 It is a flowchart of step S12 provided in Embodiment 1 of the present disclosure;

[0058] Figure 7Flowchart of a method for operating an unmanned live broadcast system provided in Embodiment 1 of the present disclosure;

[0059] Figure 8 Specific design diagram of an audio - video processing method provided in Embodiment 1 of the present disclosure;

[0060] Figure 9 Schematic diagram of modules of an audio - video processing system provided in Embodiment 2 of the present disclosure;

[0061] Figure 10 Schematic diagram of modules of an electronic device provided in Embodiment 3 of the present disclosure. Specific implementation manners

[0062] The present disclosure will be further illustrated by way of examples below, but the present disclosure is not limited to the scope of the described examples.

[0063] In the embodiments of the present disclosure, prefix words such as "first" and "second" are only used to distinguish different described objects, and have no restrictive effect on the position, order, priority, quantity, content, etc. of the described objects. The use of ordinal words and other prefix words for distinguishing described objects in the embodiments of the present disclosure does not constitute a limitation on the described objects. The statements of the described objects refer to the description in the context of the claims or embodiments, and should not constitute unnecessary limitations because of the use of such prefix words. In addition, in the description of this embodiment, unless otherwise specified, the meaning of "a plurality" is two or more.

[0064] In the embodiments of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of user personal information and other processes all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0065] Embodiment 1

[0066] Figure 1 Flowchart of an audio - video processing method provided for this embodiment. The audio - video processing method includes the following steps:

[0067] S11. Obtain a video stream and an audio stream; wherein, the video stream includes at least one video segment, and the audio stream includes at least one audio segment.

[0068] S12. Play the video stream and the audio stream simultaneously.

[0069] S13. In response to an operation on the audio stream, perform corresponding processing on the audio stream.

[0070] In this embodiment, video playback is divided into an audio stream and a video stream. The audio stream is used to play sound, and the video stream is used to play images. By playing the audio stream and the video stream simultaneously, the synchronous presentation of sound and images is achieved. The separate playback of the video stream and the audio stream also enables the user to perform separate operations on the audio stream without having to process the video stream. Therefore, the user can flexibly adjust the audio without having to adjust the video, reducing the technical difficulty and cost of video editing.

[0071] In the specific implementation process, the user can perform operations such as real-time addition, editing, and deletion of audio segments during playback, thus ensuring the flexibility of content playback. At the same time, since there is no need to process the video, the dependence on high-performance graphics rendering capabilities can be reduced or eliminated, removing the cost barrier for technology application.

[0072] Specifically, the audio-visual processing method can be applied to the scenario of unmanned live streaming. Since the video stream is used to play videos and the audio stream is used to play audio, during the live streaming process, operations can be performed only on the audio segments in the audio stream according to requirements, reducing the technical difficulty and cost of unmanned live streaming video editing.

[0073] In an alternative embodiment, the step of obtaining the video stream in step S11 specifically includes:

[0074] S111. In response to the import of at least one video material, erase the sound in the video material to obtain a corresponding video segment. In this embodiment, by erasing the sound in the video material, the video segment is only used for image display and is not affected by the audio segment.

[0075] In a specific embodiment, the sound of the video can be erased by using FFmpeg (Fast Forward Moving Picture Experts Group, an open-source multimedia processing tool).

[0076] In an alternative embodiment, as Figure 2 shown, the step of obtaining the audio stream in step S11 specifically includes:

[0077] S112. In response to the import of a text script, split the text script to obtain a plurality of text segments. In this embodiment, the content of the text script is relevant to the video material. For example, if the video material is used to display a hotel scene, the text script can be the introduction of the hotel room types and the corresponding preferential prices.

[0078] In a specific example, if the purpose of a certain live stream is to sell products in Hotel A, including calendar rooms, pre-sale rooms, meals, etc. in Hotel A, therefore, in the actual live stream, the video material only needs to include at least one of the rooms, facilities, meals, etc. in Hotel A, and does not need to correspond one by one to the products explained in the audio segments in the audio stream, and the audio segments are used to explain any of the above products in Hotel A. If the duration of the video material is less than the total duration of the audio segments in the audio stream, the video material is played repeatedly until it aligns with the total duration of the audio segments.

[0079] S113. Convert the text segment into a corresponding audio segment.

[0080] In a specific example, the text script can be reasonably split by a large language model (LLM, Large Language Model) first. For example, it can be split into segments of 20 - 50 words each, and then these segments are converted into ordered audio segments through a text-to-speech (TTS) tool.

[0081] In an optional implementation manner, step S12 specifically includes:

[0082] In response to the start time of the live broadcast room, play the video stream and the audio stream through different players. In this implementation manner, a video player can be used to play the video stream, and an audio player can be used to play the audio stream. In the specific implementation process, both the video player and the audio player can only start playing after they can both pull content. After one of them has no content to play, they both end the play at the same time, so as to ensure that the picture and the sound can appear and end synchronously.

[0083] In an optional implementation manner, as Figure 3 shown, the audio - video processing method further includes:

[0084] S14. In response to identifying that a viewer enters the live broadcast room, detect the behavior type of the viewer.

[0085] S15. Generate a reply voice according to the behavior type.

[0086] In this implementation manner, when it is detected that a viewer enters the live broadcast room or has an interactive behavior, such as consulting, liking, or purchasing a product, corresponding reply content can be generated in real time according to the behavior type of the viewer, and then converted into a reply voice.

[0087] The step of playing the audio stream in step S12 specifically includes: S121. In response to the completion of playing a target audio segment, play the reply voice; where the target audio segment is any one of the at least one audio segment.

[0088] In this embodiment, the target audio segment may be the currently playing audio segment. After the currently playing audio segment finishes playing, the generated reply voice is played, thereby realizing interaction with the audience during the unmanned live broadcast process.

[0089] In other embodiments, the target audio segment may be the first audio segment after the currently playing audio segment. In response to detecting that the relevance of the user's behavior type to the first audio segment is greater than that to other audio segments, after the first audio segment finishes playing, the generated reply voice is played, thereby enhancing the experience of audience interaction.

[0090] In an alternative embodiment, the audio playlist includes all audio segments in the audio stream, and the audio-video processing method further includes:

[0091] S16. Set the reply voice between the target audio segment and the next audio segment in the audio playlist. In this embodiment, if the target audio segment is the currently playing audio segment, the reply voice is inserted before the nearest audio segment after the currently playing audio segment in the audio playlist, and the reply voice can be played immediately after the currently playing segment ends, ensuring real-time voice interaction with the user.

[0092] S17. In response to the completion of the playback of the reply voice, delete the reply voice in the audio playlist.

[0093] In this embodiment, the reply voice can be deleted after being played and will not appear again in subsequent live broadcast cycles. Because the video screen and audio are split, no modification is required for the video screen, and only the audio part is added, modified, and deleted. These operations do not rely on high-performance graphics computing capabilities. Specifically, in a tourism scenario, since the products in the tourism scenario hardly change, the video materials and live broadcast scripts (i.e., text scripts) generated by the same merchant hardly change after being generated, and the text script of the interaction part with the user is immediately deleted after being played. Therefore, the video playlist and audio playlist can be reused, and the merchant only needs to set a cycle on the system and perform automated cyclic live broadcasts according to the cycle, and the material reuse rate can reach 100%.

[0094] In an alternative embodiment, as Figure 4 shown, step S111 specifically includes:

[0095] S1111. Erase the sound in the video material.

[0096] S1112. Transcode the video material with the sound erased according to the first target parameter to obtain the corresponding video segment;

[0097] In this embodiment, the video material can be multiple segments or one segment. Erase the sound of all video materials, and then transcode the video materials with the sound erased to the first target parameters, where the first target parameters can include target parameters such as a set frame rate, bit rate, resolution, etc. In a specific embodiment, a live video playlist can also be created according to the transcoded video segments and stored in the database.

[0098] In an alternative embodiment, as Figure 5 shown, step S113 specifically includes:

[0099] S1131. Convert the text segment into an initial audio segment. In this embodiment, the text segment undergoes natural language processing. To convert the text segment into an initial audio segment, specifically, each initial audio segment can be played for about 10 seconds.

[0100] S1132. Perform transcoding processing on the initial audio segment according to the second target parameters to obtain the corresponding audio segment.

[0101] In this embodiment, the initial audio segment can be transcoded to the second target parameters through the FFmpeg tool, where the second target parameters can include a set sampling rate and encapsulation format. In a specific embodiment, an audio playlist can also be created according to the last transcoded audio segment and stored in the database.

[0102] In an alternative embodiment, as Figure 6 shown, step S12 specifically includes:

[0103] S122. In response to the video stream playback being completed and the audio stream not being completed, loop play the video stream;

[0104] S123. In response to the audio stream playback being completed, stop playing the video stream. In this embodiment, the video segments in the video stream are played sequentially, and the audio segments in the audio stream are played sequentially; by stopping the playback of the video segments in the video stream when all the audio segments in the audio stream are played sequentially, synchronous playback of audio and video is achieved.

[0105] It should be noted that in practical applications, if the duration of the audio stream is greater than the duration of the video stream, step S122 is executed first and then step S123; if the duration of the video stream is greater than the duration of the audio stream, then step S123 is executed first and then step S122; if the durations of the video stream and the audio stream are the same, the playback of both the audio stream and the video stream is stopped simultaneously.

[0106] In a specific embodiment, Figure 7 is a flowchart of a method for operating an unattended live broadcast system:

[0107] First, material preparation:

[0108] In response to the user uploading materials (including video materials and text scripts) from the host app, the video materials are processed, and the text scripts are processed, split, and converted into voice (i.e., audio clips). Specifically, the unmanned live streaming system erases the sound of all video materials, and the video materials can be multiple segments or one segment; the text script is the content that needs to be spoken through voice during the live stream. The unmanned live streaming system will process the text script through natural language processing, split it into smaller segments, and then convert these segments into voice. Each voice segment can be played for about 10 seconds.

[0109] Create a video playlist based on the processed video materials (i.e., video clips with erased sound), and at the same time create an audio playlist based on the audio clips; the video playlist and the audio playlist constitute the live stream materials and are stored in the database.

[0110] Secondly, automatic live streaming:

[0111] In response to the user's settings, after reaching the preset time, automatic live streaming is carried out. After the live stream starts, the video playlist and the audio playlist are pushed live synchronously, forming a video stream and an audio stream respectively, ensuring that the video and audio can be started and stopped simultaneously. Specifically, the unmanned live streaming system automatically starts the live stream according to the start time set by the user when creating the live stream. When starting the live stream, it is necessary to first query the video playlist and the audio playlist generated in the material preparation stage from the database, and then use the FFmpeg tool to push them to the live streaming platform simultaneously. The viewer app can start the video player and the audio player respectively through the pull stream address of the live streaming platform to pull and play the stream. The process of pulling the stream needs to coordinate the synchronization of the two players. Only after both players can pull data will they start synchronous playback. If any one of the players fails to pull the stream or ends, the other will also end accordingly, so as to achieve simultaneous presentation and simultaneous end of the sound and the picture.

[0112] In addition, the method for operating the unmanned live streaming system further includes a method for realizing intelligent interaction with the audience. Figure 9 It is a design diagram of a specific audio-video processing method. Among them, the video playlist includes segment 1 to segment n, and the audio playlist includes s1 to sn, as well as t1 and t2. Multiple s segments form a complete product explanation, which is the fixed part of the audio playlist; s1 to s3 are the explanations of product 1, and s4 and s5 are the explanations of product 2. The t segment is a reply voice, which is the temporary part of the audio playlist. The t segment can be inserted between any two other segments and is deleted after being played. Figure 9The current playback point is s4. At this time, the response voice t2 generated according to the user's behavior type is inserted between s4 and s5 to wait for playback. The response voice t1 has been played and needs to be deleted.

[0113] In this embodiment, the audience enters the live broadcast room through the audience-side APP to watch the live broadcast, and may generate behaviors such as consulting, liking, giving gifts, and purchasing goods in the live broadcast room. These behaviors are submitted to the unmanned live broadcast system through the audience-side APP. The unmanned live broadcast system uses a large language model to generate a response script in real time according to the type and content of the audience's behavior, and then converts the script into a response voice through a text-to-speech tool. After transcoding to the same sampling rate and encapsulation format as the audio playlist, it is inserted after the currently playing audio segment. After the currently playing audio segment is played, the response voice can be played. Specifically, if the audio segments of the audio playlist are all about 10 seconds, the response voice should also be played within 10 seconds, thus realizing real-time voice interaction comparable to that of a live human broadcast.

[0114] Therefore, when the user sets the recent live broadcast plan, such as live broadcasting from 10 to 22 o'clock every day, the video playlist and audio playlist generated after the first live broadcast can be fully reused in subsequent live broadcasts, and the response voices inserted for interaction with the user will be deleted from the audio playlist after being played. Only the content in the original audio playlist will be played during the next rebroadcast, so as to be able to interact with the current audience in real time and intelligently again.

[0115] Embodiment 2

[0116] Corresponding to the foregoing embodiment 1 of the audio-video processing method, the present disclosure also provides an embodiment of an audio-video processing system.

[0117] Figure 10 A module schematic diagram of an audio-video processing system provided for this embodiment, the audio-video processing system 20 includes: an acquisition module 201, a playback module 202, and a processing module 203;

[0118] The acquisition module is used to acquire a video stream and an audio stream; wherein, the video stream includes at least one video segment, and the audio stream includes at least one audio segment;

[0119] The playback module is used to play the video stream and the audio stream simultaneously;

[0120] The processing module is used to perform corresponding processing on the audio stream in response to an operation on the audio stream.

[0121] In this embodiment, video playback is divided into an audio stream and a video stream. The audio stream is used to play sound, and the video stream is used to play images. By playing the audio stream and the video stream simultaneously, the synchronous presentation of sound and images is achieved. The separate playback of video and audio also enables users to perform separate operations on the audio without having to process the video segments in the video playlist. Therefore, users can flexibly adjust the audio without having to adjust the video, reducing the technical difficulty and cost of video editing.

[0122] For the system embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment. The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present disclosure.

[0123] Embodiment 3

[0124] Figure 10 The following is a schematic structural diagram of an electronic device shown in an exemplary embodiment of the present disclosure. The electronic device includes a memory, a processor, and a computer program stored on the memory and configured to run on the processor. When the processor executes the computer program, the audio-visual processing method described in Embodiment 1 above is implemented. Figure 10 The shown electronic device 30 is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.

[0125] As Figure 10 shown, the electronic device 30 may be presented in the form of a general-purpose computing device. For example, it may be a server device. The components of the electronic device 30 may include, but are not limited to: at least one of the above-mentioned processors 31, at least one of the above-mentioned memories 32, and a bus 33 connecting different system components (including the memory 32 and the processor 31).

[0126] The bus 33 includes a data bus, an address bus, and a control bus.

[0127] The memory 32 may include volatile memory, such as a random access memory (RAM) 321 and / or a cache memory 322, and may further include a read-only memory (ROM) 323.

[0128] The memory 32 may also include a program tool 325 (or utility) having a set (at least one) of program modules 324. Such program modules 324 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.

[0129] The processor 31 executes various functional applications and data processing by running computer programs stored in the memory 32, such as the audio-visual processing method provided in the above-mentioned Embodiment 1.

[0130] The electronic device 30 can also communicate with one or more external devices 34 (such as a keyboard, a pointing device, etc.). Such communication can be carried out through the input / output (I / O) interface 35. And, the electronic device 30 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 36. As shown in the figure, the network adapter 36 communicates with other modules of the electronic device 30 through the bus 33. It should be understood that although Figure 10 not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 30, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID (Redundant Array of Independent Disks) systems, magnetic tape drives, and data backup storage systems, etc.

[0131] It should be noted that although several units / modules or sub-units / modules of the electronic device are mentioned in the above detailed description, such a division is merely exemplary and not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more of the above-mentioned units / modules can be embodied in one unit / module. Conversely, the features and functions of one unit / module described above can be further divided and embodied by multiple units / modules.

[0132] Embodiment 4

[0133] The embodiments of the present disclosure also provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the audio-visual processing method provided in the above-mentioned Embodiment 1 is implemented.

[0134] Among them, the more specific computer-readable storage medium that can be adopted includes, but is not limited to: a portable disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0135] Embodiment 5

[0136] Embodiments of the present disclosure also provide a computer program product, including a computer program, which implements the audio-video processing method described in Embodiment 1 above when executed by a processor.

[0137] Among them, the program code for executing the computer program product of the present disclosure can be written in any combination of one or more programming languages. The program code can be executed entirely on the user device, partially on the user device, executed as an independent software package, partially on the user device and partially on a remote device, or entirely on a remote device.

[0138] Although the specific embodiments of the present disclosure have been described above, those skilled in the art should understand that this is only an example. The protection scope of the present disclosure is defined by the appended claims. Without departing from the principle and essence of the present disclosure, those skilled in the art can make various changes or modifications to these embodiments, but these changes and modifications all fall within the protection scope of the present disclosure.

Claims

1. An audio and video processing method, characterized in that: The audio and video processing method comprises the following steps: Acquire a video stream and an audio stream; wherein the video stream includes at least one video segment, and the audio stream includes at least one audio segment; Play the video stream and the audio stream simultaneously; In response to the operation on the audio stream, the audio stream is processed accordingly.

2. The audio and video processing method according to claim 1, characterized in that: The steps to obtain the video stream include: In response to the import of at least one video material, erasing the sound in the video material to obtain a corresponding video clip; and / or, The steps to obtain the audio stream include: In response to the import of the text script, split the text script to obtain multiple text segments; The text segment is converted into a corresponding audio segment.

3. The audio and video processing method according to claim 1, characterized in that: The step of simultaneously playing the video stream and the audio stream specifically includes: In response to the start time of the live broadcast room, the video stream and the audio stream are played through different players.

4. The audio and video processing method according to claim 3, characterized in that: The audio and video processing method also includes: In response to identifying that a viewer has entered the live broadcast room, detecting a behavior type of the viewer; Generate a reply voice according to the behavior type; The step of playing the audio stream specifically includes: in response to completion of playing the target audio segment, playing the reply voice; wherein the target audio segment is any audio segment of the at least one audio segment.

5. The audio and video processing method according to claim 2, characterized in that: The step of erasing the sound in the video material to obtain the corresponding video clip specifically includes: erasing sound from said video footage; Transcoding the video material from which the sound has been erased according to the first target parameter to obtain a corresponding video clip; and / or, The step of converting the text segment into the corresponding audio segment specifically includes: converting the text segment into an initial audio segment; The initial audio segment is transcoded according to the second target parameter to obtain a corresponding audio segment.

6. The audio and video processing method according to claim 1, characterized in that: The step of simultaneously playing the video stream and the audio stream specifically includes: In response to the video stream being played completely but the audio stream not being played completely, playing the video stream in a loop; In response to completion of playing of the audio stream, stop playing the video stream.

7. An audio and video processing system, characterized in that: The audio and video processing system comprises: an acquisition module, a playback module and a processing module; The acquisition module is used to acquire a video stream and an audio stream; wherein the video stream includes at least one video segment, and the audio stream includes at least one audio segment; The playing module is used to play the video stream and the audio stream simultaneously; The processing module is used for performing corresponding processing on the audio stream in response to an operation on the audio stream.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and used to run on the processor, characterized in that: When the processor executes the computer program, the audio and video processing method according to any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the audio and video processing method according to any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the audio and video processing method according to any one of claims 1 to 6 is implemented.