System and method for fade-in and fade-out processing of artificial intelligence generated content in real-time communication architecture
By introducing a fade-in/fade-out control module into the AI-generated content system, the discontinuity problem during audio switching or interruption was solved, resulting in a more natural user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2026-04-14
AI Technical Summary
In existing technologies, audio content generated by artificial intelligence lacks fade-in/fade-out processing when switching or interrupting, resulting in discontinuity and affecting user experience.
By introducing a fade-in/fade-out control module into the AI-generated content system, the gradual change in audio level is controlled to ensure continuity during audio switching or interruption.
It improves the continuity and naturalness of the user experience, reduces audio discontinuity, and provides a more pleasant auditory experience.
Smart Images

Figure CN121858059A_ABST
Abstract
Description
Cross-references to related applications
[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 706,256, filed October 11, 2024, and U.S. Patent Application No. 18 / 970,121, filed December 5, 2024. Technical Field
[0002] This invention generally relates to the field of audio control for generated content, and more specifically, to methods, computer programs, and systems for fade-in / fade-out processing of speech and other audio generated by artificial intelligence (AI) systems. The audio level is adjusted according to the stage of the speech and the presence or absence of interruptions, thereby achieving a more natural and enjoyable user experience. Background Technology
[0003] Recently, Artificial Intelligence Generated Content (AIGC) has gained increasing attention and is developing rapidly. AIGC is a technology based on machine learning and natural language processing that can automatically generate various types of content, including text, images, and audio. This content can be news articles, novels, images, music, or even software code. AIGC systems learn to mimic human creativity by analyzing large amounts of data and text, thereby generating high-quality content.
[0004] Human-computer communication, as a crucial application of artificial intelligence and computer networks, has gained widespread popularity and attention across various fields. This is thanks to advancements in AI and natural language processing technologies, which enable such systems to understand and produce human-like responses. As businesses and organizations continuously enhance customer engagement, chatbots and virtual assistants have become essential tools for providing instant support and personalized experiences. For example, Apple's intelligent assistant Siri is widely used on Apple devices, allowing users to easily access information or engage in human-like conversations.
[0005] During human-computer communication, when an AI agent begins to speak, the voice amplitude can suddenly jump from zero to a very large value, causing discontinuity. Furthermore, when a user interrupts the AI agent, the agent must stop speaking, listen to what the user has said, and then respond. Currently, providers of AI-generated content (AIGC) services simply start and stop playing AIGC-generated speech without any transition, resulting in incoherent speech and a poor audiovisual experience for users.
[0006] Therefore, providing users with a pleasant AIGC experience and reducing the discontinuity of audio content is of great value. Accordingly, this invention proposes a control system and method for fading in and out generated audio in a real-time communication system. Summary of the Invention
[0007] This system and method involve audio level control technology, primarily referring to fade-in and fade-out control of AIGC audio. This system and method can reduce discontinuities between AIGC-generated speech or other audio elements and other audio.
[0008] In some implementations, the audio stream is first received from a content generator. The content generator can be a cloud-based AI-generated content (AIGC) system. The user's local device then begins playing the audio stream. Next, an interruption event is received. An interruption occurs when the user speaks, the voice changes, or the audio stream's content violates a policy, etc. The audio stream then fades out within a fade-out time window. This window may last from 50 milliseconds to 1 second or longer.
[0009] Finally, when the amplitude of the audio stream falls below a certain threshold, a stop response flag is generated and fed back to the content generator. This threshold can be zero or below the amplitude perceived by human hearing. Fade-in and fade-out refer to changes in amplitude, which can be linear, exponential, logarithmic, or follow an S-curve. Fade-in processing can also be performed at the beginning of the audio stream. Sometimes the audio stream refers to speech. Fade-out processing can also be performed at the end of the audio stream if it has not been interrupted. In some embodiments, the audio stream is generated upon user request.
[0010] It should be noted that the various functions of the present invention described above can be used individually or in combination. These functions and other features of the present invention will be described in detail below with reference to the accompanying drawings. Attached Figure Description
[0011] To more clearly illustrate the present invention, some embodiments of the present invention will be described below with reference to the accompanying drawings, wherein:
[0012] Figure 1 This is an example block diagram of a system for transmitting artificial intelligence-generated content (AIGC) to multiple users, drawn according to an embodiment of the present invention;
[0013] Figure 2A This is an example block diagram of a system drawn according to an embodiment of the present invention, illustrating a system for generating audio with fade-in and fade-out processing;
[0014] Figure 2B This is an example block diagram of an audio encoding and transmission module drawn according to an embodiment of the present invention;
[0015] Figure 2C This is an example block diagram of a fade-in / fade-out control module drawn according to an embodiment of the present invention;
[0016] Figure 3This is a flowchart illustrating an example process for transmitting AIGC with fade-in / fade-out control, drawn according to an embodiment of the present invention; and
[0017] Figure 4A and 4B This is a schematic diagram of a computer system capable of implementing fade-in and fade-out processing according to an embodiment of the present invention. Detailed Implementation
[0018] The present invention will be described in detail below with reference to several embodiments shown in the accompanying drawings, wherein some technical details will be described to facilitate a comprehensive understanding of the embodiments of the present invention. However, those skilled in the art can also implement the embodiments of the present invention without these specific details. On the other hand, well-known technical steps and / or structures will not be described in detail to avoid unnecessarily obscuring the present invention. The functionality and advantages of the embodiments can be better understood with reference to the following drawings and descriptions.
[0019] The accompanying drawings and descriptions below will help to understand the various aspects, functions, and advantages of exemplary embodiments of the present invention. Those skilled in the art should understand that the embodiments of the present invention described herein are presented by way of example only and are not intended to limit. All functions disclosed herein can be replaced by other functions having the same or similar purpose unless explicitly stated otherwise. Therefore, other modified embodiments also fall within the scope of the invention and its equivalents as defined herein. Therefore, the imperative and / or sequential terms used in this article, such as “will,” “will not,” “should,” “should not,” “must,” “must not,” “first,” “initially,” “next,” “subsequently,” “before,” “after,” “finally,” and “end,” etc., are not intended to limit the scope of the invention, as the embodiments disclosed herein are merely illustrative examples.
[0020] This invention relates to systems and methods for generating, transmitting, and controlling fade-in and fade-out of audio elements. In some embodiments, this disclosure will focus particularly on content generated by artificial intelligence systems. Such AI-generated content (AIGC) is a particularly prominent use case and is not intended to limit the scope of this disclosure. Audio elements generated by other means independent of artificial intelligence (AI) systems can also benefit from such systems and methods. Therefore, while this disclosure will focus on AIGC, other generated audio may also be interrupted by the user.
[0021] Figure 1This is an example diagram of a system for providing AIGC content to one or more users 140a-n, with the system as a whole represented by 100. In this architecture, one or more AIGC servers 110 receive input from users 140a-n through their respective terminal devices. These terminal devices may include smartphones, smart speakers, computer systems, etc. Regardless of the form of the terminal device, they each include an audio local interface 130a-n. These local interfaces can receive input from users 140a-n and feed the input back to the AIGC server 110 via network 120. In most cases, this network consists of a cellular network and / or the Internet. However, the network also includes any wide area network (WAN) architecture, including a private WAN or a private local area network (LAN) combined with a private or public WAN.
[0022] The use cases described in this paper are based on cloud-based systems. However, with the improvement of computing power, local interfaces 130a-n may contain sufficient computing and data resources to provide AIGC without the need for a cloud-connected architecture. Therefore, although this system and method will focus on cloud-derived systems, the fade-in and fade-out control methods disclosed in this paper are equally effective if the AI content is generated locally.
[0023] In an AIGC system based on Real-Time Communication (RTC), the user's audio is first encoded and transmitted via network 120 to the AIGC server 110. After system processing, the response audio is streamed back to the user 140a-n. If the user 140a-n interrupts the AI agent, the system should continue sending the generated audio in a "fade-out" manner for a period of time. This means the system avoids immediately stopping sending responses to the user, but instead continues sending fade-out responses for a certain period, gradually reducing their energy to zero, thus maintaining continuity and enhancing the user experience. Furthermore, the encoding, transmission, decoding, and playback modules continue to operate until the system completely stops sending audio. Once the fade-out process is complete, the AIGC system receives a stop response flag, and audio encoding and transmission cease.
[0024] Figure 2A This is an example block diagram illustrating an AIGC application with the aforementioned fade-in / fade-out control in an RTC system. The system, designated 200, primarily comprises four components: a front-end processor 210, an AIGC system 110, fade-in / fade-out control 230, and an audio encoding and transmission pipeline 220.
[0025] The front-end processor 210 includes a buffer 211 for buffering input frames, a window 213 for caching input frames, and a 3A system 215. The 3A system 215 is an umbrella-shaped component capable of performing acoustic echo suppression, acoustic noise reduction, and automatic gain control. By performing front-end processing on the input speech frames, interference can be removed as early as possible.
[0026] Details of the audio encoding and transmission module 220 are as follows: Figure 2B As shown, this component includes audio encoding 221, transmission 223, and audio decoding 225, illustrating a typical audio transmission process. User speech is processed by the audio encoding and transmission system 220 and transmitted to the AIGC server 110. The AIGC system 110 processes and understands the user's speech and generates a response speech to send to the user's terminal. The user can interrupt the audio at any time, at which point the AIGC system 110 will receive a stop response flag.
[0027] Next, the voice (or other generated audio) will be transmitted to the fade-in / fade-out control module 230. Figure 2C The fade-in / fade-out control module is explained in more detail. Within this module, there are three paths: fade-in path 231, fade-out path 233, and other paths. Other paths are designated as "pass" and do not require any modification. Fade-in path 231 and fade-out module 233 are used to avoid discontinuities in the audio, thereby enhancing the user experience.
[0028] The fade-in / fade-out control module 230 has four operating scenarios. The first is when the AIGC system is ready to generate speech. In this case, at some point, the amplitude of the generated speech will suddenly jump from zero to a large value, resulting in a strong discontinuity and an unpleasant auditory experience. To mitigate this effect, fade-in technology can be used to gradually increase the speech amplitude back to its original value. The second scenario is when the user interrupts the AI agent while it is speaking. All providers of AIGC services simply stop playing the generated speech without any transition. The user experiences the sound suddenly and unnaturally disappearing. This also causes the sound to jump from a large value to zero. Fade-out technology can eliminate this effect by gradually reducing the speech amplitude to zero. It is important to note that fade-in / fade-out technology is not only needed when the user interrupts the AI agent, but also when the user switches the AI agent's voice to someone else's voice, or when the system detects that the generated content violates security or compliance standards. In other words, fade-in / fade-out technology can be widely applied in various situations. The third scenario involves applying fade-in / fade-out technology to gradually reduce the amplitude of the speech to zero when the AIGC system finishes speaking, thus eliminating any inconsistencies in the sound. The final scenario occurs when the AIGC system is still generating speech normally; in this case, the fade-in / fade-out control system performs no operation, and the generated speech simply bypasses the system. In the overall flowchart, there are two fade-in / fade-out control modules: one after the AIGC module and the other before the audio playback module. The latter is essential because it ensures that the generated audio can indeed fade in or out.
[0029] Looking at it now Figure 3 , Figure 3 This diagram illustrates an example flowchart of the fade-in / fade-out control process in AIGC real-time communication, represented by the number 300. First, the user operates using a smart speaker, smartphone, or other interface device. Typically, the user utters a trigger word, initiating voice recording. The user can ask questions or make requests to the AI system. The user's audio is processed by a front-end processor, and the recording is buffered. It then undergoes windowing processing followed by a series of acoustic processing steps, including echo cancellation, noise reduction, and automatic gain control. Next, the generated output frames are encoded, transmitted, and decoded on the AIGC server. Due to the large data and computational demands, local processing is not feasible; therefore, AIGC servers are typically cloud-based, and transmission usually occurs over the internet or other networks.
[0030] The AIGC server uses deep learning models to generate responses to user requests. These responses may include audio and other outputs (e.g., video, links and other web content, images, etc.). In some embodiments, the audio portion of the generated output may require initial fade-in / fade-out processing. It is then encoded, transmitted, and decoded at the local device and returned to the user's local device—the device from which the user receives the speech or other audio content, as shown in 310. The initial speech or other audio initially fades in from zero until it reaches the maximum amplitude in use, as shown in 320. Typically, this maximum amplitude can be adjusted by the user using volume controls. The fade-in may occur relatively quickly. In some embodiments, the fade-in may take anywhere from 50 milliseconds to a full second. In some embodiments, the fade-in may take approximately 50 milliseconds, 100 milliseconds, 200 milliseconds, 300 milliseconds, 400 milliseconds, 500 milliseconds, 600 milliseconds, 700 milliseconds, 800 milliseconds, or 900 milliseconds. "Approximately" typically refers to a deviation of up to about 20% from the stated value. In some embodiments, the fade-in may be a linear offset of amplitude within the fade-in time window. In other embodiments, the fade-in may be logarithmic, exponential, or follow an S-curve.
[0031] The audio will continue playing until interrupted by the user or other interruption event, as shown in 330. If an interruption occurs, the audio (or other audio) will fade out in the opposite manner to the fade-in, as shown in 340. The fade-out duration can be the same as the fade-in duration or a different duration. Typically, the fade-out time ranges from 50 milliseconds to 1 second. In some embodiments, the fade-in may require approximately 50 milliseconds, 100 milliseconds, 200 milliseconds, 300 milliseconds, 400 milliseconds, 500 milliseconds, 600 milliseconds, 700 milliseconds, 800 milliseconds, or 900 milliseconds.
[0032] When the amplitude of the speech approaches zero (or a volume level imperceptible to the human ear), a stop response flag is generated and sent to the AIGC server to stop content generation, as shown in 350. This concludes the example process.
[0033] However, if no interruption is encountered, the system will continue playing the content provided by the AIGC server (as shown in 360) until the content ends. After the content ends, the system can fade out the last part of the audio (as shown in 370). This audio fade-out method is basically similar to the method used when an interruption event occurs. The example process ends here.
[0034] The preceding text explained the systems and methods for fade-in and fade-out processing of AI-generated content. Now, let's look at the devices used to perform these functions in real time. For ease of discussion, Figure 4A and 4BA computer system is shown for implementing embodiments of the present invention, the computer system being referred to as system 400. Figure 4A This diagram illustrates one possible physical form of the computer system 400. Of course, the computer system 400 can have various physical forms, ranging from printed circuit boards, integrated circuits, and small handheld devices to large supercomputers. The computer system 400 may include a monitor 402, a display 404, a stand 406, a blade server including one or more storage drives 408, a keyboard 410, and a mouse 412, etc. Medium 414 is a computer-readable medium used for transmitting data to the computer system 400. Figure 4B This is an example block diagram of computer system 400. System bus 420 is connected to a variety of subsystems. Processor 422 (also called central processing unit or CPU) is adapted to storage devices (including memory 424). Memory 424 includes random access memory (RAM) and read-only memory (ROM). As is well known to those skilled in the art, ROM is used for unidirectional transfer of data and instructions to the CPU, while RAM is typically used for bidirectional transfer of data and instructions. Both types of memory can include any suitable form of computer-readable medium described below. Fixed medium 426 can also be bidirectionally adapted to processor 422 to provide additional data storage capacity and can also include any computer-readable medium described below. Fixed medium 426 is an auxiliary storage medium (e.g., a hard disk) for storing programs, data, etc., and typically operates slower than main memory. It should be noted that, where appropriate, information stored in fixed medium 426 can be incorporated into memory 424 in a standard manner as virtual memory. Removable medium 414 can take the form of any computer-readable medium described below.
[0035] Processor 422 is also compatible with various input / output devices, such as display 404, keyboard 410, mouse 412, and speaker 430. Generally, input / output devices can be any of the following: video display, trackball, mouse, keyboard, microphone, touch-sensitive display, sensor card reader, tape or paper tape reader, tablet computer, stylus, voice or handwriting recognition device, biometric reader, motion sensor, EEG reader, or other computer, etc. Processor 422 can also interconnect with another computer or telecommunications network using network interface 440. It is conceivable that processor 422 uses network interface 440 to receive information from the network, or can output information to the network during the execution of the above-described fade-in / fade-out control method. Furthermore, the method embodiments of the present invention can run independently on processor 422, or can also cooperate with a remote CPU sharing partial processing via a network such as the Internet.
[0036] Software is typically stored in non-volatile memory and / or drive units. In fact, for large programs, it may not even be possible to store the entire program in a single memory. However, it is understood that during software execution, it can be moved to a computer-readable location suitable for processing, which, for ease of explanation, is referred to as memory, if necessary. Even when the software is moved to memory for execution, the processor typically uses hardware registers and caches to store software-related values, ideally for speeding up execution. In this document, when a software program is stated to be “implemented in a computer-readable medium,” it is assumed that the software program is stored in any known or convenient location (non-volatile memory or hardware registers, etc.). A processor is considered “configured to run a program” when at least one value related to the program is stored in a processor-readable register.
[0037] During operation, computer system 400 can be controlled by operating system software (such as a media operating system) that includes a file management system. For example, Microsoft Corporation in Redmond, Washington. An operating system series and its associated file management system is essentially operating system software with a file management system. For example, the Linux operating system and its associated file management system are also operating system software with a file management system. File management systems are typically stored in non-volatile memory and / or drive units, allowing the processor to perform various operations required by the operating system, input and output data, and store data in memory, including storing files in non-volatile memory and / or drive units.
[0038] Some parts described in detail in this document may be presented in the form of algorithms and symbolic representations of operations on data bits within computer memory. These algorithmic descriptions and representations are the means by which those skilled in the art of data processing most effectively communicate their work to others skilled in the art. As defined herein, an algorithm is a series of self-consistent operational steps designed to achieve a desired result. These operations require physical manipulation of physical quantities. Typically, these physical quantities are represented as electrical or magnetic signals that can be stored, transmitted, combined, compared, or otherwise manipulated, but this is not always necessary. For common use and ease of interpretation, these signals are often referred to as bits, values, elements, symbols, characters, items, numbers, etc.
[0039] The algorithms and representations described herein are inherently independent of any particular computer or device. The procedural methods described herein can be implemented using various general-purpose systems, or specific embodiments can be designed for dedicated devices to run, thereby achieving greater convenience. The architectures required for these systems will be detailed below. Furthermore, the techniques described herein are not referred to in any particular programming language, and therefore various embodiments can be implemented using various programming languages.
[0040] In other embodiments, the machine may operate as a standalone device or may be interconnected (e.g., networked) with other machines. In a networked deployment, the machine may operate as a server or client machine in a client-server network environment, or as a peer machine in a peer-to-peer (or distributed) network environment.
[0041] The aforementioned machines can be server computers, client computers, personal computers (PCs), tablets, laptops, set-top boxes (STBs), personal digital assistants (PDAs), cellular phones, Apple iPhones, Blackberry iPhones, glasses with processors, headsets with processors, virtual reality devices, processors, distributed processors working together, telephones, network devices, network routers, switches or bridges, or any machine capable of running a set of instructions (whether sequentially or otherwise) that specify the operations that the machine needs to perform.
[0042] Although machine-readable media or machine-readable storage media are shown as a single medium in exemplary embodiments, "machine-readable media" and "machine-readable storage media" should be understood to include a single medium or multiple media (e.g., a centralized or distributed database and / or associated caches and servers) that store one or more sets of instructions. "Machine-readable media" and "machine-readable storage media" should also be understood to include any medium capable of storing, encoding, or carrying a set of instructions for machine execution and enabling the machine to operate any one or more methods of currently disclosed technologies and innovations.
[0043] Generally, routines that run to implement embodiments of the present invention may be implemented as part of an operating system or a particular application, component, program, object, module, or sequence of instructions (referred to as a "computer program"). A computer program typically includes one or more instructions written at different times in various memories and storage devices within a computer (or distributed across computers), and when one or more processing units or processors within (or across computers) read and execute these instructions, the computer can perform operations to implement the functions disclosed in the present invention.
[0044] Furthermore, while the embodiments described herein are set in the context of fully functional computers and computer systems, those skilled in the art will understand that various embodiments can be distributed in various forms of program products, and the disclosure of this invention applies equally to any particular type of machine or computer-readable medium in which the distribution is actually implemented.
[0045] The present invention has been described through several embodiments, but variations, modifications, substitutions, and equivalents that fall within the scope of the invention are still possible. While subsection headings have been used herein to aid in the description of the invention, these headings are illustrative only and are not intended to limit the scope of the invention. It should also be noted that many alternatives are available for implementing the methods and apparatus of the invention. Therefore, the appended claims should be considered to encompass all such variations, modifications, substitutions, and equivalents within the spirit and scope of the invention.
Claims
1. A computer program method for fading in and out generated audio in a real-time communication system, comprising: Receive audio streams from the content generator; Start playing the audio stream; An interruption event in the audio stream reception; Fade out the audio stream within the fade-in / fade-out time window; A stop response flag is generated when the amplitude of the audio stream falls below a threshold. as well as Send the stop response flag to the content generator.
2. The computer program method according to claim 1, wherein the content generator is an artificial intelligence-generated content (AIGC) system.
3. The computer program method according to claim 1, wherein the threshold is zero.
4. The computer program method of claim 1, wherein the threshold is below the amplitude of human hearing.
5. The computer program method of claim 1, wherein the fade-in / fade-out time window is between 50 milliseconds and 1 second.
6. The computer program method according to claim 1, wherein the fade-in / fade-out is a change in amplitude, which may be a linear, exponential, logarithmic change, or a change conforming to an S-curve.
7. The computer program method according to claim 1 further includes performing a fade-in process at the beginning of the audio stream.
8. The computer program method according to claim 1, wherein the audio stream refers to speech.
9. The computer program method of claim 1, wherein the audio stream is generated according to a user's request.
10. The computer program method of claim 1, wherein the interruption refers to a user speaking, a voice switching, or the content of an audio stream violating one or more policies.
11. A real-time communication system for performing fade-in and fade-out processing on generated audio, comprising: An encoder system configured to receive audio streams from a content generator; The local device is configured to start playing an audio stream and receive interruption events from the audio stream; The fade-in / fade-out control module in the local device is configured to fade out the audio stream within the fade-in / fade-out time window, and generate a stop response flag when the amplitude of the audio stream falls below a threshold; and The encoder system is also configured to transmit the stop response flag to the content generator.
12. The real-time communication system according to claim 11, wherein the content generator is an artificial intelligence-generated content (AIGC) system.
13. The real-time communication system according to claim 11, wherein the threshold is zero.
14. The real-time communication system of claim 11, wherein the threshold is below the amplitude of human hearing.
15. The real-time communication system of claim 11, wherein the fade-in / fade-out time window is between 50 milliseconds and 1 second.
16. The real-time communication system according to claim 11, wherein the fade-in / fade-out is a change in amplitude, which may be a linear, exponential, logarithmic change, or a change conforming to an S-curve.
17. The real-time communication system according to claim 11 further includes a fade-in process for the beginning of the audio stream.
18. The real-time communication system according to claim 11, wherein the audio stream refers to speech.
19. The real-time communication system of claim 11, wherein the audio stream is generated according to a user's request.
20. The real-time communication system of claim 11, wherein the interruption refers to a user speaking, voice switching, or the content of an audio stream violating one or more policies.