Live broadcast voice playing method and device, equipment and storage medium

By splitting the audio data to be played into audio fragments and pushing it to the live audio channel, the problem of audio discontinuity in digital people's live broadcast is solved, and the smooth and natural live sound and flexible adjustment of audio content is achieved.

CN120017872APending Publication Date: 2025-05-16SHANGHAI XULU INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510166685.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The problem of discontinuity of audio in digital people's live broadcasts has led to the intermittent sound of live broadcasts and poor user experience.

Method used

By obtaining the audio data to be played, splitting it into multiple audio fragment data, and generating corresponding fragment index files, pushing the audio fragment data to the live audio channel based on the index file, real-time interruption and interpolation of audio content is achieved.

Benefits of technology

It solves the problem of discontinuity of audio in digital live broadcasts, and realizes the smooth and natural sound of live broadcasts, making the adjustment of audio content in live broadcast scenarios more flexible.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017872A_ABST
    Figure CN120017872A_ABST
Patent Text Reader

Abstract

The invention discloses a live broadcast voice playing method and device, equipment and a storage medium. The method comprises the following steps: acquiring to-be-played audio data corresponding to a digital live broadcast object in a live broadcast process; splitting the audio data to be played to obtain a plurality of pieces of audio fragment data and fragment index files corresponding to the audio fragment data; and pushing the multiple pieces of audio fragment data to a live broadcast audio channel based on the fragment index file, and realizing live broadcast voice playing of the digital live broadcast object based on the arrangement sequence of the multiple pieces of audio fragment data, thereby realizing real-time interruption and inter-cut of audio content, meeting the requirement of flexible content adjustment in a live broadcast scene, and improving the user experience. The problem of discontinuous audio in digital human live broadcast is solved, and the live broadcast sound is smoother and more natural.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of audio processing technology, and in particular to a live voice playback method, device, equipment and storage medium. Background Art

[0002] With the rapid development of the live broadcast industry and the continuous maturity of AI technology, a new type of live broadcast that uses digital people to replace real anchors has also emerged. In traditional live broadcasts, audio can be collected by recording equipment from real anchors. But when digital people broadcast live, the audio is MP3 generated by TTS technology. Since the audio required by the live broadcast platform must be a continuous voice stream, if discrete MP3s are directly pushed, the live broadcast sound will be intermittent, resulting in a poor user experience in terms of hearing. Summary of the invention

[0003] The present invention provides a live voice playback method, device, equipment and storage medium, which realizes the real-time interruption and insertion of audio content, meets the demand for flexible adjustment of content in live broadcast scenarios, solves the problem of audio discontinuity in digital human live broadcast, and makes the live broadcast sound smoother and more natural.

[0004] According to one aspect of the present invention, a live voice playback method is provided, the method comprising:

[0005] Obtain the audio data to be played corresponding to the digital live broadcast object during the live broadcast process;

[0006] Splitting the audio data to be played to obtain a plurality of audio fragment data and fragment index files corresponding to the audio fragment data;

[0007] Based on the fragment index file, the plurality of audio fragment data are pushed into the live audio channel, and based on the arrangement order of the plurality of audio fragment data, the live voice playback of the digital live object is realized.

[0008] According to another aspect of the present invention, a live voice playback device is provided, the device comprising:

[0009] An audio data acquisition module is used to acquire the to-be-played audio data corresponding to the digital live broadcast object during the live broadcast process;

[0010] A fragment data acquisition module, used for splitting the audio data to be played to obtain a plurality of audio fragment data and fragment index files corresponding to the audio fragment data;

[0011] The live voice playback module is used to push the multiple audio fragment data to the live audio channel based on the fragment index file, and realize the live voice playback of the digital live object based on the arrangement order of the multiple audio fragment data.

[0012] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0013] at least one processor; and

[0014] a memory communicatively connected to the at least one processor; wherein,

[0015] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the live voice playback method described in any embodiment of the present invention.

[0016] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the live voice playback method described in any embodiment of the present invention when executed.

[0017] The technical solution of the embodiment of the present invention obtains the audio data to be played corresponding to the digital live broadcast object during the live broadcast process. The audio data to be played is split and processed to obtain multiple audio fragment data and fragment index files corresponding to the audio fragment data. Based on the fragment index file, the multiple audio fragment data are pushed to the live broadcast audio channel, and based on the arrangement order of the multiple audio fragment data, the live voice playback of the digital live broadcast object is realized, which solves the problem of discontinuous audio in the live broadcast of the digital human, meets the demand for flexible adjustment of audio content in the live broadcast scene, and makes the live broadcast sound of the digital human live broadcast more smooth and natural.

[0018] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0020] Figure 1 is a flow chart of a live voice playback method provided according to Embodiment 1 of the present invention;

[0021] Figure 2 is a flow chart of a live voice playback method provided according to Embodiment 2 of the present invention;

[0022] Figure 3 is a structural diagram of a live voice playback device provided according to Embodiment 3 of the present invention;

[0023] Figure 4 It is a structural schematic diagram of an electronic device that implements the live voice playback method of an embodiment of the present invention. DETAILED DESCRIPTION

[0024] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0025] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0026] Embodiment 1

[0027] Figure 1 This is a flowchart of a live voice playback method provided in Embodiment 1 of the present invention. This embodiment is applicable to the case of configuring live voice for a digital live broadcaster. The method can be executed by a live voice playback device. The live voice playback device can be implemented in the form of hardware and / or software. The live voice playback device can be configured in an electronic device. Figure 1 As shown, the method includes:

[0028] S101, obtaining the to-be-played audio data corresponding to the digital live broadcast object during the live broadcast process.

[0029] The digital live broadcast object may refer to a virtual object created by digital technology during the live broadcast process, such as a virtual character and a virtual doll. The digital live broadcast object can simulate human expressions, movements and sounds, thereby providing users with a realistic interactive experience. The audio data to be played may refer to audio data used to generate live broadcast voice, such as discrete MP3 audio clips.

[0030] In live broadcast scenarios, digital humans can appear as anchors or guests and interact with users in real time. To achieve this function, the technical support behind digital humans is crucial, including the audio stream generation service. This generation service can convert discrete MP3 audio clips into continuous voice streams and support real-time content insertion, thus meeting the audio needs of digital humans in live broadcasts.

[0031] It should be noted that the present invention is applied to the server, which provides a live audio generation service for the caller (such as a live broadcast platform or a digital human system). Before obtaining the audio data to be played corresponding to the digital live broadcast object during the live broadcast process, the server establishes a long connection with the caller to transmit and process the audio data to be played in real time. Specifically, the server can establish a long connection with the caller based on the WebSocket protocol, and the server sends an audio acquisition request to the caller, and the caller sends the audio data to be played to the server, so as to realize real-time transmission and processing of audio content, so as to improve the efficiency and real-time performance of data transmission.

[0032] S102: Split the audio data to be played to obtain a plurality of audio fragment data and fragment index files corresponding to the audio fragment data.

[0033] The audio fragment data may refer to the audio fragments obtained by splitting the audio data to be played. The fragment index file may refer to the text index file corresponding to the audio fragment data, which can be used to arrange multiple audio fragment data in sequence so that the player can download and play them continuously on demand to form a complete audio stream.

[0034] Exemplarily, the audio fragment data is TS fragments, and the fragment index file is an M3U8 index file. It should be noted that TS fragments are multiple .ts files of only a few seconds in length into which the audio and video streams in the HLS protocol are divided. These files contain part of the audio and video streams and are arranged in sequence so that the player can download and play them continuously on demand. M3U8 is an extended version of the M3U (Moving Picture Experts Group Audio Layer 3Uniform ResourceLocator) file, which is specially optimized for the Unicode character set and adds support for UTF-8 character encoding, which enables it to better handle multilingual content.

[0035] Specifically, the server uses HLS (HTTP Live Streaming) technology to split the audio data to be played into more fine-grained audio fragment data, and generates a fragment index file corresponding to the audio fragment data.

[0036] S103. Based on the fragment index file, the plurality of audio fragment data are pushed to the live audio channel, and based on the arrangement order of the plurality of audio fragment data, the live voice playback of the digital live object is realized.

[0037] Specifically, the server pushes the audio fragment data to the live audio channel of the caller through the FFmpeg tool in the order of the fragment index file, so that the caller can realize the live voice playback of the digital live object according to the arrangement order of the multiple audio fragment data.

[0038] Exemplarily, after pushing the plurality of audio fragment data to the live audio channel, the method further includes: updating the total index file according to the fragment index file.

[0039] It should be noted that a general index file is set at the calling end. After pushing multiple audio fragment data to the live audio channel, the fragment index file needs to be appended to the general index file, which can easily manage and maintain each index file in a unified manner. For example, you can easily add, delete or update the content in the index file without having to perform separate operations on each audio fragment data, which helps to reduce management costs and improve maintenance efficiency. At the same time, it also makes the update and upgrade of the audio stream simpler and more convenient.

[0040] Exemplarily, the method further includes: when the number of unplayed audio fragment data is less than a preset number threshold, continuing to obtain the audio data to be played, and realizing the live voice playback of the digital live broadcast object.

[0041] The preset quantity threshold may be set according to actual conditions or historical experience, and the present invention does not make a specific setting for this. Preferably, the preset quantity threshold may be 1.

[0042] That is to say, when the number of unplayed audio fragment data is less than the preset threshold, or there is only one sentence left in the live voice playback, the server continues to obtain the next stage of audio data to be played from the caller in advance, and repeats the above processing steps to achieve continuous playback of the live voice of the digital live object.

[0043] Exemplarily, the method further includes: in the case where the audio data to be played is not obtained, pushing blank fragment data to the live audio channel.

[0044] That is to say, if the caller fails to provide new audio data to be played in time, the server will push blank audio fragments to the audio channel to ensure that the audio stream is not interrupted, thereby improving the stability of the live broadcast and the user experience.

[0045] The technical solution of the embodiment of the present invention obtains the audio data to be played corresponding to the digital live broadcast object during the live broadcast process. The audio data to be played is split and processed to obtain multiple audio fragment data and fragment index files corresponding to the audio fragment data. Based on the fragment index file, the multiple audio fragment data are pushed to the live broadcast audio channel, and based on the arrangement order of the multiple audio fragment data, the live voice playback of the digital live broadcast object is realized, which solves the problem of discontinuous audio in the live broadcast of the digital human, meets the demand for flexible adjustment of audio content in the live broadcast scene, and makes the live broadcast sound of the digital human live broadcast more smooth and natural.

[0046] On the basis of the above embodiments, the method further includes: when an audio insertion instruction is received, obtaining the audio data to be inserted according to the audio insertion instruction; inserting the insertion fragment data obtained after splitting the audio data to be inserted into the live audio channel to realize the inserted voice playback of the digital live object.

[0047] The audio insertion instruction may refer to an instruction for inserting audio playback during the preset live audio playback process. The audio data to be inserted may refer to audio data for generating the inserted audio. The insertion fragment data may refer to audio fragments for splitting the audio data to be inserted.

[0048] Specifically, after receiving the audio insertion instruction, the server requests the caller to obtain the audio data to be inserted corresponding to the audio insertion instruction. According to the requirements of the audio insertion instruction, the server splits the audio data to be inserted into insertion fragment data and generates an insertion index file. Referring to the arrangement order of the insertion fragment data in the broadcast index file, the insertion fragment data is inserted into the caller's live audio channel to realize the insertion voice playback of the digital live object. At the same time, the insertion index file can also be used to update the total index file in the caller.

[0049] Exemplarily, before inserting the insertion fragment data into the live audio channel, it also includes: determining the insertion type corresponding to the audio insertion instruction, wherein the insertion type includes interruptive insertion and temporary insertion; when it is determined that the insertion type is an interruptive insertion, the live audio channel is cleared, and the insertion fragment data is inserted into the cleared live audio channel; when it is determined that the insertion type is a temporary insertion, based on the insertion time point corresponding to the audio insertion instruction, the insertion fragment data is inserted into the live audio channel.

[0050] Specifically, if the audio insertion instruction requires interrupting the currently playing content, the caller discards the unplayed audio fragment data in the live audio channel, clears the live audio channel, and appends the insertion index file to the general index. If the audio insertion instruction requires temporary insertion of the currently playing content, the insertion fragment data is inserted into the live audio channel according to the insertion time point required by the audio insertion instruction, and the insertion index file is appended to the general index to achieve non-interruption insertion. The technical solution of the embodiment of the present invention meets the diverse content insertion needs in live broadcast scenarios and improves the user experience by supporting interrupting the currently playing content for insertion and achieving insertion without interrupting the current content.

[0051] Embodiment 2

[0052] Figure 2 This is a flowchart of a live voice playback method provided by Example 2 of the present invention. This example is based on the above examples and is a preferred implementation scheme of the present invention. Figure 2 As shown, the method includes:

[0053] During the broadcast phase, the server establishes a long WebSocket connection with the caller.

[0054] The server sends an audio acquisition request to the caller, and the caller sends the audio data to be played to the service.

[0055] The server uses HLS technology to split the audio data to be played into audio fragment data and fragment index files.

[0056] The server stores the audio fragment data in the to-be-played list and appends the fragment index file to the total index.

[0057] The server uses the FFmpeg tool to push the audio fragment data to the audio channel for playback in the order of the total index file.

[0058] When there is only one sentence left to be played on the current audio channel, request the next audio data to be played from the caller in advance and repeat the above steps.

[0059] During the insertion phase, the caller sends an audio insertion instruction to the server.

[0060] The server splits the audio data to be inserted into insertion fragment data and insertion index files according to the requirements of the audio insertion instruction.

[0061] If the audio insertion instruction requires interrupting the current playing content, the caller discards the audio fragment data that has not been played completely, inserts the insertion fragment data into the value audio channel, and appends the insertion index file to the total index.

[0062] If the interstitial message does not need to interrupt the current content, the server only appends the interstitial index file to the total index. The server pushes the interstitial fragment data to the audio channel in the order of the updated total index file.

[0063] During the exception handling phase, if the caller fails to provide new audio data to be played in time, the server will push blank audio fragments to the audio channel to prevent streaming interruption.

[0064] Embodiment 3

[0065] Figure 3 This is a schematic diagram of the structure of a live voice playback device provided by Embodiment 3 of the present invention. Figure 3 As shown, the device comprises:

[0066] The audio data acquisition module 301 is used to acquire the to-be-played audio data corresponding to the digital live broadcast object during the live broadcast process;

[0067] The fragment data obtaining module 302 is used to split the audio data to be played to obtain a plurality of audio fragment data and fragment index files corresponding to the audio fragment data;

[0068] The live voice playback module 303 is used to push the multiple audio fragment data to the live audio channel based on the fragment index file, and realize the live voice playback of the digital live object based on the arrangement order of the multiple audio fragment data.

[0069] The technical solution of the embodiment of the present invention obtains the audio data to be played corresponding to the digital live broadcast object during the live broadcast process. The audio data to be played is split and processed to obtain multiple audio fragment data and fragment index files corresponding to the audio fragment data. Based on the fragment index file, the multiple audio fragment data are pushed to the live broadcast audio channel, and based on the arrangement order of the multiple audio fragment data, the live voice playback of the digital live broadcast object is realized, which solves the problem of discontinuous audio in the live broadcast of the digital human, meets the demand for flexible adjustment of audio content in the live broadcast scene, and makes the live broadcast sound of the digital human live broadcast more smooth and natural.

[0070] Optionally, the audio data acquisition module 301 is further used to:

[0071] When the number of unplayed audio fragment data is less than a preset number threshold, continue to obtain the audio data to be played, and implement the live voice playback of the digital live broadcast object.

[0072] Optionally, the live voice playing module 303 is further used for:

[0073] In the case where no audio data to be played is obtained, blank fragment data is pushed to the live audio channel.

[0074] Optionally, the device further comprises: a live voice insertion module.

[0075] Live voice insertion module, used for:

[0076] When receiving an audio insertion instruction, obtaining the audio data to be inserted according to the audio insertion instruction;

[0077] The insertion fragment data obtained after splitting the audio data to be inserted is inserted into the live audio channel to realize the insertion voice playback of the digital live object.

[0078] Optionally, the device further includes: an insertion type determination module.

[0079] The insertion type determination module is used to:

[0080] Before inserting the insertion fragment data into the live audio channel, determining the insertion type corresponding to the audio insertion instruction, wherein the insertion type includes interruptive insertion and temporary insertion;

[0081] When it is determined that the insertion type is an interruptive insertion, the live audio channel is cleared, and the insertion fragment data is inserted into the cleared live audio channel;

[0082] When it is determined that the insertion type is a temporary insertion, the insertion fragment data is inserted into the live audio channel based on the insertion time point corresponding to the audio insertion instruction.

[0083] Optionally, the device further comprises an index file updating module.

[0084] Index file update module, used to:

[0085] After the plurality of audio fragment data are pushed to the live audio channel, the total index file is updated according to the fragment index file.

[0086] Optionally, the audio fragment data is TS fragments, and the fragment index file is an M3U8 index file.

[0087] The live voice playback device provided in the embodiment of the present invention can execute the live voice playback method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0088] Embodiment 4

[0089] Figure 4 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.

[0090] like Figure 4 As shown, the electronic device 10 includes at least one processor 11, and a memory connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., wherein the memory stores a computer program that can be executed by at least one processor, and the processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 to the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0091] A number of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0092] The processor 11 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as a live voice playback method.

[0093] In some embodiments, the live voice playback method may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the live voice playback method described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to perform the live voice playback method in any other appropriate manner (e.g., by means of firmware).

[0094] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0095] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the computer program is executed by the processor, the functions / operations specified in the flow chart and / or block diagram are implemented. The computer program may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0096] In the context of the present invention, a computer-readable storage medium may be a tangible medium that may contain or store a computer program for use by or in combination with an instruction execution system, device or equipment. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0097] To provide interaction with a user, the systems and techniques described herein may be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).

[0098] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0099] A computing system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The client and server relationship is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services.

[0100] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and this document does not limit this.

[0101] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A live voice playback method, characterized in that: include: Obtain the audio data to be played corresponding to the digital live broadcast object during the live broadcast process; Splitting the audio data to be played to obtain a plurality of audio fragment data and fragment index files corresponding to the audio fragment data; Based on the fragment index file, the plurality of audio fragment data are pushed into the live audio channel, and based on the arrangement order of the plurality of audio fragment data, the live voice playback of the digital live object is realized.

2. The method according to claim 1, characterized in that The method further comprises: When the number of unplayed audio fragment data is less than a preset number threshold, continue to obtain the audio data to be played, and implement the live voice playback of the digital live broadcast object.

3. The method according to claim 2, characterized in that The method further comprises: In the case where no audio data to be played is obtained, blank fragment data is pushed to the live audio channel.

4. The method according to claim 1, characterized in that: The method further comprises: When receiving an audio insertion instruction, obtaining the audio data to be inserted according to the audio insertion instruction; The insertion fragment data obtained after splitting the audio data to be inserted is inserted into the live audio channel to realize the insertion voice playback of the digital live object.

5. The method according to claim 4, characterized in that Before inserting the insert fragment data into the live audio channel, the method further includes: Determine the insertion type corresponding to the audio insertion instruction, wherein the insertion type includes interruptive insertion and temporary insertion; When it is determined that the insertion type is an interruptive insertion, the live audio channel is cleared, and the insertion fragment data is inserted into the cleared live audio channel; When it is determined that the insertion type is a temporary insertion, the insertion fragment data is inserted into the live audio channel based on the insertion time point corresponding to the audio insertion instruction.

6. The method according to claim 1, characterized in that After pushing the plurality of audio fragment data to the live audio channel, the method further includes: The total index file is updated according to the fragment index file.

7. The method according to any one of claims 1 to 6, characterized in that: The audio fragment data is TS fragments, and the fragment index file is an M3U8 index file.

8. A live voice playback device, characterized in that: include: An audio data acquisition module is used to acquire the to-be-played audio data corresponding to the digital live broadcast object during the live broadcast process; A fragment data acquisition module, used for splitting the audio data to be played to obtain a plurality of audio fragment data and fragment index files corresponding to the audio fragment data; The live voice playback module is used to push the multiple audio fragment data to the live audio channel based on the fragment index file, and realize the live voice playback of the digital live object based on the arrangement order of the multiple audio fragment data.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the live voice playback method described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the live voice playback method described in any one of claims 1-7 when executed.