Electronic device, method for providing content, storage medium, and program product

By providing an electronic device that includes a controller and a player, and utilizing machine learning models to generate high-quality media content, the problem of insufficient manpower and weak interactive capabilities for local businesses in live streaming is solved, thereby improving the quality and effectiveness of live streaming.

CN121967735APending Publication Date: 2026-05-01BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2026-03-13
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Local businesses lack professional training in live streaming, have limited language skills, and weak interactive abilities, making it difficult to support high-frequency, long-duration live streams, which affects the effectiveness of the live streams.

Method used

An electronic device is provided, including a controller and a player, capable of playing media content in response to object selection instructions, assisting businesses in live streaming, reducing reliance on human labor, and generating high-quality media content using machine learning models.

Benefits of technology

Using electronic devices to assist merchants in live streaming reduces reliance on manpower, improves the quality and interactivity of the live stream, and enhances the overall effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121967735A_ABST
    Figure CN121967735A_ABST
Patent Text Reader

Abstract

An electronic device, a method for providing content, a storage medium, and a program product are provided. An electronic device presented herein includes a controller configured to provide a playback instruction for media content to a player in response to a selection instruction for at least one object, the media content including an explanation voice for the at least one object; and a player configured to play the media content in response to the play instruction. In this way, the electronic device can play the media content to assist the user in live broadcast, the dependence of the live broadcast process on manpower is reduced, and the live broadcast quality is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic devices, methods for providing content, storage media, and program products Technical Field

[0001] The examples in this article generally relate to the field of computers, and in particular to electronic devices, methods for providing content, computer-readable storage media, and computer program products. Background Technology

[0002] Live streaming is the real-time transmission of content (such as video, audio, images, etc.) over the internet, allowing users to watch live content instantly. With its immediacy, interactivity, and immersive experience, it has rapidly gained widespread application in various fields such as entertainment and e-commerce. For example, in live-streaming e-commerce scenarios, hosts effectively increase users' willingness to buy and conversion rates by showcasing product details in real time, answering user questions, and interacting with viewers, making live streaming a crucial bridge connecting products and consumers. Summary of the Invention

[0003] In a first aspect, an electronic device is provided. The electronic device includes: a controller configured to provide a playback instruction to a player for media content, including narration of the at least one object, in response to selection of at least one object; and a player configured to play the media content in response to the playback instruction.

[0004] In a second aspect, a method for providing content is provided. The method includes: receiving a selection instruction for at least one object, the at least one object being associated with a live event; and playing media content, the media content including narration of the at least one object.

[0005] In a third aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores computer-executable instructions that can be executed by a processor to implement the method of the second aspect.

[0006] In a fourth aspect, a computer program product is provided, which is tangibly stored in a computer storage medium and includes computer-executable instructions that, when executed by a device, cause the device to perform the method of the second aspect.

[0007] In this way, media content can be played through electronic devices to assist users in live streaming, reducing reliance on human intervention and helping to improve the quality of the live stream.

[0008] It should be understood that the content described in this section is not intended to limit the key or important features of the examples in this article, nor is it intended to restrict the scope of the solution. Other features will become readily apparent from the following description. Attached Figure Description

[0009] The above and other features, advantages, and aspects of the various examples herein will become more apparent when taken in conjunction with the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein: Figure 1 shows a schematic diagram of an example environment; Figure 2A shows a schematic diagram of an example architecture for providing content according to some scenarios; Figure 2B shows a schematic diagram of an example architecture of an electronic device according to some scenarios; Figure 3 shows a schematic diagram of an example architecture for providing content according to other scenarios; Figure 4 shows a flowchart of a signaling flow for providing content according to some scenarios; Figure 5 shows a flowchart of an example process for providing content according to some scenarios; and Figure 6 shows a block diagram of an electronic device according to other scenarios. Detailed Implementation

[0010] The examples in the text will now be described in more detail with reference to the accompanying drawings. While some examples are shown in the drawings, it should be understood that solutions can be implemented in various forms and should not be construed as limited to the examples presented herein. Rather, these examples are provided to provide a more thorough and complete understanding of the solutions. It should be understood that the drawings and examples in this document are for illustrative purposes only and are not intended to limit the scope of protection of the solutions.

[0011] It should be noted that the headings of any section / subsection provided herein are not restrictive. Various examples are described throughout this document, and examples of any type may be included under any section / subsection. Furthermore, examples described in any section / subsection may be combined in any way with any other examples described in the same section / subsection and / or different sections / subsections.

[0012] In the description of the examples in this document, the term "including" and similar terms should be understood as open inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "an example" or "the example" should be understood as "at least one example". The term "some examples" should be understood as "at least some examples". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0013] The examples in this document may involve user data, data acquisition, and / or use. All of these aspects comply with relevant laws, regulations, and provisions. In the examples presented herein, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, when implementing each example, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained through appropriate means, in accordance with relevant laws and regulations. The specific methods of notification and / or authorization can vary depending on the actual situation and application scenario; the scope of the solution is not limited in this regard.

[0014] In this manual and the sample solutions, any processing of personal information will be conducted only under legal grounds (such as obtaining the consent of the data subject or being necessary for the performance of a contract) and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information beyond what is necessary for basic functions will not affect the user's use of basic functions.

[0015] As used in this paper, the term "model" refers to a model that learns the relationship between inputs and outputs from training data, enabling it to generate corresponding outputs for a given input after training. Model generation can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs using multiple layers of processing units. A neural network model is an example of a deep learning-based model. In this paper, "model" may also be referred to as a "machine learning model," "learning model," "machine learning network," or "learning network," and these terms are used interchangeably.

[0016] A neural network is a machine learning network based on deep learning. A neural network processes input and provides a corresponding output, typically consisting of an input layer, an output layer, and one or more hidden layers between the input and output layers. Neural networks used in deep learning applications often include many hidden layers, thus increasing the network's depth. The layers of a neural network are connected sequentially, so that the output of the previous layer is provided as the input to the next layer. The input layer receives the input to the neural network, while the output layer's output serves as the final output. Each layer of a neural network includes one or more nodes (also called processing nodes or neurons), each node processing the input from the layer above.

[0017] As mentioned above, in live-streaming e-commerce, hosts effectively increase users' willingness to buy and conversion rates by showcasing product details in real time, answering user questions, and interacting with viewers, making live-streaming a crucial bridge connecting products and consumers. However, local businesses (such as small and medium-sized shops) generally lack professional live-streaming training, have limited language skills, and relatively weak interactive abilities. Furthermore, local businesses often need to manage their physical stores, making it difficult to support high-frequency, long-duration live streams, thus impacting the overall effectiveness of the broadcasts.

[0018] An electronic device for live streaming is proposed herein. In this scheme, the electronic device includes a controller and a player. The controller can provide the player with a playback instruction for media content in response to a selection instruction for at least one object. The at least one object is associated with a live streaming event, and the media content includes narration of the at least one object. The player can play the media content in response to the playback instruction.

[0019] In this way, users (such as merchants) can be assisted in live streaming by playing media content through devices, reducing the reliance on human labor in the live streaming process and improving the quality of the live stream.

[0020] The following describes various examples of this scheme in further detail with reference to the accompanying drawings.

[0021] Figure 1 illustrates a schematic diagram of example environment 100. As shown in Figure 1, in example environment 100, one or more users 110-1, 110-2, 110-3, ..., 110-N can watch live streams and participate in live stream sessions (also referred to as "live stream events") through their respective associated terminal devices 120-1, 120-2, ..., 120-N. User 110 can also be referred to as a participant. For ease of discussion, users 110-1, 110-2, ..., 110-N can be collectively referred to as user 110 or individually, and terminal devices 120-1, 120-2, ..., 120-N can be collectively referred to as terminal device 120 or individually. Among all users 110, one or more users, such as user 110-1, can establish a live stream session and initiate a live stream through their associated terminal device 120-1.

[0022] In some scenarios, the user 110-1 who initiates the live stream can also be referred to as the host of the live stream session, a host participant, a live stream user, or a live stream administrator. One or more other users 110 may include guest participants 110-N, representing users who participate in the live stream interaction initiated by the host (e.g., live chat). In addition to guest participants in the live stream interaction, other users 110 may also include viewers, listeners, or viewers of the live stream session.

[0023] In some cases, terminal device 120 may have applications installed that can provide live streaming services, or may have access to websites that can provide live streaming services. User 110 can operate terminal device 120 to access the corresponding applications or websites. Accordingly, terminal device 120 can display a corresponding live streaming interface, which can provide media content of the live event, such as audio content or video media content.

[0024] In some cases, terminal device 120 can also communicate with server 130 via network 132 to provide live streaming services. Server 130 can also provide functions such as application or website management, configuration, and maintenance. Server 130 can also store data generated during the live streaming session, including live interaction data generated during live interaction.

[0025] Terminal device 120 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable gaming terminals, VR / AR devices, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some cases, terminal device 120 may also support any type of user-facing interface (such as "wearable" circuitry).

[0026] Server 130 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. Server 130 may include, for example, computing systems / servers such as mainframes, edge computing nodes, computing devices in cloud environments, etc.

[0027] A communication connection can be established between server 130 and terminal device 120. This communication connection can be established via wired or wireless means. The communication connection can include, but is not limited to, Bluetooth, mobile network, Universal Serial Bus (USB), and Wireless Fidelity (WiFi) connections. In some cases, server 150 and terminal device 120 can exchange signaling information through their communication connection.

[0028] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.

[0029] Figure 2A illustrates a schematic diagram of an example architecture 200 for providing content under certain circumstances, and Figure 2B illustrates a schematic diagram of an example architecture 210 for an electronic device under certain circumstances. For ease of discussion, the example architecture 200 shown in Figure 2A and the electronic device 210 shown in Figure 2B will be described in conjunction with the example environment 100 shown in Figure 1, but it should be understood that this is merely exemplary. As shown in Figures 2A and 2B, the electronic device 210 may include a controller 212, a display 214, and a speaker 216. The controller 212, the display 214, and the speaker 216 may be communicatively connected via, for example, a bus or other means. The display 214 and the speaker 216 are sometimes referred to as output devices or players. The electronic device 210 (e.g., controller 212) may receive a selection instruction for at least one object, which may instruct the provision of media content about the at least one object, which may be associated with a live event. In response to the selection instruction, the controller 212 may provide a playback instruction for the media content to the player (e.g., speaker 216). The player may play the media content in response to the playback instruction.

[0030] In some examples, electronic device 210 may receive a first request for a live event. This first request may be used to request electronic device 210 to provide media content for the live event. Electronic device 210 may use display 214 to present candidate objects related to the live event; these candidate objects may include objects suitable for live streaming during the live event. User 110-1 may select at least one object from the candidate objects (for ease of description, the object selected by user 110-1 will be referred to as the "live object" below). The live object may include various virtual or physical objects, such as product objects.

[0031] In some cases, electronic device 210 can serve as smart hardware to assist live streaming, such as assisting the broadcaster in playing and controlling audio narration during a live stream. For ease of description, electronic device 210 may also be referred to herein as a live streaming auxiliary device, live streaming equipment, smart hardware, or any other suitable name.

[0032] In some examples, as shown in Figure 2B, electronic device 110 may also include input device 222, which may be communicatively connected to controller 212 via, for example, a bus or other means. Input device 222 can be used to input data or instructions to electronic device 210. Input device 222 may include any suitable hardware device, such as, but not limited to, touch devices, mice, keyboards, buttons, image capture devices, or audio capture devices. User 110-1 can operate input device 222 to send a first request to controller 212. Controller 212 may respond to the first request by controlling display 214 to present candidate objects. It should be understood that input device 222 here and input device 650 hereinafter may be implemented as the same or different input devices.

[0033] Media content may include live audio or live video containing audio. If electronic device 210 receives a selection of at least one live object from the candidate list, electronic device 210 may play live audio using speaker 216. For example, controller 212 may send a playback command for live audio to speaker 216 in response to the selection of a live object from the candidate list. Speaker 216 may play live audio in response to the playback command. The live audio described herein may be associated with a live event, for example, it may be used as media content. Live audio includes narration of at least one live object. Terminal device 120-1 may collect media content, and terminal device 120-1 may provide live streaming services by communicating with server 130. For example, terminal device 120-1 may push the collected media content to server 130. In this way, electronic device 210 can play media content to assist user 110-1 (e.g., a merchant) in live streaming, reducing the reliance on human intervention in the live streaming process and improving the quality of the live stream.

[0034] In some cases, live audio playback can also be accompanied by video. For example, electronic device 210 can also synchronously display video on display 214, which may include a virtual anchor acted by a virtual object. The virtual anchor's lip movements can be synchronized with the live audio.

[0035] In some examples, referring to Figure 2B, electronic device 210 may also include transceiver 218, through which electronic device 210 can communicate with server 220 to provide auxiliary live streaming services. Transceiver 218 may include any suitable hardware unit with communication capabilities, such as the communication unit described below. In response to a first request, electronic device 210 may use transceiver 218 to send a request (sometimes referred to herein as a "fourth request") to server 220 to request the server to provide candidate objects. This request may include an identifier of the live event. The identifier may include any suitable information capable of identifying the live event, such as, but not limited to, the live event number, the live event name, the live account number, the live account username, etc. Server 220 may, in response to this request, determine the candidate objects.

[0036] Candidate objects can include any suitable object. In some examples, user 110-1 can initiate a live stream targeting a product. In this case, candidate objects can include one or more products. Alternatively and / or additionally, user 110-1 can also initiate live streams targeting other areas, such as entertainment, games, knowledge, or lifestyle services. In this case, candidate objects can include, for example, game objects, knowledge points, or service items. It is understood that candidate objects may differ when live streaming targets different areas. This document does not limit the live streaming area or the type of candidate objects.

[0037] Electronic device 210 can receive an indication of a candidate object from server 220 using transceiver 218. This indication may include any appropriate information related to the candidate object, including but not limited to the candidate object's name, number, description, or image. For example, if the candidate object includes a product, the indication may include, but is not limited to, the product name, product number, product description, product image, or product link. Upon receiving the indication of a candidate object, electronic device 210 can present the candidate object. In this way, electronic device 210 can accurately determine the candidate object using server 220.

[0038] In some examples, controller 212 may receive a selection instruction for the at least one object in response to receiving a first question. The first question includes a question posed to at least one object during a live event. For example, the first question may include a question asked by a guest or viewer about a product or service during a live event. In this case, electronic device 210 may provide live content for answering the question, such as audio and / or video for answering the question.

[0039] In some examples, controller 210 may receive a selection instruction for at least one object in response to receiving a fifth request. The fifth request may instruct live streaming of at least one object during a live event. In this case, electronic device 210 may automatically play live content to assist user 110-1 in live streaming. For example, electronic device 210 may be used to control a live event, and user 110-1 may select objects for the live event through electronic device 210. Electronic device 210 may generate a fifth request in response to the selection of objects to be live or to be live streamed. Also, for example, electronic device 210 may communicate with terminal device 120-1 or server 130 to receive a fifth request from terminal device 120-1 or server 130.

[0040] Continuing with Figure 2A, in some examples, electronic device 210 can use transceiver 218 to send a request (sometimes referred to herein as a "third request") to server 220. The third request requests media content and includes an object identifier for at least one live-streaming object. This object identifier indicates the corresponding live-streaming object and may include any appropriate information, such as, but not limited to, a number, name, a link to the object's display page, etc. Server 220 can generate media content based on this object identifier. Electronic device 210 can receive the media content from server 220 and then play it. In this way, the computing resources of server 220 can be used to generate media content, which helps improve the quality of the media content. Of course, media content can also be generated by electronic device 210 or other devices. Furthermore, electronic device 210 can also provide other visual content (such as text, images, or videos, etc.) to assist user 110-1 in live-streaming. In this case, server 220 can also generate visual content based on the object identifier of the live-streaming object and provide the visual content to electronic device 210.

[0041] In some examples, electronic device 210 can utilize machine learning model 230 (which may include one or more machine learning models, such as machine learning model 230-1, machine learning model 230-2, ..., machine learning model 230-N, etc., where N is a positive integer. For ease of description, one or more machine learning models are collectively referred to as machine learning model 230 in this document) to provide auxiliary live streaming services. As an example, server 220 can utilize machine learning model 230 to generate media content to improve the quality and efficiency of media content generation.

[0042] Machine learning model 230 can be of different types. In some examples, machine learning model 230 can be built based on a language model (LM). The machine learning model used is a content-generative model, capable of generating corresponding outputs based on model inputs. A language model-based machine learning model can receive text-modal model inputs (e.g., natural language and / or machine language) and / or non-text-modal model inputs (e.g., images, speech, video, etc.), and can generate the desired output based on the model inputs and prompts. Prompts can indicate user needs, guiding the machine learning model to generate model outputs that match user needs. In other examples, machine learning model 230 can be built based on a multimodal model. In still other examples, machine learning model 230 can be a speech-related model, including an automatic speech recognition (ASR) model or a text-to-speech (TTS) model. The input of an ASR model is speech, and the output is text. The input of a TTS model is text, and the output is the corresponding speech. It should be understood that the above machine learning model 230 is merely exemplary. In practical applications, any suitable model architecture can be used. It should also be noted that although the machine learning model 230 is deployed on server 220 in Figures 2A and 2B, the machine learning model 230 can also be deployed on electronic device 210, terminal device 120-1, edge device, or other devices. This document does not impose any restrictions on the deployment location of the machine learning model 230.

[0043] Figure 3 illustrates a schematic diagram of an example architecture 300 for assisting live streaming under certain scenarios. As shown in Figure 3, the electronic device 310 includes a player 306. The player 306 may include the display 214 and / or speaker 216 shown in Figure 2A. Figure 4 illustrates a flowchart of a signaling flow 400 for user-provided content under certain scenarios. The signaling flow 400 involves the electronic device 210 and a server 220, which may include an interface 312, a service node 314, and a queue 316. The example architecture 300 will be described below in conjunction with the signaling flow 400.

[0044] As shown in Figures 3 and 4, electronic device 210 can receive (402) a first request 302 for a live event. Electronic device 210 can use transceiver 218 to send (404) a fourth request to interface 312 to request server 220 to provide candidate objects. The fourth request may include an identifier of the live event. Interface 312 can determine candidate objects associated with the live event based on the identifier. Electronic device 210 can use transceiver 218 to receive (406) an indication of candidate objects from interface 312. Subsequently, electronic device 210 can use display 214 to present (408) the candidate objects.

[0045] In some cases, interface 312 can determine the products displayed in the live stream event (e.g., products located on the shelf of the live stream event) based on the live stream event number. Then, interface 312 can provide feedback on the product information (i.e., indication of candidate items) to electronic device 210. Alternatively and / or additionally, interface 312 can also determine the products displayed in the live stream account's product showcase based on the live stream account's account information. Of course, the above candidate items are merely exemplary. The method of determining candidate items may differ in different live streaming scenarios.

[0046] In some cases, the display 214 of the electronic device 210 can present an interface. This interface can display candidate objects. As shown in Figure 2B, the user 110-1 can operate the input device 222 to send a first request 302 to the controller 212. The first request 302 can be initiated by touching the corresponding virtual control displayed on the display 214, or by using physical controls set on the electronic device 210. Alternatively or additionally, the first request 302 can also be initiated by providing a voice command or text command to the electronic device 210. In response to the first request 302, the controller 212 can control the display 214 to present candidate objects.

[0047] In some situations, the display 214 can be used to display an interface related to the live event. For example, the display 214 can display a control interface for controlling the live event. This control interface can be used to control the flow or content of the live event, such as object display, object switching, or other live-related operations.

[0048] Returning to Figure 4, user 110-1 can select at least one live stream object from the candidate objects. Electronic device 210 can receive (410) the selection of the live stream object using transceiver 218 and send (412) a third request to interface 312. The third request may include the object identifier of the live stream object. Interface 312 can send (414) a generation request to service node 314, which may include the object identifier of the live stream object and can be used to request service node 314 to generate media content. Service node 314 can obtain (416) the explanatory text of the live stream object based on the object identifier of the live stream object. The explanatory text is used to explain the corresponding live stream object. For example, the explanatory text may include the explanation content of products, knowledge points, or service items.

[0049] In some scenarios, as shown in Figure 2B, electronic device 210 can communicate with server 220 via transceiver 218. User 110-1 can select a live streaming object via input device 222, and controller 212 can respond to the selection of the live streaming object by sending a third request to interface 312 via transceiver 218.

[0050] In other scenarios, as shown in Figure 3, service node 314 can determine information related to the live-streaming object (e.g., information related to products, knowledge points, or service items) based on the object identifier of the live-streaming object. Then, service node 314 can request explanatory text from machine learning model 230-1 based on the information related to the live-streaming object. Machine learning model 230-1 can generate explanatory text based on the information related to the live-streaming object. Service node 314 can then obtain the explanatory text from machine learning model 230-1. Of course, service node 314 is not limited to using machine learning model 230-1 to generate explanatory text. In practical applications, service node 314 can also generate explanatory text based on information related to the live-streaming object and an explanatory template. This paper does not restrict the method of generating explanatory text.

[0051] Continuing with Figure 4, service node 314 can obtain (418) the narration audio of the live stream object based on the narration text. As an example, service node 314 can request machine learning model 230-2 (e.g., a TTS model) to generate the narration audio based on the narration text. Afterwards, service node 314 can receive the narration audio from machine learning model 230-2. Of course, service node 314 is not limited to using machine learning model 230-2 to generate narration audio, but can also generate narration audio through any other appropriate method. This article does not impose any restrictions on this.

[0052] Regarding the acquisition of narration audio, in some scenarios, service node 314 can acquire the narration audio of one live stream object at a time. For example, service node 314 can provide the narration text of one live stream object to machine learning model 230-2 at a time and acquire the narration audio generated by machine learning model 230-2. In other scenarios, service node 314 can acquire the narration audio of multiple live stream objects simultaneously. For example, service node 314 can provide the narration text of multiple live stream objects to machine learning model 230-2 at a time and acquire multiple narration audios from machine learning model 230-2.

[0053] In some examples, server 220 (e.g., service node 314) can send narration audio to electronic device 210. Electronic device 210 can play the narration audio via, for example, speaker 216, to allow user 110-1 to listen to the generated narration audio. If user 110-1 is satisfied with the generated quality of the narration audio, they can trigger a confirmation of the narration audio. If electronic device 210 receives the confirmation of the narration audio, it can send a confirmation instruction to server 220, instructing server 220 to generate media content (e.g., live audio) based on the narration audio. If user 110-1 is not satisfied with the generated quality of the narration audio, they can trigger an update request for the narration audio. If live streaming device 210 receives the update request for the narration audio, it can send an update instruction for the narration audio to server 220. In this case, service node 314 can update the narration audio. Of course, if user 110-1 is not satisfied with the generated quality of the narration audio, they can provide feedback to electronic device 210. Server 220 can update the narration audio based on the feedback. For example, service node 314 can re-acquire the narration text and narration audio based on the feedback.

[0054] As shown in Figure 4, upon receiving the narration audio, service node 314 can send (420) the narration audio to queue 316. Queue 316 generates media content based on the narration audio. Electronic device 210 can receive (422) the media content (e.g., live audio) from queue 316 using transceiver 218. Afterward, electronic device 210 can play (422) the media content.

[0055] In some cases, service node 314 can also provide explanatory text to queue 316, and service node 314 can generate live text corresponding to the media content based on the explanatory text. Electronic device 210 can also receive live text from queue 316. Then, electronic device 210 can synchronously play the media content and display the live text. This method helps improve the service quality of the auxiliary live streaming service. In some examples, a dual-channel can be established between electronic device 210 and queue 316, and electronic device 210 can receive media content and live text from queue 316 based on the dual-channel.

[0056] As an example, electronic device 210 can receive media content and live text from queue 316 via transceiver 218. Controller 212 can invoke speaker 216 to drive it to play media content. Controller 212 can also control display 214 to display live text.

[0057] In some examples, the media content may include narration content from one or more live stream subjects, which may include at least one of the subject's narration audio and video. Alternatively and / or additionally, the media content may also include introductory content (also referred to as, for example, an "opening"), which precedes the narration content of at least one subject. The introductory content may include at least one of introductory audio and introductory visual content. Introductory content can quickly attract the attention of viewers or listeners, thus stimulating their willingness to watch or listen. For example, server 220 may generate introductory audio and corresponding introductory text based on information related to the live stream event (such as information related to the live stream account, information related to the live stream event, etc.). Of course, server 220 may pre-build a first set, which may include multiple introductory contents and corresponding multiple closing texts. Upon receiving a third request, server 220 may also select introductory content and closing text from the first set.

[0058] Alternatively and / or additionally, the media content may also include concluding content (also referred to as, for example, a "closing statement"), which follows the narration of at least one object. The concluding content may include at least one of concluding audio and concluding visual content. This improves the completeness of the media content's presentation and enhances the viewer's or listener's memory, thereby improving the quality of the live stream. As an example, server 220 may generate concluding audio and corresponding concluding text based on information related to the live event, using, for example, machine learning models 230-1 and 230-2. Alternatively, server 220 may pre-build a second set, which may include multiple concluding audios and corresponding multiple concluding texts. Upon receiving a third request, server 220 may select concluding audio and concluding text from the second set.

[0059] Alternatively and / or additionally, the media content may also include related content (also referred to as, for example, "transitional speech" or "transitional speech") located between adjacent explanatory content. Setting related content between adjacent explanatory content improves the coherence of the media content. Related content may include at least one of related speech and related visual content (e.g., text, images, or videos). For example, server 220 may pre-build a third set. The third set may include multiple related speech and corresponding multiple related texts. Upon receiving a third request, server 220 can select related speech and corresponding related text from the third set. Of course, server 220 may also use a model to generate related speech and related text. This document does not limit the method of obtaining related speech and related text.

[0060] Alternatively and / or additionally, the media content may also include interactive content, which is used to respond to interactions related to the live event (also referred to as, for example, "live interaction"). Live interaction here may include any appropriate interactive action, such as, but not limited to, viewers or listeners paying attention to the live event, viewers or listeners participating in (or entering) the live event, etc.

[0061] As an example, if it is determined that viewers are paying attention to a live event, server 220 can generate interactive content corresponding to the attention (i.e., live interaction), which can be used to express gratitude to viewers for their attention. Server 220 can add this interactive content to queue 316. Queue 316 can determine the order of introductory content, at least one explanatory piece of content, interactive content, and concluding content, and generate media content based on the determined order. Electronic device 210 can receive the media content from queue 316.

[0062] As another example, if it is determined that a viewer has entered a live broadcast event, server 220 can generate interactive content corresponding to the entry, which can be used to welcome the viewer. Then, server 220 can provide media content to electronic device 210 based on the interactive and explanatory content. In this way, the interaction between the host and the audience or guests in a real-world scenario can be simulated, improving the interactivity of the auxiliary live broadcast service.

[0063] In some examples, the electronic device 210 can also receive a second request for a task. Based on the task type, the electronic device 210 can either execute the task or continue playing the media content. In this way, task execution and media content playback can be coordinated based on the task type, avoiding conflicts. As an example, the electronic device 210 can determine a scheduling policy corresponding to the task type. Then, based on the scheduling policy, the electronic device 210 can choose to execute the task or continue playing the media content. As another example, the electronic device 210 can determine the task priority based on the task type. Based on the task priority and the media content priority, it can either execute the task or continue playing the media content. If the task priority is higher than the media content priority, the electronic device 210 can pause media content playback. The electronic device 210 can then execute the task, playing task-related prompts (e.g., audio or visual prompts). If the task priority is lower than the media content priority, the electronic device 210 can block the second request for the task and continue playing the media content.

[0064] It is understood that electronic device 210 is not limited to scheduling task execution and media content playback based on priority. In some examples, electronic device 210 can determine the type of task. If the task type is a first-class task, electronic device 210 can pause media content playback. Then, electronic device 210 can play the first prompt content associated with the first-class task. In some cases, the priority of the first-class task may be higher than the priority of playing media content. If the task type is determined to be a first-class task, meaning that the task has a higher priority than the priority of playing media content, electronic device 210 can pause media content playback. Electronic device 210 can also execute the task and play the first prompt content based on the task execution result.

[0065] For example, electronic device 210 can receive an order reconciliation request. This reconciliation request may include, for example, a reconciliation code associated with the order, which can be used to request electronic device 210 to reconcile the order based on the reconciliation code. Electronic device 210 can determine that the order reconciliation task is a first-class task, meaning that the reconciliation task has a higher priority than playing media content. Electronic device 210 can pause the media content playback and play a prompt audio related to the reconciliation task (i.e., the first prompt content) to indicate the task status of the reconciliation task (e.g., whether the order reconciliation was successful or failed).

[0066] In some examples, electronic device 210 can resume playback of media content once the initial prompt has finished playing. For instance, electronic device 210 can resume playback of media content once the prompt audio for the verification task has finished playing.

[0067] In some examples, electronic device 210 can determine the type of task based on a second request. If the task type is a second type of task, electronic device 210 can pause playback of media content. Electronic device 210 can play mixed media content. This mixed media content can be synthesized based on at least a portion of a second cue content and media content, the second cue content being related to the second type of task. Here, the second cue content can include cue audio or visual cue content (e.g., text, images, or videos, etc.), and the mixed media content can also include mixed audio, visual content with audio, or spliced ​​video content, etc. For example, if the media content includes live audio, the second cue content includes cue audio, and the mixed media content can include mixed audio. As another example, if the media content includes live video with audio, the second cue content includes cue video, and the mixed media content can include spliced ​​mixed video.

[0068] In some situations, the priority of a second-type task can be higher than the priority of playing media content. If a task is determined to be a second-type task, it means that the task has a higher priority than playing media content. The electronic device 210 can pause the playback of media content and execute this task. Afterward, the electronic device 210 can play mixed media content.

[0069] In other cases, the media content may include live audio, and the second prompt content may include prompt audio. Electronic device 210 can output live audio through the first audio channel. If the task type is a second type of task, electronic device 210 can perform the second type of task and output prompt audio through the second audio channel. Electronic device 210 can perform mixing processing on the live audio and prompt audio to obtain mixed audio. Then, electronic device 210 can play the mixed audio.

[0070] As an example, electronic device 210 receives a reconciliation request for an order, which requests electronic device 210 to perform the order reconciliation task. Electronic device 210 can determine that the order reconciliation task is a second type of task. In this case, electronic device 210 can perform audio mixing processing on the media content and the audio prompt for the reconciliation task (i.e., the second prompt content) to obtain mixed audio. By playing this mixed audio, electronic device 210 can simultaneously provide explanations of the live broadcast subject and prompts for the task status, which helps to improve the realism and richness of the live broadcast event.

[0071] In some examples, electronic device 210 can determine the task type based on the second request. If the task type is a third-type task, electronic device 210 can block or ignore the second request and continue playing the media content. For example, electronic device 210 can receive a question voice through an audio collector, requesting electronic device 210 to perform a voice question-and-answer task to play the response voice. Electronic device 210 can determine that the question-and-answer task is a third-type task, and can discard the second request to abandon the execution of the question-and-answer task. At the same time, electronic device 210 can continue playing the live audio to avoid the response voice mixing into the live audio, which helps maintain the live broadcast effect.

[0072] It should be noted that the tasks and task types mentioned above are not exemplary. In practical application scenarios, electronic device 210 can execute any appropriate task and can also select any appropriate scheduling strategy according to actual needs. This document does not impose any restrictions on this.

[0073] Figure 5 shows a flowchart of an example process 500 for providing content under certain circumstances. Process 500 can be implemented at electronic device 210. Process 500 will now be described with reference to Figure 1.

[0074] In box 510, electronic device 210 receives a selection instruction for at least one object, which is associated with a live event.

[0075] In frame 520, electronic device 210 plays media content, which includes narration of at least one object.

[0076] In some examples, process 500 also includes: receiving a second request for requesting the execution of a task; and, based on the type of task, executing the task or maintaining playback of the media content.

[0077] In some examples, the tasks performed include: pausing media content playback in response to the task being of type 1; and playing first prompt content related to the type 1 task.

[0078] In some examples, process 500 also includes: resuming playback of media content in response to the end of the first prompt content playback.

[0079] In some examples, the task includes: pausing playback of media content in response to the task being of type 2; and playing mixed media content, which is synthesized based on at least a portion of a second cue content and media content, the second cue content being related to the type 2 task.

[0080] In some examples, maintaining media content playback includes: maintaining media content playback in response to a task of type 3.

[0081] In some examples, process 500 further includes: sending a third request to the server for requesting media content, the third request including at least one object identifier indicating at least one object; and receiving media content from the server.

[0082] In some examples, the selection instruction is received in the following manner: in response to a first request, candidate objects are presented, which are related to the live event; and in response to the selection of at least one of the candidate objects, a playback instruction is provided to the player.

[0083] In some examples, presenting candidate objects involves: sending a fourth request to the server to request candidate objects, the fourth request including an identifier for the live event; receiving an indication of the candidate objects from the server; and presenting the candidate objects in response to the indication.

[0084] In some examples, the selection instruction is received in the following manner: in response to receiving a first question, the selection instruction is received, the first question including a question posed to at least one object in the live event.

[0085] In some examples, the selection instruction is received in the following manner: in response to receiving a fifth request, the selection instruction is received, which instructs that a live broadcast be performed on at least one object during the live broadcast event.

[0086] In some examples, the media content also includes at least one of the following: a guiding voice, which precedes the narration of at least one object; an ending voice, which follows the narration of at least one object; an associated voice, which is located between adjacent narrations; or an interactive voice, which is used to respond to interactions related to the live event.

[0087] Figure 6 shows a block diagram of an electronic device 600 according to other scenarios. It should be understood that the electronic device 600 shown in Figure 6 is merely exemplary and should not constitute any limitation on the functionality and scope of the examples described herein. The electronic device 600 shown in Figure 6 may be implemented as the same or different electronic device as the electronic device 210 discussed above.

[0088] As shown in Figure 6, the electronic device 600 is in the form of a general-purpose electronic device. Components of the electronic device 600 may include, but are not limited to, one or more processing units or processors 610, memory 620, storage devices 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. The processor 610 may be a physical or virtual processor and is capable of performing various processes according to programs stored in the memory 620. In a multiprocessor system, multiple processors execute computer-executable instructions in parallel to improve the parallel processing capability of the electronic device 600.

[0089] Electronic device 600 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 600, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 620 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof). Storage device 630 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 600.

[0090] Electronic device 600 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 6, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks may be provided. In these cases, each drive may be connected to a bus (not shown) via one or more data media interfaces. Memory 620 may include computer program product 625 having one or more program modules configured to perform various methods or actions of various examples.

[0091] The communication unit 640 enables communication with other electronic devices via a communication medium. Additionally, the functionality of the components of the electronic device 600 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the electronic device 600 can operate in a networked environment using logical connections to one or more other servers, networked personal computers, or another network node.

[0092] Input device 650 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 660 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 600 can also communicate with one or more external devices (not shown) via communication unit 640 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 600, or with any device that enables electronic device 600 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).

[0093] A computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. A computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.

[0094] The flowcharts and / or block diagrams of the methods, apparatus, devices, and computer program products referred to herein describe various aspects. It should be understood that each block of the flowcharts and / or block diagrams, as well as combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0095] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0096] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0097] The flowcharts and block diagrams in the accompanying figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products under various scenarios. In this respect, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the figures. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0098] Various examples have been described above. The foregoing descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. An electronic device, comprising: The controller is configured to provide a playback instruction for media content to the player in response to a selection instruction for at least one object, the at least one object being associated with a live event, the media content including narration of the at least one object; The player is configured to play the media content in response to the playback command.

2. The electronic device of claim 1, wherein the controller is further configured to: receive a second request for requesting the execution of a task; and, based on the type of the task, execute the task or maintain playback of the media content.

3. The electronic device of claim 2, wherein the player is further configured to: pause playback of the media content in response to the task being of the first type of task; and play first prompt content related to the first type of task.

4. The electronic device according to claim 3, wherein the player is further configured to: resume playback of the media content in response to the end of playback of the first prompt content.

5. The electronic device of claim 2, wherein the player is further configured to: pause playback of the media content in response to the task being a second type of task; and play mixed media content, the mixed media content being synthesized based on at least a portion of a second prompt content and the media content, the second prompt content being related to the second type of task.

6. The electronic device of claim 2, wherein the player is further configured to: maintain playback of the media content in response to the task being a third type of task.

7. The electronic device of claim 1, further comprising a transceiver configured to: send a third request to a server, the third request being for requesting the media content, the third request including at least one object identifier indicating the at least one object; and receive the media content from the server.

8. The electronic device of claim 1, further comprising a display configured to: present candidate objects in response to a first request, the candidate objects being related to the live event; and wherein the controller is further configured to: provide the playback instruction to the player in response to selection of at least one of the candidate objects.

9. The electronic device of claim 8, further comprising a transceiver configured to: send a fourth request to a server, the fourth request being for requesting the candidate object, the fourth request including an identifier of the live event; and receive an indication of the candidate object from the server; and wherein the display is further configured to, in response to the indication, present the candidate object.

10. The electronic device of claim 1, wherein the controller is further configured to: receive the selection instruction in response to receiving a first question, the first question comprising a question posed to the at least one object during the live event.

11. The electronic device of claim 1, wherein the controller is further configured to: receive the selection instruction in response to receiving a fifth request, the fifth request indicating that a live broadcast be performed for the at least one object in the live broadcast event.

12. The electronic device of claim 1, wherein the media content further comprises at least one of the following: guiding voice, the guiding voice being located before the explanatory voice of the at least one object, ending voice, the ending voice being located after the explanatory voice of the at least one object, associated voice, the associated voice being located between adjacent explanatory voices, or interactive voice, the interactive voice being used to respond to interactions related to the live event.

13. A method for providing content, comprising: Receive a selection instruction for at least one object, said at least one object being associated with a live event; And playing media content, the media content including narration of the at least one object.

14. A computer-readable storage medium having stored thereon computer-executable instructions that can be executed by a processor to implement the method of claim 13.

15. A computer program product tangibly stored in a computer storage medium and comprising computer-executable instructions that, when executed by a device, cause the device to perform the method according to claim 13.