Method and electronic device for live presentation
By receiving and sequentially presenting different descriptions of multiple live stream segments during the live stream, and using a machine learning model to generate live stream content, the problem of rigid and inaccurate live stream scripts generated by machine learning is solved, thereby improving the convenience and richness of live streaming.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING ZITIAO NETWORK TECH CO LTD
- Filing Date
- 2026-03-13
- Publication Date
- 2026-05-29
AI Technical Summary
In existing live streaming solutions, the scripts generated by machine learning are rather rigid and have poor accuracy, which affects the user experience.
During the live stream, multiple live segments are received, each associated with a different object, and presented in sequence. A machine learning model is used to generate different descriptions associated with each object.
It reduces the reliance on human resources for live streaming, increases the convenience and richness of live streaming, and enhances the user experience.
Smart Images

Figure CN122120480A_ABST
Abstract
Description
Technical Field
[0001] The examples in this article generally relate to the field of computer science, and in particular to methods and electronic devices used for live streaming presentation. Background Technology
[0002] With the development of computer technology, various forms of live streaming can provide people with information, education, entertainment, and other content. Live streams can be associated with multiple objects, allowing viewers to learn about various products based on the content. Conventional solutions typically require the host to manually write the live stream script and conduct the broadcast. There is a desire to provide richer and more diverse live streaming options. Summary of the Invention
[0003] In a first aspect, a method for live streaming is provided. The method includes: during a live stream, receiving multiple live stream segments associated with multiple objects, wherein at least two live stream segments associated with a first object include different descriptions of the first object; and presenting the multiple live stream segments sequentially.
[0004] In a second aspect, an electronic device for live streaming is provided. The electronic device includes: a receiving device configured to receive, during a live stream, a plurality of live stream segments associated with a plurality of objects, wherein at least two live stream segments associated with a first object include different descriptions of the first object; and an output device configured to present the plurality of live stream segments sequentially.
[0005] This approach can improve the convenience and reduce the difficulty of live streaming. While ensuring the accuracy of content related to the live stream, it can also enrich the live streaming methods and improve the user experience.
[0006] It should be understood that the content described in this section is not intended to limit the key or important features of the examples in this article, nor is it intended to restrict the scope of the solution. Other features will become readily apparent from the following description. Attached Figure Description
[0007] The above and other features, advantages, and aspects of the various examples herein will become more apparent when taken in conjunction with the accompanying drawings and the following detailed description. In the accompanying drawings, the same or similar reference numerals denote the same or similar elements, wherein: Figure 1 A schematic diagram of the example environment is shown; Figure 2A A schematic diagram of an example architecture for assisting live streaming is shown, based on some examples; Figure 2B A schematic diagram of an example architecture of an electronic device is shown, depending on several scenarios. Figure 3A schematic diagram of an example architecture for live streaming presentation is shown, based on some examples; Figure 4 The following are some example signaling streams used for live streaming presentation; Figure 5 Flowcharts showing some examples of sample methods for live streaming presentation are provided; and Figure 6 A block diagram of an electronic device capable of implementing multiple illustrative examples is shown. Detailed Implementation
[0008] The examples in the text will now be described in more detail with reference to the accompanying drawings. While some examples are shown in the drawings, it should be understood that solutions can be implemented in various forms and should not be construed as limited to the examples presented herein. Rather, these examples are provided to provide a more thorough and complete understanding of the solutions. It should be understood that the drawings and examples in this document are for illustrative purposes only and are not intended to limit the scope of protection of the solutions.
[0009] In the description of examples in this document, the term "including" and similar terms should be understood as open inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "an example" or "the example" should be understood as "at least one example". The term "some examples" should be understood as "at least some examples". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. The term "trigger" refers to one or more interactive actions by a user on a terminal device. Further, these interactive actions may be triggered within the same user interface / pop-up window or within different user interfaces / pop-up windows. There is no limitation in this respect. Other explicit and implicit definitions may also be included below. It should be noted that, in this document, unless explicitly stated otherwise, performing a step in "response to A" does not mean that the step is performed immediately after "A", but may include one or more intermediate steps.
[0010] The examples in this article may involve user data, data acquisition, and / or use. All of these aspects comply with relevant laws, regulations, and rules. In the examples presented here, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, when implementing each example, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained through appropriate means, in accordance with relevant laws and regulations. The specific methods of notification and / or authorization can vary depending on the actual situation and application scenario; the scope of the solution is not limited in this regard.
[0011] In this manual and the sample solutions, any processing of personal information will be conducted only under legal grounds (such as obtaining the consent of the data subject or being necessary for the performance of a contract) and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information beyond what is necessary for basic functions will not affect the user's use of basic functions.
[0012] As used in this paper, the term "model" refers to a model that learns the relationship between inputs and outputs from training data, enabling it to generate corresponding outputs for a given input after training. Model generation can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs using multiple layers of processing units. A neural network model is an example of a deep learning-based model. In this paper, "model" may also be referred to as a "machine learning model," "learning model," "machine learning network," or "learning network," and these terms are used interchangeably.
[0013] A neural network is a machine learning network based on deep learning. A neural network processes input and provides a corresponding output, typically consisting of an input layer, an output layer, and one or more hidden layers between the input and output layers. Neural networks used in deep learning applications often include many hidden layers, thus increasing the network's depth. The layers of a neural network are connected sequentially, so that the output of the previous layer is provided as the input to the next layer. The input layer receives the input to the neural network, while the output layer's output serves as the final output. Each layer of a neural network includes one or more nodes (also called processing nodes or neurons), each node processing the input from the layer above.
[0014] As mentioned above, live streaming can be associated with multiple objects, allowing viewers to learn about these products based on the content. Conventional solutions typically require the host to manually write the live stream script and conduct the broadcast. While some methods have been proposed to generate live stream scripts using machine learning models, these methods often produce rigid scripts with repetitive descriptions of the same object. Furthermore, the accuracy of existing scripts generated using machine learning models is relatively poor, which negatively impacts the user experience for live stream viewers.
[0015] In view of this, an improved live streaming presentation scheme is proposed. The scheme includes: during the live stream, receiving multiple live stream segments, each associated with a multiple object, wherein at least two live stream segments associated with a first object include different descriptions of the first object; and presenting the multiple live stream segments sequentially.
[0016] This method allows devices to automatically present live stream segments related to the subject, reducing reliance on human intervention and improving the convenience and difficulty of live streaming. It also ensures the accuracy of the live stream content while enriching the streaming experience and enhancing user experience.
[0017] The following sections, in conjunction with accompanying diagrams, further illustrate various examples of solutions used for live streaming.
[0018] Figure 1 A schematic diagram of example environment 100 is shown. (e.g.) Figure 1 As shown, in example environment 100, one or more users 110-1, 110-2, 110-3, ..., 110-N can watch live streams and participate in live stream sessions (also referred to as "live stream events") through their respective associated terminal devices 120-1, 120-2, ..., 120-N. User 110 can also be referred to as a participant. For ease of discussion, users 110-1, 110-2, ..., 110-N can be collectively referred to as user 110 or individually, and terminal devices 120-1, 120-2, ..., 120-N can be collectively referred to as terminal device 120 or individually. Among all users 110, one or more users, such as user 110-1, can establish a live stream session and initiate a live stream through their associated terminal device 120-1.
[0019] In some scenarios, the user 110-1 who initiates the live stream can also be referred to as the host of the live stream session, a host participant, a live stream user, or a live stream administrator. One or more other users 110 may include guest participants 110-N, representing users who participate in the live stream interaction initiated by the host (e.g., live chat). In addition to guest participants in the live stream interaction, other users 110 may also include viewers, listeners, or viewers of the live stream session.
[0020] In some examples, terminal device 120 may have an application installed that can provide live streaming services, or access a website that can provide live streaming services. User 110 can operate terminal device 120 to access the corresponding application or website. Accordingly, terminal device 120 can display a corresponding live streaming interface, which can provide live streaming content of the live event, such as audio content or video live streaming content.
[0021] In some examples, terminal device 120 can also communicate with server 130 via network 132 to provide live streaming services. Server 130 can also provide functions such as application or website management, configuration, and maintenance. Server 130 can also store data generated during the live streaming session, including live interaction data generated during live interaction.
[0022] Terminal device 120 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable gaming terminals, VR / AR devices, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some examples, terminal device 120 may also support any type of user-facing interface (such as "wearable" circuitry).
[0023] Server 130 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. Server 130 may include, for example, computing systems / servers such as mainframes, edge computing nodes, computing devices in cloud environments, etc.
[0024] A communication connection can be established between server 130 and terminal device 120. This communication connection can be established via wired or wireless means. The communication connection can include, but is not limited to, Bluetooth, mobile network, Universal Serial Bus (USB), and Wireless Fidelity (WiFi) connections. In some examples, server 150 and electronic device 110 can interact via a communication connection.
[0025] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.
[0026] Figure 2A A schematic diagram of an example architecture 200A for providing content is shown, based on some examples. Figure 2B A schematic diagram of an example architecture of an electronic device 210 according to some scenarios is shown. For ease of discussion, it will be combined with... Figure 1 Example environment 100 is shown to describe Figure 2A The example architecture shown is 200A and Figure 2BThe electronic device 210 shown is merely illustrative. Figure 2A and Figure 2B As shown, electronic device 210 may include controller 212, display 214, and speaker 216. Controller 212, display 214, and speaker 216 may be communicatively connected via, for example, a bus or other means. In some examples, display 214 may be used to display an interface related to the live event. For example, display 214 may display a control interface for controlling the live event. This control interface may be used to control the flow or content of the live event, such as object display, object switching, or other live-related operations.
[0027] It should be noted that electronic device 210 can be independent of terminal device 120-1, or it can be integrated into terminal device 120-1 as a piece of hardware. When electronic device 210 is integrated into terminal device 120-1 as hardware, server 220 and server 130 can be the same server. In some examples, when electronic device 210 is independent of terminal device 120-1, the two do not directly interact via a communication connection. Unless otherwise specified, example architecture 200 is described using the example of electronic device 210 being independent of terminal device 120-1.
[0028] like Figure 2A As shown, electronic device 210 can receive requests for live events, such as requests for live content of the event. In some examples, electronic device 210 can present candidate objects related to the live event; these candidate objects may include objects suitable for live streaming during the event. User 110-1 can select multiple objects from the candidate objects (these objects may also be referred to as live objects). Objects may include various virtual or physical objects, such as product objects.
[0029] In some examples, electronic device 210 can serve as smart hardware for assisting live streaming, such as assisting the broadcaster in playing and controlling audio narration during a live stream. For ease of description, electronic device 210 may also be referred to herein as a live streaming assisting device, live streaming equipment, smart hardware, or any other suitable name.
[0030] In some examples, such as Figure 2BAs shown, electronic device 110 may further include input device 222, which can be communicatively connected to controller 212 via, for example, a bus or other means. Input device 222 can be used to input data or instructions to electronic device 210. Input device 222 can include any suitable hardware device, such as, but not limited to, touch devices, mice, keyboards, buttons, image capture devices, or audio capture devices. User 110-1 can operate input device 222 to send requests to controller 212. Controller 212 can respond to the request by controlling display 214 to present candidate objects. It should be understood that input device 222 here and input device 650 below can be implemented as the same or different input devices.
[0031] If electronic device 210 receives selections of multiple objects from the candidate objects, electronic device 210 can provide at least multiple live stream segments associated with these multiple objects, each object may have at least one live stream segment associated with it. In some examples, electronic device 210 may also provide other content, such as opening content, ending content, transition content for different live stream segments, etc. It is understood that multiple live stream segments, opening content, ending content, transition content, etc., may all include at least one of text and audio. Electronic device 210 may, for example, display text via display 214 and play audio corresponding to the displayed text via speaker 216. For example, controller 212 may send a playback instruction for the live stream segment of the selected object to speaker 216 in response to the selection of one of the multiple objects. Speaker 216 may play the audio corresponding to the object in response to the playback instruction.
[0032] In some examples, terminal device 120-1 can capture content provided by electronic device 210 (i.e., multiple live stream segments, opening content, ending content, transition content, etc.). For example, terminal device 120-1 can directly capture audio played by electronic device 210. In other examples, terminal device 120-1 can also capture content provided by electronic device 210 via a user. For example, if electronic device 210 provides text, terminal device 120-1 can capture video or audio of the broadcaster explaining the text. Terminal device 120-1 can provide live streaming services by communicating with server 130. For example, terminal device 120-1 can push the captured content as live stream content to server 130.
[0033] In this way, electronic devices 210 can provide content to assist users 110-1 (such as merchants) in conducting live broadcasts, reducing reliance on human labor and lowering the barrier to entry for live broadcasting. Furthermore, it helps improve the quality of live broadcasts, thereby enhancing the overall effect.
[0034] In some examples, the playback of live audio can also be accompanied by video. For example, electronic device 210 can also synchronously display video on display 214, which may include a virtual anchor acted by a virtual object. The virtual anchor's lip movements can be synchronized with the live audio. In some examples, combined with... Figure 2B As shown, electronic device 210 may also include transceiver 218, which can communicate with server 220 (e.g., obtain multiple live stream segments, opening content, ending content, transition content, etc. from server 220) to provide auxiliary live streaming services. The transceiver 218 may include any suitable hardware unit with communication capabilities; for example, the transceiver 218 may include the communication unit described below.
[0035] In some examples, server 220 may utilize machine learning model 230 (which may include one or more machine learning models, such as machine learning model 230-1, machine learning model 230-2, ..., machine learning model 230-N, etc., where N is a positive integer. For ease of description, one or more machine learning models are collectively referred to as machine learning model 230 in this document) to provide auxiliary live streaming services. As an example, server 220 may utilize machine learning model 230 to generate multiple live streaming segments, opening content, ending content, transition content, etc., to improve the generation quality and efficiency of this content.
[0036] Machine learning model 230 can be of different types. In some examples, machine learning model 230 can be built based on a language model (LM). The machine learning model used is a content-generative model, capable of generating corresponding outputs based on model inputs. A language model-based machine learning model can receive text-modal model inputs (e.g., natural language and / or machine language) and / or non-text-modal model inputs (e.g., images, speech, video, etc.), and can generate the desired output based on the model inputs and prompts. Prompts can indicate user needs, guiding the machine learning model to generate model outputs that match user needs. In other examples, machine learning model 230 can be built based on a multimodal model.
[0037] In other examples, machine learning model 230 can be a speech-related model, including an Automatic Speech Recognition (ASR) model or a Text-to-Speech (TTS) model. An ASR model takes speech as input and outputs text. A TTS model takes text as input and outputs the corresponding speech. It should be understood that the machine learning model 230 described above is merely exemplary. In practical applications, any suitable model architecture can be used. It should also be noted that although machine learning model 230 is deployed on server 220 in Figure 2, it can also be deployed on electronic device 210, terminal device 120-1, edge devices, or other devices. This document does not limit the deployment location of machine learning model 230.
[0038] The following is combined with Figure 3 and Figure 4 To describe an example of auxiliary live streaming. Figure 3 A schematic diagram of an example architecture 300 for live streaming presentation is shown, based on some examples. Figure 4 Example signaling stream 400 for live streaming is shown. Electronic device 310 in example architecture 300 can be a terminal device 120-1 integrating electronic device 210, or it can exist independently of terminal device 120-1. Correspondingly, server 320 in example architecture 300 can be server 130 or server 220. The text generation model 330 and speech synthesis model 340 in example architecture 300 can be machine learning models in machine learning model 230. Server 320 may include interface 322, service node 324, and queue 326.
[0039] Electronic device 310 may have a hardware button associated with live streaming / live streaming assistance or present operation controls associated with live streaming / live streaming assistance. Electronic device 310 may determine (411) that it has received an instruction associated with live streaming in response to a triggering of the hardware button or operation control. For example, if electronic device 310 is terminal device 120-1, electronic device 310 may receive a live streaming instruction. If electronic device 310 is electronic device 210, electronic device 310 may receive a live streaming assistance instruction.
[0040] Electronic device 310 may, in response to a live broadcast instruction or a live broadcast assistance instruction, send a (412) request to the interface of server 320. This request is used to obtain object information for each candidate object, and may include, for example, an identifier of the live broadcast event. This identifier may include any appropriate information capable of identifying the live broadcast event, such as, but not limited to, the live broadcast event number, the live broadcast event name, the live broadcast account number, the live broadcast account username, etc. Candidate objects may include any number of appropriate objects. In some examples, candidate objects may include multiple products. It should be noted that candidate objects may differ when broadcasting to different fields. This document does not limit the type of live broadcast field or candidate object. For example, if the live broadcast field is the gaming field, candidate objects may include multiple game objects; if the live broadcast field is life services, candidate objects may include multiple service items, etc.
[0041] Interface 322 can determine candidate objects associated with the live event based on the identifier of the received live event. Electronic device 310 can receive (413) object information of each candidate object from interface 322. The object information of each candidate object can include at least the identifier information, attribute information, and source information of the candidate object. Taking a product as an example, the identifier information of the product can include the name, number, icon, image, link, etc. of the product, and the attribute information of the product can include the size, weight, price, instructions for use, precautions, etc. of the product.
[0042] As an example, interface 322 can determine the products displayed in the live stream event (e.g., products located on the shelf of the live stream event) based on the event's ID. Then, interface 322 can feed back the product's object information to electronic device 310. Optionally and / or additionally, interface 322 can also determine the products displayed in the live stream account's product showcase based on the account information of the live stream account. Of course, the above candidate objects are merely exemplary. The method of determining candidate objects may differ in different live streaming scenarios.
[0043] Electronic device 310 can present object information for each candidate object. Not all candidate objects need to be explained; in this case, electronic device 310 can receive (414) the user's selection of multiple objects from the candidate objects. For example, such as... Figure 3As shown, electronic device 310 can present interface 312 and display object information of each candidate object via interface 312. A user (e.g., user 110-1) can select multiple objects from the candidate objects. In response to receiving the selection of multiple objects, electronic device 310 can send (415) the identifiers of each of these multiple objects to interface 322. Interface 322 can then send (416) a request to service node 324, which may, for example, request service node 324 to obtain content associated with the live event. This content may include, for example, multiple live segments associated with the multiple objects selected by the user, the opening content of the live event, the ending content of the live event, transition content for transitioning between two live segments, etc., wherein each object may be associated with at least one live segment. In some examples, at least one of the multiple objects (e.g., the first object) is associated with at least two live segments. The request may, for example, include at least the identifiers of each of the multiple objects received by interface 322.
[0044] Service node 324 can provide the acquired content to queue 326. Electronic device 310 can obtain (422) the content from queue 326. Thus, the computing power of server 320 can be used to obtain content associated with the live event, which is beneficial to improving the quality of the content. It should be noted that the content associated with the live event can include text and / or audio. For example, a live segment associated with an object can include descriptive text and / or descriptive audio of the object, and the opening content can include opening text and / or opening audio, etc.
[0045] If the content associated with the live event includes text, service node 324 can obtain (417) the text and provide (418) the obtained text to queue 326. Service node 324 can obtain the text in any appropriate manner; for example, service node 324 can use text generation model 330 to obtain the text. Service node 324 can provide at least several object identifiers to text generation model 330 to instruct text generation model 330 to generate text associated with the live event. Service node 324 can obtain the text generated by text generation model 330 and provide it to queue 326. It is understood that the text generated by text generation model 330 can include various types of text, including but not limited to opening text 332, closing text 334, text associated with objects 336, etc.
[0046] If the content associated with the live event includes audio, service node 324 can further perform the steps shown in box 419. In this case, service node 324 can acquire (420) the audio and provide (421) the acquired audio to queue 326. Service node 324 can acquire the audio in any suitable manner; as an example only, service node 324 can utilize speech synthesis model 340 to acquire the audio. Specifically, service node 324 can provide the text generated by text generation model 330 to speech synthesis model 340 to instruct speech synthesis model 340 to generate audio that matches the text. Service node 324 can then acquire the audio generated by speech synthesis model 340 and provide it to queue 326.
[0047] As an example, a dual-channel can be established between electronic device 310 and queue 316, and electronic device 310 can receive text and / or audio from queue 326 based on the dual-channel. After acquiring the text and / or audio, electronic device 310 can provide (423) the text and / or audio. Specifically, electronic device 310 can have a display, which can present interface 312, and electronic device 310 can present text through interface 312. Electronic device 310 can have a speaker 314 and can play audio through speaker 314. Taking electronic device 310 as terminal device 120-1 as an example, terminal device 120-1 can directly present the acquired text in a specific area of the live broadcast interface or directly play the acquired audio through speaker 314. It can be understood that the text presented or the audio played by terminal device 120 can also be different depending on the object currently presented in the live broadcast. Taking electronic device 310 as electronic device 210 as an example, electronic device 310 can directly provide text or audio, and terminal device 120 can collect the text or audio provided by electronic device 310 and broadcast the collected content as live broadcast content.
[0048] It should be noted that steps 411 to 421 described above can be performed before the live stream begins or during the live stream. That is, in some examples, electronic device 310 can provide server 320 with information related to the live stream (which may include at least information related to the live stream event, object information for at least one object, etc.) before the live stream begins, instructing server 320 to pre-generate the content related to the live stream event as shown above before the live stream begins. In some examples, electronic device 310 can obtain the content related to the live stream event as shown above from server 320 before the live stream begins and store it locally. In some examples, electronic device 310 can also obtain the content related to the live stream event as shown above from server 320 in real time during the live stream.
[0049] As mentioned earlier, the text obtained by service node 324 may include at least the opening text 332 of the live stream, the ending text 334 of the live stream, and text 336 associated with each of the multiple objects. In some examples, the text may also include other text 338, which may include any appropriate text.
[0050] The text 336 associated with an object may include descriptive text or structured text for the object. For each of the multiple objects, server 320 may generate the descriptive text for that object in any appropriate manner. In some examples, server 320 may utilize a text generation model to directly generate multiple descriptive texts for the object based on its object information, wherein the different descriptive texts are at least partially different from each other.
[0051] In some examples, server 320 can utilize a text generation model to generate at least one structured text for the object based on its object information. If multiple structured texts are generated, the different structured texts within these multiple structured texts are at least partially different. Each structured text may include at least one placeholder. Each placeholder may indicate the type of attribute information to be filled in. That is, when generating the structured text for the object, the model can be instructed to leave some information blank (i.e., by adding placeholders). This information will be filled in later.
[0052] The structured text of an object can be used to determine the object's descriptive text. As an example, since placeholders in the structured text indicate the type of attribute information to be filled, server 320 can determine the complete descriptive text by filling at least one placeholder in the structured text with information (which may include, for example, attribute values of the object's attribute information corresponding to the structured text). It is understood that each structured text can be filled with information to generate at least one descriptive text. Taking, for example, an object's structured text containing a placeholder indicating that attribute information of the component type should be filled, if the component type attribute information indicates that the object's components include one important component 1 and two special components 2 and 3, that is, the attribute values of the attribute information include component 1, component 2, and component 3, then three descriptive texts for the object can be obtained by filling the placeholders with these three attribute values respectively.
[0053] The object's attribute information and its values can be provided by the object's provider. For example, server 320 can obtain this information from the object's provider. Taking a product as an example, the product's attribute information could include its size, price, weight, ingredients, precautions, usage instructions, etc. Therefore, based on the type of attribute information indicated by each placeholder, the corresponding attribute value is filled into each placeholder. This ensures that the object's attribute values are always correct, preventing errors in important information due to model illusions during content generation.
[0054] In some examples, other text 338 may include transition text, interactive text, etc. In some scenarios, transition text may also be referred to as transition words. In some examples, server 320 may pre-generate multiple transition texts and send them to electronic device 310. It is important to note that each transition text is not associated with a specific object. In some examples, interactive text may include structured text associated with the interaction. Similar to the structured text of the object mentioned above, the structured text associated with the interaction may also include at least one placeholder, each placeholder indicating the type of information to be filled. It can be understood that transition text may be pre-generated before the live broadcast begins, while interactive text is generated in real time during the live broadcast based on the interactions of users watching the live broadcast.
[0055] In some examples, the interaction may include at least one of a user joining the live stream and the user interacting with objects in the live stream. Objects in the live stream may be multiple objects selected above (e.g., multiple products), or user objects (e.g., the streamer) or user groups (e.g., shops, businesses, etc.) corresponding to the live stream. User interactions with objects in the live stream may include browsing, adding to cart, placing orders, favorites, forwarding, etc., of any of the multiple objects, or following and commenting on user objects or user groups. Different interactions may correspond to different structured texts, and each interaction may correspond to multiple structured texts. For example, during the live stream, the electronic device 310 or server 320 may, based on real-time interactions, populate information into the corresponding structured text. For example, if the interaction is a user joining the live stream, the placeholder in the corresponding structured text may indicate that the information to be filled is a user identifier (e.g., user name), which may be, for example, "Welcome XX to the live stream." The electronic device 310 or server 320 can generate descriptive text matching the real-time interaction by filling the user identifier corresponding to the interaction into the "XX" position of the structured text (e.g., the descriptive text may be "Welcome User A to follow the streamer").
[0056] For example, if the interaction is a user's interaction with an object in the live stream, the placeholders in the corresponding structured text can indicate the information to be filled in as the user identifier (e.g., user name), the object being interacted with, and the interaction content. For example, it could be "Thank you AA for CC'ing BB". Taking the specific interaction as a user following the streamer as an example, the electronic device 310 or server 320 can generate descriptive text that matches the real-time interaction by filling the user identifier corresponding to the interaction into the "AA" position of the structured text, filling the streamer identifier into the "BB" position of the structured text, and filling the interaction details (e.g., the text "follow") into the "CC" position of the structured text. (For example, the descriptive text could be "Thank you user A for following the streamer").
[0057] It should be noted that the text 336 associated with the object obtained by the electronic device 310 from the server 320 can be the object's descriptive text or the object's structured text, and other text 338 obtained may include interactive text or structured text corresponding to the interaction. In some examples, if the electronic device 310 obtains structured text, the electronic device 310 can pad the structured text to determine the complete descriptive text or interactive text.
[0058] In some examples, after receiving text, electronic device 310 can automatically generate corresponding audio based on the text. Of course, in some examples, electronic device 310 can also directly obtain audio from server 320. In some examples, electronic device 310 can stream content from server 320. Taking only object description text or description audio as an example, server 320 can send the generated description text or description audio to electronic device 310 in response to the generation of a description text or description audio for an object. In this case, electronic device 310 can obtain one description text or description audio for one object at a time. In some examples, electronic device 310 can obtain content from server 320 in batches. For example, electronic device 310 can obtain a predetermined number of texts or audios from server 320 each time.
[0059] In some examples, electronic device 310 can receive user feedback on content associated with a live event. If the user is satisfied with the content quality, they can trigger a confirmation. In response to receiving the confirmation, electronic device 310 can send a confirmation command to server 320. If the user is dissatisfied with the content quality, they can trigger an update request. In response to receiving this update request, electronic device 310 can send an update command to server 320. In this case, service node 324 can update the content and resend the updated content to electronic device 310. Of course, if the user is not satisfied with the content quality, they can also provide feedback. Server 220 can update the content based on this feedback.
[0060] It should be noted that although the foregoing examples illustrate how electronic device 310 obtains text or audio from server 320, in real-world scenarios, electronic device 310 can also generate text or audio itself. For instance, electronic device 310 may have a locally deployed machine learning model, which it can use to generate text or audio. The examples in this article do not limit this to such cases.
[0061] Figure 5 A flowchart of an example method 500 for live streaming presentation is shown, based on some examples. Method 500 can be implemented at electronic device 310.
[0062] In box 510, during live streaming, electronic device 310 receives multiple live stream segments associated with multiple objects, wherein at least two live stream segments associated with a first object include different descriptions of the first object. These multiple objects may, for example, be multiple objects selected by the user from candidate objects. In some examples, each object may have at least one associated live stream segment, and a subset of the multiple objects (e.g., including at least the first object) may have at least two associated live stream segments. It is understood that if the same object has at least two associated live stream segments, the descriptions of the objects included in any two different live stream segments associated with that object are at least partially different. The description of each object may include descriptive text and / or descriptive audio. As mentioned above, electronic device 310 may receive multiple live stream segments either streaming or in batches.
[0063] In frame 520, electronic device 310 presents multiple live stream segments in sequence.
[0064] In some examples, electronic device 310 can cyclically present multiple live segments based on multiple objects. In each cycle, electronic device 310 can present one live segment corresponding to each object. In some examples, in the first cycle, electronic device 310 can sequentially present the live segments associated with each of the multiple objects. This order can be based on the order of the multiple objects. In the second cycle, for an object among the multiple objects that has at least two associated live segments (e.g., the first object), in some examples, electronic device 310 can present the first live segment in response to the first live segment of the at least two live segments associated with that object not being presented. In other examples, electronic device 310 can present the second live segment of the at least two live segments in response to the at least two live segments associated with that object having been previously presented, the second live segment being presented earlier than at least one other live segment among the at least two live segments.
[0065] Therefore, in the current cycle, the electronic device 310 can prioritize presenting live stream segments that have not been presented before for each object, or prioritize presenting the earliest live stream segment that has been presented for that object. Thus, the live stream segments for the same object are different in two adjacent cycles. This achieves dynamic changes and diversity in the live stream script, effectively simulating the randomness and richness of a real broadcaster, and solving the problem of repetitive and boring long-term live stream content.
[0066] Regarding the specific method of sequentially presenting multiple live stream segments associated with each object, in some examples, the electronic device 310 may display the text of each of the multiple live stream segments in sequence. The text of each live stream segment may, for example, include descriptive text of the object associated with that live stream segment. In other examples, the electronic device 310 may also play audio corresponding to the text of each live stream segment in sequence. The audio of each live stream segment may, for example, include descriptive audio of the object associated with that live stream segment. Of course, in some examples, for each live stream segment, the electronic device 310 may simultaneously display the text of that live stream segment and play the audio of that live stream segment.
[0067] In some examples, to help users quickly understand the objects corresponding to the current live stream segment, the electronic device 310 can respond to the presentation of the live stream segment by providing objects associated with the presented live stream segment in the corresponding live stream interface. For example, still using a product as an example, the electronic device 310 can respond to the presentation of a product's live stream segment by displaying the product's product card, link, or image in the corresponding live stream interface.
[0068] In some examples, electronic device 310 may also receive content other than multiple live stream segments, including but not limited to opening content, ending content, transition content, interactive content, etc. Each of these contents may include at least one of text and audio. Electronic device 310 may, for example, present opening content in response to the start of a live stream. Electronic device 310 may, for example, present ending content in response to a trigger to end the live stream. Electronic device 310 may, for example, determine that an end trigger to the live stream has been received in response to receiving a user's touch operation on the end control. Electronic device 310 may also determine that an end trigger to the live stream has been received in response to the live stream duration reaching a predetermined duration, the number of loops of multiple live stream segments reaching a predetermined number, the current live stream time reaching a predetermined time point, etc.
[0069] In some examples, electronic device 310 can insert transition content between any two live segments. In some examples, the transition content inserted between different live segments may differ, and these different transition contents may all come from multiple pre-configured transition contents. Similar to live segments, in some examples, electronic device 310 can acquire multiple transition contents and prioritize the presentation of transition contents that have not yet been presented, or prioritize the presentation of the earliest presented transition content among those that have already been presented. This avoids repeatedly presenting the same transition content, thereby improving the user experience of watching live streams.
[0070] In some examples, during the presentation of multiple live stream segments, in response to interactions occurring during the live stream, electronic device 310 can insert the presentation of live stream segments related to the interaction. These interaction-related live stream segments may include interaction-related descriptions, which may include at least one of interactive text and interactive audio. As mentioned earlier, interactions may include at least one of a user joining the live stream and a user interacting with objects in the live stream. As mentioned earlier, the interactive text may be text generated in real-time by populating pre-generated structured text with interactive information. Therefore, by presenting live stream segments that match real-time interactions, the real-time nature and engagement of the live stream presentation can be improved, enhancing the user experience of watching the live stream.
[0071] Regarding the specific methods for presenting interactive live stream segments, in some examples, each of multiple live stream segments related to multiple objects can include multiple statements. Taking electronic device 310 currently presenting a live stream segment as an example, electronic device 310 can respond to the end of the presentation of the first statement currently being presented in that live stream segment by presenting the interactive live stream segment. Electronic device 310 can also respond to the end of the presentation of the interactive live stream segment by continuing to present other statements after the first statement of that live stream segment. In this way, the real-time presentation of interactive live stream segments can be guaranteed without affecting the presentation of other live stream segments.
[0072] Figure 6 A block diagram of an electronic device 600 in which one or more examples may be implemented is shown. It should be understood that... Figure 6 The electronic device 600 shown is merely exemplary and should not be construed as limiting the functionality and scope of the examples described herein. Figure 6 The illustrated electronic device 600 can be used to implement the electronic device 310 discussed above.
[0073] like Figure 6 As shown, electronic device 600 is in the form of a general-purpose electronic device. Components of electronic device 600 may include, but are not limited to, one or more processing units or processors 610, memory 620, storage device 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. Processor 610 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 620. In a multiprocessor system, multiple processors execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 600.
[0074] Electronic device 600 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 600, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 620 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof). Storage device 630 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 600.
[0075] Electronic device 600 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 6 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 620 may include computer program product 625 having one or more program modules configured to perform various methods or actions of various examples.
[0076] The communication unit 640 enables communication with other electronic devices via a communication medium. Additionally, the functionality of the components of the electronic device 600 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the electronic device 600 can operate in a networked environment using logical connections to one or more other servers, networked personal computers, or another network node.
[0077] Input device 650 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 660 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 600 can also communicate with one or more external devices (not shown) via communication unit 640 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 600, or with any device that enables electronic device 600 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).
[0078] A computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. A computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.
[0079] The flowcharts and / or block diagrams of the methods, apparatus, devices, and computer program products referred to herein describe various aspects. It should be understood that each block of the flowcharts and / or block diagrams, as well as combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0080] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0081] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0082] The flowcharts and block diagrams in the accompanying figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to some examples. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the figures. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0083] Various examples have been described above. The foregoing descriptions are exemplary and not exhaustive, nor are they limited to the proposed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the proposed implementations.
Claims
1. A method for live streaming presentation, comprising: During the live stream, multiple live stream segments are received, each segment being associated with a multiple object. At least two live stream segments associated with a first object include different descriptions of that first object. The multiple live stream segments are presented in sequence.
2. The method according to claim 1, further comprising: In response to the start of the live stream, opening content is presented, wherein the multiple live stream segments are presented after the opening content; as well as In response to the termination of the live stream, end content is displayed.
3. The method according to claim 1, further comprising: Transitional content is inserted between the presentation of two live stream segments.
4. The method of claim 3, wherein the transition content inserted between the two different live stream segments is different.
5. The method of claim 4, wherein the different transition content is derived from a plurality of pre-configured transition content.
6. The method according to claim 1, wherein the plurality of live streaming segments are a plurality of first live streaming segments, and the method further comprises: During the presentation of the multiple first live broadcast segments, In response to an interaction occurring during the live stream, a second live stream segment is inserted, the second live stream segment including a description related to the interaction.
7. The method of claim 6, wherein the presentation of inserting the second live stream segment comprises: In response to the end of the presentation of the first statement, the second live stream segment is presented, wherein the first statement is a part of the first live stream segment presented. The method further includes: In response to the end of the presentation of the second live segment, the remaining statements following the first statement in the first live segment will continue to be presented.
8. The method according to claim 1, further comprising: In response to the presentation of a live segment among the plurality of live segments, a second object is provided in the interface corresponding to the live stream, the second object being associated with the presented live segment.
9. The method of claim 1, wherein presenting the plurality of live stream segments in sequence comprises: In the first loop, the live stream segments associated with each of the multiple objects are presented sequentially; as well as In the second loop, for the first object, In response to the first live segment not being displayed in the at least two live segments, the first live segment is displayed, or In response to the fact that the at least two live segments have been previously presented, a second live segment of the at least two live segments is presented, the presentation time of the second live segment being earlier than at least one other live segment of the at least two live segments.
10. The method of claim 1, wherein the plurality of live stream segments respectively comprise at least one of text and audio, and wherein presenting the plurality of live stream segments in sequence comprises at least one of the following: Display the text of the multiple live stream segments in sequence; or Play the audio corresponding to the displayed text in sequence.
11. The method of claim 1, wherein the at least two live stream segments associated with the first object are generated in the following manner: Based on the object information of the object, at least two texts are generated, the at least two texts being at least partially different; and Based on the at least two texts, generate the at least two live stream segments.
12. The method of claim 11, wherein generating the at least two texts comprises: Based on the object information of the object, at least two structured texts of the object are generated, wherein the different structured texts are at least partially different, and each structured text includes at least one placeholder, the at least one placeholder indicating the type of attribute information; as well as Based on the attribute values of the object's attribute information, placeholders in the at least two structured texts are filled to obtain the at least two texts.
13. The method of claim 12, wherein the attribute information of the object is provided by the provider of the object.
14. An electronic device, comprising: A receiving device is configured to receive multiple live stream segments during a live stream, the multiple live stream segments being associated with multiple objects, wherein at least two live stream segments associated with a first object include different descriptions of the first object; and The output device is configured to present the multiple live stream segments in sequence.
15. The device of claim 14, wherein the plurality of live stream segments each comprise at least one of text and audio, and the output device comprises at least one of: The display is configured to sequentially display the text of the plurality of live stream segments; or The speaker is configured to play audio corresponding to the displayed text in sequence.