Live broadcast method and device, electronic equipment and storage medium

By pre-creating candidate live streaming rooms and matching connections when users visit, the problem of excessively long live streaming room creation time in one-to-one live streaming rooms is solved, improving user experience and saving server resources.

CN121603700APending Publication Date: 2026-03-03BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511909543.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-17
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

In one-on-one live streaming rooms, the existing technology takes too long to create the room, resulting in long waiting times for users and a poor user experience.

Method used

Multiple candidate live streaming rooms are pre-created and matched and connected when a user visits the target page, reducing the waiting time for live streaming room creation.

Benefits of technology

By pre-creating candidate live streaming rooms, the time spent creating live streaming rooms is reduced, the user experience is improved, and server resources are saved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121603700A_ABST
    Figure CN121603700A_ABST
Patent Text Reader

Abstract

The invention provides a live broadcast method and device, electronic equipment and a storage medium, and relates to the technical field of data processing, in particular to the fields of intelligent search, information flow and the like. According to the specific implementation scheme, in response to an access request received from a client, page information for a target page in the access request is determined; in response to the detection that a plurality of pre-created candidate live broadcasting rooms comprise candidate live broadcasting rooms matched with the page information, taking the candidate live broadcasting rooms matched with the page information as target live broadcasting rooms; any candidate live broadcast room in the plurality of candidate live broadcast rooms is a live broadcast room for performing one-to-one interaction with the client through the virtual image; establishing a connection between the client and the target live broadcast room; and sending a video stream for the target live broadcast room to the client which establishes the connection with the target live broadcast room.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of data processing technology, and in particular to the fields of intelligent search and information flow. More specifically, this disclosure provides a live streaming method, apparatus, electronic device, storage medium, and computer program product. Background Technology

[0002] In scenarios such as recommendation, advertising, and e-commerce, using one-on-one live streaming to interact with users can enhance the user experience.

[0003] Unlike one-to-many live streaming rooms, where users typically enter a live stream after it has started by browsing the feed or clicking on it, and multiple users can join the same room, one-to-one live streaming rooms serve only one user. Therefore, users usually click on an option when they have interactive needs such as asking questions or learning more, and then the server creates a separate live streaming room for that user and renders a virtual avatar to interact with them.

[0004] In the architecture of real-time rendering live streaming services, the initialization process, such as the creation of one-to-one live streaming rooms, takes a long time, resulting in an excessively long overall waiting time for users from accessing the service to establishing a connection with the live streaming room, leading to a poor user experience. Summary of the Invention

[0005] This disclosure provides a live streaming method, apparatus, electronic device, storage medium, and computer program product.

[0006] According to one aspect of this disclosure, a live streaming method is provided, comprising: in response to receiving an access request from a client, determining page information for a target page in the access request; in response to detecting that a plurality of pre-created candidate live streaming rooms include a candidate live streaming room that matches the page information, selecting the candidate live streaming room that matches the page information as the target live streaming room; any one of the plurality of candidate live streaming rooms being a live streaming room for one-to-one interaction with the client through a virtual avatar; establishing a connection between the client and the target live streaming room; and sending a video stream for the target live streaming room to the client after establishing a connection with the target live streaming room.

[0007] According to another aspect of this disclosure, a live streaming method is provided, comprising: in response to detecting that an object accesses a target page, sending an access request to a server, the access request including page information of the target page; establishing a connection with a target live streaming room among a plurality of pre-created candidate live streaming rooms; matching the target live streaming room with the page information; any one of the plurality of candidate live streaming rooms being a live streaming room for one-to-one interaction with a client through a virtual avatar; acquiring a video stream from the established target live streaming room; and displaying the video stream.

[0008] According to another aspect of this disclosure, a live streaming device is provided, comprising: a page information determination module, a live streaming room determination module, a first connection module, and a sending module. The page information determination module is used to determine page information for a target page in response to receiving an access request from a client. The live streaming room determination module is used to select the candidate live streaming room matching the page information as the target live streaming room in response to detecting that a plurality of pre-created candidate live streaming rooms include a candidate live streaming room matching the page information. Any one of the plurality of candidate live streaming rooms is a live streaming room used for one-to-one interaction with the client through a virtual avatar. The first connection module is used to establish a connection between the client and the target live streaming room. The sending module is used to send a video stream for the target live streaming room to the client after establishing a connection with the target live streaming room.

[0009] According to another aspect of this disclosure, a live streaming device is provided, comprising: an access request sending module, a second connection module, a video stream acquisition module, and a video stream display module. The access request sending module is used to send an access request to a server in response to detecting that an object is accessing a target page. The access request includes page information of the target page. The second connection module is used to establish a connection with a target live streaming room among a plurality of pre-created candidate live streaming rooms. The target live streaming room is matched with the page information. Any one of the plurality of candidate live streaming rooms is a live streaming room used for one-to-one interaction with a client through a virtual avatar. The video stream acquisition module is used to acquire the video stream of the target live streaming room with the established connection. The video stream display module is used to display the video stream.

[0010] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the methods provided in this disclosure.

[0011] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause a computer to perform the methods provided in this disclosure.

[0012] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the methods provided in this disclosure.

[0013] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0014] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0015] Figure 1 This is a schematic diagram illustrating an application scenario of the live streaming method and apparatus according to embodiments of this disclosure;

[0016] Figure 2 This is a schematic flowchart of a live streaming method according to an embodiment of the present disclosure;

[0017] Figure 3 This is a schematic diagram illustrating the principle of creating a room pool according to an embodiment of this disclosure;

[0018] Figure 4 This is a schematic diagram illustrating the principle of a live streaming method according to an embodiment of this disclosure;

[0019] Figure 5 This is a schematic flowchart of a live streaming method according to an embodiment of the present disclosure;

[0020] Figure 6 This is a schematic diagram illustrating the generation and transmission of data packets according to embodiments of this disclosure;

[0021] Figure 7 This is a schematic diagram illustrating the linkage of key areas according to an embodiment of this disclosure;

[0022] Figure 8 This is a schematic diagram illustrating the background area linkage according to an embodiment of the present disclosure;

[0023] Figure 9 This is a schematic structural block diagram of a live streaming device according to an embodiment of the present disclosure;

[0024] Figure 10 This is a schematic structural block diagram of a live streaming device according to an embodiment of the present disclosure; and

[0025] Figure 11 This is a structural block diagram of an electronic device used to implement the live streaming method of the embodiments of this disclosure. Detailed Implementation

[0026] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0027] The collection, storage, use, processing, transmission, provision, and disclosure of any type of information, such as user personal information, in this technical solution comply with relevant laws and regulations and do not violate public order and good morals. In this technical solution, user authorization or consent is obtained before acquiring or collecting user personal information.

[0028] In some technical solutions, a static digital avatar image can be used as a placeholder on the front-end page before the live stream is fully created, along with prompts such as "Connecting...", to alleviate user anxiety while waiting. However, this solution only alleviates user anxiety through static images and prompts; it does not actually solve the problem of slow live stream loading.

[0029] In other technical solutions, the creation process for a live stream can be initiated in advance on the parent page (e.g., the live stream list page) when a user enters the live stream. This way, the live stream is ready when the user enters, saving the user's waiting time for room creation. However, this solution is only suitable for one-to-many live streams, not one-to-one live streams. The reason is that in a one-to-one live stream scenario, a separate live stream room needs to be created for each user. Creating the live stream room in advance on the parent page would result in the creation of a large number of empty live stream rooms, severely consuming server resources.

[0030] This disclosure aims to provide a live streaming method that pre-creates candidate live streaming rooms and allocates one of these rooms as the target live streaming room to the client when interaction is needed, instead of creating a new room only when user interaction is required. This saves time in creating live streaming rooms and improves user experience by reducing user waiting time. Furthermore, in some embodiments, the estimated usage of candidate live streaming rooms can be determined based on historical data, and then candidate live streaming rooms can be created according to the estimated usage, thus avoiding excessive consumption of server resources.

[0031] The technical solutions provided in this disclosure will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0032] Figure 1 This is a schematic diagram illustrating an application scenario of the live streaming method and apparatus according to embodiments of this disclosure.

[0033] It is important to note that Figure 1 The examples shown are merely examples of system architectures that can be applied to the embodiments of this disclosure, in order to help those skilled in the art understand the technical content of this disclosure, but do not mean that the embodiments of this disclosure cannot be used in other devices, systems, environments or scenarios.

[0034] like Figure 1As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, and 103, a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between the terminal devices 101, 102, and 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0035] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Terminal devices 101, 102, and 103 can be various electronic devices with displays and web browsing capabilities, including but not limited to smartphones, tablets, laptops, and desktop computers.

[0036] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using terminal devices 101, 102, and 103 (for example only). The backend management server can analyze and process data such as received user requests and feed the processing results back to the terminal devices.

[0037] For example, server 105 receives an access request from terminal device 101 and determines the page information for the target page in the access request. Then, server 105 determines whether any of the pre-created candidate live stream rooms include one that matches the page information. If so, the candidate live stream room matching the page information is selected as the target live stream room. Afterwards, terminal device 101 establishes a connection with the target live stream room, and server 105 sends a video stream for the target live stream room to the connected terminal device 101. Terminal device 101 then acquires and displays the live video stream.

[0038] It should be noted that the live streaming method provided in this embodiment can generally be executed by server 105 or terminal devices 101, 102, and 103. Accordingly, the live streaming device provided in this embodiment can generally be set in server 105 or terminal devices 101, 102, and 103.

[0039] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0040] Figure 2 This is a schematic flowchart of a live streaming method according to an embodiment of the present disclosure.

[0041] like Figure 2 As shown, the live streaming method 200 can be executed by the server, and the live streaming method 200 includes operations S210 to S240.

[0042] In operation S210, in response to receiving an access request from the client, the page information for the target page in the access request is determined.

[0043] For example, when a client detects that an object is accessing a target page, it sends an access request to the server. Upon receiving the access request, the server extracts the page information of the target page from it. The target page can be an advertisement page, a resource page, etc., and the page information can be the URL (Uniform Resource Locator) of the target page.

[0044] In operation S220, in response to detecting that multiple pre-created candidate live rooms include candidate live rooms that match the page information, the candidate live room that matches the page information is selected as the target live room; any one of the multiple candidate live rooms is a live room used for one-to-one interaction with the client through a virtual avatar.

[0045] For example, several candidate live streaming rooms can be pre-created, and a correspondence between these rooms and the target page can be established. One page can correspond to one or more candidate live streaming rooms. A candidate live streaming room is a one-to-one live streaming room, primarily consisting of a virtual avatar and a scene. The virtual avatar can be a digital human, and the scene can include the environment in which the virtual avatar is located. It can be determined whether, among the multiple candidate live streaming rooms, there exists a candidate live streaming room that corresponds to the target page and is in an idle state. If so, that candidate live streaming room is used as the target live streaming room.

[0046] Using S230, establish a connection between the client and the target live streaming room.

[0047] For example, the server returns the identifier of the target live stream to the client, and then the client sends a connection request to the server, which includes the identifier of the target live stream. Upon receiving the connection request, the server's live streaming service establishes a connection with the client.

[0048] In operation S240, a video stream for the target live room is sent to the client after a connection has been established with the target live room.

[0049] For example, the server sends the live video stream from the target room to the client in real time.

[0050] According to the live streaming method provided in this disclosure, the server pre-creates multiple candidate live streaming rooms, which are used for one-to-one interaction between the client and the virtual avatar. After the object accesses the target page, the client sends an access request to the server, which includes the page information of the target page. The server then detects whether the pre-created multiple candidate live streaming rooms include a candidate live streaming room that matches the page information. If so, the matching candidate live streaming room is selected as the target live streaming room. A connection is then established between the client and the target live streaming room, and a video stream for the target live streaming room is sent to the client. The client then displays the video stream of the target live streaming room.

[0051] As can be seen, because multiple candidate live stream rooms are created in advance, and when a user's intention to enter a one-on-one live stream room is detected, the system first checks if the target live stream room exists among the multiple candidate live stream rooms. If so, a connection is directly established with the target live stream room. In this way, the target live stream room is created in advance, rather than when interaction with the user is needed, thus saving the time spent creating the live stream room and improving the user experience by reducing user waiting time.

[0052] Figure 3 This is a schematic diagram illustrating the principle of creating a room pool according to an embodiment of this disclosure.

[0053] In this embodiment, multiple candidate live streaming rooms 303 can be pre-created. During the actual creation process, for any candidate category among at least one candidate category, the estimated usage 302 for the target time period can be determined based on the historical usage 301 of the live streaming rooms of the candidate category within a predetermined time period. Then, based on the estimated usage 302 and the live streaming room configuration information corresponding to the candidate category, candidate live streaming rooms 303 for the candidate category are created.

[0054] For example, candidate categories can be divided according to pages. One page for users to access a live stream might correspond to one candidate category, and there could be one or more live streams within the same candidate category. Alternatively, candidate categories can be divided according to advertisers, products, or other dimensions. For example, one page for an advertiser might correspond to one candidate category. This embodiment does not limit the method of dividing candidate categories. A live stream within a candidate category can support one digital human form, voice, and scene, rather than being universally applicable to different forms, voices, and scenes.

[0055] Taking a candidate category as an example, the process of creating candidate live stream rooms 303 for that category is explained below. For instance, to estimate the estimated usage 302 of candidate live stream rooms 303 for the target time period of 2:00 AM to 2:15 PM, we can first read the historical usage 301 of the live stream rooms for that candidate category during the past day or more during the 2:00 AM to 2:15 PM time period. For example, the historical usage 301 for that time period in the past day might be 100. If the historical usage 301 cannot be directly read, we can calculate the historical usage 301 for a certain time period using indicators such as the creation time and usage duration of the live stream rooms for that candidate category. Then, we can perform weighted calculations or other calculations on the historical usage 301 for each of the past 2:00 AM to 2:15 PM time periods, and use the calculation result as the estimated usage 302 for the future 2:00 AM to 2:15 PM time period. For example, the estimated usage 302 might be 100. Next, we can create 100 candidate live stream rooms 303 for the candidate categories using the live stream room configuration information corresponding to the candidate categories. The configuration information for the live streaming room includes, for example, the digital human form, voice, and live streaming room scene.

[0056] The creation method for candidate live stream rooms 303 in other candidate categories can be found above. After creating candidate live stream rooms 303 for each candidate category, each candidate live stream room 303 can be added to the room pool 304.

[0057] In this embodiment, candidate live streaming rooms 303 are created in categories, and the estimated usage 302 for a future target time period is calculated based on the historical usage 301 of a predetermined time period. Then, candidate live streaming rooms 303 for that category are created according to the estimated usage 302. It is understood that the historical usage 301 of a candidate category reflects the number of times users have used this type of live streaming room in the past. Determining the estimated usage 302 through historical usage 301 allows for a relatively accurate determination of user usage during the future target time period. Therefore, creating candidate live streaming rooms 303 according to the estimated usage 302 avoids resource waste due to an excessive number of candidate live streaming rooms 303, and also avoids insufficient candidate live streaming rooms 303 to meet the needs of a large number of users watching live streams.

[0058] Figure 4 This is a schematic diagram illustrating the principle of a live streaming method according to an embodiment of this disclosure.

[0059] In this embodiment, an object can access the target page 401 by clicking or other operations. When the client detects that the object is accessing the target page 401, it can send an access request to the server. The access request includes the page information 402 of the target page 401.

[0060] The server extracts page information 402 of the target page 401 from the access request. Then it determines whether the multiple candidate live rooms pre-created in the room pool include a candidate live room that matches page information 402.

[0061] For example, the server can map page information 402 to live room configuration based on a first mapping relationship between page information 402 and live room configuration. Based on a second mapping relationship between live room configuration and live room identifier, the live room configuration is mapped to at least one live room identifier. Then, from at least one candidate live room indicated by the at least one live room identifier, it is determined whether there is an idle candidate live room. If so, an idle candidate live room is randomly selected as the candidate live room matching page information 402, and then this candidate live room can be reused. That is, this candidate live room is used as the target live room 403.

[0062] If no candidate live room matches page information 402 in the room pool, a new live room can be created based on page information 402. For example, the client first loads the necessary code, components, and other resources. After loading, the client calls the server's interface for creating a live room. Then, the server can create a new live room. For instance, the server can map page information 402 to live room configurations such as digital human form, voice, and scene based on the first mapping information, and then create a new live room according to the configuration, using the new live room as the target live room 404.

[0063] Then a connection can be established between the client and the target live room (one of the target live room 403 and target live room 404 mentioned above), and the live video stream can be sent to the client after the connection is established.

[0064] In this embodiment, if the pre-created room pool includes a target live room that matches the page information, the target live room is directly reused to reduce user waiting time. If the room pool does not include candidate live rooms that match the page information, a new live room is created based on the page information as the target live room, thereby providing live streaming services to users.

[0065] According to another embodiment of this disclosure, among a plurality of candidate live streaming rooms, the candidate live streaming room that is in an idle state and meets the predetermined matching conditions can be used as the candidate live streaming room that matches the page information, that is, as the target live streaming room.

[0066] For example, the pre-defined matching criteria include: the virtual avatar corresponding to the page information is consistent with the virtual avatar in the candidate live stream. Another example is: the pre-defined matching criteria include: the voice timbre corresponding to the page information is the same as the voice timbre used by the virtual avatar in the candidate live stream. Yet another example is: the pre-defined matching criteria include: the live stream scene corresponding to the page information is consistent with the scene in the candidate live stream.

[0067] In this embodiment, the matching between candidate live streaming rooms and page information is determined by virtual avatars, voice timbre, and scene. This allows the user to be provided with a live streaming room suitable for the target page after visiting the target page.

[0068] Figure 5 This is a schematic flowchart of a live streaming method according to an embodiment of the present disclosure.

[0069] like Figure 5 As shown, the live streaming method 500 can be executed by the client, and the live streaming method 500 includes operations S510 to S540.

[0070] When operating S510, in response to detecting that an object is accessing the target page, an access request is sent to the server. The access request includes the page information of the target page.

[0071] For example, an object can access a target page through clicking or other operations. When the client detects that the object is accessing the target page, it can send an access request to the server, which includes the page information of the target page.

[0072] For example, the server pre-creates multiple candidate live streaming rooms. After receiving an access request, the server extracts the page information of the target page from the access request. Then, it determines whether the pre-created multiple candidate live streaming rooms include a candidate live streaming room that matches the page information. If so, the candidate live streaming room that matches the page information is selected as the target live streaming room.

[0073] In operation S520, a connection is established with the target live room among multiple pre-created candidate live rooms; the target live room is matched with the page information; any one of the multiple candidate live rooms is a live room used for one-to-one interaction with the client through a virtual avatar.

[0074] For example, the server returns the identifier of the target live stream to the client, and then the client sends a connection request to the server, which includes the identifier of the target live stream. Upon receiving the connection request, the server's live streaming service establishes a connection with the client.

[0075] When operating the S530, acquire the video stream of the target live room with the established connection.

[0076] Operating the S540 to display a video stream.

[0077] For example, the client pulls a real-time video stream from the target live stream room and then performs front-end rendering.

[0078] According to the live streaming method provided in this disclosure, since multiple candidate live streaming rooms are created in advance, and when it is detected that a user intends to enter a one-on-one live streaming room, the method first checks whether a target live streaming room exists among the multiple candidate live streaming rooms. If so, a connection is directly established with the target live streaming room. In this way, the target live streaming room is created in advance, rather than being created when interaction with the user is required, saving the time spent creating the live streaming room and improving the user experience by reducing the user's waiting time.

[0079] According to another embodiment of this disclosure, users can ask questions while watching a live stream, and the live stream can answer these questions. For example, a user inputs question information through a front-end page. Upon receiving the input question information, the client sends a question request, including the question information, to the server. Next, the server, upon receiving the question request, generates response information based on a large model, addressing the question information in the request. Furthermore, the server generates explanatory segments based on the response information, such as driving a digital human to verbally recite the response information, or driving the digital human to perform corresponding actions, lip movements, and facial expressions. The server then pushes the explanatory segments to the live stream.

[0080] It's worth noting that some livestream avatars have rather simplistic interaction modes. For example, some avatars simply convert text into audio and broadcast it, lacking interaction during the explanation process. This mode relies solely on voice and limited facial expressions to convey information, which can be tedious, especially when explaining long texts. It's also inefficient at conveying complex or abstract concepts. Therefore, it's necessary to improve the interaction methods in livestreams to enhance the user experience.

[0081] To address the aforementioned issues, in this embodiment, the server can first obtain some visual materials, response information, and summary information, and then return this information to the client. In this way, the client can not only perform text broadcasting (e.g., voice broadcasting of response information), but also display the aforementioned visual materials, response information, and summary information, thereby optimizing the interaction between the live stream and the user and improving the user experience.

[0082] Figure 6 This is a schematic diagram illustrating the generation and transmission of data packets according to embodiments of this disclosure.

[0083] like Figure 6 As shown, the user inputs question information 601 through the front-end page. Upon receiving the question information 601, the client 609 sends a question request, including the question information 601, to the server. Next, the server, upon receiving the question request, generates response information 603 based on the large model 602, addressing the question information 601 in the question request.

[0084] After receiving response information 603, the server generates a narration segment based on the response information 603, and then pushes the narration segment to the live broadcast room, where the client 609 can display the narration segment.

[0085] After receiving response information 603, the server retrieves at least one visual material 607 related to response information 603 from the material library 606, and then outputs at least one visual material 607 to the client 609. Visual material 607 can be an image, GIF, video, etc. In this way, the client 609 can display visual material 607 while showing explanatory segments in the video stream.

[0086] In this embodiment, before the virtual avatar in the live stream begins its explanation, visual material 607 related to the response information 603 is retrieved and returned to the client 609. The visual material 607 is displayed synchronously while the virtual avatar is explaining, allowing the user to visually and intuitively perceive the information related to the response information 603. The live stream presents the response to the user through both visual and auditory means, making the virtual avatar's explanation more specific and efficient, thus enhancing the user experience.

[0087] In one example, the server can use a large model 602 to perform intent analysis or entity extraction on the response information 603, thereby extracting key entities 605 from the response information 603. For example, key entity 605 could be "xx scenic spot" or "xx travel guide". Then, at least one visual material 607 that satisfies a predetermined similarity relationship with key entity 605 is retrieved from the material library 606. The predetermined similarity relationship is, for example, that the similarity between visual material 607 and key entity 605 is greater than a threshold. In this embodiment, key entity 605 is first extracted from the response information 603, and then visual material 607 is retrieved through key entity 605. Since key entity 605 is a relatively important part of the response information 603, the retrieved visual material 607 is more relevant to the response information 603, thus providing the user with more information to aid understanding.

[0088] Furthermore, after receiving response information 603, the server can also detect whether the number of characters in response information 603 is greater than or equal to a first predetermined number, such as 50. If so, a summary information 604 is generated for response information 603, where the number of characters in summary information 604 is less than or equal to a second predetermined number, which is less than the first predetermined number, such as 10. If not, summary information 604 is not generated. Then, the server outputs summary information 604 to client 609. Client 609 can display summary information 604 as needed. In this embodiment, summary information 604 is generated when response information 603 is long text. Subsequently, client 609 can display summary information 604, allowing users to understand the summary of the content through summary information 604, thereby better understanding the response information 603 output by the virtual avatar in the live broadcast room.

[0089] The server can send the response information 603, visual material 607, and summary information 604 to the client 609 separately, or it can assemble the response information 603, visual material 607, and summary information 604 into a data packet 608 and then send the data packet 608 to the client 609.

[0090] According to another embodiment of this disclosure, after receiving a data packet, the client can extract information from the data packet, such as response information, visual materials, and summary information, and then display the extracted information. The client can determine a display area based on the quantity and category of at least one visual material, and then display at least one visual material in the display area.

[0091] It's important to note that the live stream primarily displays two parts: a video stream and a data packet. The first part is the video stream itself. For example, the server drives the virtual avatar's lip movements, voice, actions, and facial expressions to generate a video stream. The client then pulls this real-time video stream from the established live stream connection for playback. The second part is the data packet the client receives from the server, which includes at least one of the following: response information, visual materials, and summary information. During video stream playback, the client can also display the response information, visual materials, and summary information in layers.

[0092] In this embodiment, the client obtains visual materials from the server. However, during the display process, a uniform display mode is not adopted. Instead, the display areas in the live broadcast room are dynamically determined based on the quantity and category of the materials. This makes the display effect more suitable for the current quantity and category of materials, thereby improving the display effect of the visual materials.

[0093] Figure 7 This is a schematic diagram illustrating the linkage of key areas according to an embodiment of this disclosure.

[0094] like Figure 7 As shown, for example, if the number of materials is less than or equal to a predetermined number, and at least one of the visual materials does not contain any video material, the display area can be determined as a key area. For example, if there are few visual materials and they are all images, the display area can be determined as a key area. The key area can be a local area above the virtual avatar in the live broadcast room. If the display area is a key area above the virtual avatar, at least one visual material can be displayed in the key area. For example, before the user asks a question, the client can display regular content 701; after the user asks a question, the client can display explanatory content 702. After the virtual avatar finishes explaining, the client can display regular content 701 again.

[0095] In this embodiment, the key area above the virtual avatar is more visually appealing to users. When the virtual avatar begins explaining long text and detects related supplementary comprehension content in the data packet, the content in the key area is replaced with this supplementary comprehension content. This supplementary comprehension content can include at least one of visual materials and summary information. A fade-in / fade-out animation effect can be used to further attract and maintain the user's visual attention. This embodiment makes full use of the key area, placing visual information in a superior visual and interactive zone, allowing users to clearly perceive the visual materials and summary information, thereby using this information to help users understand the virtual avatar's explanation. After the virtual avatar finishes explaining the response, the key area of ​​the live stream can automatically switch back to the previously displayed information, restoring the previous state.

[0096] Figure 8 This is a schematic diagram illustrating the principle of background area linkage according to an embodiment of the present disclosure.

[0097] like Figure 8 As shown, for example, if the number of materials exceeds a predetermined number, or if at least one visual material includes a video-type visual material, the display area is determined to be the background area behind the virtual avatar. For example, if there are many visual materials or if some of the visual materials are videos, the display area is determined to be the background area. If the display area is the background area, at least one visual material can be displayed there, for example, multiple visual materials can be rotated in the background area without displaying summary information. For example, before a user asks a question, the client can display regular content 801; after the user asks a question, the client displays explanatory content 802. After the virtual avatar finishes its explanation, the client displays regular content 801 again.

[0098] In this embodiment, during the virtual avatar's explanation, the background area of ​​the live broadcast room is switched to visual materials related to the explanation content. This makes the virtual avatar's explanation more immersive and specific, and improves the user experience.

[0099] Figure 9This is a schematic structural block diagram of a live streaming device according to an embodiment of the present disclosure.

[0100] like Figure 9 As shown, the live streaming device 900 may include a page information determination module 910, a live streaming room determination module 920, a first connection module 930, and a sending module 940.

[0101] The page information determination module 910 is used to determine the page information for the target page in the access request in response to receiving an access request from the client.

[0102] The live stream room determination module 920 is used to respond to the detection of multiple pre-created candidate live stream rooms, including candidate live stream rooms that match the page information, and to select the candidate live stream room that matches the page information as the target live stream room. Any one of the multiple candidate live stream rooms is a live stream room used for one-to-one interaction with the client through a virtual avatar.

[0103] The first connection module 930 is used to establish a connection between the client and the target live streaming room.

[0104] The sending module 940 is used to send a video stream to the target live room to the client after a connection has been established with the target live room.

[0105] According to another embodiment of this disclosure, multiple candidate live streaming rooms are created through the following modules: an estimated usage determination module and a first creation module. The estimated usage determination module is used to determine, for any candidate category among at least one candidate category, an estimated usage for a target time period based on the historical usage of live streaming rooms in that candidate category within a predetermined time period. The first creation module is used to create candidate live streaming rooms for the candidate categories based on the estimated usage and the live streaming room configuration information corresponding to the candidate categories.

[0106] According to another embodiment of this disclosure, the live stream room determination module includes a live stream room sub-determination module, used to select a candidate live stream room that is idle and meets predetermined matching conditions from a plurality of candidate live stream rooms as the target live stream room. The predetermined matching conditions include at least one of the following: the virtual avatar corresponding to the page information is consistent with the virtual avatar in the candidate live stream room; the timbre corresponding to the page information is consistent with the timbre used by the virtual avatar in the candidate live stream room; the live stream room scene corresponding to the page information is consistent with the scene in the candidate live stream room.

[0107] According to another embodiment of this disclosure, it further includes: a second creation module, configured to, in response to the absence of detection of multiple candidate live rooms including candidate live rooms matching the page information, create a new live room as the target live room based on the page information.

[0108] According to another embodiment of this disclosure, the system further includes: a response generation module, a retrieval module, and a first output module. The response generation module is used to generate response information based on a large model to the question information in the question request in response to receiving a question request. The retrieval module is used to retrieve at least one visual material related to the response information from a material library. The first output module is used to output at least one visual material to the client so that the client can display at least one visual material during the presentation of an explanatory segment in the video stream. The explanatory segment is generated based on the response information.

[0109] According to another embodiment of this disclosure, the retrieval module includes: a response generation module, a retrieval module, and a first output module. The extraction submodule is used to extract key entities from the response information. The retrieval submodule is used to retrieve at least one visual material from a material library that satisfies a predetermined similarity relationship with the key entities.

[0110] According to another embodiment of this disclosure, the system further includes a digest generation module and a second output module. The digest generation module is used to generate a digest of the response information after generating the response information, in response to detecting that the number of characters in the response information is greater than or equal to a first predetermined number. The number of characters in the digest information is less than or equal to a second predetermined number, and the second predetermined number is less than the first predetermined number. The second output module is used to output the digest information to the client.

[0111] Figure 10 This is a schematic structural block diagram of a live streaming device according to an embodiment of the present disclosure.

[0112] like Figure 10 As shown, the live streaming device 1000 may include an access request sending module 1010, a second connection module 1020, a video stream acquisition module 1030, and a video stream display module 1040.

[0113] The access request sending module 1010 is used to send an access request to the server in response to detecting that an object is accessing the target page. The access request includes the page information of the target page.

[0114] The second connection module 1020 is used to establish a connection with a target live room among a plurality of pre-created candidate live rooms. The target live room is matched with page information. Any one of the candidate live rooms is a live room used for one-to-one interaction with the client through a virtual avatar.

[0115] The video stream acquisition module 1030 is used to acquire the video stream of the target live room with an established connection.

[0116] The video stream display module 1040 is used to display video streams.

[0117] According to another embodiment of this disclosure, the system further includes: a question request sending module, a material acquisition module, and a material display module. The question request sending module is used to send a question request, including the question information, in response to receiving question information input by an object. The material acquisition module is used to acquire at least one visual material related to the response information, which is generated based on a large model in response to the question information in the question request. The material display module is used to display at least one visual material during the presentation of an explanatory segment in the video stream. The explanatory segment is generated based on the response information.

[0118] According to another embodiment of this disclosure, the material display module includes an area determination submodule and a material display submodule. The area determination submodule is used to determine a display area based on the quantity and category of at least one visual material. The material display submodule is used to display at least one visual material in the display area during the presentation of explanatory segments in the video stream.

[0119] According to another embodiment of this disclosure, the material display submodule includes: a first display unit and a second display unit. The first display unit is configured to display at least one visual material in the key area when the display area is detected to be a key area located above the virtual avatar. The second display unit is configured to display at least one visual material in the background area when the display area is detected to be a background area located behind the virtual avatar.

[0120] According to another embodiment of this disclosure, the region determination submodule includes: a first region determination unit and a second region determination unit. The first region determination unit is configured to determine the display area as a key region in response to detecting that the number of materials is less than or equal to a predetermined number, and that at least one visual material does not contain visual material of the video category. The second region determination unit is configured to determine the display area as a background region in response to detecting that the number of materials is greater than the predetermined number, or that at least one visual material includes visual material of the video category.

[0121] According to another embodiment of this disclosure, the first display unit includes: a display subunit, configured to display summary information and at least one visual material in the key area in response to detecting that the display area is a key area located above the virtual image and receiving summary information for the response information.

[0122] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product. Figure 11This is a structural block diagram of an electronic device used to implement the live streaming method of embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0123] like Figure 11 As shown, device 1100 includes a computing unit 1101, which can perform various appropriate actions and processes according to a computer program stored in read-only memory (ROM) 1102 or a computer program loaded into random access memory (RAM) 1103 from storage unit 1108. The RAM 1103 may also store various programs and data required for the operation of device 1100. The computing unit 1101, ROM 1102, and RAM 1103 are interconnected via bus 1104. Input / output (I / O) interface 1105 is also connected to bus 1104.

[0124] Multiple components in device 1100 are connected to I / O interface 1105, including: input unit 1106, such as keyboard, mouse, etc.; output unit 1107, such as various types of monitors, speakers, etc.; storage unit 1108, such as disk, optical disk, etc.; and communication unit 1109, such as network card, modem, wireless transceiver, etc. Communication unit 1109 allows device 1100 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0125] The computing unit 1101 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1101 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1101 performs the various methods and processes described above, such as the live streaming method. For example, in some embodiments, the live streaming method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1108. In some embodiments, part or all of the computer program may be loaded and / or installed on device 1100 via ROM 1102 and / or communication unit 1109. When the computer program is loaded into RAM 1103 and executed by the computing unit 1101, one or more steps of the live streaming method described above may be performed. Alternatively, in other embodiments, the computing unit 1101 may be configured to perform the live streaming method by any other suitable means (e.g., by means of firmware).

[0126] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0127] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0128] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0129] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0130] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0131] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other.

[0132] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0133] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A live streaming method, comprising: In response to receiving an access request from a client, determine the page information for the target page in the access request; In response to detecting that a plurality of pre-created candidate live streaming rooms include a candidate live streaming room that matches the page information, the candidate live streaming room that matches the page information is selected as the target live streaming room; any one of the plurality of candidate live streaming rooms is a live streaming room used for one-to-one interaction with the client through a virtual avatar; Establish a connection between the client and the target live streaming room; as well as Send a video stream for the target live stream to the client that has established a connection with the target live stream.

2. The method according to claim 1, wherein, The multiple candidate live streaming rooms were created in the following way: For any candidate category among at least one candidate category, determine the estimated usage for the target time period based on the historical usage of the live streaming rooms of the candidate category within a predetermined time period; as well as Based on the estimated usage and the live stream configuration information corresponding to the candidate category, candidate live stream rooms for the candidate category are created.

3. The method according to claim 1, wherein, The response to detecting that multiple pre-created candidate live stream rooms include candidate live stream rooms that match the page information, and selecting the candidate live stream room that matches the page information as the target live stream room includes: The candidate live room that is idle and meets the predetermined matching conditions among the multiple candidate live rooms is selected as the target live room. The predetermined matching conditions include at least one of the following: The virtual avatar corresponding to the page information is consistent with the virtual avatar in the candidate live broadcast room; The timbre corresponding to the page information is consistent with the timbre used by the virtual avatar in the candidate live stream room; and The live stream scene corresponding to the page information is consistent with the scene in the candidate live stream rooms.

4. The method according to claim 1, further comprising: In response to the absence of a candidate live room matching the page information among the plurality of candidate live rooms, a new live room is created as the target live room based on the page information.

5. The method according to any one of claims 1 to 4, further comprising: In response to receiving a question request, the system generates response information based on the large model to address the question information in the question request. Retrieve at least one visual material related to the response information from the material library; as well as The at least one visual material is output to the client so that the client can display the at least one visual material while displaying the explanatory segment in the video stream; wherein the explanatory segment is generated based on the response information.

6. The method according to claim 5, wherein, The step of retrieving at least one visual material related to the response information from the material library includes: Extract key entities from the response information; and Retrieve at least one visual material from the material library that satisfies a predetermined similarity relationship with the key entity.

7. The method according to claim 5, further comprising: After generating the response information In response to detecting that the number of characters in the response information is greater than or equal to a first predetermined number, a summary information for the response information is generated; The number of characters in the summary information is less than or equal to a second predetermined number, and the second predetermined number is less than the first predetermined number; as well as The summary information is output to the client.

8. A live streaming method, comprising: In response to detecting that an object is accessing a target page, an access request is sent to the server, the access request including page information of the target page; A connection is established with a target live room among a plurality of pre-created candidate live rooms; the target live room is matched with the page information; any one of the plurality of candidate live rooms is a live room for one-to-one interaction with the client through a virtual avatar; Acquire the video stream of the target live streaming room with the established connection; and The video stream is displayed.

9. The method according to claim 8, further comprising: In response to receiving question information input by the object, a question request including the question information is sent; Acquire at least one visual element related to the response information; The response information is generated based on the question information in the question request using a large model. as well as The at least one visual material is displayed during the presentation of the explanatory segment in the video stream; wherein the explanatory segment is generated based on the response information.

10. The method according to claim 9, wherein, The process of displaying at least one visual material during the presentation of the explanatory segment in the video stream includes: The display area is determined based on the quantity and category of the at least one visual material; and During the presentation of explanatory segments in the video stream, at least one visual element is displayed in the presentation area.

11. The method according to claim 10, wherein, Displaying the at least one visual material in the display area includes: In response to detecting that the display area is a key area located above the virtual avatar, at least one visual material is displayed in the key area; and In response to detecting that the display area is a background area located behind the virtual avatar, at least one visual material is displayed in the background area.

12. The method according to claim 11, wherein, The step of determining the display area based on the quantity and category of the at least one visual material includes: In response to detecting that the number of materials is less than or equal to a predetermined number, and that none of the at least one visual material is of the video category, the display area is determined to be the key area; and In response to detecting that the number of materials is greater than the predetermined number, or that the at least one visual material includes visual materials of the video category, the display area is determined to be the background area.

13. The method according to claim 11, wherein, The step of displaying at least one visual material in the key area in response to detecting that the display area is a key area located above the virtual avatar includes: In response to detecting that the display area is a key area located above the virtual avatar, and receiving summary information for the response information, the summary information and the at least one visual material are displayed in the key area.

14. A live streaming device, comprising: The page information determination module is used to determine the page information for the target page in the access request in response to receiving an access request from the client. The live stream room determination module is used to respond to the detection that a plurality of pre-created candidate live stream rooms include a candidate live stream room that matches the page information, and to select the candidate live stream room that matches the page information as the target live stream room; any one of the plurality of candidate live stream rooms is a live stream room for one-to-one interaction with the client through a virtual avatar; The first connection module is used to establish a connection between the client and the target live streaming room; as well as The sending module is used to send a video stream to the target live room to the client after establishing a connection with the target live room.

15. A live streaming device, comprising: An access request sending module is used to send an access request to the server in response to detecting that an object is accessing a target page. The access request includes page information of the target page. The second connection module is used to establish a connection with a target live room among a plurality of pre-created candidate live rooms; the target live room is matched with the page information; any one of the plurality of candidate live rooms is a live room used for one-to-one interaction with the client through a virtual avatar; The video stream acquisition module is used to acquire the video stream of the target live streaming room with an established connection; as well as The video stream display module is used to display the video stream.

16. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 13.

17. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 13.

18. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 13.