Video rendering method, device and electronic equipment for live broadcast scene
By using voice recognition to generate and render virtual characters based on the anchor's speech, the problem of anchors having limited topics was solved. This enabled interaction between the anchor and the virtual characters, improving the interactivity of the live stream and the viewer experience.
Patent Information
- Application Number
- CN202311708249.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-12
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2043-12-12
AI Technical Summary
In live streaming scenarios, if the host's topics are limited and it is difficult to interact with the audience or other hosts, it will affect the live streaming effect and the audience experience.
By performing voice recognition on the live streamer's audio, the popularity of the live stream topic is determined, and under certain conditions, reply text information is generated to render a virtual character, enabling the live streamer and the virtual character to interact.
Increase the popularity of live stream topics, enliven the atmosphere of the live stream, and enhance interactive effects.
Smart Images

Figure CN117812375B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of artificial intelligence, in particular to the field of live broadcast and the field of large models. The present disclosure specifically relates to a video rendering method and device for a live broadcast scenario, an electronic device and a storage medium. BACKGROUND
[0002] In a live broadcast scenario, for example, a scenario in which a host sells goods for a certain product or the host chats with other hosts or audiences, sometimes the host's topic is single and it is difficult to chat with the audience or other hosts. Over a long period of time, this affects the live broadcast effect and the experience of the host and other audiences. SUMMARY
[0003] The present disclosure provides a video rendering method and device for a live broadcast scenario, an electronic device and a storage medium.
[0004] According to an aspect of the present disclosure, a video rendering method for a live broadcast scenario is provided, comprising:
[0005] live broadcast recording of a host to obtain a first video stream;
[0006] voice recognition of live broadcast voice in the first video stream to obtain first text information;
[0007] determination of live broadcast topic heat based on audience response information in the live broadcast recording process and the first text information;
[0008] determination of corresponding reply text information based on the first text information in a case where the live broadcast topic heat meets a first set condition;
[0009] rendering of a virtual character based on the reply text information to obtain a second video stream;
[0010] generation of a third video stream of the host chatting with the virtual character based on the first video stream and the second video stream.
[0011] According to another aspect of the present disclosure, a live broadcast device is provided, comprising:
[0012] a video recording module configured to live broadcast record a host to obtain a first video stream;
[0013] a voice recognition module configured to perform voice recognition on live broadcast voice in the first video stream to obtain first text information;
[0014] a heat determination module configured to determine live broadcast topic heat based on audience response information in the live broadcast recording process and the first text information;
[0015] a reply text determination module configured to determine corresponding reply text information based on the first text information, in a case where the live topic heat meets a first set condition;
[0016] a virtual character rendering module configured to render a virtual character based on the reply text information, to obtain a second video stream;
[0017] a video generation module configured to generate a third video stream in which the anchor and the virtual character chat with each other, based on the first video stream and the second video stream.
[0018] According to another aspect of the present disclosure, an electronic device is provided, comprising:
[0019] at least one processor; and
[0020] a memory in communication with the at least one processor; wherein
[0021] the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform any of the live scene-oriented video rendering methods in embodiments of the present disclosure.
[0022] According to another aspect of the present disclosure, a non-transitory computer readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to perform any of the live scene-oriented video rendering methods in embodiments of the present disclosure.
[0023] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements any of the live scene-oriented video rendering methods in embodiments of the present disclosure.
[0024] According to the technology of the present disclosure, a virtual character is set for an anchor in a live room, the live voice of the anchor is recognized to obtain first text information, then based on the audience response information and the first text information, the live topic heat can be determined, when the anchor has no topic to chat about, the corresponding reply text information is determined based on the first text information, and the virtual character is rendered based on the reply text information to obtain a second video stream, the first video stream obtained by live recording of the anchor and the second video stream are mixed to generate a third video stream in which the anchor and the virtual character chat with each other, thereby realizing a live scene in which the virtual character and the anchor interact on topics. In this way, when the anchor has no topic to chat about, the virtual character can find a topic to chat with the anchor, thereby livening up the atmosphere of the live room and improving the heat of the live topic.
[0025] It should be appreciated that the description set forth in this section is not intended to identify key or essential features of embodiments of the disclosure, and the scope of the disclosure is not to be limited to any of the features set forth in this section. Other features, aspects, and advantages of the disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0026] The accompanying drawings are included to provide a further understanding of the present application, and are incorporated in and constitute a part of this specification. Illustrations in the drawings are for purposes of illustrating an implementation of the present application and are not intended to limit the present application.
[0027] Figure 1 is a flowchart of a live scene-oriented video rendering method according to an embodiment of the present disclosure;
[0028] Figure 2 is a scene diagram of anchor and virtual character mutual chatting according to an embodiment of the present disclosure;
[0029] Figure 3 is a scene diagram of anchor and virtual character mutual chatting according to an embodiment of the present disclosure;
[0030] Figure 4 is a schematic diagram of a live companion chatting method according to an embodiment of the present disclosure;
[0031] Figure 5 is a schematic diagram of a live microphone method according to an embodiment of the present disclosure;
[0032] Figure 6 is a structural block diagram of a live device according to an embodiment of the present disclosure;
[0033] Figure 7 is a structural block diagram of a live device according to another embodiment of the present disclosure;
[0034] Figure 8 is a block diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0035] Exemplary embodiments of the present disclosure are described herein with reference to the accompanying drawings, which are included to provide a further understanding of the present disclosure and are incorporated in and constitute a part of this specification. The descriptions of the embodiments of the present disclosure are intended to be illustrative, and not to be the only
[0036] Figure 1is a flowchart of a video rendering method for a live streaming scenario according to an embodiment of the present disclosure. The method can be applied to an electronic device. The electronic device is, for example, a terminal, a server, or other processing device, where the terminal can be a desktop computer, a mobile device, a PDA (Personal Digital Assistant), a handheld device, a computing device, a vehicle-mounted device, a wearable device, or the like UE (User Equipment). In some implementations, the electronic device can implement the video rendering method for a live streaming scenario according to an embodiment of the present disclosure by invoking computer-readable instructions stored in a memory through a processor.
[0037] As shown in Figure 1 The video rendering method for a live streaming scenario can include the following steps.
[0038] S110, live streaming recording of an anchor is performed to obtain a first video stream;
[0039] S120, speech recognition is performed on live streaming speech in the first video stream to obtain first text information;
[0040] S130, live streaming topic heat is determined based on audience response information during live streaming recording and the first text information;
[0041] S140, in a case where the live streaming topic heat meets a first set condition, corresponding reply text information is determined based on the first text information;
[0042] S150, a virtual character is rendered based on the reply text information to obtain a second video stream;
[0043] S160, a third video stream of an anchor and a virtual character chatting with each other is generated based on the first video stream and the second video stream.
[0044] It can be understood that the video rendering method for a live streaming scenario according to an embodiment of the present disclosure can be performed by a local streaming device, or the live streaming can be recorded and uploaded to the cloud by the local streaming device, the cloud performs steps S120 to S160, the cloud returns the third video stream to the streaming device, and the streaming device pushes the third video stream to each client for playing. For example, the video stream is distributed through a CDN.
[0045] It can be understood that live streaming recording of an anchor includes receiving voice information of the anchor through a microphone and obtaining a camera picture of a live streaming room centered on the anchor through a camera, thereby obtaining the first video stream. The first video stream includes multimedia data of video and voice.
[0046] It can be understood that the live streaming speech is voice information from the anchor.
[0047] It can be understood that the live voice is recognized by using a voice recognition technology (Automatic Speech Recognition, ASR) to obtain first text information.
[0048] It can be understood that the audience response information can include audience comment information in a comment area, and reply information of other anchors or audiences to the microphone with the anchor (the reply voice of the other anchors or audiences is recognized to obtain the reply information).
[0049] It can be understood that the virtual character can be a virtual character with a specific style. For example, the anchor selects a virtual character with a first style in a style selection interface, starts rendering by using model data of the virtual character with the first style to obtain an initial video stream of the virtual character with the first network, wherein the initial video stream includes an initial picture and an initial voice.
[0050] It can be understood that based on the reply text information, the virtual character in the initial video stream is rendered to obtain a second video stream. For example, by using the reply text information, corresponding lip shape data and body movement data are determined, and the lip shape and body movement of the virtual character in the initial video stream are adjusted by using the lip shape data and the body movement data, so that the lip shape of the virtual character matches the lip shape in the lip shape data, and the body movement matches the body movement in the body movement data.
[0051] It can be understood that the first video stream and the second video stream are mixed to obtain a third video stream. For example, the mixing includes inserting a video frame in the second video stream into the first video stream as a video frame in the first video stream, or merging the video frame in the second video stream with the video frame in the first video stream into the same frame, etc. For another example, the mixing can also include deleting part of the frames in the first video stream, and deleting part of the frames in the second video stream.
[0052] Figure 2 and Figure 3 is a scene diagram of anchor and virtual character mutual chatting in an embodiment of the disclosure. As shown in Figure 2 , the virtual character as an assistant role and the anchor chat with each other in the same live room, and the virtual character and the anchor exist in the same picture and there is no chat box. As shown in Figure 3 , the virtual character as an audience or an anchor in other live rooms and the anchor in the live room exist in the live picture in different chat boxes. Of course, in Figure 3 , real audiences or other anchors can also exist in the live picture in the form of chat boxes together with the chat boxes of the virtual characters.
[0053] According to the above embodiment, a virtual character is set for the host in the live room, speech recognition is performed on the live speech of the host to obtain first text information, and then based on the audience response information and the first text information, the live topic heat can be determined. When the host has no words to chat, the corresponding reply text information is determined by using the first text information, and the virtual character is rendered by using the reply text information to obtain a second video stream. The first video stream obtained by live recording for the host is mixed with the second video stream to generate a third video stream of the host and the virtual character chatting with each other, and the live scene of the virtual character and the host interacting with each other on the topic is realized. In this way, when the host has no words to chat, the virtual character can find a topic to chat with the host, thereby activating the atmosphere of the live room and improving the heat of the live topic.
[0054] In an embodiment, based on the audience interaction information in the process of live recording and the first text information, the live topic heat is determined, comprising: extracting at least one first text segment from the first text information; for each first text segment, finding a second text segment in the audience response information which can form a key-value pair with the first text segment, and counting the number of key-value pair groups; and determining the live topic heat based on the number of key-value pair groups.
[0055] It can be understood that the first text segment can be a text segment as key information. For example, a question sentence or a statement sentence, etc.
[0056] It can be understood that the second text segment can be a text segment as value information, for example, an answer sentence or a rhetorical question sentence, etc.
[0057] It can be understood that the audience response information can include a second text segment which forms a key-value pair with the first text segment, or can not include a second text segment which forms a key-value pair with the first text segment.
[0058] Exemplarily, each text segment in the audience response information includes a timestamp, and each text segment in the first text information also has a timestamp. For the first text segment, a second text segment which forms a key-value pair with the first text segment and has a timestamp matching the timestamp of the first text segment is found in the audience response information. The two matching timestamps can be two timestamps within the same time range.
[0059] Exemplarily, there is a preset key-value pair database, the first text segment is used to find a target key-value pair in the key-value pair database, the target key-value pair includes the first text segment and a third text segment, the third text segment is extracted from the target key-value pair, and then a second text segment identical or similar to the third text segment is found in the audience response information. Finally, the first text segment and the second text segment form a new key-value pair.
[0060] It can be understood that the more the number of key-value pairs of the newly formed group, the higher the live topic heat. For example, a linear function can be used to calculate the number of key-value pairs of the group to obtain the live topic heat.
[0061] According to the above embodiment, the number of key-value pairs composed of the text segment in the first text information from the live voice and the text segment in the audience response information can be used to accurately determine the live topic heat.
[0062] In one embodiment, in the case where the live topic heat meets the first set condition, the corresponding reply text information is determined based on the first text information, including: in the case where the live topic heat meets the first set condition, extracting keywords from the first text information to obtain a plurality of keywords; topic classification is performed on the plurality of keywords to obtain at least one topic set; based on the number of topic set collections, the live topic repetition degree is determined; in the case where the live topic repetition degree meets the second set condition, the corresponding reply text information is determined based on the first text information.
[0063] For example, in the case where the live topic heat does not meet the first set condition, or in the case where the live topic heat meets the first set condition and the live topic repetition degree does not meet the second set condition, it means that the anchor has topics to chat and the interaction heat between the anchor and the audience is high, so it is not necessary to generate the reply text information to render the virtual character.
[0064] It can be understood that the first set condition is that the live topic heat is less than a set heat threshold.
[0065] For example, a topic classification model can be used to classify the plurality of keywords to obtain the topic type of each keyword, and then the plurality of keywords are grouped using the topic type of each keyword to obtain at least one topic set. The topic classification model is a model obtained by pre-training according to topic data samples.
[0066] It can be understood that the more the number of topic set collections, the more the topics in the first text information from the live voice, that is, the higher the live topic repetition degree of the anchor. For example, a linear function can be used to calculate the number of topic set collections to obtain the live topic repetition degree.
[0067] It can be understood that the text generation model can be used to process the first text information to obtain the corresponding reply text information. The text generation model can be a natural language processing model, for example, a large model. The text generation model is a model obtained by pre-training according to text data samples.
[0068] It can be understood that the second set condition is that the live topic repetition degree is less than a set repetition degree threshold.
[0069] According to the above embodiment, when the live topic heat is too low, the number of topics in the first text information is counted. If the number of topics is also too low, it means that the host has nothing to talk about, and the corresponding reply text information needs to be generated to enable the virtual character to chat with the host.
[0070] In an embodiment, the virtual character includes N virtual characters, N being a positive integer greater than 1, and determining the corresponding reply text information based on the first text information includes: determining a corresponding target text generation model in M text generation models based on the style of the first virtual character in the N virtual characters, where M is a positive integer greater than 1; inputting the first text information into the target text generation model corresponding to the first virtual character to obtain the reply text information of the first virtual character; and performing the following operations on the i-th virtual character in the N virtual characters: determining a corresponding target text generation model in the M text generation models based on the style of the i-th virtual character, where i is a positive integer greater than 1; inputting the first text information and the reply text information from the first virtual character to the i-1-th virtual character into the target text generation model corresponding to the i-th virtual character to obtain the reply text information of the i-th virtual character.
[0071] It can be understood that the value of N can be the same as the value of M, or can not be the same.
[0072] It can be understood that the styles of the virtual characters in the N virtual characters are different, and some of the virtual characters can have the same style.
[0073] Exemplarily, the style of the virtual character can be a role style such as a news host, an entertainment host, a reporter, an audience, etc. The style of the virtual character can also be a style such as male, female, child, middle school student, etc.
[0074] Exemplarily, the host can set the number of virtual characters and the style of each virtual character in a virtual character setting interface. Alternatively, the number of virtual characters and the style of each virtual character can be set according to the live content in the live room.
[0075] It can be understood that any two virtual characters of different styles correspond to different target text generation models. Any two virtual characters of the same style correspond to the same target text generation model.
[0076] It can be understood that the styles corresponding to each text generation model in the above M text generation models are different. Each text generation model is obtained by training a model using text data samples of the corresponding style.
[0077] Exemplarily, starting from an initial value of i being 2, the following operations are performed for the i-th virtual character in the N virtual characters one by one: based on the style of the i-th virtual character, determining a corresponding target text generation model in the M text generation models, where i is a positive integer greater than 1; inputting the first text information and the reply text information from the first virtual character to the i-1-th virtual character into the target text generation model corresponding to the i-th virtual character to obtain the reply text information of the i-th virtual character.
[0078] When the initial value of i is 2, the first text information and the reply text information of the first virtual character are input into the target text generation model corresponding to the second virtual character to obtain the reply text information of the second virtual character.
[0079] In actual application, if there are multiple virtual characters simultaneously chatting with the host, it is necessary to consider whether the chatting content between each virtual character and the host can be connected and mutually connected.
[0080] Therefore, in the present example, the reply text information of the next virtual character is generated according to the first text information spoken by the host and the generated reply text information of all virtual characters, so that the chatting content between each virtual character and the host can be connected and mutually connected.
[0081] In some embodiments, the reply text information of the next virtual character can also be generated only according to the reply text information of the last virtual character. Or the reply text information of the next virtual character can be generated according to the first text information and the reply text information of the last virtual character.
[0082] Exemplarily, based on the first text information, the corresponding reply text information is determined, including: based on the style of the first virtual character in the N virtual characters, determining a corresponding target text generation model in the M text generation models; inputting the first text information into the target text generation model corresponding to the first virtual character to obtain the reply text information of the first virtual character; starting from an initial value of i being 2, the following operations are performed for the i-th virtual character in the N virtual characters: based on the style of the i-th virtual character, determining a corresponding target text generation model in the M text generation models; inputting the reply text information of the i-1-th virtual character into the target text generation model corresponding to the i-th virtual character to obtain the reply text information of the i-th virtual character.
[0083] Illustratively, based on the first text information, the corresponding reply text information is determined, including: based on the style of the first virtual character in the N virtual characters, determining the corresponding target text generation model in the M text generation models; inputting the first text information into the target text generation model corresponding to the first virtual character to obtain the reply text information of the first virtual character; starting from the initial value of i being 2, performing the following operations on the i-th virtual character in the N virtual characters: based on the style of the i-th virtual character, determining the corresponding target text generation model in the M text generation models; inputting the first text information and the reply text information of the i-1-th virtual character into the target text generation model corresponding to the i-th virtual character to obtain the reply text information of the i-th virtual character.
[0084] According to the above-mentioned embodiments, the reply text information of the next virtual character is generated according to the first text information spoken by the host and / or the reply text information of one or more virtual characters that have been generated, so that the chat content between each virtual character and the host can be connected and continued.
[0085] In an embodiment, based on the reply text information, the virtual characters are rendered to obtain a second video stream, including: based on the reply text information of each virtual character, respectively rendering each virtual character to obtain a second video stream of each virtual character.
[0086] It can be understood that each virtual character can be a virtual character with a corresponding style. For example, the host selects a virtual character of a first style in a style selection interface, and starts rendering using model data of the virtual character of the first style to obtain an initial video stream of the virtual character of the first network, wherein the initial video stream includes an initial picture and an initial voice.
[0087] It can be understood that based on the reply text information of each virtual character, the virtual character in the initial video stream of each virtual character is rendered respectively to obtain a second video stream of each virtual character. For example, the corresponding lip shape data and body movement data are determined using the reply text information, and the lip shape and body movement of the virtual character in the initial video stream are adjusted using the lip shape data and body movement data, so that the lip shape of the virtual character matches the lip shape in the lip shape data, and the body movement matches the body movement in the body movement data.
[0088] According to the above-mentioned embodiments, in the case where a plurality of virtual characters exist, each virtual character is rendered using the reply text information of each virtual character to obtain a second video stream of each virtual character, which facilitates subsequent mixing.
[0089] In an embodiment, based on the first video stream and the second video stream, a third video stream of the anchor chatting with the virtual characters is generated, including: based on the generation order of the reply text information of each virtual character, mixing the first video stream and the second video stream of each virtual character to obtain the third video stream of the anchor chatting with each virtual character.
[0090] It can be understood that the generation order of the reply text information of each virtual character in the N virtual characters is the same as the arrangement order of each virtual character in the N virtual characters, and the above i arrangement is used as a reference.
[0091] For example, the first video stream and the second video stream of the first virtual character in the N virtual characters are mixed to obtain a first mixed result; starting from the initial value of i being 2, for the i-th virtual character in the N virtual characters, the (i-1)-th mixed result and the second video stream of the i-th virtual character are mixed to obtain the i-th mixed result. Finally, the N-th mixed result is taken as the third video stream of the anchor chatting with each virtual character.
[0092] Among them, each mixed result is a video stream.
[0093] It can be understood that mixing is inserting one or more video frames in one video stream between two video frames in another video stream. Alternatively, one or more video frames in one video stream are spliced with one or more video frames in another video stream.
[0094] In actual application, if there are multiple virtual characters chatting with the anchor at the same time, the order of each virtual character chatting with the anchor needs to be considered to avoid multiple virtual characters speaking at the same time.
[0095] Therefore, by using the above embodiment, it can be avoided that multiple virtual characters chat with the anchor at the same time, and the situation of too mixed voice can be avoided.
[0096] In an embodiment, based on the first text information, the corresponding reply text information is determined, including: based on the style of the virtual character, processing the first text information to obtain second text information; based on a text generation model, processing the second text information to obtain the reply text information of the virtual character.
[0097] In actual application, only one virtual character can be set to chat with the anchor, and the style of the virtual character can be set by the anchor or automatically set according to the live content. Multiple virtual characters with different styles can also be set to chat with the anchor. At the same time, only one text generation model can also be set, and multiple text generation models can also be set. These text generation models can be large language models.
[0098] Understandably, the second text information is the stylized version of the first text information.
[0099] In this example, the first text information is stylized before being input into the text generation model. In this way, different styles of reply text information can be generated using only one text generation model.
[0100] Figure 4 This is a schematic diagram of a live chat method according to an embodiment of the present disclosure.
[0101] like Figure 4 As shown, under normal circumstances, during live streaming, the video comes from the camera of the streaming device (computer / phone), and the audio comes from the microphone of the streaming device. The broadcaster's audio data is sent to a speech recognition module (locally deployed or implemented via cloud service) to convert the broadcaster's speech into text. Then, the text is transformed into reasonable text information through a propmt (text) stylization module. The specific style depends on the broadcaster's needs and can be configured by the broadcaster. The stylized text information is input to a large model (generative AI) service in the cloud to obtain the response text provided by the large model. Locally, the response text is converted into speech and merged with the local rendered video, such as a cartoon character opening and closing its mouth, into a generative AI response video. The broadcaster's own captured video and the generative AI response video are mixed into a live stream and pushed to the live streaming server for viewers to consume.
[0102] Figure 5 This is a schematic diagram of a live microphone interaction method according to an embodiment of the present disclosure.
[0103] The broadcaster pre-selects a digital avatar with a style appropriate to their content. After selection, the cloud creates a corresponding digital avatar rendering instance and a large-scale model inference instance, performs stylistic initialization settings, and begins rendering the digital avatar's visuals and corresponding audio. During live broadcasting, the streaming device retrieves the rendered digital avatar visuals from the cloud. When the broadcaster needs to interact, they speak to the streaming device; their voice signal is copied and uploaded to the cloud for speech-to-text conversion. This text is then input into different large-scale model inference instances. Different large models generate different styles of response text, which are then input into the digital avatar rendering instance to render the corresponding speaking visuals and audio. The rendered visuals are retrieved by the streaming device in real-time. The streaming device mixes the live broadcaster's visuals and audio with the audio from the cloud-based digital avatar visuals into a single video and audio stream, pushes it to the live streaming service, and distributes it via CDN for viewers to watch.
[0104] Figure 6 This is a structural block diagram of a live streaming device according to an embodiment of the present disclosure.
[0105] As shown in Figure 6 , a live streaming device can include:
[0106] a video recording module 610 configured to record a live streaming of an anchor to obtain a first video stream;
[0107] a speech recognition module 620 configured to perform speech recognition on live streaming speech in the first video stream to obtain first text information;
[0108] a hotness determination module 630 configured to determine a live streaming topic hotness based on audience response information in the process of the live streaming recording and the first text information;
[0109] a reply text determination module 640 configured to determine corresponding reply text information based on the first text information in a case where the live streaming topic hotness meets a first set condition;
[0110] a virtual character rendering module 650 configured to render a virtual character based on the reply text information to obtain a second video stream;
[0111] a video generation module 660 configured to generate a third video stream of the anchor and the virtual character chatting with each other based on the first video stream and the second video stream.
[0112] Figure 7 is a structural block diagram of a live streaming device according to another embodiment of the present disclosure. Figure 7 The video recording module 710, the speech recognition module 720, the hotness determination module 730, the reply text determination module 740, the virtual character rendering module 750, and the video generation module 760 in Figure 6 The video recording module 610, the speech recognition module 620, the hotness determination module 630, the reply text determination module 640, the virtual character rendering module 650, and the video generation module 660 in
[0113] In an implementation, the hotness determination module 730 includes:
[0114] a text segment extraction module 731 configured to extract at least one first text segment from the first text information;
[0115] a key-value pair statistics module 732 configured to, for each of the first text segments, respectively find a second text segment in the audience response information that can form a key-value pair with the first text segment, and count a number of pairs of the key-value pairs;
[0116] a hotness determination unit 733 configured to determine the live streaming topic hotness based on the number of pairs of the key-value pairs.
[0117] In an implementation, the reply text determination module 740 includes:
[0118] a keyword extraction unit 741 configured to extract keywords from the first text information to obtain a plurality of keywords, when the live topic heat meets a first set condition;
[0119] a topic classification unit 742 configured to perform topic classification on the plurality of keywords to obtain at least one topic set;
[0120] a repetition degree determination unit 743 configured to determine a live topic repetition degree based on a number of the topic set;
[0121] a reply text determination unit 744 configured to determine corresponding reply text information based on the first text information, when the live topic repetition degree meets a second set condition.
[0122] In an implementation, the virtual characters include N virtual characters, N is a positive integer greater than 1, and the reply text determination unit is specifically configured to:
[0123] determine a target text generation model corresponding to a first virtual character among the N virtual characters based on a style of the first virtual character, where M is a positive integer greater than 1;
[0124] input the first text information into the target text generation model corresponding to the first virtual character to obtain reply text information of the first virtual character;
[0125] perform the following operations on an i-th virtual character among the N virtual characters:
[0126] determine a target text generation model corresponding to the i-th virtual character based on a style of the i-th virtual character, where i is a positive integer greater than 1;
[0127] input the first text information and the reply text information from the first virtual character to an (i-1)-th virtual character into the target text generation model corresponding to the i-th virtual character to obtain reply text information of the i-th virtual character.
[0128] In an implementation, the virtual character rendering module is specifically configured to:
[0129] render each virtual character based on the reply text information of the virtual character to obtain a second video stream of the virtual character.
[0130] In an implementation, the video generation module is specifically configured to:
[0131] According to the generation sequence of the reply text information of each virtual character, the first video stream and the second video stream of each virtual character are mixed to obtain a third video stream in which the anchor and each virtual character chat with each other.
[0132] In an implementation, the reply text determination unit is specifically configured to:
[0133] According to the style of the virtual character, the first text information is processed to obtain second text information;
[0134] According to a text generation model, the second text information is processed to obtain the reply text information of the virtual character.
[0135] The specific functions and examples of each module and sub-module of the apparatus of the embodiments of the present disclosure are described in the related description of the corresponding steps in the above method embodiments, which will not be described here.
[0136] In the technical solutions of the present disclosure, the acquisition, storage and application of user personal information comply with relevant laws and regulations and do not violate public order and good customs.
[0137] According to the embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium and a computer program product.
[0138] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present disclosure described and / or claimed in this document.
[0139] As shown in Figure 8 The device 800 includes a computing unit 801 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded into a random access memory (RAM) 803 from a storage unit 808. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0140] A plurality of components in the device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the device 800 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.
[0141] The computing unit 801 can be various general and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 801 performs various methods and processes described above, such as a live-scene-oriented video rendering method. For example, in some embodiments, a live-scene-oriented video rendering method can be implemented as a computer software program, which is tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded to the RAM 803 and executed by the computing unit 801, one or more steps in a live-scene-oriented video rendering method described above can be performed. Alternatively, in other embodiments, the computing unit 801 can be configured to perform a live-scene-oriented video rendering method by other any appropriate means, such as by means of firmware.
[0142] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0143] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces the functions / operations specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a standalone software package, or entirely on a remote machine or server.
[0144] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0145] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0146] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0147] The computer system can include clients and servers. This relationship can be. The servers are typically remote from the clients with the interactions between them occurring over a communication network. The relationship between client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The servers can be cloud servers, servers of a distributed system, or servers incorporating blockchain.
[0148] It should be understood that the various forms of flow shown above can be re-ordered, added to, or deleted from without departing from the scope of the present disclosure. For example, the steps recited in the present disclosure can be performed in parallel, in series, or in a different order, without departing from the desired results of the technology disclosed in the present disclosure, and are not limited herein.
[0149] The specific embodiments described above are not intended to be limiting. One of skill in the art will understand that various modifications, combinations, sub-combinations, and alternatives can be made to the specific embodiments described above without departing from the principles of the present disclosure. Any such modifications, equivalents, and alternatives are intended to be included within the scope of the present disclosure.
Claims
1. A video rendering method for live streaming scenarios, comprising: Record the live stream of the broadcaster to obtain the first video stream; Speech recognition is performed on the live audio in the first video stream to obtain the first text information; Determining the popularity of a live stream topic based on audience response information during the live stream recording process and the first text information includes: extracting at least one first text segment from the first text information; for each first text segment, searching the audience response information for a second text segment that can form a key-value pair with the first text segment and whose timestamp matches the timestamp of the first text segment, and counting the number of key-value pair pairs; and determining the popularity of the live stream topic based on the number of key-value pair pairs. When the popularity of the live topic meets the first set condition, based on the first text information, the corresponding reply text information is determined, including: based on the style of the i-th virtual character among N virtual characters, the corresponding target text generation model is determined among M text generation models; the first text information and the reply text information from the 1st virtual character to the (i-1th)th virtual character are input into the target text generation model corresponding to the i-th virtual character to obtain the reply text information of the i-th virtual character, where N, M and i are positive integers greater than 1; Based on the reply text information, the virtual character is rendered to obtain a second video stream; Based on the first video stream and the second video stream, a third video stream is generated showing the interaction between the anchor and the virtual character.
2. The method according to claim 1, wherein, When the popularity of the live stream topic meets a first predetermined condition, the corresponding reply text information is determined based on the first text information, including: When the popularity of the live broadcast topic meets the first set condition, keywords are extracted from the first text information to obtain multiple keywords; The multiple keywords are categorized into topics to obtain at least one topic set; The repetition rate of live-stream topics is determined based on the number of topics in the set. If the repetition rate of the live broadcast topic meets the second set condition, the corresponding reply text information is determined based on the first text information.
3. The method according to claim 1 or 2, wherein, The step of determining the corresponding reply text information based on the first text information includes: Based on the style of the first virtual character among the N virtual characters, the corresponding target text generation model is determined among the M text generation models; The first text information is input into the target text generation model corresponding to the first virtual character to obtain the reply text information of the first virtual character.
4. The method according to claim 3, wherein, The step of rendering the virtual character based on the reply text information to obtain a second video stream includes: Based on the reply text information of each virtual character, each virtual character is rendered to obtain a second video stream of each virtual character.
5. The method according to claim 4, wherein, The step of generating a third video stream based on the first video stream and the second video stream, whereby the anchor and the virtual character interact, includes: Based on the generation order of the reply text information of each virtual character, the first video stream and the second video streams of each virtual character are mixed to obtain a third video stream in which the anchor chats with each virtual character.
6. The method according to claim 1 or 2, wherein, The step of determining the corresponding reply text information based on the first text information includes: Based on the style of the virtual character, the first text information is processed to obtain the second text information; Based on the text generation model, the second text information is processed to obtain the virtual character's reply text information.
7. A live streaming device, comprising: The video recording module is used to record the live stream of the broadcaster and obtain the first video stream; The speech recognition module is used to perform speech recognition on the live audio in the first video stream to obtain the first text information; A popularity determination module is used to determine the popularity of a live stream topic based on audience response information during the live stream recording process and the first text information. The reply text determination module is used to determine the corresponding reply text information based on the first text information when the popularity of the live topic meets the first preset condition. The virtual character rendering module is used to render the virtual character based on the reply text information to obtain a second video stream; The video generation module is used to generate a third video stream based on the first video stream and the second video stream, in which the anchor and the virtual character interact. The virtual characters include N virtual characters, where N is a positive integer greater than 1. The reply text determination module is specifically used to: determine the corresponding target text generation model among the M text generation models based on the style of the i-th virtual character, where i is a positive integer greater than 1; input the first text information and the reply text information from the 1st virtual character to the (i-1th)th virtual character into the target text generation model corresponding to the i-th virtual character to obtain the reply text information of the i-th virtual character. The heat determination module includes: A text segment extraction unit is used to extract at least one first text segment from the first text information; The key-value pair counting unit is used to search for a second text segment in the audience response information for each of the first text segments, which can form a key-value pair with the first text segment and whose timestamp matches the timestamp of the first text segment, and to count the number of key-value pairs. The popularity determination unit is used to determine the popularity of the live broadcast topic based on the number of pairs of the key-value pairs.
8. The apparatus according to claim 7, wherein, The response text determination module includes: The keyword extraction unit is used to extract keywords from the first text information to obtain multiple keywords when the popularity of the live topic meets a first set condition. A topic classification unit is used to classify the multiple keywords into topics to obtain at least one topic set; The repetition determination unit is used to determine the repetition of live topics based on the set size of the topic collection. The reply text determination unit is used to determine the corresponding reply text information based on the first text information when the repetition rate of the live topic meets the second preset condition.
9. The apparatus according to claim 8, wherein, The virtual characters include N virtual characters, where N is a positive integer greater than 1. The reply text determination unit is specifically used for: Based on the style of the first virtual character among the N virtual characters, the corresponding target text generation model is determined from the M text generation models, where M is a positive integer greater than 1; The first text information is input into the target text generation model corresponding to the first virtual character to obtain the reply text information of the first virtual character.
10. The apparatus according to claim 9, wherein, The virtual character rendering module is specifically used for: Based on the reply text information of each virtual character, each virtual character is rendered to obtain a second video stream of each virtual character.
11. The apparatus according to claim 10, wherein, The video generation module is specifically used for: Based on the generation order of the reply text information of each virtual character, the first video stream and the second video streams of each virtual character are mixed to obtain a third video stream in which the anchor chats with each virtual character.
12. The apparatus according to claim 8, wherein, The reply text determination unit is specifically used for: Based on the style of the virtual character, the first text information is processed to obtain the second text information; Based on the text generation model, the second text information is processed to obtain the virtual character's reply text information.
13. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-6.
14. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-6.
15. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-6.
Citation Information
Patent Citations
Virtual robot multi-mode interaction method and system applied to video live-broadcasting platform
CN107423809A
Robot interaction method based on real-time video chatting scene
CN108416286A
Virtual robot interaction method and device and storage medium
CN116756285A