Intelligent generation method and device of live video stream in network live broadcast scenario
By automatically determining live streaming content data through intelligent generation equipment, the problem of low intelligence level in live streaming rooms under network live streaming scenarios is solved, and a more efficient live streaming room setup process is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NANJING SILICON INTELLIGENCE TECH CO LTD
- Filing Date
- 2023-04-18
- Publication Date
- 2026-04-24
AI Technical Summary
In current online live streaming scenarios, the level of intelligence in setting up live streaming rooms is low, the setup efficiency is slow, and a large amount of manual intervention is required.
The system acquires content demand data from service recipients through intelligent generation devices, automatically determines live streaming content data, and sends the live streaming content data to the live streaming data generation device. This data is then combined with the data from the process-driven device to generate a live streaming video stream.
It improves the intelligence level of the live streaming room construction process, reduces the complexity of data organization, and increases the efficiency of setup.
Smart Images

Figure CN118828031B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and in particular to an intelligent method and device for generating live video streams in a network live streaming scenario. Background Technology
[0002] With the development of live streaming technology, more and more users are choosing live streaming to achieve various purposes, such as live streaming games, videos, or selling products.
[0003] In the current online live streaming scenario, the entire process of planning and execution of a live streaming room requires human intervention. For example, the live streaming process is designed manually, the live streaming team is assembled manually, the background of the live streaming room is set up manually, the images and voices of the streamer are collected manually, etc. The level of intelligence is low and the setup efficiency is slow. Summary of the Invention
[0004] This invention provides an intelligent method and device for generating live video streams in online live streaming scenarios, to solve the problems of low intelligence and slow setup efficiency in existing live streaming room construction processes. Specifically, the embodiments of this application disclose the following technical solutions:
[0005] Firstly, this application provides an intelligent generation method for live video streams in a network live streaming scenario. This method is applied to an intelligent generation device in a live streaming system, which also includes a live data generation device connected to the intelligent generation device, and the live data generation device is further connected to a process-driven device. The method includes: the intelligent generation device acquiring content requirement data of a service recipient; the intelligent generation device determining live content data corresponding to the content requirement characteristics based on the content requirement data; and the intelligent generation device sending the live content data to the live data generation device, so that the live data generation device generates a live video stream for live streaming based on the live content data, as well as live process data and live rule data from the process-driven device.
[0006] In one possible implementation of the first aspect, the content requirement data includes image feature data and audio feature data; the intelligent generation device determines the live content data corresponding to the content requirement features based on the content requirement data, including: the intelligent generation device determines live image data matching the image feature data from an image library and determines live audio data matching the audio feature data from a speech library; the image library includes the correspondence between various image feature data and multiple live image data, and the speech library includes the correspondence between various audio feature data and multiple live audio data.
[0007] In one possible implementation of the first aspect, the content requirement data includes live stream types; the intelligent generation device determines live stream content data corresponding to the content requirement characteristics based on the content requirement data, including: the intelligent generation device determines live stream image data matching the live stream type from an image library and live stream audio data matching the live stream type from a voice library; the image library includes the correspondence between multiple live stream types and multiple live stream image data, and the voice library includes the correspondence between multiple live stream types and multiple live stream audio data.
[0008] In one possible implementation of the first aspect, the method further includes: the intelligent generation device receiving a first operation from the object; the intelligent generation device responding to the first operation to determine content adjustment information corresponding to the first operation; and the intelligent generation device adjusting the live content data according to the content adjustment information.
[0009] In one possible implementation of the first aspect, the live video stream includes live image data and live audio data, and the live image data includes the anchor's facial data; the first operation of the intelligent generation device receiving the object includes: the intelligent generation device receiving the service object's first selection operation of a first face image among multiple preset face images; the intelligent generation device responding to the first operation determines the content adjustment information corresponding to the first operation, including: the intelligent generation device responding to the first selection operation extracting facial features from the first face image to obtain facial feature data; the intelligent generation device adjusting the live content data according to the content adjustment information, including: the intelligent generation device acquiring the anchor's facial data from the live image data; the intelligent generation device fusing the facial feature data into the anchor's facial data to obtain fused target facial data, and replacing the anchor's facial data in the live video stream with the target facial data.
[0010] In one possible implementation of the first aspect, the live video stream includes live image data and live audio data. The intelligent generation device receives a first operation from the target, including: the intelligent generation device receives a second selection operation from the target among multiple preset timbres for a first timbre; the intelligent generation device responds to the first operation by determining content adjustment information corresponding to the first operation, including: the intelligent generation device responds to the second selection operation by extracting timbre features from the first timbre to obtain timbre feature data; the intelligent generation device adjusts the live content data according to the content adjustment information, including: the intelligent generation device fuses the live audio data and the timbre feature data to obtain fused target audio data, and replaces the live audio data in the live video stream with the target audio data.
[0011] Secondly, this application also provides an intelligent generation device for a live streaming scenario, comprising: an acquisition module for acquiring content demand data of a service object; a processing module for determining live streaming content data corresponding to the content demand data based on the content demand data acquired by the acquisition module; and a sending module for sending the live streaming content data obtained by the processing module to the live streaming data generation device, so that the live streaming data generation device can generate a live streaming video stream for live streaming based on the live streaming content data, as well as live streaming process data and live streaming rule data from the process-driven device.
[0012] Thirdly, this application provides an electronic device including a processor and a memory; the memory is used to store computer instructions, and when the electronic device is running, the processor executes the computer instructions stored in the memory to enable the electronic device to implement the intelligent generation method of live video stream in the network live streaming scenario provided in the first aspect above.
[0013] Fourthly, this application also provides a computer-readable storage medium storing instructions that, when executed on an electronic device, cause the electronic device to implement the intelligent generation method for live video streams in a network live streaming scenario as provided in the first aspect above.
[0014] Fifthly, this application also provides a computer program product containing computer instructions, which, when executed on an electronic device, enables the electronic device to implement the intelligent generation method for live video streams in a network live streaming scenario as provided in the first aspect above.
[0015] In this application, the aforementioned names do not limit the devices or functional modules themselves. In actual implementation, these devices or functional modules may appear under other names. As long as the functions of each device or functional module are similar to those in this application, they fall within the scope of the claims of this application and their equivalents.
[0016] These or other aspects of this application will become more readily apparent in the following description.
[0017] In the technical solution provided in this application, the intelligent generation device acquires the content requirement data of the service recipient; the intelligent generation device determines the live content data corresponding to the content requirement characteristics based on the content requirement data; the intelligent generation device sends the live content data to the live data generation device, so that the live data generation device generates a live video stream for live streaming based on the live content data, as well as the live process data and live rule data from the process-driven device. Applying the technical solution of this disclosure, the intelligent generation device automatically determines the live content data required for the live video stream based on the content requirement data of the service recipient, thereby enabling the live data generation device to smoothly combine the live process data and live rule data to generate a live video stream for live streaming. This technical solution does not require the service recipient to organize the live content data; instead, the intelligent generation device automatically organizes and generates it, improving the intelligence of the live streaming room construction process in existing network live streaming scenarios. Furthermore, this technical solution also reduces the complexity of data organization in the existing live streaming room construction process and improves data organization efficiency. Attached Figure Description
[0018] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a schematic diagram of the structure of an implementation environment provided in an embodiment of this application;
[0020] Figure 2 This application provides a schematic diagram of the structure of a live streaming system according to an embodiment of the present application.
[0021] Figure 3 A flowchart illustrating an intelligent generation method for live video streams in a network live streaming scenario provided in this application embodiment. Figure 1 ;
[0022] Figure 4 A flowchart illustrating an intelligent generation method for live video streams in a network live streaming scenario provided in this application embodiment. Figure 2 ;
[0023] Figure 5 A flowchart illustrating an intelligent generation method for live video streams in a network live streaming scenario provided in this application embodiment. Figure 3 ;
[0024] Figure 6 A flowchart illustrating an intelligent generation method for live video streams in a network live streaming scenario provided in this application embodiment. Figure 4 ;
[0025] Figure 7 This application provides a schematic diagram of a scenario for selecting a broadcaster.
[0026] Figure 8 A flowchart illustrating an intelligent generation method for live video streams in a network live streaming scenario provided in this application embodiment. Figure 5 ;
[0027] Figure 9 A schematic diagram illustrating a face selection operation scenario provided in an embodiment of this application;
[0028] Figure 10 A flowchart illustrating an intelligent generation method for live video streams in a network live streaming scenario provided in this application embodiment. Figure 6 ;
[0029] Figure 11 A schematic diagram illustrating a timbre selection operation as provided in an embodiment of this application;
[0030] Figure 12 This is a schematic diagram of the structure of an intelligent generation device provided in an embodiment of this application;
[0031] Figure 13 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0032] To enable those skilled in the art to better understand the technical solutions in the embodiments of this application, and to make the above-mentioned objectives, features, and advantages of the embodiments of this application more apparent, the technical solutions in the embodiments of this application will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are all within the protection scope of this application.
[0033] To make the objectives and implementation methods of this application clearer, the exemplary implementation methods of this application will be clearly and completely described below with reference to the accompanying drawings of the exemplary embodiments of this application. Obviously, the exemplary embodiments described are only some embodiments of this application, and not all embodiments.
[0034] It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.
[0035] The terms "first," "second," "third," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar or related objects or entities, and do not necessarily imply a specific order or sequence, unless otherwise specified. It should be understood that such terms are interchangeable where appropriate.
[0036] The terms “comprising” and “having”, and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device that includes a range of components is not necessarily limited to all of the components that are clearly listed, but may include other components that are not clearly listed or that are inherent to such product or device.
[0037] The term "module" refers to any known or subsequently developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code that is capable of performing the functions associated with that element.
[0038] The term "live video stream" refers to a sequence of digitally encoded data used to transmit live video from a live streaming system to any live streaming platform or terminal without significant delay. Therefore, as used in this disclosure, a live video stream includes data transmission from one computing device to another (or to multiple computing devices) (including servers or groups of servers). A live video stream also includes unidirectional broadcasting of data from the live streaming system (or group of servers) to the live streaming platform or user terminal (or to multiple viewer devices).
[0039] In the current online live streaming scenario, the entire process from planning to execution of the live stream requires human intervention, resulting in a low level of automation.
[0040] To address the aforementioned issues, this application provides an intelligent generation method for live video streams in a network live streaming scenario. This method is applied to an intelligent generation device within a live streaming system. The live streaming system also includes a live data generation device connected to the intelligent generation device, and the live data generation device is further connected to a process-driven device. In this method, the intelligent generation device acquires content requirement data of the service recipient; the intelligent generation device determines live content data corresponding to the content requirement characteristics based on the content requirement data; the intelligent generation device sends the live content data to the live data generation device, enabling the live data generation device to generate a live video stream for live streaming based on the live content data, as well as live process data and live rule data from the process-driven device.
[0041] By applying the technical solution disclosed herein, an intelligent generation device automatically determines the live content data required for the live video stream based on the content needs data of the service recipient. This enables the live data generation device to smoothly combine live process data and live rule data to generate the live video stream for live streaming. This technical solution eliminates the need for the service recipient to organize the live content data; instead, the intelligent generation device automatically organizes and generates it, improving the intelligence of the live streaming room construction process in existing online live streaming scenarios. Furthermore, this technical solution reduces the complexity of data organization in existing live streaming room setup processes, improving data organization efficiency.
[0042] The technical solutions provided in the embodiments of this application can be applied to, for example... Figure 1 The implementation environment shown may include a live streaming system 01, a live streaming platform 02, and a user terminal 03. The live streaming system 01 and the live streaming platform 02 can communicate via wired or wireless communication methods, and the live streaming platform 02 and the user terminal 03 can also communicate via wired or wireless communication methods.
[0043] The live streaming system 01 is mainly used to generate a live video stream, which is then pushed to the live streaming platform 02 for live streaming. In this embodiment, there may be one or more live streaming platforms 02. Figure 1 This example uses only one live streaming platform, 02, and does not impose any specific restrictions.
[0044] The live streaming platform 02 is mainly used to push the live streaming data stream to the user terminal 03 according to specific rules after receiving the live video stream from the live streaming system 01, so that the user terminal 03 can play the live data corresponding to the live video stream for the user to watch. In this embodiment, there may be one or more user terminals 03. The user terminal 03 can be an electronic device with a live streaming application corresponding to the live streaming platform 02 installed.
[0045] Reference Figure 2 As shown, the live streaming system 01 provided in this embodiment may specifically include a process-driven device 011, a live streaming data generation device 012, and an intelligent generation device 013. The process-driven device 011 and the live streaming data generation device 012 can communicate via wired or wireless means. The live streaming data generation device 012 and the intelligent generation device can also communicate via wired or wireless means.
[0046] The process-driven device 011 can acquire user process requirement data and determine the corresponding live streaming process data and live streaming rule data. Then, the process-driven device 011 can send the live streaming process data and live streaming rule data to the live streaming data generation device 012, so that the live streaming data generation device 012 can generate a live video stream for live streaming based on the live streaming process data, live streaming rule data, and live streaming content data from the intelligent generation device.
[0047] It should be noted that in this application, the process-driven device 011, the live data generation device 012, and the intelligent generation device 013 can be separate devices, different parts of the same device, or modules in a data center that implement different functions, depending on actual needs. This application does not impose any specific restrictions on these.
[0048] Based on the aforementioned live streaming system, referring to Figure 3 As shown, this application embodiment provides an intelligent generation method for live video streams in a network live streaming scenario, which may include S301-S303:
[0049] S301, Intelligent generation device acquires content requirement data of service objects.
[0050] In this embodiment, the content demand data can be characteristic content that reflects the service recipient's demand for content in the live stream. Examples include "selling goods," "baby products," "milk powder," "humorous host," "mature female voice," and "warm lighting."
[0051] In one possible implementation, the service object can directly provide content requirement data to the process-driven device. For example, the service object can input content requirement data on its own service terminal or process-driven device, so that the process-driven device can obtain the content requirement data. Of course, the service object can input the content requirement data in any feasible way, such as keyboard input, voice input, etc. This application does not impose specific restrictions on this.
[0052] In another possible implementation, refer to Figure 2 As shown, the live streaming system may also include a data sensing device 014. The intelligent generation device can obtain the content requirement data of the service recipients through the data sensing device.
[0053] In some embodiments, the data sensing device may include an order system. Service recipients can publish content requirement descriptions on the order system, such as, "I want live-streaming content selling maternity and baby products, preferably featuring a beautiful female host with a gentle, intellectual voice, and ideally in a maternity and baby product store setting." The data sensing device can obtain the service recipient's content requirement description through the order system. After obtaining the service recipient's content requirement description, the data sensing device can preprocess the description to obtain the corresponding content requirement data.
[0054] Typically, service recipients can input their content requirements in the order system in several different ways. For example, they can input the content requirements in text format, or in audio format, or by selecting from multiple requirement options provided by the order system.
[0055] Since service recipients can input their content requirements in the order system in various formats, the preprocessing method of the data sensing device will also be adjusted according to the format of the input content requirements. For example, if the service recipient inputs the content requirements in text format, the data sensing device can perform semantic analysis on the text and output the corresponding semantic result, thus obtaining the content requirements data. As another example, if the service recipient inputs the content requirements in audio format, the data sensing device can first convert the audio to text, then perform semantic analysis on the text and output the corresponding semantic result, thus obtaining the content requirements data. Yet another example is when the service recipient selects a content requirements description from multiple options provided by the order system; the data sensing device can then obtain the content requirements data based on the selected option.
[0056] For example, a service recipient might input a content requirement description like, "I want a live stream process for selling maternity and baby products, with a female host who has a mature, sophisticated style and preferably a gentle voice." The data sensing device preprocesses this content requirement description, outputting content requirement data such as "live stream sales room," "maternity and baby products," "mature, sophisticated host," and "gentle voice." After obtaining this content requirement data, it can be sent to an intelligent generation device, which then determines the corresponding live stream content data based on this data.
[0057] In some possible embodiments, the content requirement description of the service object can be directly the content requirement data. In this case, the preprocessing of the data sensing device can be no processing or deduplication and other processing operations that do not affect the content.
[0058] S302. The intelligent generation device determines the live content data corresponding to the content requirement characteristics based on the content requirement data.
[0059] The live stream content data includes live image data and live audio data. This data helps enrich the live stream content, thereby attracting more users to stay. For example, in addition to live image and audio data, the live stream content data can also include all the specific content required by the target audience for the live stream. This includes, for example, the host's information, script information, live stream background information, and product information.
[0060] In this embodiment of the application, to avoid the personal influence of the broadcaster on the live stream from being affected by the external environment, the broadcaster can be a digital human broadcaster. The broadcaster information can include the digital human broadcaster's image information and voice information. The image information can include the digital human broadcaster's expressions, movements, clothing, body characteristics, facial features, etc. The voice information can include the digital human broadcaster's voice characteristics (e.g., timbre).
[0061] The script information can include the scripts spoken by the digital human anchor during different live stream segments and the scripts corresponding to the interaction rules. For example, the scripts corresponding to the live stream segments can specifically include opening scripts, warm-up scripts, scripts to get viewers excited, scripts to pamper fans, and closing scripts. The scripts corresponding to the interaction rules can include: promotional scripts (used to promote products), guiding scripts (used to guide users to buy products), explanation scripts (used to explain products), interactive scripts (used to enhance the atmosphere of the live stream), bullet screen response scripts (used to respond to bullet screen questions, etc.), and event response scripts (used to respond to specific live stream events), etc. Taking interactive scripts as an example, when viewers in the live stream request gifts, the digital human anchor can say something like, "Please like this post a lot. The more likes, the more gifts the anchor can apply for!"
[0062] The background information of a live stream can include the background features of the entire live stream. For example, the background information of the live stream may include: specific background content (such as a textured kitchen background), background lighting parameters (brightness, orientation, etc.), background music, and parameters of the decorations in the background (position, size, etc.).
[0063] Product information can include content related to the products for sale, such as corresponding sales pitches and product materials (product images, product videos, product inspection reports, etc.).
[0064] After acquiring content requirement data, the intelligent generation device will determine how it acquires live content data based on the different content requirement data. Specifically, this can include the following two methods:
[0065] In the first feasible approach, if the content requirement data includes image feature data and audio feature data, then combine... Figure 3 , refer to Figure 4 As shown, S302 may specifically include S302A:
[0066] S302A, the intelligent generation device determines live image data that matches the image feature data from the image library, and determines live audio data that matches the audio feature data from the voice library.
[0067] Image feature data can be data indicating the anchor and scene in the live stream, such as "mature and sophisticated style," "kitchen background," and "warm sunlight." Audio feature data can be data indicating the audio during the live stream, such as "mature and sophisticated voice," "gentle tone," and "enthusiastic." Content requirement data can include multiple image feature data and multiple audio feature data.
[0068] In this embodiment of the application, in order to facilitate the intelligent generation device to obtain the process library and rule library based on image feature data and audio feature data, the intelligent generation device may include a process library and a rule library.
[0069] The image library includes multiple live-stream image data sets, such as image information for digital avatar anchors in various styles (e.g., mature woman style, chef outfit, suit), as well as images of baby and maternity products, kitchenware, baby and maternity product stores, and kitchen utensil stores. The library also includes various mappings between image feature data and live-stream image data. For example, the image feature data "mature woman style" corresponds to the image information of a digital avatar anchor in that style, while the image feature data "baby and maternity products" corresponds to scene images of baby and maternity product stores.
[0070] The voice library can include various types of live audio data. For example, it can contain voice data for a mature, sophisticated woman; live audio data for sales scripts on maternal and infant products; live audio data for sales scripts on kitchen utensils; live audio data for warm-up conversations; live audio data for responses to comments on live streams, and so on. The voice library also includes the correspondence between various audio feature data and multiple live audio data sets. For example, the audio feature data "mature, sophisticated woman" corresponds to the voice data for a mature, sophisticated woman; the audio feature data "maternal and infant product sales" corresponds to live audio data for sales scripts on maternal and infant products, and so on.
[0071] In some embodiments, the voice library may also include various types of live text data (or live speech data) and the correspondence between various audio feature data and various types of live text data. After the intelligent generation device determines the live text data based on the audio feature data, it can also use a preset text-to-speech model to convert the live text data into live audio data. The specific content included in the voice library can be determined according to actual needs, and this application does not impose specific limitations on it.
[0072] It should be noted that the content included in the image library can cover the image content in the anchor information, live room background information, and product information provided in the aforementioned embodiments, and the content rules included in the voice library can cover the text and / or audio content in the anchor information, script information, live room background information, and product information provided in the aforementioned embodiments.
[0073] In other embodiments, when a service recipient needs to create a live streaming room, the intelligent generation device can provide the service recipient with all the content required to create the live streaming room, such as multiple live streaming content data (including live streaming image data and live streaming audio data, or including anchor information, script information, live streaming room background information, product information, etc.). Users can filter target live streaming content data from the multiple live streaming content data provided by the intelligent generation device according to their needs. Then, the live streaming data generation device can generate the final live streaming video stream based on the filtered target live streaming content data and the live streaming rule data and live streaming process data provided by the process device.
[0074] Based on the technical solution corresponding to S032A, the intelligent generation device can obtain live content data (i.e., live image data and live audio data) that meets the needs of the service recipients, thereby improving data support for the subsequent generation of live video streams.
[0075] In the second feasible approach, if the content requirement data includes the live streaming type, then it is combined with... Figure 3 , refer to Figure 5 As shown, S302 may specifically include S302B:
[0076] The S302B intelligent generation device determines the live image data matching the live type from the process library and the live audio data matching the live type from the rule library, based on the live type.
[0077] The image library includes multiple live image datasets. The specific content of the live image datasets can be found in the descriptions in the preceding embodiments, and will not be repeated here. The image library also includes the correspondence between various live streaming types and the multiple live image datasets.
[0078] The audio library may include various types of live audio data. The specific content of the live audio data can be found in the relevant descriptions in the foregoing embodiments, and will not be repeated here. The image library also includes the correspondence between various live streaming types and multiple live audio data sets.
[0079] Typically, the types of live streaming in the market can include casual chat live streams, game live streams (used to showcase playing games), e-commerce live streams (used to sell and convert products), and knowledge-sharing live streams (used to explain knowledge and skills in a certain field, share on a certain topic, or sell a certain book).
[0080] Since the live content of different live streaming types varies slightly, the live image data and live audio data of several common types of live streaming can be sorted out in advance. Then, the live image data of multiple live streaming types and the live audio data of each live streaming type can be stored in the image library, and the live audio data of multiple live streaming types and the live audio data of each live streaming type can be stored in the audio library.
[0081] When the intelligent generation device needs to determine the live image data and live audio data corresponding to the content requirement data, it can select the live image data and live audio data corresponding to the live type from the image library and the audio library, based on the live type included in the content requirement data. For example, the image library may include at least live image data corresponding to everyday chat-type live streams, game live streams, e-commerce live streams, and knowledge-sharing live streams. The audio library may include at least live audio data corresponding to everyday chat-type live streams, game live streams, e-commerce live streams, and knowledge-sharing live streams.
[0082] For example, if the live streaming type included in the content requirement data is "e-commerce sales live streaming", the intelligent generation device can determine the live streaming image data corresponding to the e-commerce sales live streaming from the image library. Specifically, it can include: the image information of the sexy female anchor, the background information of the sales live streaming, etc.
[0083] The intelligent generation device can also identify the corresponding live audio data for e-commerce sales-oriented live streams from the voice library, such as the voice timbre of a mature and sophisticated woman and audio data of sales-oriented speech.
[0084] Based on the technical solution corresponding to S032B, and based on the image library and voice library, the intelligent generation device can obtain live image data and live audio data that meet the needs of the service recipients, providing data support for the subsequent generation of live video streams.
[0085] S303. The intelligent generation device sends live content data to the live data generation device, so that the live data generation device can generate a live video stream for live streaming based on the live content data, as well as the live process data and live rule data from the process-driven device.
[0086] Following S303, the live streaming data generation device can generate a live video stream for live streaming based on live content data, as well as live streaming process data and live streaming rule data from the process-driven device (specifically, it can push the live video stream to a live streaming platform, so that the live streaming platform can push the live video stream to the user terminal for live streaming, thereby displaying the live room to the viewing user). The live video stream can include live image data and live content data.
[0087] In the technical solution provided in this application embodiment, an intelligent generation device automatically determines the live content data required for the live video stream based on the content demand data of the service recipient. This enables the live data generation device to smoothly combine live process data and live rule data to generate the live video stream for live streaming. This technical solution eliminates the need for the service recipient to organize the live content data; instead, the intelligent generation device automatically organizes and generates it, improving the intelligence of the live streaming room construction process in existing network live streaming scenarios. Furthermore, this technical solution reduces the complexity of data organization in the existing live streaming room construction process and improves data organization efficiency.
[0088] In some embodiments, after the intelligent generation device determines the live content data, it can display the live content data in any manner and allow the service recipient to adjust the live content data. Subsequently, the live content data can be adjusted based on the adjustment operation to improve the user experience. Based on this, referring to... Figure 6 As shown, the intelligent generation method for live video streams in network live streaming scenarios provided in this application embodiment further includes S601-S603:
[0089] S601, The intelligent generation device receives the first operation from the service object.
[0090] For example, the first operation may include a content deletion operation, a content adjustment operation, and a content addition operation. For instance, a content deletion operation may involve deleting a first piece of content (e.g., background content) from the live stream content data. A content adjustment operation may involve adjusting a second piece of content (e.g., lighting content) from the live stream content data (e.g., adjusting the lighting angle). A content addition operation may involve adding a third piece of content (e.g., interactive dialogue / audio) from the live stream content data.
[0091] S602, the intelligent generation device responds to the first operation and determines the content adjustment information corresponding to the first operation.
[0092] The content adjustment information can be the specific content corresponding to the first operation. For example, if the first operation is to delete the first content (e.g., background content) in the live content data, then the content adjustment information would be the deletion of the first content in the live content data, or a related instruction (e.g., the delete instruction) indicating the deletion of the first content in the live content data.
[0093] S603, the intelligent generation device adjusts information and data for live streaming content based on the content.
[0094] For example, if the content adjustment information is to delete the first content in the live content data, then S603 specifically means that the intelligent generation device deletes the first content (e.g., background content) according to the content adjustment information, that is, deletes the background content.
[0095] like Figure 7 As shown, the intelligent generation device receives a streamer selection operation from the service recipient, which specifically involves selecting a live streamer for the live broadcast room. For example, the service recipient selects from... Figure 7 From the digital human anchor A and digital human anchor B shown, select the target digital human anchor A. In response to the anchor selection operation of the service recipient, determine the content adjustment information corresponding to that anchor selection operation. Based on the content adjustment information, adjust the live anchor in the live broadcast room. For details on the selectable objects for the anchor selection operation, see [link to relevant documentation]. Figure 7 The content shown.
[0096] Based on the above technical solution, users can arbitrarily change and adjust the live content data according to their own needs, so that the live content data in the live video stream better meets the needs of the users and improves the user experience.
[0097] In some embodiments, to improve the user experience of the live streaming system, the intelligent generation device can pre-store multiple preset facial images. Whether before or after generating the live video stream, users can select a face from these preset images to adjust the appearance of the live stream host, making the adjusted host more attractive to viewers. Based on this, in Figure 6 Based on the illustrated embodiment, referring to Figure 8 As shown, S601 may specifically include S601a, S602 may specifically include S602a, and S603 may specifically include S6031 and S6032:
[0098] S601a, The intelligent generation device receives the first selection operation of the first face image from multiple preset face images from the service object.
[0099] The first operation mentioned above refers to the service recipient's first selection operation of the first face image among multiple preset face images.
[0100] For example, see Figure 9 As shown, the user can select the first face image 901 from multiple preset face images provided by the intelligent generation device as the face image they need.
[0101] S602a, the intelligent generation device responds to the first selection operation by extracting facial features from the first face image to obtain facial feature data.
[0102] Upon receiving a selection operation from a service recipient for a first facial image, the intelligent generation device can respond to this selection operation by extracting facial features from the first facial image to obtain facial feature data. The extracted facial feature data can then be integrated into the live streamer's image.
[0103] For example, facial feature extraction of a first face image to obtain facial feature data can be achieved by inputting the first face image into a facial feature extraction model, and the facial feature extraction model outputting the facial feature data corresponding to the first face image. In this embodiment, the facial feature extraction model can be obtained by training a convolutional neural network based on multiple face images and their facial feature data. The specific training method can be any feasible method, and this embodiment does not impose any specific limitations on it.
[0104] S6031, Intelligent generation device acquires the anchor's facial data from live broadcast image data.
[0105] In this embodiment of the application, the live video stream includes live image data and live audio data.
[0106] The method of obtaining the anchor's facial data can be any feasible method, and this application does not impose any specific restrictions on it.
[0107] S6032, The intelligent generation device integrates facial feature data into the anchor's facial data to obtain the integrated target facial data, and uses the target facial data to replace the anchor's facial data in the live video stream.
[0108] After obtaining the facial feature data of the streamer's face image and the first face image, the facial feature data of the streamer's face image and the first face image can be fused to obtain the fused target face data. The fused target face data is more in line with user expectations than the live streamer's face image.
[0109] Based on the above technical solution, users can arbitrarily change and adjust the face of the streamer in the live content data according to their own needs, so that the image of the live streamer in the final live video stream better meets the user's needs and improves the user experience.
[0110] In some embodiments, to improve the user experience of the live streaming system, the intelligent generation device can pre-store multiple timbre data sets. Whether before or after generating the live video stream, users can select a specific timbre data set to adjust the voice of the live stream host, making the adjusted voice more appealing to viewers. Based on this, in Figure 6 Based on the illustrated embodiment, referring to Figure 10 As shown, S601 may specifically include S601b, S602 may specifically include S602b, and S603 may specifically include S603b:
[0111] S601b, the intelligent generation device receives a second selection operation of the first timbre from multiple preset timbres by the service object.
[0112] The second selection operation of the service recipient on the first timbre among multiple preset timbres is the aforementioned first operation.
[0113] For example, see Figure 11 As shown, the user can select the first timbre 111 from the multiple preset timbres provided by the intelligent generation device as the face image they need.
[0114] S602b, the intelligent generation device responds to the second selection operation by extracting timbre features from the first timbre to obtain timbre feature data.
[0115] When a service user wants to optimize the timbre of the live audio data generated by the intelligent generation device using a first timbre, the intelligent generation device can, upon receiving a second selection operation from the service user regarding the first timbre, extract timbre features from the first timbre in response to this selection operation. The extracted timbre features are then fused into the voice of the live streamer.
[0116] For example, timbre feature extraction of the first timbre can be achieved by extracting its frequency domain features and time domain features separately, thereby obtaining timbre feature data corresponding to the first timbre. Specifically, the MFCC algorithm can be used to extract timbre features from the first timbre to obtain timbre feature data. Alternatively, the first timbre can be input into a timbre extraction model, and the timbre extraction model can output the timbre feature data corresponding to the first timbre. In this embodiment, the timbre extraction model can be obtained by training a neural network model based on multiple timbres and their timbre feature data. The specific training method can be any feasible method, and this embodiment does not impose any specific limitations on it.
[0117] The S603b intelligent generation device fuses live audio data and timbre feature data to obtain fused target audio data, and uses the target audio data to replace the live audio data in the live video stream.
[0118] After obtaining the timbre feature data of the first timbre, the live audio data and the timbre feature data can be fused to obtain the fused target audio data. The fused target audio data is more in line with user expectations than the timbre of the anchor in the original live video stream.
[0119] Based on the above technical solution, users can arbitrarily change and adjust the anchor's voice in the live broadcast content data according to their own needs, so that the live broadcast audio in the final live video stream better meets the user's needs and improves the user experience.
[0120] The foregoing mainly describes the solutions provided by the embodiments of this application from a methodological perspective. To achieve the above functions, it includes corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should readily recognize that, based on the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0121] This application embodiment can divide the intelligent generation device into functional modules according to the above method example. For example, each function can be divided into its own functional module, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware or as a software functional module. It should be noted that the module division in this application embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.
[0122] This application embodiment can also divide the intelligent generation device mentioned in the foregoing embodiments into functional modules according to different functions of the device. For example, each function can be divided into its own functional module, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware or as a software functional module. It should be noted that the module division in this application embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.
[0123] This application provides an intelligent generation device. This intelligent generation device is used to implement the aforementioned video data stream generation method. The functions of this intelligent generation device can be implemented through hardware or through hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions.
[0124] When intelligent generation devices are divided into functional modules, Figure 12 This illustrates one possible scenario for an intelligent generation device. (Refer to...) Figure 12 As shown, the intelligent generation device includes an acquisition module 121, a processing module 122, and a sending module 123.
[0125] The acquisition module 121 is used to acquire content requirement data of the service object; the processing module 122 is used to determine the live content data corresponding to the content requirement data based on the content requirement data acquired by the acquisition module 121; and the sending module 123 is used to send the live content data obtained by the processing module 122 to the live data generation device, so that the live data generation device can generate a live video stream for live streaming based on the live content data, as well as the live process data and live rule data from the process driving device.
[0126] Optionally, the content requirement data includes image feature data and audio feature data; the processing module 122 is specifically used to: determine live image data that matches the image feature data from the image library, and determine live audio data that matches the audio feature data from the voice library; the image library includes the correspondence between various image feature data and multiple live image data, and the voice library includes the correspondence between various audio feature data and multiple live audio data.
[0127] Optionally, the content requirement data includes the live stream type; the processing module 122 is specifically used to: determine the live stream image data that matches the live stream type from the image library, and determine the live stream audio data that matches the live stream type from the audio library; the image library includes the correspondence between multiple live stream types and multiple live stream image data, and the audio library includes the correspondence between multiple live stream types and multiple live stream audio data.
[0128] Optionally, the acquisition module 121 is further configured to receive a first operation of the object; the processing module 122 is further configured to respond to the first operation received by the acquisition module 121, determine the content adjustment information corresponding to the first operation; the processing module 122 is further configured to adjust the live content data according to the content adjustment information.
[0129] Optionally, the live video stream includes live image data and live audio data, and the live image data includes the anchor's facial data; the acquisition module 121 is specifically used to receive the selection operation of the service object on the first face image from multiple preset face images; the processing module 122 is specifically used to extract facial features from the first face image in response to the first selection operation received by the acquisition module 121 to obtain facial feature data; the processing module 122 is also specifically used to acquire the anchor's facial data in the live image data, fuse the facial feature data into the anchor's facial data to obtain the fused target facial data, and replace the anchor's facial data in the live video stream with the target facial data.
[0130] Optionally, the live video stream includes live image data and live audio data. The acquisition module 121 is specifically used to receive a second selection operation of the service object on the first timbre among multiple preset timbres. The processing module 122 is specifically used to: in response to the second selection operation received by the acquisition module 121, extract timbre features from the first timbre to obtain timbre feature data; fuse the live audio data and the timbre feature data to obtain fused target audio data, and replace the live audio data in the live video stream with the target audio data.
[0131] The effects that the aforementioned intelligent generation device can achieve can be referred to the effects of the intelligent generation method for live video streams in the network live streaming scenario in the foregoing embodiments, and will not be repeated here.
[0132] Figure 13 A structural diagram of an electronic device provided in this disclosure embodiment, such as... Figure 13 As shown, the electronic device 1300 includes one or more processors 1301 and a memory 1302. This electronic device can be the intelligent generation device described in the foregoing embodiments, or it can be a device that carries a live streaming system.
[0133] The processor 1301 may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 1300 to perform desired functions.
[0134] The memory 1302 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 1301 may execute the program instructions to implement the intelligent generation method for live video streams in network live streaming scenarios according to the various embodiments of this disclosure described above.
[0135] In one example, the electronic device 1300 may also include an input device 1303 and an output device 1304, which are interconnected via a bus system and / or other forms of connection mechanism (not shown).
[0136] Of course, for the sake of simplicity, Figure 13 Only some of the components of the electronic device 1300 relevant to this disclosure are shown, omitting components such as buses, input / output interfaces, etc. In addition, the electronic device 1300 may include any other suitable components depending on the specific application.
[0137] In addition to the methods and devices described above, embodiments of this disclosure may also be computer program products, which include computer program instructions that, when executed by a processor, cause the processor to perform the intelligent generation method for live video streams in the network live streaming scenario described in the foregoing embodiments.
[0138] Computer program products can be written in any combination of one or more programming languages to perform the operations of embodiments of this disclosure. The programming languages include object-oriented programming languages such as Java and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on a user's computing device, partially on a user's computing device, as a standalone software package, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0139] Furthermore, embodiments of this disclosure may also be computer-readable storage media storing computer program instructions, which, when executed by a processor, cause the processor to perform the intelligent generation method for live video streams in the network live streaming scenario described in the foregoing embodiments.
[0140] Computer-readable storage media may take the form of any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may, for example, include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0141] The basic principles of this disclosure have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.
[0142] The block diagrams of devices, apparatuses, devices, and systems disclosed herein are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.
[0143] It should also be noted that in the apparatus, devices, and methods of this disclosure, the components or steps can be disassembled and / or recombined. These disassemblies and / or recombinations should be considered as equivalent solutions to this disclosure.
[0144] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to be carried out within the widest scope consistent with the principles and novel features disclosed herein.
[0145] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.
Claims
1. A method for intelligently generating live video streams in a network live streaming scenario, applied to an intelligent generation device in a live streaming system, wherein the live streaming system further includes a live data generation device connected to the intelligent generation device, and the live data generation device is further connected to a process driving device, characterized in that, The live streaming system includes a process-driven device, a live streaming data generation device, and an intelligent generation device; wherein the process-driven device and the live streaming data generation device communicate with each other via wired or wireless means; and the live streaming data generation device and the intelligent generation device communicate with each other via wired or wireless means. The process-driven device is used to acquire user process requirement data and generate live streaming process data and live streaming rule data based on the user process requirement data. The live streaming data generation device is used to generate a live video stream for live streaming based on the live streaming process data, the live streaming rule data, and the live streaming content data. The method includes: The intelligent generation device acquires content requirement data of the service recipients; the content requirement data includes image feature data and audio feature data. The intelligent generation device determines live image data matching the image feature data from an image library and live audio data matching the audio feature data from a voice library based on the image feature data. The image library includes the correspondence between various image feature data and multiple live image data, and the voice library includes the correspondence between various audio feature data and multiple live audio data. The intelligent generation device sends multiple live content data to the live data generation device. The live content data includes live image data and live audio data, so that the live data generation device can filter out target live content data from the multiple live content data according to user needs. Based on the target live content data, as well as the live process data from the process driving device and the live rule data, a live video stream is generated for live streaming. The intelligent generation device receives the first operation from the object; The intelligent generation device responds to the first operation and determines the content adjustment information corresponding to the first operation; The intelligent generation device adjusts the live stream content data based on the content adjustment information; The live video stream includes live image data and live audio data. The intelligent generation device receives a first operation from the object, including: the intelligent generation device receives a second selection operation from the service object on a first timbre among multiple preset timbres. The intelligent generation device responds to the first operation and determines the content adjustment information corresponding to the first operation, including: the intelligent generation device responds to the second selection operation and extracts timbre features from the first timbre to obtain timbre feature data; The intelligent generation device adjusts the live content data according to the content adjustment information, including: the intelligent generation device fuses the live audio data and the timbre feature data to obtain fused target audio data, and uses the target audio data to replace the live audio data in the live video stream; The first operation of the intelligent generation device receiving the object includes: the intelligent generation device receiving the service object's first selection operation of a first face image among multiple preset face images; The intelligent generation device responds to the first operation and determines the content adjustment information corresponding to the first operation, including: the intelligent generation device responds to the first selection operation and extracts facial features from the first face image to obtain facial feature data; The intelligent generation device adjusts the live broadcast content data according to the content adjustment information, including: the intelligent generation device acquiring the anchor's facial data in the live broadcast image data; the intelligent generation device fusing the facial feature data into the anchor's facial data to obtain fused target facial data, and using the target facial data to replace the anchor's facial data in the live broadcast video stream.
2. The method according to claim 1, characterized in that, The content requirement data includes the live stream type; the intelligent generation device determines the live stream content data corresponding to the content requirement data based on the content requirement data, including: The intelligent generation device determines live image data matching the live stream type from an image library and live audio data matching the live stream type from a voice library, based on the live stream type. The image library includes a correspondence between multiple live stream types and multiple live image data, and the voice library includes a correspondence between multiple live stream types and multiple live audio data.
3. An intelligent generation device for online live streaming scenarios, applied to a live streaming system, wherein the live streaming system further includes a live streaming data generation device connected to the intelligent generation device, and the live streaming data generation device is further connected to a process driving device, characterized in that... The intelligent generation device is used to execute the intelligent generation method for live video streams in a network live streaming scenario as described in claim 1; the intelligent generation device includes: The acquisition module is used to acquire content requirement data of the service object; the content requirement data includes image feature data and audio feature data; The processing module is used to determine, based on the image feature data, live image data matching the image feature data from an image library, and live audio data matching the audio feature data from a speech library; the image library includes the correspondence between various image feature data and multiple live image data, and the speech library includes the correspondence between various audio feature data and multiple live audio data; The sending module is used to send multiple live content data to the live data generation device. The live content data includes live image data and live audio data, so that the live data generation device can filter out target live content data from the multiple live content data according to user needs, and generate a live video stream for live streaming based on the target live content data, as well as live process data and live rule data from the process driving device.
4. An electronic device, characterized in that, include: A memory and a processor, wherein the memory is used to store computer programs; The processor is used to cause the electronic device to perform the intelligent generation method for live video streams in a network live streaming scenario as described in claim 1 or 2 when executing a computer program.
5. A computer-readable storage medium, characterized in that, It includes at least one instruction that, when executed on an electronic device, causes the electronic device to perform the intelligent generation method for live video streams in a network live streaming scenario as described in claim 1 or 2.
6. A live streaming system, characterized in that, This includes the intelligent generation device as described in claim 3 or the electronic device as described in claim 4.
Citation Information
Patent Citations
Live broadcast method and system based on artificial intelligence
CN113873286A