Live broadcast risk control false touch prevention method, device and equipment and product
Patent Information
- Application Number
- CN202410286589.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-13
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2044-03-13
AI Technical Summary
[0004]虚拟直播活动中的视频流,通常是基于一定时长的素材视频进行驱动生成的,而素材视频中的图像帧总是有限的,当长时间生成直播所需的视频流时,由于风控系统在技术中的滞后性,风控系统容易将虚拟直播活动中的视频流误识别为机器人代播行为,对虚拟直播活动做出错误的干涉,例如将其停播或者禁播等,导致电商直播活动无法正常进行,造成电商店铺和平台方的巨大经济损失,也破坏了用户体验,同时限制了虚拟人直播这一新技术的发展,因而,有必要结合风控系统所存在的技术缺陷,改进虚拟直播活动相关技术
[0019] Compared to existing technologies, this application, when a live streamer starts a live stream room to execute a virtual live stream activity, first calls the script list from the script server through the script generation service to obtain the script texts corresponding to each business link of the same live stream business process. Then, the video generation service drives the video server to instantly generate the human voice video corresponding to each script text based on the template video that has reached the preset regulation duration. Finally, the virtual camera driving service pushes these human voice videos to the e-commerce live stream room to implement the virtual live stream activity.
Smart Images

Figure CN118158450B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of e-commerce information security technology, and in particular to a method, device, equipment and product for preventing accidental touches during live streaming risk control. Background Technology
[0002] E-commerce platforms all deploy risk control systems to identify various non-compliant behaviors of e-commerce users, enabling timely regulatory actions to maintain platform information security. With the development of live streaming technology, e-commerce platforms and live streaming have become deeply integrated. E-commerce platforms are gradually strengthening security controls for live streaming activities, such as identifying merchants using bots for repetitive, mechanical broadcasting, to promptly detect illegal broadcasting and ensure the quality of live streaming content.
[0003] Virtual human live streaming, also known as digital human live streaming, is increasingly widely used in online live streaming. Unlike robot-assisted broadcasting, virtual human live streaming uses text or voice to drive the generation of video streams for live streaming. This video stream contains image content corresponding to information spoken by a specific person. The information to be spoken can be pre-defined or determined in real time. Therefore, virtual human live streaming is a relatively real-time dynamic video generation technology. When the video stream corresponding to the virtual human is pushed to the live streaming room for playback, a virtual live streaming activity can be executed.
[0004] The video stream in virtual live streaming events is usually generated based on a certain length of source video. However, the number of image frames in the source video is always limited. When generating the video stream required for the live stream for a long time, due to the technological lag of the risk control system, the system may easily misidentify the video stream in the virtual live streaming event as robot-assisted broadcasting, and make incorrect interventions, such as stopping or banning the broadcast. This can cause e-commerce live streaming events to be unable to proceed normally, resulting in huge economic losses for e-commerce stores and platforms, damaging the user experience, and limiting the development of virtual human live streaming as a new technology. Therefore, it is necessary to improve the relevant technologies for virtual live streaming events by addressing the technical deficiencies of the risk control system. Summary of the Invention
[0005] The purpose of this application is to provide a method, device, equipment and product for preventing accidental touches during live streaming risk control.
[0006] According to one aspect of this application, a method for preventing accidental touches during live streaming risk control is provided, comprising:
[0007] In response to the virtual live streaming start command, launch the e-commerce live streaming room on the e-commerce platform used to execute virtual live streaming activities;
[0008] The script generation service obtains a script list from the script server. The script list contains script texts corresponding to different business stages in the same live streaming business process.
[0009] The video generation service calls the video server to generate a human-image spoken video corresponding to each of the spoken texts based on the template video. The template video reaches the preset regulation duration, and each of its image frames contains a facial image captured based on the same person.
[0010] Using a virtual camera-driven service, the corresponding human-image spoken video of each script text in the script list is pushed to the e-commerce live streaming room in accordance with the live streaming business process to carry out the virtual live streaming activity.
[0011] According to another aspect of this application, a live streaming risk control and accidental touch prevention device is provided, comprising:
[0012] The live streaming response module is configured to respond to virtual live streaming start commands and launch e-commerce live streaming rooms on the e-commerce platform used to execute virtual live streaming activities.
[0013] The script acquisition module is configured to obtain a script list from the script server through the script generation service. The script list contains script texts corresponding to different business stages in the same live broadcast business process.
[0014] The video acquisition module is configured to call the video server through the video generation service to generate a human image broadcast video corresponding to each of the spoken texts based on the template video. The template video reaches the preset regulation duration, and each of its image frames contains a facial image captured based on the same person.
[0015] The live streaming module is configured to push the human-image spoken video corresponding to each script text in the script list to the e-commerce live streaming room to implement the virtual live streaming activity, according to the live streaming business process, through a virtual camera-driven service.
[0016] According to another aspect of this application, a computer device is provided, including a central processing unit and a memory, wherein the central processing unit is used to invoke and run a computer program stored in the memory to perform the steps of the live streaming risk control and accidental touch prevention method described in this application.
[0017] According to another aspect of this application, a non-volatile readable storage medium is provided, which stores a computer program implemented according to the live streaming risk control and accidental touch prevention method in the form of computer-readable instructions. When the computer program is invoked by a computer, it executes the steps included in the method.
[0018] According to another aspect of this application, a computer program product is provided, including a computer program / instructions, which, when executed by a processor, perform the steps of the live streaming risk control and accidental touch prevention method.
[0019] Compared to existing technologies, this application, when a live streamer starts a live stream room to execute a virtual live stream activity, first calls the script list from the script server through the script generation service to obtain the script texts corresponding to each business link of the same live stream business process. Then, the video generation service drives the video server to instantly generate the human voice video corresponding to each script text based on the template video that has reached the preset regulation duration. Finally, the virtual camera driving service pushes these human voice videos to the e-commerce live stream room to implement the virtual live stream activity.
[0020] This demonstrates the advantages of this application in several aspects, including: in terms of dynamic recognition mechanism, template videos that comply with the prescribed duration can enrich and generalize the action features of facial images in the corresponding human voice broadcast videos of various script texts, reducing the frequency of the risk control system detecting repeated action features; in terms of static recognition mechanism, the virtual camera-driven service avoids the inherent defects of the risk control system that mechanically judges robot broadcasting behavior based on data sources.
[0021] It is evident that, through the combined efforts of these two aspects, this application effectively prevents virtual live-streaming activities from being mistakenly identified as robot broadcasting behavior by the e-commerce platform's risk control system, both in terms of dynamic and static identification mechanisms. This reduces the frequency of the risk control system's erroneous intervention in the virtual live-streaming activities of this e-commerce live-streaming room, improves the stability and security of virtual live-streaming activities, avoids unnecessary economic losses for the e-commerce stores to which the live-streaming room belongs, and also safeguards the application of virtual human live-streaming in the e-commerce field, removing obstacles to its application. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 The network architecture of the e-commerce platform is an example of that used in this application;
[0024] Figure 2 This application provides an example of a network architecture that supports a digital broadcast control system.
[0025] Figure 3This is a flowchart illustrating the live streaming risk control and accidental touch prevention method in the embodiments of this application;
[0026] Figure 4 This is a schematic diagram illustrating the process of inserting a human-voiced video in response to a dynamic insertion command in an embodiment of this application;
[0027] Figure 5 This is a flowchart illustrating the process of determining the inserted text based on user input information in an embodiment of this application.
[0028] Figure 6 This is a schematic diagram illustrating the process of creating a template video in an embodiment of this application;
[0029] Figure 7 This is a schematic diagram illustrating the process of creating template videos based on the interference rate of the risk control system in an embodiment of this application.
[0030] Figure 8 This is a schematic diagram illustrating the process by which the video server generates corresponding human-image spoken video based on the spoken text in an embodiment of this application.
[0031] Figure 9 This is a schematic diagram of the live streaming risk control and accidental touch prevention device in the embodiments of this application;
[0032] Figure 10 This is a schematic diagram of the structure of the computer device in the embodiments of this application. Detailed Implementation
[0033] like Figure 1 In the network architecture shown, e-commerce platform 82 is deployed on the Internet to provide corresponding services to its users. Similarly, the devices 80 of merchant users and 81 of consumer users on e-commerce platform 82 are also connected to the Internet to use the services provided by the e-commerce platform. For example, the e-commerce platform can configure a live streaming system to open live streaming services to merchant users of various online stores on the e-commerce platform. For the live streaming service, the merchant users are the broadcasters. Broadcasters can start the e-commerce live streaming room to carry out live streaming activities, including live streaming with real people or virtual live streaming activities based on virtual humans or digital humans. Consumer users can also act as viewers of the live streaming service, entering the e-commerce live streaming room to receive the video and information streams of the live streaming activities and interact bidirectionally with the broadcasters to participate in the live streaming activities.
[0034] An exemplary e-commerce platform 82 provides supply and demand matching of products and / or services to the general public through the Internet infrastructure. In e-commerce platform 82, products and / or services are provided as commodity information. For the sake of simplicity, the concepts of commodity and product are used in this application to refer to the products and / or services in e-commerce platform 82. Specifically, these may be physical products, digital products, tickets, service subscriptions, other offline services, etc.
[0035] In reality, various entities can access e-commerce platform 82 as users and utilize its online services to participate in the business activities facilitated by the platform. These entities can be natural persons, legal persons, or social organizations. Corresponding to the two types of entities in business activities—merchants and consumers—e-commerce platform 82 has two corresponding categories of users: merchant users and consumer users. Entities involved in the product distribution chain in business activities, including manufacturers, sellers, retailers, and logistics providers, can all use online services on e-commerce platform 82 as merchant users. Similarly, consumers in business activities, including actual or potential consumers, can use online services on e-commerce platform 82 as consumer users. In actual business activities, the same entity can operate as both a merchant user and a consumer user; this should be interpreted flexibly.
[0036] The infrastructure used to deploy the e-commerce platform 82 mainly includes the backend architecture and frontend devices. The backend architecture runs various online services through a service cluster, including middleware or frontend services for the platform, services for consumers, and services for merchants, to enrich and improve its service functions. The frontend devices mainly cover the terminal devices used by users as clients to access the e-commerce platform 82, including but not limited to various mobile terminals, personal computers, and point-of-sale devices. For example, merchant users can use their terminal device 80 to enter product information for their online stores or use the interfaces opened by the e-commerce platform to generate their product information; consumer users can use their terminal device 81 to access the webpage of the online store implemented by the e-commerce platform 82, trigger the shopping process by clicking the shopping button provided on the webpage, and call various online services provided by the e-commerce platform 82 during the shopping process to achieve the purpose of placing an order.
[0037] In some embodiments, the e-commerce platform 82 may be implemented via a processing facility including a processor and memory, which stores a set of instructions that, when executed, cause the e-commerce platform 82 to perform the e-commerce and support functions as described in this application. The processing facility may be part of a server, client, network infrastructure, mobile computing platform, cloud computing platform, fixed computing platform, or other computing platform, and may provide electronic components, merchant devices, payment gateways, application developers, marketing channels, transportation providers, customer devices, point-of-sale devices, etc., for the e-commerce platform 82.
[0038] E-commerce platform 82 can provide online services such as cloud computing services, Software as a Service (SaaS), Infrastructure as a Service (IaaS), Platform as a Service (PaaS), Desktop as a Service (DaaS), Hosted Software as a Service, Mobile Backend as a Service (MBaaS), and Information Technology Management as a Service (ITMaaS). In some embodiments, the various functional components of e-commerce platform 82 can be implemented to operate on various platforms and operating systems. For example, for an online store, its administrator user enjoys the same or similar functions regardless of whether it is on iOS, Android, HomonyOS, or a web page.
[0039] E-commerce platform 82 enables merchants to create their own independent websites to run their online stores. It provides merchants with corresponding business management engine instances, allowing them to establish, maintain, and operate one or more online stores across these independent websites. The business management engine instance can be used for content management, task automation, and data management for one or more online stores. It can be configured through interfaces or built-in components to support various specific business processes in the online store, supporting business activities. Independent websites are the infrastructure of e-commerce platform 82, which offers cross-border services. Merchants can maintain their online stores relatively independently and centrally based on these independent websites. Independent websites typically have dedicated domain names and storage space, and different independent websites are relatively independent. E-commerce platform 82 can provide standardized or customized technical support for a large number of independent websites, allowing merchants to customize a business management engine instance that suits their needs and use it to maintain one or more online stores.
[0040] Online stores can be configured and maintained in the backend by merchant users logging into their Business Management Engine instance as administrators. Supported by the various online services provided by the e-commerce platform 82's infrastructure, merchant users can configure various functions within their online stores and view various data as administrators. For example, merchant users can manage various aspects of their online stores, such as viewing recent online store activities, updating the online store's product catalog, managing orders, recent visit activity, and total order activity. Merchant users can also view more detailed information about their business and visitors to their online store by obtaining reports or metrics, such as displaying a sales summary of the merchant's overall business, specific sales and engagement data from promotional sales and marketing channels, etc.
[0041] E-commerce platforms 82 can provide communication facilities and associated merchant interfaces for electronic communication and marketing. For example, they can utilize electronic messaging aggregation facilities to collect and analyze communication interactions between merchants, consumers, merchant devices, customer devices, point-of-sale devices, etc., aggregating and analyzing communications to increase the potential for product sales. For instance, a consumer may have product-related questions, which could lead to a dialogue between the consumer and the merchant (or an automated processor-based agent representing the merchant), where the communication facilities handle the interaction and provide the merchant with analysis on how to increase the probability of a sale.
[0042] In some embodiments, applications suitable for installation on terminal devices can be provided to serve the access needs of different users, enabling various users to access the e-commerce platform 82 by running the application on their terminal devices. Examples include the merchant backend module of online stores within the e-commerce platform 82. During the process of conducting business activities through these functions, the e-commerce platform 82 can implement various functions related to business activities as middleware or online services and expose corresponding interfaces. Then, toolkits corresponding to the interface access functions are embedded into the application to achieve functional expansion and task completion. The business management engine can include a series of basic functions and expose these functions to online services and / or applications via APIs. Online services and applications use the corresponding functions by remotely calling the corresponding APIs.
[0043] With the support of various components of the Business Management Engine instance, the e-commerce platform 82 can provide online shopping functionality, enabling merchants to connect with customers in a flexible and transparent manner. Consumers can select items online, create orders, provide delivery addresses in the orders, and complete payment confirmation. Merchants can then review and fulfill or cancel orders. The review component included with the Business Management Engine instance ensures compliant use of business processes, guaranteeing that orders are suitable for fulfillment before actual execution. Orders may sometimes be fraudulent and require verification (e.g., ID checks). Payment methods that require merchants to wait to ensure receipt of funds can mitigate this risk, and so on. Order risks may arise from fraud detection tools submitted by third parties through order risk APIs, etc. Before fulfillment, merchants may need to obtain or wait to receive payment information to mark the order as paid before preparing to deliver the product. Such situations can all be subject to appropriate review. The review process can be implemented by the fulfillment component. Merchants can leverage fulfillment components to review and adjust operations, and trigger related fulfillment services. These include: manual fulfillment services, used when merchants select and pack products into boxes, purchase shipping labels and enter tracking numbers, or simply mark items as fulfilled; custom fulfillment services, which can define email notifications; API fulfillment services, which can trigger third-party applications to create fulfillment records; legacy fulfillment services, which can trigger custom API calls from the Commerce Management Engine to third parties; and gift card fulfillment services, which can generate and activate gift cards. Merchants can use an order printer application to print shipping documents. The fulfillment process can be executed once items are packed and ready for shipment, tracked, delivered, and verified by the consumer.
[0044] E-commerce platform 82 can also deploy a risk control system to perform security checks on the network access behavior of merchants and / or consumers during e-commerce or live streaming activities, promptly detect non-compliant operations, and implement corresponding technical intervention measures to ensure the healthy operation of the e-commerce platform. The computer devices of the merchants (i.e., the broadcasters) in e-commerce live streaming can run a computer program product implemented according to the live streaming risk control and accidental touch prevention method of this application, serving as a digital broadcast control system. This prevents the risk control system from mistakenly identifying the merchant's actions during the live streaming activity as non-compliant, reducing the frequency of the risk control system interfering with the live streaming activity.
[0045] To support merchant users, i.e., broadcaster users, in running the computer program product described in this application, the following can be provided: Figure 2The illustrated digital broadcast control system network architecture includes a script server, a video server, a live streaming server on the e-commerce platform 82 to support live streaming, and the terminal devices used by the broadcasters, i.e., live streaming devices. The script server provides relevant text, such as script text and insert text, to the broadcasters' devices. The video server helps the broadcasters generate corresponding voice-over videos. The e-commerce platform's live streaming server sends videos, such as voice-over videos, pushed by the broadcasters' devices to the e-commerce live streaming room, allowing viewers to receive and play the corresponding voice-over videos, enabling the broadcasters to achieve their live streaming goals. When broadcasting through their e-commerce live streaming room, broadcasters can conduct live events with real people or virtual live events based on virtual humans (digital humans). During virtual live events, they interact with the script server, video server, and live streaming server to push specified scripts and corresponding voice-over videos to the e-commerce live streaming room.
[0046] Please see Figure 3 In some embodiments, the live streaming risk control and accidental touch prevention method of this application can be implemented as a computer program product, running on the terminal device of the live streamer, to construct a digital broadcast control system, helping the live streamer to implement playback control of virtual live streaming activities in the e-commerce live streaming room of an e-commerce platform. The method includes:
[0047] Step S5100: Respond to the virtual live streaming start command and start the e-commerce live streaming room in the e-commerce platform used to execute virtual live streaming activities;
[0048] After running the computer program product of this application on their terminal device, the live streamer can enter the e-commerce live streaming room they have registered on the e-commerce platform through a preset control method to start a live stream. For example, the live streamer can click the live stream start button on the management page of their online store registered on the e-commerce platform, thereby triggering a virtual live stream start command. The background process of the computer program product of this application, after running, responds to the virtual live stream start command and interacts with the live streaming server of the e-commerce platform according to its default business logic to start the live streamer's e-commerce live streaming room, so as to execute virtual live streaming activities through the e-commerce live streaming room.
[0049] Step S5200: Obtain a script list from the script server through the script generation service. The script list contains script texts corresponding to different business links in the same live broadcast business process.
[0050] To achieve efficient operation, the computer program product of this application runs a script generation service, which communicates with a script server. Thus, before starting to push the virtual live broadcast event, the script generation service can first obtain the list of scripts corresponding to the implementation of the virtual live broadcast event from the script server.
[0051] The script server has the capability to generate text corresponding to the scripts needed for virtual live streaming activities for broadcast users. It can provide the script generation service with corresponding texts through various implementation methods, pre-storing these texts in a database. When the script generation service needs to access the scripts, it retrieves the matching text from the database and returns it to the service. For example:
[0052] In one embodiment, the script server can maintain a live streaming business process database. This database stores script texts corresponding to different business stages within various live streaming business processes. The broadcaster can retrieve a specific list of scripts corresponding to a particular live streaming business process from the script generation service. The script server simply needs to retrieve the script texts corresponding to each business stage from the database to construct the script list and push it to the script generation service. The live streaming business process consists of multiple business stages. Each stage's corresponding script outputs the expected information type for that stage, achieving phased information dissemination. The information disseminated at each stage constitutes the overall information disseminated throughout the live streaming business process, achieving the intended dissemination purpose. The dissemination role of each business stage in the live streaming business process can be flexibly customized. For example, typical live streaming scripts related to recommending products from online stores include: an opening introduction, a pain point explanation, a product selling point presentation, a purchase facilitation stage, and an attention-attracting stage. Each business stage can pursue its expected dissemination effect through corresponding scripts, and the scripts used for different business stages are generally different. Multiple script texts can be set for each business process. When constructing the script list for a live business process, the script server can randomly select one script text from the multiple script texts corresponding to each business process and add it to the script list.
[0053] In another embodiment, the script server can also maintain an advertising script database. After the broadcaster's terminal device submits product information to be promoted to the script server, the script server matches one or more corresponding promotional script texts from the advertising script database based on the product information to form a script list, which is then returned to the broadcaster. This advertising-related script list can be activated by the broadcaster inputting the product to be advertised. Once the broadcaster confirms the product to be advertised, they submit it to the script server either through a script generation service or directly. The script server retrieves the product information of the product to be advertised from the e-commerce platform's product database, including image information and / or text information, and then matches the corresponding script texts from the advertising script database based on the product information to form the script list.
[0054] Similarly, in another embodiment, the script server can also maintain a customer service Q&A database. When a viewer provides input text related to a question in the e-commerce live stream, the script server matches one or more corresponding response texts from the customer service Q&A database based on the input text, forming a script list, which is then returned to the broadcaster. This script list can also be obtained by the script generation service detecting the viewer's input text in the chat history of the e-commerce live stream and interacting with the script server, or by the live stream server interacting with the script server to determine the input text corresponding to the user's question, then obtaining it from the script server and pushing it to the broadcaster to drive the generation of the corresponding human-image voiceover video.
[0055] In another embodiment, for the various scripts required by the broadcaster, the script server can use a well-trained neural network model, such as various large language models, to generate a corresponding script list based on the basic material text and prompt text submitted by the broadcaster to the script server, and directly return it to the script server. This method can provide not only script lists required for the live streaming business process, but also script lists corresponding to advertisements and customer service Q&A. The script list can contain single script texts or multiple script texts, depending on the results generated by the neural network model and specific business needs.
[0056] In embodiments that flexibly adapt the above various embodiments, the script list obtained by the terminal device from the script server through the script generation service can not only contain a pre-defined script list corresponding to the same live broadcast business process, which only contains the script text corresponding to each business link in the live broadcast business process, but also embed script text that serves as an advertisement in the script list corresponding to the same live broadcast business process, and regard the latter as a business link in the live broadcast business process, that is, an advertising insertion link. In other words, the live broadcast business process includes an advertising insertion link, so the script list also contains the script text corresponding to the advertising business link. This can promote a more natural and smooth advertising insertion effect when generated based on the script list.
[0057] In a further enriched embodiment, the script server can respond to a script generation service request to obtain a script list. First, it determines the script list corresponding to the live-streaming business process specified in the request, as well as the product information of the product to be advertised. Then, it uses each script text in the script list as a reference text, calls a preset prompt text template, fits the product information into it to become the model prompt text, inputs the reference text and the model prompt text into the large language model, controls the large language model to associate the reference text to generate the script text for the product to be advertised, and determines the position of the script text in the script list of the live-streaming business process. Thus, the final obtained script list also includes the script text corresponding to the advertising insertion segment. Furthermore, with the help of the large language model's ability to analyze contextual semantics, this script text and the script text of each business segment in the original live-streaming business process can transition more smoothly in terms of natural semantics, presenting viewers with the feeling of advertising information being presented in natural conversation, playing a role similar to soft advertising, and further optimizing the user experience.
[0058] Step S5300: Through the video generation service, call the video server to generate the corresponding human image broadcast video based on the template video. The template video reaches the preset regulation duration, and each of its image frames contains a facial image captured based on the same person.
[0059] To achieve efficient operation, the computer program product of this application runs a video generation service, which communicates with a video server to drive the video server to generate corresponding human-image spoken video based on texts directly or indirectly submitted by the broadcaster user, such as script texts and insert texts. After obtaining each script text from the script list, the digital broadcast control system of the broadcaster user's terminal device submits each script text to the video server in the network architecture through the video generation service. The video server then calls a template video preset by the broadcaster user and generates a human-image spoken video corresponding to the script text based on the template video.
[0060] Specifically, a voice-action driven model can be deployed in the video server. For example, this model first converts the text received by the video server, such as scripted text or inserted text, into a speech sequence using an acoustic model to obtain the corresponding audio data and determine the corresponding audio duration. Then, based on the audio duration, it randomly selects a video image frame sequence corresponding to the duration from the template video preset by the broadcaster. Next, based on temporal alignment, it uses the speech sequence to correct the mouth shape of the face image in each image frame of the video image frame sequence, obtaining a corrected image frame sequence. Finally, based on temporal alignment, the audio data and the corrected image frame sequence are encoded into a human-image spoken video. After the video server generates the human-image spoken video corresponding to each text, it can push it to the broadcaster's terminal device. The terminal device can then download the corresponding human-image spoken video locally and establish an association mapping with the corresponding scripted text or inserted text for subsequent use.
[0061] Template videos can be pre-collected by broadcasters and uploaded to a video server for storage. In this application, a regulated duration sufficient to ensure the diversity of human actions in the template video is pre-set. Broadcasters collect source videos based on a reference duration determined by this regulated duration. After the broadcaster completes the collection of source videos, if the total duration of the source videos still does not reach the regulated duration, the terminal device or video server can extend the source videos to meet the regulated duration through data augmentation. Once the duration reaches the regulated duration, it can be used as a template video.
[0062] In one embodiment, the broadcaster user can preset multiple template videos. The people in different template videos can be the same staff member or different staff members, and the different template videos can have different styles, such as different spatial environments, different clothing styles, different desktop arrangements, etc. When the video server needs to generate a human-image broadcast video, it randomly selects any one of the template videos to capture a sequence of video image frames to generate the human-image broadcast video.
[0063] When livestreamers collect source videos, they do so by recording upper-body images of the same person in front of the camera. The person in front of the camera may or may not speak, but can have various expressions and movements to enrich the action features of the same source video for final transfer to the template video. Following this requirement, all image frames in the source video contain facial images of the person; therefore, the template video also contains corresponding facial images of the same person, which can serve as the basis for voice-driven lip-syncing.
[0064] Step S5400: Using the virtual camera driver service, push the human portrait video corresponding to each script text in the script list to the e-commerce live broadcast room according to the live broadcast business process to implement the virtual live broadcast activity.
[0065] To further reduce the probability of misjudgment in the risk control system, the computer program product corresponding to the digital broadcast control system of this application can automatically install a virtual camera driver service on the terminal device of the broadcaster. After obtaining the script list and the corresponding human image broadcast video of the script text, the virtual camera driver service is called to push the human image broadcast video corresponding to each script text in the script list to the e-commerce live broadcast room one by one, thereby implementing the virtual live broadcast activity. Since the scripts in the activity list are organized in an orderly manner according to the various business links of the same live streaming business process, the corresponding voice-over videos can be called and pushed to the live streaming server of the e-commerce platform according to the natural order of the scripts in the activity list. After receiving the video stream of the voice-over videos, the live streaming server pushes it to the terminal devices of each viewer in the e-commerce live streaming room, so that each viewer can watch the voice-over video and play the corresponding script audio content through the playback window of the corresponding page in the e-commerce live streaming room. Since the voice-over videos are driven by the corresponding scripts to correct lip movements, from the viewer's perspective, the voice and lip movements of the person in the voice-over videos are coordinated and correspond to each other, making it more natural.
[0066] It is not difficult to see that when a live user starts a live room to execute a virtual live event, this application first calls the script list from the script server through the script generation service to obtain the script texts corresponding to each business link of the same live business process. Then, the video generation service drives the video server to generate human voice videos corresponding to each script text in real time based on the template video that has reached the preset regulation duration. Finally, the virtual camera driver service pushes these human voice videos to the e-commerce live room to implement the virtual live event.
[0067] This demonstrates the advantages of this application in several aspects, including: in terms of dynamic recognition mechanism, template videos that comply with the prescribed duration can enrich and generalize the action features of facial images in the corresponding human voice broadcast videos of various script texts, reducing the frequency of the risk control system detecting repeated action features; in terms of static recognition mechanism, the virtual camera-driven service avoids the inherent defects of the risk control system that mechanically judges robot broadcasting behavior based on data sources.
[0068] It is evident that, through the combined efforts of these two aspects, this application effectively prevents virtual live-streaming activities from being mistakenly identified as robot broadcasting behavior by the e-commerce platform's risk control system, both in terms of dynamic and static identification mechanisms. This reduces the frequency of the risk control system's erroneous intervention in the virtual live-streaming activities of this e-commerce live-streaming room, improves the stability and security of virtual live-streaming activities, avoids unnecessary economic losses for the e-commerce stores to which the live-streaming room belongs, and also safeguards the application of virtual human live-streaming in the e-commerce field, removing obstacles to its application.
[0069] Based on any embodiment of the method in this application, please refer to Figure 4 The virtual live streaming activity is implemented by pushing the corresponding human-image spoken video to the e-commerce live streaming room according to the live streaming business process, including:
[0070] Step S5410: Respond to the dynamic insertion command triggered in the e-commerce live broadcast room and determine the corresponding insertion text;
[0071] At any point during the virtual live streaming event, a dynamic insertion command can be triggered in the e-commerce live streaming room. The digital broadcast control system responds to this dynamic insertion command by determining the corresponding insertion text. Of course, the insertion text is also a type of scripted text, which is used to generate a corresponding human-image spoken video for insertion. For example:
[0072] In one embodiment, the broadcaster can trigger dynamic insertion commands through the digital broadcast control system on their terminal device. For example, the digital broadcast control system can provide a function button; after the broadcaster touches the function button, they can input the insertion text and submit it to trigger the dynamic insertion command. Alternatively, the broadcaster can directly send a chat message to the e-commerce live broadcast room, which includes the input text to trigger the dynamic insertion command, thereby triggering and confirming the insertion text. These methods facilitate broadcasters inserting advertisements or other similar content during virtual live broadcasts.
[0073] In another embodiment, viewers can also trigger corresponding dynamic insertion commands by sending chat messages in the e-commerce live stream. For example, a viewer might send a chat message containing input text indicating the question, and the digital broadcast control system can match this input text to the corresponding reply text as the insertion text. This allows viewers to ask questions about products on sale in the e-commerce live stream, and the digital broadcast control system can then automatically respond, providing an automated customer service function.
[0074] Step S5420: Using the video generation service, call the video server to generate a video of the human image speaking based on the template video;
[0075] Similar to step S5300, after the inserted text is determined, the digital broadcast control system submits the inserted text to the video server through its video generation service. After receiving the inserted text, the video server converts the inserted text into a speech sequence using an acoustic model, determines the corresponding video image frame sequence from the selected template video, and then calls the speech action driving model to correct the mouth shape of the face image in each image frame of the video image frame sequence according to the speech sequence, thereby obtaining a corrected image frame sequence. Then, using the temporal alignment relationship, the speech sequence and the corrected image frame sequence are encoded into a human voice-over video and sent back to the video generation service.
[0076] Step S5430: Insert the video of the person speaking in the background corresponding to the inserted text after the video of the person speaking in the background that is being pushed to the e-commerce live broadcast room.
[0077] After obtaining the corresponding human-image spoken video for the inserted text through its video generation service, the digital broadcast control system can insert it into the ongoing virtual live broadcast activity in the e-commerce live broadcast room. This allows the human-image spoken video corresponding to the inserted text to be naturally embedded during the process of playing the human-image spoken videos corresponding to each text in the script list according to the live broadcast business process.
[0078] Since virtual live streaming events follow a pre-defined live streaming workflow—that is, the order of the scripts in the script list—when pushing corresponding voice-over videos to the live streaming server in real time, to maintain the smoothness of the virtual live streaming playback, the voice-over video currently being pushed to the e-commerce live streaming room (i.e., the live streaming server) can be determined first. The corresponding voice-over video for the insertion text can then be inserted after the currently pushed voice-over video before being pushed. Specifically, it can be inserted either immediately after the currently pushed voice-over video or one or two voice-over videos after it, depending on the response time of the dynamic insertion command. For example, in one embodiment, when the dynamic insertion command is used to answer questions asked by viewers via chat messages, it can be inserted immediately after the currently pushed voice-over video; in another embodiment, when the dynamic insertion command is used to insert advertisements, it can be inserted several voice-over videos later than the currently pushed voice-over video.
[0079] As can be seen from the above embodiments, during the virtual live streaming event, dynamic insertion commands can be triggered in various ways to determine the corresponding insertion text. Based on this, corresponding human-image spoken video is generated and inserted into a pre-determined script list for playback. This can be used for advertising integration and also for automatically providing viewers with automated Q&A services. The entire process is relatively smooth, making the human-to-human interaction in the virtual live streaming event more realistic and natural, significantly improving the user experience for viewers. Furthermore, because dynamic insertion commands can enrich the diversity of human-image spoken video in the script list, the risk of the virtual live streaming event being identified as mechanically repetitive robot broadcasting behavior can be further reduced, thus lowering the false trigger rate of the e-commerce platform's risk control system and ensuring the stable operation of the virtual live streaming event.
[0080] Based on any embodiment of the method in this application, please refer to Figure 5 In response to a dynamic insertion command triggered in the e-commerce live stream, the corresponding insertion text is determined, including:
[0081] Step S5411: Detect user input information in the e-commerce live streaming room, and perform intent recognition on the user input information to determine whether it carries the intent to insert content.
[0082] When it is necessary to identify dynamic insertion instructions through chat messages in an e-commerce live streaming room, the digital broadcast control system of this application can be responsible for detecting each chat message in the public screen message flow of the e-commerce live streaming room, obtaining the chat messages of each user, including the host user and the audience user, that is, the user input information submitted by the corresponding user. Then, a pre-trained intent recognition model is used to identify the intent of the user input information to determine whether the user input message carries an insertion intent. When it is determined that it carries an insertion intent, the corresponding dynamic insertion instruction can be triggered.
[0083] Step S5412: When carrying an insertion intention, trigger the dynamic insertion instruction corresponding to the insertion intention, and determine the insertion type and input text corresponding to the insertion intention based on the user input information;
[0084] Once the system recognizes that the user's input information contains an intention to insert content and triggers the corresponding dynamic insertion command, the digital broadcast control system can further determine the corresponding insertion type and input text based on the user's input information and its implied insertion intention.
[0085] For example, if a user submits the input information "Please help me generate an advertisement for model XX sports shoes", the intent recognition model can determine that there is an intent to insert an advertisement based on this user input information. After removing invalid characters and emojis from the user input information, the model can obtain the input text used by the user to express their meaning effectively.
[0086] For example, if a viewer submits the user input message "Host, what material are your sneakers made of?", the intent recognition model can determine that there is an intention to insert the message in response to the question based on this user input. Similarly, after removing invalid characters and emoticons from the user input message, the user's input text that can effectively express their meaning can be obtained.
[0087] Step S5413: Determine the inference text corresponding to the input text from the database corresponding to the insertion type as the insertion text.
[0088] Once the type of insertion is determined, the digital broadcast control system knows which interface to call. For example, for inserting advertisements, the interface provided by the advertising system can be called; for answering questions, the interface provided by the customer service system can be called. By calling the corresponding interface, the insertion text corresponding to the input text can be further obtained.
[0089] For example, in an embodiment corresponding to interstitial advertising, the digital broadcast control system extracts product keywords from the user's input text, such as "XX model sports shoes," and then submits it to the interface provided by the advertising system. The interface of the advertising system, based on the product keywords, matches the target product corresponding to the product keywords from the database of online stores belonging to the anchor user on the e-commerce platform, i.e., the product database, and obtains the product information of the target product, including but not limited to any one or more of the following: product images, product titles, product detail text, and product attribute data, depending on the preset business logic. Then, the product information is input into a preset advertising copy generation model to expand and generate plain text-based reasoning text, which is essentially generating the corresponding advertising copy. This advertising copy is then used as the interstitial text.
[0090] For example, in a corresponding implementation of a question-and-answer session, the digital broadcast control system can directly submit user-input text to the customer service system via a live call to its interface. This interface uses the semantic vector of the user-input text to match semantically similar basic questions from the customer service system's database (i.e., the question-and-answer database). It then retrieves a corresponding answer text from the question-and-answer database as the inference text for the input text, which the digital broadcast control system can then use as the inserted broadcast text.
[0091] The above embodiments demonstrate that the digital broadcast control system of this application possesses highly intelligent features. It can utilize user input information in e-commerce live streaming rooms for intent recognition, determine the corresponding insertion intent and its input text, and generate inferred text corresponding to the input text by calling the interface corresponding to the insertion intent. This inferred text is then used as the insertion text. The entire process is automated, eliminating the need for complex user operations. It automatically generates corresponding human-image voiceover videos based on user input information. These videos are played during virtual live streaming activities, responding to user input and creating the effect of a person in the live stream receiving a task from the host or responding to audience requests, significantly improving the user experience. Furthermore, since the insertion text is dynamically generated based on user input, the content of each generated insertion text is generally different, resulting in different human-image voiceover videos, further reducing the probability of accidentally triggering the e-commerce platform's risk control system.
[0092] Based on any embodiment of the method in this application, inserting the human-image voiceover video corresponding to the inserted text after the human-image voiceover video being pushed to the e-commerce live stream includes:
[0093] Step S5431: According to the order of the business links in the live broadcast business process, the human image broadcast video corresponding to each script text in the script list is ordered into the cache queue, and the human image broadcast video ordered out of the cache queue is pushed to the live broadcast server.
[0094] To improve the memory efficiency of terminal devices and ensure the smoothness of virtual live streaming activities at the data level, in this embodiment, the digital broadcast control system of this application can be further optimized by incorporating caching technology. Accordingly, the digital broadcast control system can load the voice-over videos corresponding to the script texts of each business segment in the script list of the live streaming business process into a cache queue, according to the order of each business segment in the preset live streaming business process. Then, in conjunction with the queue scheduling principle, the system controls each voice-over video to be dequeued from the cache queue in the order of its corresponding business segment. When a voice-over video is dequeued from the cache queue for consumption, the consumption thread is responsible for encoding and pushing the voice-over video to the live streaming server. After receiving the corresponding video stream, the live streaming server decodes and re-encodes it and pushes it to the terminal devices of each viewer. Each viewer's terminal device decodes and plays the corresponding video stream, thus viewing the corresponding voice-over video.
[0095] Step S5432: Detect and determine the head position of the human image voiceover video being pushed out of the cache queue, and insert the human image voiceover video corresponding to the inserted text into the next position after the head position.
[0096] To integrate the accompanying video with the inserted text into the existing live streaming workflow, caching technology can be used. First, identify the video with the inserted text currently being consumed and pushed out of the cache queue. Determine the position of this video as the head of the cache queue. Then, insert the video with the inserted text immediately after the head of the queue, ensuring that it follows the video currently being consumed. Once the video being consumed finishes its push, the video with the inserted text can then be consumed immediately, thus ensuring the real-time playback of the video with the inserted text.
[0097] As can be understood from the above embodiments, by combining caching technology to process the playback order relationship between the video of the inserted text and the video of the inserted text in the existing live broadcast process, the playback of the video of the inserted text becomes a business link in the live broadcast process, playing naturally within the live broadcast process. Due to the use of caching technology, this process transitions naturally, the transmission is stable and smooth, the robustness of the digital broadcast control system is better, and placing the video of the inserted text after the first position in the queue can ensure its immediate playback. When responding to user input information, its efficiency advantage is obvious.
[0098] Based on any embodiment of the method in this application, please refer to Figure 6 Before obtaining the script list from the script server through the script generation service, the following steps are included:
[0099] Step S4100: Obtain a video clip with a duration of at least half of the regulated duration, wherein the video clip contains facial images captured based on the same person;
[0100] As disclosed above, the regulated duration in this application is preset to ensure sufficient diversity of human facial movements in the template video to effectively reduce the false trigger rate of the e-commerce platform's risk control system. In theory, when generating a template video, it is only necessary to record a template video of a specific person according to this regulated duration; however, this method incurs high time costs. Therefore, when creating a template video, a recording duration can be set. This duration can be any value greater than or equal to half the regulated duration, prioritizing time savings. Under this condition, the digital broadcast control system can open a recording program to record source videos of the person. During recording, the camera can be aimed at the upper body of the person to ensure that the person's face is visible in every frame of the source video. It should be noted that audio data recording is not required for the source video.
[0101] Step S4200: Arrange the image frames in the source video in reverse order to form the corresponding reverse video;
[0102] After acquiring the source video, the digital broadcast control system reverses the order of the video frames, transforming it into a reversed version of the source video. When played back in this way, the reversed video visually enriches the variety of human movements, especially facial expressions. Therefore, reversed video is essentially a superior method of data enhancement of the source video, thus eliminating the need for users to record long videos.
[0103] Step S4300: Combine the source video with its corresponding reverse video to form the template video used to create the portrait voiceover video.
[0104] Once the source video and the reverse video are determined, the reverse video can be stitched after the source video, creating a combined video that can be used directly as a template video. For easier future use, the template video can be associated with the corresponding identifier of its live streaming workflow and stored. Later, based on the identifier of the live streaming workflow, the corresponding template video can be called to create the persona-driven audio video. It's easy to understand that the total duration of the resulting template video will inevitably reach the preset duration. Since the first frame of the reverse video is also the last frame of the source video, and the reverse video is stitched after the source video, the movements of the person in the entire template video are relatively smooth. Because the movements of the person in the video frame sequence extracted from the middle of the template video will also be smooth. Similarly, since the last frame of the reverse video is also the first frame of the source video, when it's necessary to extract a video frame sequence at the end of the template video that exceeds the total duration, it can loop back to the first frame of the template video to continue extracting, ensuring that the movements of the person in the resulting video frame sequence are also smooth.
[0105] In some embodiments, if the total duration of the template video is much greater than the preset regulation duration, the template video can be cropped to control storage space. In this case, the duration to be cropped can be divided by two to obtain the duration to be deleted. Image frames corresponding to the duration to be deleted are deleted from both the beginning and end of the template video. This ensures that the beginning and end image frames of the template video respond to each other, and that the human movements are still smooth when the video image frame sequence is cyclically extracted based on the template video.
[0106] As can be seen from the above embodiments, although the digital broadcast control system sets a regulated duration, the data augmentation method adopted in this embodiment can effectively reduce the duration of the recorded source video. By using the method of generating a reverse video based on the source video to expand the source video to obtain the template video, the time cost of recording source video is reduced. At the same time, it can also ensure that when the video image frame sequence corresponding to the script text / interlude text is obtained based on the template video, the characters' movements in the sequence are still relatively smooth and natural. This will not make the audience feel abrupt, nor will it easily trigger the alarm of the e-commerce platform's risk control system, thus improving the robustness of the virtual live broadcast event.
[0107] Based on any embodiment of the method in this application, please refer to Figure 7 Before acquiring source video footage with a duration at least half of the regulated duration, the following procedures are included:
[0108] Step S3100: Calculate the interference rate of the e-commerce platform's risk control system for each virtual live streaming activity executed in response to the virtual live streaming start command in the e-commerce live streaming room.
[0109] During the process of live streamers conducting virtual live streaming activities through the digital broadcast control system of this application, the digital broadcast control system can be responsible for recording the situations in which each virtual live streaming activity triggers the intervention of the e-commerce platform's risk control system. Based on this, the total number of virtual live streaming activities and the number of interventions corresponding to the intervention events that trigger the risk control system can be obtained. Dividing the number of interventions by the total number of activities will yield the intervention rate corresponding to the intervention behavior of the risk control system.
[0110] Interference actions implemented by the risk control system include, but are not limited to, any punitive or warning actions such as time-limited suspension of broadcasting, banning broadcasting for the current session, deducting points from the broadcaster user, and sending alarm notifications to the broadcaster user. Each such action is considered an interference event, and the digital broadcast control system will record it accordingly for subsequent statistical purposes.
[0111] The timing for determining the interference rate in a digital broadcast control system can be implemented each time a broadcaster initiates a virtual live broadcast activity and triggers the corresponding virtual live broadcast start command, so as to ensure the smooth progress of the current virtual live broadcast activity in a timely manner.
[0112] Step S3200: Determine whether the interference rate has reached a preset threshold. When the preset threshold is reached, trigger the material video expansion instruction to update the regulation duration by adding a fixed duration. Use the updated regulation duration as the recording duration.
[0113] The digital broadcast control system also has a preset threshold, which can be an empirical or experimental threshold. This threshold is used to determine the interference rate. When the interference rate is determined to reach the threshold, a video material expansion instruction is triggered to guide the broadcaster to re-record a longer video material to generate a new template video. When the threshold is not reached, no further processing is required. The recommended threshold can be any value between 20% and 30%. That is, if relying on historical template videos for virtual live streaming results in 20% to 30% of the activities being interfered with by the risk control system, the user can be guided to recreate the template video.
[0114] The increased interference rate of the risk control system may be due to two reasons. Firstly, the diversity of the characters' movements in the template video may not be sufficient to circumvent the dynamic detection mechanism's condition parameters. Secondly, the risk control system may have increased the requirements for these condition parameters. In either case, extending the total duration of the template video can ensure the diversity of the characters' movements. Accordingly, while triggering the material video expansion command, this embodiment also updates the previously used regulatory duration. Specifically, a preset fixed duration can be used, and the original regulatory duration is added to this fixed duration to form the new regulatory duration. The material video is then recorded using the new regulatory duration. This fixed duration can be a preset value, such as 5 minutes, 10 minutes, etc.
[0115] Step S3300: In response to the material video expansion instruction, start the video recording program to record material video using the updated regulation duration for use in creating the template video.
[0116] After the instruction to expand the material video is triggered, the digital broadcast control system starts the video recording program in response to the instruction. It uses the updated regulation duration to create a new template video. Specifically, the new template video can be created according to the process of steps S3100 to S3300 in the previous embodiment.
[0117] In this embodiment, the digital broadcast control system uses the statistically obtained interference rate corresponding to the intervention of the risk control system in the virtual live broadcast activity to intelligently decide whether to remake the template video. When the interference rate is found to be high, the system will appropriately extend the regulation duration by superimposing a fixed duration and guide the user to remake the template video with the duration of the regulation duration throughout the process. This ensures that the diversity of the character action features in the template video can reduce the interference rate of the risk control system on the virtual live broadcast activity, effectively reduce the false trigger rate of the risk control system, and improve the robustness of the virtual live broadcast activity.
[0118] Based on any embodiment of the method in this application, please refer to Figure 8 The process includes: 1) Calling the video server to generate a video with the corresponding spoken text based on the template video; 2) Calling the video server to generate a video with the corresponding spoken text based on the template video; 3) Calling the video server to generate a video with the corresponding spoken text based on the template video; 4) Calling the video server to generate a video with the corresponding spoken text based on the template video.
[0119] Step S6100: The video server obtains the script text / interlude text of the voice-over video to be generated, calls the acoustic model to generate the audio data of the script text / interlude text, and determines the corresponding script duration of the audio data.
[0120] Since inserted text is also a type of scripted text, and both are analogous, this embodiment mainly uses scripted text to describe the process of the video server generating a human-image spoken video. Those skilled in the art should know that it is equally applicable to inserted text.
[0121] Once the video server receives the scripted text or insert text submitted by the digital broadcast control system for generating voice-over video, it can call a pre-defined acoustic model to convert the scripted text or insert text into audio data. This audio data can be represented as a speech sequence for easy intermediate retrieval. The acoustic model can be any mature and known model, which can be directly implemented by those skilled in the art from existing technologies.
[0122] Once the acoustic model generates the corresponding audio data through its text-to-speech reasoning ability, it determines the duration of the audio data, which can be used as the duration of the speech.
[0123] Step S6200: The video server extracts a sequence of video image frames corresponding to the duration of the speech from the template video;
[0124] The video server then extracts an image frame corresponding to the duration of the speech from a template video with a duration that meets the regulated duration, based on the speech duration. This image frame is used as a video image frame sequence corresponding to the speech duration, ensuring that the video image frame sequence maintains temporal alignment with the speech sequence of the audio data.
[0125] When a video server extracts a sequence of video image frames from a template video, it can either extract image frames from different positions in the same template video sequentially to form the corresponding video image frame sequence when processing different speech texts multiple times, or it can randomly locate and obtain the corresponding video image frame sequence from the template video for each speech text processed.
[0126] Step S6300: The video server calls the voice action driving model and, based on the audio data generated according to the corresponding speech text / interlude text, corrects the mouth movements of the face images in the video image frame sequence to obtain the corrected image frame sequence.
[0127] After the video server obtains the two data streams corresponding to the speech duration (i.e., audio data and video image frame sequence), it can call its preset speech action driving model. This model can be any mature model, which can be flexibly selected by those skilled in the art. Using this model, based on the speech data in the speech sequence, the mouth shape of the face image in each image frame in the video image frame sequence is corrected, thereby realizing the correction of the mouth shape action of the face image in the entire video image frame sequence, thus obtaining the corrected image frame sequence.
[0128] Step S6400: After the video server performs time-series alignment of the audio data and the corrected image frame sequence, it generates the human-image spoken video corresponding to the spoken text / interlude text.
[0129] After obtaining the corrected image frame sequence, the video server encodes the corrected image frame sequence with the speech sequence representing the audio data according to the temporal alignment relationship, thereby generating a human-image-based spoken video corresponding to the spoken text or inserted text. The video server then pushes this human-image-based spoken video to the digital broadcast control system in the broadcaster's terminal device, which can then be used to implement virtual live broadcast activities.
[0130] In this embodiment, the video server is responsible for centrally processing the generation of human-image-based spoken videos corresponding to various text-based scripts. Its business logic is centralized and reusable, resulting in high computational efficiency. Because the video server can utilize acoustic models and speech-to-speech correction models to extract material from template videos to generate human-image-based spoken videos, the generated videos are of high quality. When these videos are used in virtual live-streaming events, they help reduce the false trigger rate of e-commerce platform risk control systems by leveraging the diverse characteristics of human actions within the videos.
[0131] Please see Figure 9 Another embodiment of this application provides a live streaming risk control and accidental touch prevention device, which includes a live streaming response module 5100, a script acquisition module 5200, a video acquisition module 5300, and a live streaming push module 5400. The live streaming response module 5100 is configured to respond to a virtual live streaming start command and start an e-commerce live streaming room on an e-commerce platform for executing virtual live streaming activities. The script acquisition module 5200 is configured to obtain a script list from a script server through a script generation service. The script list contains script texts corresponding to different business stages in the same live streaming business process. The video acquisition module 5300 is configured to call a video server through a video generation service to generate human-image spoken videos corresponding to each script text based on a template video. The template video reaches a preset regulated duration, and each image frame contains a facial image captured based on the same person. The live streaming push module 5400 is configured to push the human-image spoken videos corresponding to each script text in the script list to the e-commerce live streaming room to implement the virtual live streaming activity according to the live streaming business process through a virtual camera driver service.
[0132] Based on any embodiment of the device in this application, the live streaming module 5400 includes: an insertion response module, configured to respond to a dynamic insertion command triggered in the e-commerce live streaming room and determine the corresponding insertion text; an insertion generation module, configured to call a video server through a video generation service to generate a human voice-over video corresponding to the insertion text based on the template video; and a video insertion module, configured to insert the human voice-over video corresponding to the insertion text after the human voice-over video being pushed to the e-commerce live streaming room.
[0133] Based on any embodiment of the device in this application, the insertion response module includes: an input detection module, configured to detect user input information in the e-commerce live streaming room, perform intent recognition on the user input information to determine whether it carries an insertion intent; an information extraction module, configured to trigger a dynamic insertion instruction corresponding to the insertion intent when it carries an insertion intent, and determine the insertion type and input text corresponding to the insertion intent based on the user input information; and a text determination module, configured to determine the inferred text corresponding to the input text from the database corresponding to the insertion type as the insertion text.
[0134] Based on any embodiment of the device in this application, the video insertion module includes: a cache scheduling module, configured to sequentially load the human-image spoken videos corresponding to each script text in the script list into a cache queue according to the order of business links in the live broadcast business process, and push the human-image spoken videos sequentially dequeued from the cache queue to the live broadcast server; and a positioning insertion module, configured to detect and determine the head position of the human-image spoken video being dequeued and pushed in the cache queue, and insert the human-image spoken video corresponding to the insertion text into the position after the head position.
[0135] Based on any embodiment of the device in this application, prior to the operation of the script acquisition module 5200, the live broadcast risk control and accidental touch prevention device of this application includes: a material acquisition module, configured to acquire material videos with a duration of at least half of the regulated duration, the material videos including facial images captured based on the same person; a reverse playback expansion module, configured to arrange the image frames in the material videos in reverse order to form their corresponding reverse playback videos; and a template creation module, configured to splice the material videos and their corresponding reverse playback videos to form the template video used to create the portrait voiceover video.
[0136] Based on any embodiment of the device in this application, prior to the operation of the material acquisition module, the live streaming risk control and accidental touch prevention device of this application includes: an interference statistics module, configured to count the interference rate corresponding to the interference of the risk control system of the e-commerce platform when a virtual live streaming activity is executed in response to the virtual live streaming start command in the e-commerce live streaming room; a duration update module, configured to determine whether the interference rate reaches a preset threshold, and when the preset threshold is reached, trigger a material video expansion command to update the regulated duration by adding a fixed duration, and use the updated regulated duration as the recording duration; and a video re-recording module, configured to respond to the material video expansion command, start a video recording program, and record material video using the updated regulated duration for use in creating the template video.
[0137] Based on any embodiment of the device in this application, the video acquisition module 5300 / the insertion generation module includes: a generation preparation module, configured to have a video server acquire the script text / insertion text of the human portrait audio video to be generated, call an acoustic model to generate audio data of the script text / insertion text, and determine the corresponding script duration of the audio data; a template extraction module, configured to have a video server extract a video image frame sequence corresponding to the script duration from the template video; a lip-shape correction module, configured to have a video server call a voice action driving model to correct the lip-shape action of the face image in the video image frame sequence according to the audio data generated corresponding to the script text / insertion text, and obtain a corrected image frame sequence; and a video generation module, configured to have the video server perform time-series alignment of the audio data and the corrected image frame sequence, and generate the human portrait audio video corresponding to the script text / insertion text.
[0138] Based on any embodiment of this application, please refer to Figure 10 Another embodiment of this application also provides a computer device, such as... Figure 10 The diagram shows the internal structure of a computer device. The computer device includes a processor, a computer-readable storage medium, a memory, and a network interface connected via a system bus. The computer-readable storage medium stores an operating system, a database, and a computer program encapsulated with computer-readable instructions. The database may store a sequence of control information. When the computer-readable instructions are executed by the processor, the processor can implement a live-streaming risk control method for preventing accidental touches. The processor of the computer device provides computing and control capabilities to support the operation of the entire computer device. The memory of the computer device may store computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor can execute the live-streaming risk control method for preventing accidental touches of this application. The network interface of the computer device is used for communication with a terminal. Those skilled in the art will understand that… Figure 10The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0139] In this embodiment, the processor is used to execute... Figure 9 The system contains the specific functions of each module and its sub-modules. The memory stores the program code and various data required to execute these modules or sub-modules. The network interface is used for data transmission between the user terminal and the server. In this embodiment, the memory stores the program code and data required to execute all modules / sub-modules in the live streaming risk control and accidental touch prevention device of this application. The server can call the server's program code and data to execute the functions of all sub-modules.
[0140] This application also provides a storage medium storing computer-readable instructions, which, when executed by one or more processors, cause the one or more processors to perform the steps of the live streaming risk control and accidental touch prevention method described in any embodiment of this application.
[0141] This application also provides a computer program product, including a computer program / instructions, which, when executed by one or more processors, implement the steps of the live streaming risk control and accidental touch prevention method described in any embodiment of this application.
[0142] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. This computer program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. The aforementioned storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0143] The above description is only a partial embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
[0144] In summary, this application enables virtual live streaming activities to effectively avoid being misidentified as robot broadcasting by e-commerce platform risk control systems, both dynamically and statically. This reduces the frequency of accidental intervention by risk control systems targeting virtual live streaming activities in the e-commerce live streaming room, improves the stability and security of virtual live streaming activities, avoids unnecessary economic losses for the e-commerce stores to which the live streaming room belongs, and also safeguards the application of virtual human live streaming in the e-commerce field, removing obstacles to its application.
Claims
1. A method for preventing accidental touches during live streaming risk control, characterized in that, include: In response to the virtual live streaming start command, launch the e-commerce live streaming room on the e-commerce platform used to execute virtual live streaming activities; The statistics show the interference rate among e-commerce live streaming rooms where virtual live streaming activities executed in response to virtual live streaming start commands are interfered with by the e-commerce platform's risk control system. Determine whether the interference rate reaches a preset threshold. When the preset threshold is reached, trigger the material video expansion instruction to update the regulation duration by adding a fixed duration. Use the updated regulation duration as the recording duration. In response to the aforementioned material video expansion command, the video recording program is started to record the material video using the updated regulation duration; Acquire video footage with a duration at least half of the regulated duration, the video footage containing facial images captured based on the same person; The image frames in the source video are arranged in reverse order to form the corresponding reversed video; The source video and its corresponding reversed video are spliced together to form a template video for creating a human voiceover video; The script generation service obtains a script list from the script server. The script list contains script texts corresponding to different business stages in the same live streaming business process. The video generation service calls the video server to generate a human-image spoken video corresponding to each script text based on the template video. The template video reaches the preset regulation duration, and each of its image frames contains a facial image captured based on the same person. Using a virtual camera-driven service, the corresponding human-image spoken video of each script text in the script list is pushed to the e-commerce live streaming room in accordance with the live streaming business process to carry out the virtual live streaming activity.
2. The live streaming risk control and accidental touch prevention method according to claim 1, characterized in that, According to the live streaming business process, the corresponding human-image spoken video for each script text in the script list is pushed to the e-commerce live streaming room to implement the virtual live streaming activity, including: In response to the dynamic insertion command triggered in the e-commerce live broadcast room, determine the corresponding insertion text; The video generation service calls the video server to generate a video of the human image speaking based on the template video; The video of the person speaking in the inserted text is inserted after the video of the person speaking in the inserted text that is being pushed to the e-commerce live broadcast room.
3. The live streaming risk control and accidental touch prevention method according to claim 2, characterized in that, In response to a dynamic insertion command triggered in the e-commerce live stream, determine the corresponding insertion text, including: Detect user input information in the e-commerce live streaming room, perform intent recognition on the user input information to determine whether it carries the intent to insert content; When a user expresses an intention to insert content, a dynamic insertion instruction corresponding to that intention is triggered. The insertion type and input text corresponding to the intention are determined based on the user input information. The inference text corresponding to the input text is determined from the database corresponding to the insertion type as the insertion text.
4. The live streaming risk control and accidental touch prevention method according to claim 2, characterized in that, Inserting the video of the person speaking in the background corresponding to the inserted text after the video of the person speaking in the background being pushed to the e-commerce live stream includes: According to the order of the business links in the live streaming business process, the human voice videos corresponding to each script text in the script list are sequentially loaded into the cache queue, and the human voice videos that are sequentially dequeued from the cache queue are pushed to the live streaming server. The system detects and determines the head position of the human-image voiceover video being pushed out of the cache queue, and inserts the human-image voiceover video corresponding to the inserted text into the position after the head position.
5. The live streaming risk control and accidental touch prevention method according to any one of claims 1 to 4, characterized in that, The video server is invoked to generate corresponding spoken video clips for each script text based on the template video, including: The video server obtains the script text of the human voice video to be generated, calls the acoustic model to generate the audio data of the script text, and determines the corresponding script duration of the audio data. The video server extracts a sequence of video image frames corresponding to the duration of the speech from the template video; The video server calls the voice action driven model and, based on the audio data generated according to the corresponding speech text, corrects the mouth movements of the facial images in the video image frame sequence to obtain the corrected image frame sequence; After the video server performs time-series alignment of the audio data and the corrected image frame sequence, it generates a video of the human voice corresponding to the spoken text.
6. The live streaming risk control and accidental touch prevention method according to any one of claims 2 to 4, characterized in that, The process of calling the video server to generate a video with the human image and voiceover corresponding to the inserted text based on the template video includes: The video server obtains the text to be inserted into the video of the human voice to be generated, calls the acoustic model to generate the audio data of the inserted text, and determines the corresponding speech duration of the audio data. The video server extracts a sequence of video image frames corresponding to the duration of the speech from the template video; The video server calls the voice action driven model and, based on the audio data generated corresponding to the inserted text, corrects the mouth movements of the face images in the video image frame sequence to obtain the corrected image frame sequence; After the video server performs time-series alignment of the audio data and the corrected image frame sequence, it generates a video of the human voice corresponding to the inserted text.
7. A live streaming risk control and accidental touch prevention device, characterized in that, include: The live streaming response module is configured to respond to virtual live streaming start commands and launch e-commerce live streaming rooms on the e-commerce platform used to execute virtual live streaming activities. The interference statistics module is set to count the interference rate of the e-commerce platform's risk control system when a virtual live streaming activity is executed in response to a virtual live streaming start command in an e-commerce live streaming room. The duration update module is configured to determine whether the interference rate reaches a preset threshold. When the preset threshold is reached, a material video expansion instruction is triggered to update the regulated duration by adding a fixed duration. The updated regulated duration is then used as the recording duration. The video re-recording module is configured to respond to the material video expansion command and start the video recording program to record the material video using the updated regulation duration; The material acquisition module is configured to acquire video materials with a duration of at least half of the regulated duration, wherein the video materials contain facial images captured based on the same person; The reverse playback expansion module is configured to arrange the image frames in the source video in reverse order to form the corresponding reverse playback video. The template creation module is configured to stitch the source video with its corresponding reverse video to create a template video for creating a human voiceover video. The script acquisition module is configured to obtain a script list from the script server through the script generation service. The script list contains script texts corresponding to different business links in the same live broadcast business process. The video acquisition module is configured to call the video server through the video generation service to generate a human image broadcast video corresponding to each script text based on the template video. The template video reaches the preset regulation duration, and each of its image frames contains a facial image captured based on the same person. The live streaming module is configured to push the human-image spoken video corresponding to each text in the script list to the e-commerce live streaming room to implement the virtual live streaming activity, according to the live streaming business process, through a virtual camera-driven service.
8. The live streaming risk control and accidental touch prevention device according to claim 7, characterized in that, The live streaming module includes: The insertion response module is configured to respond to dynamic insertion commands triggered in the e-commerce live streaming room and determine the corresponding insertion text; The insertion generation module is configured to use a video generation service to call a video server to generate a human-voiced video corresponding to the insertion text based on the template video. The video insertion module is configured to insert the human-image voiceover video corresponding to the inserted text after the human-image voiceover video being pushed to the e-commerce live streaming room.
9. The live streaming risk control and accidental touch prevention device according to claim 8, characterized in that, The insertion response module includes: The input detection module is configured to detect user input information in the e-commerce live streaming room, perform intent recognition on the user input information, and determine whether it carries the intent to insert content. The information extraction module is configured to trigger a dynamic insertion instruction corresponding to the insertion intention when it carries an insertion intention, and determine the insertion type and input text corresponding to the insertion intention based on the user input information; The text determination module is configured to determine the inference text corresponding to the input text from the database corresponding to the insertion type as the insertion text.
10. A computer device comprising a central processing unit and a memory, characterized in that, The central processing unit is used to invoke and run a computer program stored in the memory to perform the steps of the method as described in any one of claims 1 to 6.
11. A computer program product, characterized in that, Includes a computer program / instruction, which, when executed by a processor, performs the steps of the method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Virtual character video generation method and device, computer equipment and storage medium
CN114998489A
Multilingual lip language data generation method and system based on virtual 2d digital human
CN116934930A