Automated live streaming method

By combining pre-recorded videos generated at headquarters with artificial intelligence, the problem of inconsistent live streaming quality and lack of interactivity in recorded broadcasts for large enterprise branches has been solved. This has enabled low-cost, personalized, and highly interactive automated live streaming, enhancing brand image and audience experience.

CN120812332BActive Publication Date: 2026-02-27BEIJING SKYFLY INTERACTIVE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511299468.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-12
Publication Date
2026-02-27
Estimated Expiration
2045-09-12

AI Technical Summary

Technical Problem

Large corporate branches conducting their own live streams result in inconsistent live stream quality and difficulty in unifying brand image. Furthermore, existing recording and broadcasting technologies lack interactivity and realism, and are costly.

Method used

An automated interactive live streaming method based on pre-recorded video is adopted. The original materials are generated through the headquarters server and combined with artificial intelligence real-time interactive technology to realize the generation and streaming of personalized video and audio. Using technologies such as green screen keying, face replacement, and audio voice changing, personalized video frame sequences and audio streams are generated, and audience questions are captured and answers are generated in real time during the live stream.

Benefits of technology

It enables low-cost, large-scale live streaming with real-time interactivity, ensuring brand image consistency, reducing training and management costs, and improving live streaming quality and audience experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120812332B_ABST
    Figure CN120812332B_ABST
Patent Text Reader

Abstract

The application belongs to the live broadcast technical field, and particularly relates to an automatic live broadcast method. Specifically, a headquarters server records original audio and video materials in a green screen environment at one time; after a foreground anchor object is extracted from the green screen, personalized video frame sequences and personalized audio streams of "thousands of stores and thousands of faces" are generated in batches in combination with local scenes, staff facial features and voiceprint information collected by each branch store; the above content is encapsulated into a task package according to a preset live broadcast schedule and is differentially issued to each branch store unmanned live broadcast mobile phone. After starting the broadcast, the mobile phone end pushes a stream in real time and answers questions of audiences in real time by using a natural language processing model during the live broadcast process, so that an interactive experience that is not different from that of a real person is realized through voice synthesis and frame-level lip correction. The application significantly reduces labor cost, guarantees brand content unification, improves audience participation and conversion rate, and is suitable for large-scale, multi-store and all-weather automatic live broadcast operation.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of live broadcast, and particularly relates to an automatic live broadcast method. BACKGROUND

[0002] At present, live broadcast marketing has become an important means for enterprises to promote products. For large enterprises with many branch agencies, if personnel of the branch agencies live broadcast by themselves, the differences in personal abilities such as eloquence and image of the personnel will lead to uneven live broadcast quality, which not only affects the promotion effect, but also may damage the unified brand image. In addition, the human and time costs required for training and managing a large number of anchors are also very high. SUMMARY

[0003] The application aims to provide an automatic interactive live broadcast method and system based on pre-recorded videos, and aims to solve the technical problems of high cost and uneven quality of real person live broadcast, lack of interactivity and authenticity of traditional recording and broadcasting, and insufficient authenticity of pure virtual person live broadcast in the prior art. The application combines pre-recorded videos with artificial intelligence real-time interaction technology, realizes automatic live broadcast with real-time interaction ability at low cost and large scale on the premise of ensuring high-quality live broadcast content and unified brand image.

[0004] To achieve the above purpose, the application adopts the following technical solutions:

[0005] To achieve the above purpose, the application provides an automatic interactive live broadcast method based on pre-recorded videos, which comprises the following steps:

[0006] The headquarters server generates original video materials and original audio materials, and the original video materials are recorded by an anchor in a green screen environment;

[0007] The original video materials are subjected to green screen matting processing to extract a foreground anchor object and generate transparent channel information;

[0008] Local real scene backgrounds of a plurality of branch stores and face feature information and voiceprint information of corresponding store clerks are collected;

[0009] For each branch store, the foreground anchor object is subjected to face replacement and voice change based on the corresponding face feature information and voiceprint information, to obtain a personalized video frame sequence and a personalized audio stream;

[0010] According to a preset live broadcast schedule, the personalized video frame sequence and the personalized audio stream are packaged as a task package and delivered to an unmanned live broadcast automatic live broadcast mobile phone located at the branch store;

[0011] When the unmanned live broadcast automatic live broadcast mobile phone reaches the starting time, the personalized video frame sequence and the personalized audio stream are pushed in real time to perform unattended live broadcast;

[0012] During the live broadcast, the audience question text is captured in real time, and a trained natural language processing model is called to generate a reply text;

[0013] The reply text is converted into synthetic speech, and frame-level lip correction is performed on the mouth region in the personalized video frame sequence to generate a third video stream synchronized with the synthetic speech;

[0014] The third video stream is encoded and pushed to the live broadcast platform until the preset stop time is reached, completing an automatic live broadcast.

[0015] In some embodiments, the task package carries:

[0016] Live broadcast scheduling indication data corresponding to the live broadcast scheduling;

[0017] A hash value for integrity verification;

[0018] Metadata indicating the number of loop plays.

[0019] In some embodiments, based on the corresponding facial feature information and voiceprint information, the foreground anchor object is replaced with a face and the audio is changed to obtain a personalized video frame sequence and a personalized audio stream; comprising:

[0020] Replace the background image of the personalized video frame sequence with the corresponding site image of the branch.

[0021] In some embodiments, after capturing the audience question text in real time, it further comprises:

[0022] Classify the audience question text;

[0023] Count the number of questions of each type;

[0024] Prioritize questions with high numbers;

[0025] And after replying, the number of questions of this type is cleared.

[0026] In some embodiments, it further comprises:

[0027] Based on the live broadcast field, a text library is pre-constructed;

[0028] The text library includes: different categories of audience question texts and their corresponding reply texts.

[0029] In some embodiments, the task package is issued using a differential update mechanism, only pushing video segments or audio segments that have changed relative to the last task package, to reduce network bandwidth occupancy.

[0030] In some embodiments, it further comprises:

[0031] During the push flow, the uplink network bandwidth is monitored in real time, and the encoding rate or frame rate is dynamically adjusted according to the bandwidth change to maintain the smoothness of live streaming.

[0032] In some embodiments, further comprising:

[0033] After the live broadcast ends, a log file containing audience interaction data and live broadcast quality indicators is generated and returned to the headquarters server for subsequent model iteration and operation analysis.

[0034] The above technical solutions are adopted in the present application, and the following beneficial effects are at least achieved:

[0035] The present method generates original video materials (recorded by professional anchors in a green screen environment) through the headquarters server, and distributes personalized content to each branch through a standardized processing flow (green screen image extraction, face replacement, audio voice change, etc.), completely eliminating the dependence on the personal live broadcast ability of branch marketing personnel. This mechanism ensures that all branch live broadcast content comes from professionally recorded materials, avoiding the problem of uneven live broadcast effects caused by individual talent and image differences, effectively maintaining the consistency and professionalism of the enterprise brand image. On the one hand, there is no need for large-scale live broadcast training of branch marketing personnel, reducing training time and human cost; on the other hand, through the preset live broadcast scheduling and task package distribution mechanism, centralized control of all branch live broadcast activities is achieved, avoiding problems such as live broadcast time conflicts, content duplication or chaos, greatly improving the management efficiency and resource utilization rate of the enterprise's live broadcast business. In view of the defects of traditional recording and broadcasting methods, such as rigidity and inability to respond to audience interaction, the present method optimizes the audience experience through the following technical innovations: a trained natural language processing model is introduced to capture and respond to audience questions in real time, generating accurate replies; combining speech synthesis technology and frame-level lip correction technology, the audio of the reply content is synchronized with the anchor's lip shape in the video, simulating the natural effect of real-time conversation; through green screen image extraction, superimposing real branch scenes, face replacement for branch employee images, and audio matching corresponding voice prints, the audience has an immersive experience of "the anchor being in the local branch", enhancing the realism and appeal of the live broadcast. The present method supports differential processing of multiple branches (generating personalized video and audio based on branch scenes and employee characteristics), which can achieve both large-scale live broadcast operation through unified management by the headquarters and consideration of the localization characteristics of each branch, meeting the experience needs of audiences in different regions, and realizing the organic combination of "large-scale unified management" and "personalized scene adaptation".

[0036] In summary, the method of the present application effectively solves the pain points of existing live broadcast technology through technical innovation, and has significant advantages in protecting brand image, reducing cost, improving efficiency, and optimizing audience experience, providing an efficient and feasible technical solution for enterprises to carry out automated and intelligent live broadcast marketing.

[0037] It should be understood that the above general description and the following detailed description are only exemplary and explanatory and are not restrictive of the application. BRIEF DESCRIPTION OF DRAWINGS

[0038] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.

[0039] Figure 1 The automatic live broadcast method flowchart provided for an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0040] In order to make the purposes, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be described in detail below. Obviously, the described embodiments are only some of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0041] The scheme of the present application is to provide an unmanned live broadcast automatic live broadcast mobile phone and a corresponding automatic live broadcast method, which are as follows:

[0042] The headquarters professional anchor records original video and audio material in a green screen environment, the server performs green screen matting on the original video, extracts the foreground anchor object and generates transparent channel information; at the same time, the local real scene background of each branch and the face features and voiceprint information of the corresponding staff are collected.

[0043] For each branch, based on the face features and voiceprint information of the staff, the foreground anchor object is replaced with a face and the audio is changed, and combined with the real scene background of the branch, a personalized video frame sequence and an audio stream are generated; through a task setting system, a task package encapsulating the above content is distributed to the unmanned live broadcast automatic live broadcast mobile phone of each branch according to the preset live broadcast schedule.

[0044] During live broadcast, the mobile phone pushes personalized content at a specified time to realize unmanned live broadcast; during the live broadcast process, with the help of an artificial intelligence module, the audience's questions are captured and analyzed in real time through natural language processing technology, a reply text is generated, which is then converted into synthesized speech, and the lip shape is corrected at the frame level at the same time, so that the video and the voice are synchronized and then pushed to the live broadcast platform until the preset stop time completes the live broadcast.

[0045] The scheme solves many problems of traditional live broadcast and improves the professionalism, consistency, interactivity and realism of live broadcast through efficient video distribution, centralized task management, intelligent interaction and multimedia synthesis.

[0046] After introducing the basic principles of the application, various non-limiting embodiments of the application will be specifically introduced with reference to the drawings.

[0047] Figure 1 An automatic live broadcast method flowchart is provided for an exemplary embodiment of the application. The method runs in a "headquarters-store" two-level architecture, in which the headquarters is responsible for content production and task distribution, and the store side only deploys unmanned live broadcast automatic live broadcast mobile phones. The steps are described in detail below.

[0048] Step S11, the headquarters server generates original video material and original audio material, the original video material being recorded by the anchor in a green screen environment;

[0049] The headquarters uniformly organizes professional anchors to record videos in a preset green screen background, and the recording content includes standardized content such as product introduction, brand promotion, and activity explanation. The setting of the green screen environment provides a basis for background separation in subsequent video processing, ensuring high contrast between the anchor subject and the background. During the recording process, the original audio material of the anchor (i.e. the anchor's explanation sound) is synchronously collected, and the original video material and the original audio material are kept in time axis synchronization, which are collectively used as the basis material for subsequent personalized processing. The headquarters server checks and stores the recorded original material to ensure the integrity and availability of the material.

[0050] Step S12, performing green screen matting processing on the original video material to extract the foreground anchor object and generate transparent channel information;

[0051] The headquarters server calls a video processing module to perform green screen matting processing on the original video material. Specifically, based on a chroma key control technology (such as HSV color space threshold segmentation), the green background area is identified and removed, and only the foreground information of the anchor's body contour, action and expression, i.e. the "foreground anchor object", is retained. At the same time, the transparent channel (Alpha channel) information corresponding to the video frame is generated, which is used to mark the edge and transparent area of the foreground anchor object, providing accurate layer fusion basis for superimposing the real scene background of the store in the subsequent process, and ensuring the natural transition of the foreground anchor and the new background.

[0052] Step S13, collecting the local real scene background of a plurality of stores and the face feature information and voiceprint information of the corresponding store clerks;

[0053] For each branch, the actual operating scene (such as the store environment, the commodity display area, etc.) of the branch is photographed by the image acquisition device (such as a fixed camera) deployed in the branch to generate a "local real scene background" video or image sequence, which is used to replace the green screen background of the original video and enhance the regional authenticity and scene immersion of the live broadcast. At the same time, the facial feature information (by shooting multi-angle facial images through a high-definition camera, extracting key points, texture, etc.) and voiceprint information (by collecting voice samples of the staff through a recording device, extracting voiceprint spectrum features) of the specified staff of each branch are collected, and these information are associated with the corresponding branch and uploaded to the headquarters server to form a "branch-staff feature library", which provides data support for subsequent personalized processing.

[0054] Step S14, for each branch, the foreground anchor object is subjected to face replacement and audio voice conversion based on the corresponding facial feature information and voiceprint information, to obtain a personalized video frame sequence and a personalized audio stream;

[0055] The headquarters server performs branch-level personalized processing on the foreground anchor object extracted in step S12 according to the "branch-staff feature library":

[0056] Face replacement: based on the facial feature information of the staff corresponding to the branch, the facial region of the foreground anchor object is replaced with the facial image of the staff by using a deep learning model such as a generative adversarial network (GAN), while keeping the body movements, gestures and expression dynamics of the anchor consistent with the original video, ensuring that the replaced face is natural and has no sense of strangeness, forming a "personalized video frame sequence".

[0057] Audio voice conversion: based on the voiceprint information of the staff corresponding to the branch, the timbre and intonation of the original audio material are adjusted to match the voiceprint of the staff by using voice synthesis and conversion technology (such as spectrum mapping based on vocoder), while keeping the voice content (text information) unchanged, to generate a "personalized audio stream". Through the above processing, each branch will obtain exclusive live broadcast materials presented in the form of local staff images and voices, realizing "unified content and personalized form".

[0058] Step S15, according to the preset live broadcast schedule, the personalized video frame sequence and the personalized audio stream are packaged as a task package and delivered to the unmanned live broadcast automatic live broadcast mobile phone located in the branch;

[0059] The task management module of the headquarters server generates a corresponding live broadcast task for each unmanned live broadcast automatic live broadcast mobile phone of a branch store according to a preset live broadcast plan (such as the starting time of each branch store, the live broadcast duration, the content priority, and the like). In addition to the personalized video frame sequence and the personalized audio stream generated in step S14, the task package also encapsulates control information such as live broadcast scheduling instruction data (the starting time and the stopping time), a complete verification hash value (for verifying whether the data transmission is complete), and loop playing parameters (such as the number of times of repeated playing if necessary). The task package is delivered to the live broadcast mobile phone of each branch store through an encrypted transmission protocol (such as HTTPS), and the live broadcast mobile phone stores the task package in the local storage module after receiving the task package and waits for execution.

[0060] In step S16, the unmanned live broadcast automatic live broadcast mobile phone pushes the personalized video frame sequence and the personalized audio stream in real time when the starting time is reached, to perform unattended live broadcast.

[0061] The unmanned live broadcast automatic live broadcast mobile phone is internally provided with a timing trigger module. When the system time reaches the preset starting time in the task package, the live broadcast pushing program is automatically started: the personalized video frame sequence and the personalized audio stream stored locally are read, real-time encoding is performed through a hardware encoding and decoding module (such as an H.265 encoder), the audio and video streams are pushed to a designated live broadcast platform according to a live broadcast platform protocol (such as RTMP or HLS), and unattended live broadcast is started. During the pushing process, the mobile phone monitors the material playing progress in real time to ensure audio-visual synchronization, without the need for manual intervention.

[0062] In step S17, the audience question text is captured in real time during the live broadcast, and a trained natural language processing model is called to generate a reply text.

[0063] During the live broadcast, the unmanned live broadcast automatic live broadcast mobile phone captures the question text (such as the pop-up window, the comment area message) sent by the audience in the live broadcast room in real time through the open interface of the live broadcast platform. The question text is input into the built-in natural language processing (NLP) model (which is trained with a large amount of industry corpus and product knowledge and has the ability of intent recognition and question and answer matching). The model quickly analyzes the question intent (such as the product price, the activity rules, the after-sales policy, and the like), generates a reply content from a preset knowledge base or dynamically, outputs a “reply text”, and ensures the accuracy and pertinence of the answer.

[0064] In step S18, the reply text is converted into a synthesized speech, and frame-level lip correction is performed on the mouth area in the personalized video frame sequence, to generate a third video stream synchronized with the synthesized speech.

[0065] The voice synthesis module of the unmanned live automatic live mobile phone converts the reply text generated in step S17 into synthesized voice (uses text-to-speech TTS technology, and the tone is consistent with the voiceprint of the store clerk). At the same time, the lip correction module is called, and based on the phoneme duration and rhythm (such as syllable, pause) of the synthesized voice, the anchor mouth area of the current playing frame in the personalized video frame sequence is adjusted at the frame level: through key point detection (such as lip contour point) and image generation technology, the opening and closing degree and shape of the mouth are corrected, so that the mouth movement is accurately synchronized with the pronunciation rhythm of the synthesized voice. The corrected video stream and the synthesized voice together constitute the "third video stream", realizing the synchronization of "question and answer voice - lip shape - picture".

[0066] Step S19, encode and push the third video stream to the live platform until the preset stop time is reached, and complete an automatic live broadcast.

[0067] The unmanned live automatic live mobile phone encodes the third video stream in real time (such as using a low-delay encoding standard), and pushes it to the live platform in the form of a barrage voice or an embedded voice in the picture, and presents it to the audience, realizing real-time interaction with the audience. During the live broadcast, the system continuously monitors the preset stop time, and when the stop time is reached, it automatically stops streaming, cleans up local temporary files, and feeds back key data (such as viewing time, interaction times, etc.) of this live broadcast to the headquarters server, completing an unmanned automatic live broadcast process.

[0068] In some embodiments, the task package carries:

[0069] Live scheduling indication data corresponding to the live scheduling;

[0070] The data contains detailed information such as the start time, stop time, live duration of each store live mobile phone, and the playback content segment index corresponding to different time periods. Through these data, the live mobile phone can accurately know when to start live broadcast and when to end live broadcast, and what specific content should be played at different time points during live broadcast, so as to strictly follow the live broadcast plan formulated by the headquarters and execute in order, ensuring that the live broadcast activities of each store are consistent with the overall marketing arrangement.

[0071] Hash value for integrity check;

[0072] In the task package generation process, the headquarters server performs a hash operation on all data in the task package (including personalized video frame sequences, personalized audio streams, and various types of indication information, etc.), generates a unique hash value, and sends it together with the task package. When the live broadcast mobile phone of the branch store receives the task package, it re-performs a hash operation on the received data and compares the result with the hash value carried by the task package. If they are consistent, it indicates that the task package has not been lost or tampered with during transmission, ensuring the integrity and reliability of the data; if they are not consistent, the live broadcast mobile phone will request the headquarters server to resend the task package to ensure that the subsequent live broadcast can be based on complete and accurate data.

[0073] Metadata indicating the number of looped plays;

[0074] The metadata specifies the number of looped plays of the personalized video frame sequences and the personalized audio stream during the live broadcast process. For example, when the live broadcast time of a branch store is relatively long and the play time of a single personalized content is relatively short, by setting the number of looped plays, the live broadcast mobile phone can automatically repeat the relevant content within the preset live broadcast period without manual intervention. This setting is especially suitable for scenarios where the product information is relatively fixed and needs to be continuously displayed to the audience for a long time, saving the cost of the headquarters producing a large amount of repetitive content and ensuring the continuity and sustainability of the live broadcast.

[0075] In some embodiments, the face replacement and voice change of the foreground anchor object based on the corresponding facial feature information and voiceprint information to obtain the personalized video frame sequence and the personalized audio stream; comprising:

[0076] Replacing the background image of the personalized video frame sequence with the corresponding site image of the branch store.

[0077] In some embodiments, the process of face replacement and voice change of the foreground anchor object based on the corresponding facial feature information and voiceprint information to obtain the personalized video frame sequence and the personalized audio stream also includes replacing the background image of the personalized video frame sequence with the corresponding site image of the branch store, and the specific implementation is as follows:

[0078] After the face replacement and voice change processing are completed, the headquarters server calls the video synthesis module to perform layer fusion on the transparent channel information generated in step S12 and the local real scene background (i.e., the site image) of the branch store collected in step S13. Specifically, the foreground anchor object after face replacement is accurately superimposed on the preset position (such as in front of the store counter, in the center of the product display area, etc.) of the branch store site image according to the edge information of the foreground anchor object marked by the transparent channel, and the size and light and shadow effect of the foreground anchor object are adjusted according to the perspective relationship and lighting conditions of the site image, to ensure that the foreground anchor and the branch store background are visually naturally integrated, as if the anchor is actually in the branch store scene.

[0079] For example, when the branch store site image is a shelf display area, the foreground anchor object will be adjusted to a reasonable position in front of the shelf, and the shadow direction will be consistent with the natural light or light direction in the site, avoiding the appearance of a floating feeling or light and shadow conflict. After background replacement, the personalized video frame sequence not only associates with the branch store in terms of the character image (face, voice), but also completely matches the actual environment of the branch store in terms of scene presentation, further strengthening the authenticity and regional relevance of the live broadcast, making the audience have the cognition of "local store staff live in the store in real time", and improving the trust and acceptance of the live broadcast content.

[0080] Through this processing, the personalized video frame sequence and the personalized audio stream together constitute a "person-voice-scene" three-in-one branch store exclusive live broadcast material, which not only retains the professionalism of the unified content of the headquarters, but also realizes deep personalization through scene adaptation, effectively solving the problems of fixed recording and broadcasting scene and lack of regional targeting in traditional recording and broadcasting.

[0081] After capturing the audience question text in real time, the method further includes: classifying the audience question text; counting the number of questions of each type; preferentially replying to the question with a high number; and after replying, clearing the number of questions of this type.

[0082] Classifying the audience question text: the automatic live broadcast mobile phone question and answer processing module calls a preset classification model (such as an intent classification model based on text keyword matching or deep learning) to classify the captured audience question text by type. The classification dimension can be pre-set according to the live broadcast scene, for example, in a product sales live broadcast, it can be divided into "price consultation type" (such as "how much is this product"), "specification parameter type" (such as "what is the size"), "after-sales policy type" (such as "does it support return and exchange"), "activity rule type" (such as "how to participate in the full-price reduction activity"), etc. Through classification, the structure of a large number of audience questions is combed, laying a foundation for subsequent efficient processing.

[0083] Count the number of each type of question: the system assigns an independent counter to each type of question, and the counter automatically increments each time a new question of that type is received. For example, if 8 viewers ask about the product price within 10 minutes, the "price consultation" counter is 8. The counting process is updated in real time to ensure that the data reflects the current audience's core concerns.

[0084] Prioritize questions with high numbers: the Q&A processing module sorts the questions by the number of each type of question and lists the type with the highest number as the priority. During the live broadcast, the system will prioritize the use of natural language processing models to generate reply text for that type of question, and will present it to the audience through speech synthesis and lip correction. For example, when the "price consultation" question has the highest number, the system will prioritize answering price-related questions to ensure that the most frequently asked questions are answered in a timely manner, improving interaction efficiency and audience satisfaction.

[0085] Clear the number of that type of question after replying: after completing the reply to a high-priority question type, the system automatically resets the counter value corresponding to that type to zero to avoid repeated replies to solved core questions. At the same time, it continues to count the number of newly added questions in real time and dynamically updates the number of each type of question, forming a "capture - classification - statistics - priority reply - clear" cycle to ensure that the audience's current most concerned content is always focused on during the live broadcast, achieving precision and timeliness in interaction.

[0086] Through the above steps, the application can intelligently identify the core needs of the audience without human intervention, allocate interaction resources reasonably, and avoid ineffective responses to repeated or low-priority questions, improving interaction efficiency and enhancing the audience's sense of participation and acceptance of the live content.

[0087] In some embodiments, it also includes: based on the live broadcast field, pre-constructing a text library; the text library includes: different categories of viewer question texts and their corresponding reply texts.

[0088] In some embodiments, to further improve the accuracy and response speed of the artificial intelligence automatic answering function, the application also includes the step of pre-constructing a text library based on the live broadcast field, as follows:

[0089] The text library is a structured Q&A database specially constructed for a specific live broadcast field (such as beauty sales, home appliance shopping guides, and food promotion, etc.), and its core content includes different categories of viewer question texts and their corresponding standard reply texts. During the construction process, the headquarters technical team will combine industry characteristics, product information, common marketing scenarios, and historical live broadcast interaction data to sort out the types of questions that viewers frequently ask about, and match each type of question with a professionally reviewed reply.

[0090] For example, in the field of live streaming of makeup products, the “product efficacy category” in the text library can contain the question text “Is this cream suitable for sensitive skin?” and the corresponding answer text “This cream has been tested by dermatologists for sensitive skin and does not contain irritating ingredients such as alcohol and fragrance. Sensitive skin users can use it with confidence. It is recommended to test it on the back of the ear before first use.” The “use method category” can contain the question text “How to apply foundation more smoothly” and the corresponding answer text “It is recommended to first moisturize the skin, take the appropriate amount of foundation, and apply it to the forehead, cheeks, nose tip, and chin. Use a damp makeup egg to gently spread it from the inside out. Pay attention to the transition at the edge. You can add another layer to enhance the concealer effect as needed.”

[0091] The text library is dynamically maintained according to the updates of live streaming content (such as new product launches and activity adjustments) and changes in audience questions: when new high-frequency questions appear and there is no corresponding record in the text library, the system will prompt the operator to supplement the input; when product information or activity rules change, the relevant answer text will be updated simultaneously to ensure the timeliness and accuracy of the content in the library.

[0092] During live streaming, when the artificial intelligence module captures the audience question text, it will first match with the question text in the text library (such as based on keyword similarity algorithm), if it matches successfully, it will directly call the corresponding answer text for reply, without the need to generate it again, greatly improving the speed of answering; if it does not match, it will call the natural language processing model to generate the answer text, and at the same time, the new question and the corresponding answer will be supplemented to the text library (after being manually reviewed, it will take effect), realizing the self-iteration optimization of the text library.

[0093] By pre-constructing and dynamically maintaining the text library, the invention not only improves the efficiency and accuracy of automatic answering, but also reduces the dependence on natural language processing models for real-time generation of answers, reduces the problem of declining interactive experience caused by model calculation delay or error, and further ensures the stability and professionalism of the interactive process in the process of unmanned live streaming.

[0094] In some embodiments, the task package delivery adopts a differential update mechanism, only pushing video segments or audio segments that have changed relative to the last task package, to reduce network bandwidth occupancy.

[0095] In some embodiments, to further optimize the efficiency of task package delivery and reduce network bandwidth occupancy, the task package delivery process of the invention adopts a differential update mechanism, the specific implementation is as follows:

[0096] The headquarters server does not re-encapsulate and issue the complete personalized video frame sequence and personalized audio stream every time when generating a new task package. Instead, by comparing the current task package with the historical task package received by the branch store last time, the differences between the two are accurately identified, including but not limited to newly added video segments, modified audio passages, adjusted live streaming scheduling instruction data, etc. For video segments that do not change (such as repeatedly used commodity basic introduction segments), audio streams (such as fixed brand promotion voice audio), and configuration information, they are no longer pushed repeatedly.

[0097] Specifically, the differential update mechanism is implemented through the following steps:

[0098] The headquarters server compares the historical task package and the new task package, generates a "difference file", which only contains the video frame sequence segments, audio stream segments and corresponding update instructions (such as replacement position, insertion time point, etc.) of the changed parts;

[0099] The difference file is encapsulated into a lightweight differential task package, which has a much smaller data volume than the complete task package;

[0100] After receiving the differential task package, the unmanned live streaming automatic live streaming mobile phone of the branch store fuses the difference part with the unchanged content in the locally stored historical task package according to the update instructions, and reconstructs the complete new task package without the need to download all the data again.

[0101] For example, when a branch store adds a 3-minute explanation video of a promotional product to the original commodity introduction the next day, the headquarters server only needs to issue the 3-minute new video segment and the corresponding scheduling adjustment instruction. After receiving it, the live streaming mobile phone inserts it into the corresponding time node of the original task package to complete the task update.

[0102] Through this mechanism, the data transmission volume in the task package issuing process is significantly reduced. Especially in the case of unstable network environment or limited bandwidth of the branch store, the network congestion risk can be effectively reduced to ensure that the task package is quickly and stably transmitted to each live streaming mobile phone, while reducing the bandwidth pressure and data flow cost of the headquarters server, and improving the operation efficiency of the entire system.

[0103] In some embodiments, the uplink network bandwidth is also monitored in real time during the push streaming process, and the encoding code rate or frame rate is dynamically adjusted according to the bandwidth change to maintain the smoothness of live streaming.

[0104] In some embodiments, to ensure the stability and smoothness of the live streaming push streaming process, the unmanned live streaming automatic live streaming mobile phone of the present application also includes the step of monitoring the uplink network bandwidth and dynamically adjusting the encoding parameters during the push streaming process, which is as follows:

[0105] The mobile phone with built-in network monitoring module automatically monitors the live stream. During the live stream (including the personalized audio and video stream in step S16 and the third video stream in step S19), the module collects uplink network bandwidth data in real time at preset time intervals (such as once every 2 seconds), including key indicators such as current available bandwidth, bandwidth fluctuation range and network latency.

[0106] Simultaneously, the phone's encoding control module works in conjunction with the network monitoring module to dynamically adjust the video encoding bitrate and frame rate based on real-time bandwidth data.

[0107] When the uplink bandwidth is detected to be sufficient and stable (e.g., the actual bandwidth is more than 1.2 times the current encoding bitrate), the encoding control module automatically increases the encoding bitrate (e.g., from 2Mbps to 4Mbps) and frame rate (e.g., from 25fps to 30fps) to enhance the clarity and smoothness of the video and improve the viewer's visual experience.

[0108] When the uplink bandwidth is detected to be insufficient or fluctuating significantly (e.g., the actual bandwidth is less than 0.8 times the current encoding bitrate), the encoding control module prioritizes reducing the encoding bitrate (e.g., from 4Mbps to 1.5Mbps), and if necessary, appropriately reduces the frame rate (e.g., from 30fps to 20fps). By reducing the amount of data transmitted, the module adapts to the bandwidth limitation and avoids video stuttering, screen tearing, or streaming interruption due to insufficient bandwidth.

[0109] When the bandwidth is detected to have returned to a stable state, the encoding control module gradually increases the bit rate and frame rate to ensure a smooth transition in picture quality and avoid sudden changes in parameters from impacting the viewer's experience.

[0110] For example, during peak live streaming periods, if a branch's network uplink bandwidth drops sharply from 5Mbps to 1.8Mbps due to a surge in users, the system will quickly reduce the encoding bitrate from 3Mbps to 1.5Mbps while maintaining a frame rate of 25fps to ensure a continuous and stable video stream. Once the bandwidth recovers to 4Mbps, the bitrate will be gradually increased back to 2.5Mbps to balance image quality and smoothness.

[0111] Through the aforementioned dynamic adjustment mechanism, this invention can adapt to fluctuations in different network environments, prioritizing uninterrupted live streaming when bandwidth is limited, and improving picture quality when bandwidth is sufficient. This ensures a consistently high level of live streaming smoothness under complex network conditions, reduces viewer loss due to network issues, and further enhances the practicality and reliability of the unmanned live streaming system.

[0112] In some embodiments, the method further includes: after the live broadcast ends, generating a log file containing audience interaction data and live broadcast quality indicators, and sending the log file back to the headquarters server for subsequent model iteration and operational analysis.

[0113] In some embodiments, to achieve quantitative evaluation of the live streaming effect and continuous system optimization, the present invention further includes the step of generating a log file and sending it back to the headquarters server after the live streaming ends, as detailed below:

[0114] After the live stream ended, the automated live streaming phone automatically triggered a log generation program to summarize and organize key data from the live stream, forming a structured log file. This log file contains two main categories of core information:

[0115] Audience interaction data: This includes the total number of questions asked by viewers during the live stream, the specific number of questions of each type (e.g., 20 price inquiries and 15 after-sales inquiries), the total number of audience comments, peak interaction periods (e.g., 40% of interactions occur within 10-20 minutes after the start of the broadcast), and the distribution of audience dwell time (e.g., an average viewing time of 5 minutes, with 30% of viewers watching for more than 10 minutes), comprehensively reflecting audience participation and focus.

[0116] Live streaming quality metrics include total live streaming duration, actual streaming duration (excluding possible interruptions), number and duration of network fluctuations (e.g., 3 brief pauses due to insufficient bandwidth, totaling 5 seconds), encoding parameter adjustment records (e.g., the time and duration of bitrate adjustment from 2Mbps to 1.5Mbps), AI automatic response latency (e.g., average latency of 0.8 seconds, maximum latency of 1.5 seconds), and lip-sync accuracy (based on audience feedback or system self-check matching score), objectively recording the technical performance during the live stream.

[0117] After the log file is generated, the unmanned live streaming mobile phone encrypts and uploads it to the headquarters server via the network communication module. The headquarters server centrally stores and analyzes the log files returned by each branch:

[0118] For model iteration: The technical team can optimize the training data of the natural language processing model, adjust the model parameters, and improve the accuracy of automatic answers based on the high-frequency question types in the audience interaction data and the feedback of AI answers (such as whether there are irrelevant answers). At the same time, based on the network adaptation in the live broadcast quality indicators, the dynamic coding algorithm can be improved to enhance the system's adaptability to complex network environments.

[0119] For operation analysis: operation personnel optimize live broadcast scheduling and content strategy (such as increasing the live broadcast frequency in the high-interactive period) by analyzing the live broadcast data of different branches (such as the interaction of A branch is twice that of B branch, which may be related to the branch scene display); adjust the original video materials recorded by the headquarters according to the audience's attention (such as increasing the explanation time for the high-frequency asked product characteristics), and improve the overall live broadcast effect.

[0120] Through this step, the application constructs a closed-loop mechanism of "live broadcast execution - data feedback - optimization iteration", so that the system can continuously adapt to the needs of actual application scenarios, continuously improve the intelligent level and operation efficiency of unmanned live broadcast, and further consolidate its advantages in the live broadcast technology field.

[0121] In summary, in the scheme provided by the present application, first, through advanced Internet technology, the video of product introduction recorded by professional anchors in the headquarters is efficiently and accurately distributed to each live broadcast mobile phone distributed in different branches. It is no longer necessary for each branch marketing personnel to conduct live broadcast independently, which greatly guarantees the professionalism and consistency of live broadcast content and avoids poor live broadcast effect and damage to brand image caused by personal differences.

[0122] Secondly, through the carefully designed task setting system, the live broadcast time and playing content of each live broadcast mobile phone can be accurately arranged, ensuring the orderly progress of live broadcast and avoiding time conflicts and content confusion. This system enables the company to centrally manage and uniformly plan all live broadcast activities, greatly improving the operation efficiency.

[0123] More importantly, in order to overcome the rigidity and inflexibility of recording and broadcasting and the defect of being unable to respond to audience questions in time, the application innovatively develops the function of artificial intelligence automatically answering audience questions in live broadcast room. Based on deep learning and natural language processing technology, this function can analyze audience questions in real time and give accurate and appropriate text answers. Moreover, it can instantly convert the text answers into audio for playing, simulating the effect of real person conversation, greatly enhancing the interactivity and flexibility of live broadcast, and improving the audience's participation and experience.

[0124] At the same time, in order to further avoid the rigidity of recording and broadcasting, the application also integrates advanced face replacement technology. For example, after the headquarters anchor records the video, this technology can replace the anchor's face with the faces of 1000 employees in the branch, and can also replace the audio with different voices of these 1000 people. Although the text content is the same, the effect presented is as if 1000 different people are conducting live broadcast, each person's speaking, gestures and movements are consistent, only the face and voice are different, greatly enriching the diversity and individualization of live broadcast.

[0125] And, the application requires the anchor of the headquarters to stand in front of a green curtain when recording the video. This makes it easy to do green screen matting in the post-production, which separates the anchor's image from the background, and then superimposes the real scenes of each branch to create the visual effect that the anchor seems to be live streaming in different branches, further enhancing the realism and attractiveness of live streaming.

[0126] In addition, the application also introduces lip correction technology. When playing the audio converted from the text of the automatic reply, this technology is used to correct the anchor's mouth shape, making it look real and natural, as if the anchor is really talking. This greatly improves the realism of live streaming.

[0127] At the same time, in terms of hardware and software optimization, the live streaming mobile phone of the application has excellent performance and stability, can run stably for a long time, and ensures the smooth progress of live streaming, without being disturbed by network fluctuations and other factors.

[0128] In some embodiments, the unmanned live streaming automatic live streaming mobile phone of the application mainly includes the following parts: a high-performance processor for quickly processing video data and running related software programs; a large-capacity storage device for storing video, audio files and related configuration information issued by the headquarters; an advanced network communication module to ensure stable and high-speed reception of video data and instructions from the headquarters; a high-definition display screen to clearly display live streaming content; an artificial intelligence module built-in for functions such as automatically answering audience questions, text-to-audio conversion and lip correction; a video synthesis module for green screen matting, face replacement and scene superimposition operations.

[0129] In actual application, the professional anchor of the headquarters performs video recording in front of a green curtain, and the recorded video is uploaded to the server. The server preprocesses the video, including extracting key frames, separating audio, etc. At the same time, the facial features, actions and lip shapes of the anchor are analyzed and modeled using deep learning algorithms.

[0130] When live streaming is needed, the server distributes the processed video data to each live streaming mobile phone according to the preset task. After receiving the data, the live streaming mobile phone first maintains connection with the server through the network communication module to obtain the latest instructions and data in real time. Then, the video synthesis module performs green screen matting according to the preset rules to separate the anchor's image from the original background and superimpose the real scenes of each branch. At the same time, the face replacement technology replaces the anchor's face with the faces of the branch employees, and the audio is replaced with the corresponding employee's voice.

[0131] When the audience raises questions in the live room, the artificial intelligence module quickly analyzes the questions and generates a written answer. The written answer is converted into audio through voice synthesis technology, and the lip correction technology adjusts the anchor's lip shape to match the generated audio, which is finally played out through the live mobile phone, giving the audience a natural and real interactive experience.

[0132] Throughout the process, hardware and software work together to ensure that the live mobile phone can run stably and efficiently, unaffected by network fluctuations and other interference factors.

[0133] Specifically, the scheme provided by the present application has an efficient video distribution mechanism to ensure that the videos recorded at the headquarters reach each live mobile phone accurately and quickly. Advanced artificial intelligence automatic answering and text-to-audio technology enable real-time interaction. Innovative face swapping, green screen cutout, and lip correction technology enhance the realism and diversity of live streaming. The entire unmanned live automatic live mobile phone system architecture and workflow. Key technologies such as artificial intelligence automatic answering, text-to-audio, face swapping, green screen cutout, and lip correction are involved.

[0134] The unmanned live automatic live mobile phone of the present application brings many significant benefits. First, it greatly reduces the labor cost of enterprises in live streaming. Instead of requiring numerous marketing personnel at each branch to prepare and operate live streaming individually, only a few professional anchors at the headquarters are needed to record high-quality content, saving a lot of training and preparation time.

[0135] Second, it significantly improves the quality and consistency of live streaming. Professional anchor recording ensures the professionalism and standardization of live streaming content, avoiding uneven live streaming effects caused by individual ability differences, and effectively maintaining the company's brand image.

[0136] Third, it enhances the interactivity and flexibility of live streaming. The artificial intelligence automatic answering function for audience questions, as well as the real-time generated audio and corrected lip shape, make the audience feel more real interaction, improving audience engagement and retention.

[0137] At the same time, through face swapping technology and green screen cutout, real scenes are superimposed to provide viewers with a rich and more realistic live streaming experience, increasing the appeal and interest of live streaming.

[0138] In addition, stable and efficient performance ensures smooth live streaming, unaffected by factors such as network, reducing live streaming interruptions or lag, and improving user viewing satisfaction.

[0139] It can be understood that the same or similar parts in the above embodiments can be mutually referenced, and the contents not described in detail in some embodiments can be referred to the same or similar contents in other embodiments.

[0140] It should be noted that, in the description of the present application, the terms "first", "second", etc. are used only for descriptive purposes and cannot be understood as indicating or implying relative importance. In addition, in the description of the present application, the meaning of "a plurality of" or "multiple" is at least two, unless otherwise specified.

[0141] It should be understood that when an element is referred to as being "fixed to" or "set on" another element, it can be directly on the other element or a middle element can be present at the same time; when an element is referred to as "connected" to another element, it can be directly connected to the other element or a middle element can be present at the same time, in addition, "connected" used herein can include wireless connection; the phrase "and / or" used includes any unit and all combinations of the associated listed items.

[0142] Any process or method descriptions in flow charts or described elsewhere herein can be understood as representing one or more modules, segments, or portions of code that include executable instructions for performing specific logical functions or steps in the process, and the preferred embodiments of the present application include additional implementations in which the functions are performed in different orders, in substantially simultaneous fashion, or in reverse order, depending on the functionality involved, as will be understood by those skilled in the art.

[0143] It should be understood that parts of the present application can be implemented in hardware, software, firmware or a combination thereof. In the above-described embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented by any one or a combination of the following technologies known in the art: discrete logic circuit with logic gate circuit for implementing logic functions on data signals, application specific integrated circuit with suitable combination logic gate circuit, programmable gate array (PGA), field programmable gate array (FPGA), etc.

[0144] Those skilled in the art of the present technology can understand that all or part of the steps carried out by the above-mentioned embodiment method can be completed by program instructions to the relevant hardware, and the program can be stored in a computer readable storage medium, which includes one or a combination of steps of the method embodiments when executed.

[0145] In addition, each of the function units in each embodiment of the present application can be integrated in one processing module, or each unit can be physically present separately, or two or more units can be integrated in one module. The integrated module can be realized in the form of hardware, or in the form of a software function module. When the integrated module is realized in the form of a software function module and sold or used as an independent product, it can also be stored in a computer readable storage medium.

[0146] The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc.

[0147] In the description of the present specification, the description referring to the terms "one embodiment", "some embodiments", "an example", "a specific example" or "some examples" etc. means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the present specification, the illustrative description of the above terms does not necessarily mean the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0148] Although the embodiments of the present application have been shown and described above, it should be understood that the above embodiments are exemplary and should not be construed as limiting the present application, and those skilled in the art can make changes, modifications, replacements and variations to the above embodiments within the scope of the present application.

Claims

1. An automatic live streaming method, characterized in that, include: The headquarters server generates the original video and audio footage, which is recorded by the broadcaster in a green screen environment. Perform green screen keying on the original video footage to extract the foreground anchor object and generate alpha channel information; Collect real-world background images of multiple branches, as well as facial features and voiceprints of the corresponding staff. For each branch, based on the corresponding facial feature information and voiceprint information, the foreground anchor object is subjected to face replacement and audio voice changing to obtain a personalized video frame sequence and a personalized audio stream. According to the preset live broadcast schedule, the personalized video frame sequence and personalized audio stream are packaged into a task package and sent to the unmanned live broadcast automatic live broadcast mobile phone located in the branch. When the start time for the unmanned live streaming mobile phone arrives, it pushes the personalized video frame sequence and personalized audio stream in real time to perform unattended live streaming. During the live stream, the text of audience questions is captured in real time, and a trained natural language processing model is used to generate response text. The response text is converted into synthesized speech, and frame-level lip-sync correction is performed on the mouth region in the personalized video frame sequence to generate a third video stream synchronized with the synthesized speech; The third video stream is encoded and pushed to the live streaming platform until the preset stop time is reached, completing one automatic live stream.

2. The automatic live streaming method according to claim 1, characterized in that, The task package carries: Live broadcast schedule indicator data corresponding to the live broadcast schedule; The hash value used for integrity verification; Metadata indicating the number of times the playback will loop.

3. The automatic live streaming method according to claim 1, characterized in that, The process of performing face replacement and voice changing on the foreground anchor object based on corresponding facial feature information and voiceprint information to obtain a personalized video frame sequence and a personalized audio stream includes: Replace the background image of the personalized video frame sequence with the corresponding venue image of the branch.

4. The automatic live streaming method according to claim 1, characterized in that, After capturing the audience's question text in real time, the method also includes: Classify the text of audience questions; Count the number of questions for each type; Prioritize responding to questions with a high volume of responses; After replying, the number of questions of that type will be reset to zero.

5. The automatic live streaming method according to claim 4, characterized in that, Also includes: Based on the live streaming field, a text library is pre-built. The text library includes: different categories of audience question texts, and their corresponding response texts.

6. The automatic live streaming method according to claim 1, characterized in that, The task package is delivered using a differential update mechanism, which only pushes video or audio segments that have changed compared to the previous task package, in order to reduce network bandwidth usage.

7. The automatic live streaming method according to claim 1, characterized in that, Also includes: During the streaming process, the uplink network bandwidth is monitored in real time, and the encoding bitrate or frame rate is dynamically adjusted according to the bandwidth changes to maintain the smoothness of the live stream.

8. The automatic live streaming method according to claim 1, characterized in that, Also includes: After the live stream ends, a log file containing audience interaction data and live stream quality metrics is generated and sent back to the headquarters server for subsequent model iteration and operational analysis.

Citation Information

Patent Citations

  • Live broadcast method and device based on artificial intelligence, equipment and storage medium

    CN111010586A

  • Live broadcast autonomous interaction method and device, and computer readable medium

    CN117539986A