Digital human live broadcast method and system

By acquiring the voice of a real anchor through a local module and transmitting it to a remote module for processing, the problem of risk control for unmanned live streaming systems is solved, and flexible control of digital humans and real human voices is realized, thereby improving the live streaming effect.

CN120658895BActive Publication Date: 2025-12-16SHENZHEN CLOUDSKY TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511162507.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-19
Publication Date
2025-12-16
Estimated Expiration
2045-08-19

AI Technical Summary

Technical Problem

Existing unmanned live streaming systems cannot effectively prevent live streaming accounts from being subject to risk control, thus affecting the live streaming effect.

Method used

The local module acquires the voice information of the real anchor and transmits it to the remote module. After the remote module detects the access of the real anchor, it blocks or processes the voice information of the digital human, realizing flexible control of the digital human and the real human voice, and merging and processing to generate adapted live audio.

Benefits of technology

To avoid live streaming accounts being subject to risk control measures, meet user needs, and improve the live streaming experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120658895B_ABST
    Figure CN120658895B_ABST
Patent Text Reader

Abstract

The application discloses a digital person live broadcast method and system, relates to the technical field of artificial intelligence, and solves the technical problem that the existing unmanned live broadcast system cannot well meet the use demand of users. The system comprises a local module and a remote module, the local module is in communication connection with the remote module through a remote connection protocol, the local module is used for acquiring second voice information generated by a real person anchor and transmitting the second voice information to the remote module, and the remote module is used for playing automatically generated first voice information or the received second voice information. The application transmits the voice generated by a local microphone to a live broadcast partner to process the input and output of the voice, flexibly controls the voice, realizes the coexistence of the automatically generated first voice information of the digital person and the externally input second voice information, avoids the live broadcast account being controlled, and meets the user demand.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and in particular to a digital person live broadcast method and system. BACKGROUND

[0002] With the continuous development of Internet technology, unmanned live broadcast has gradually become one of the trends in the live broadcast industry. Unmanned live broadcast is completed by a virtual host (i.e. AI) simulating the image and action of a human being, which has high entertainment and ornamental value.

[0003] In the process of unmanned live broadcast, the voice of the virtual host is generally not transmitted into the live broadcast partner through a microphone, but audio is automatically generated in the background, which makes the live broadcast account at risk of being controlled. The existing unmanned live broadcast system cannot well process live broadcast audio and cannot avoid the live broadcast account being controlled, which easily affects the live broadcast effect.

[0004] In the process of implementing the present application, the inventors have found that at least the following problems exist in the prior art:

[0005] The existing unmanned live broadcast system cannot well meet the use requirements of users. SUMMARY

[0006] The present application aims to provide a digital person live broadcast method and system to solve the technical problem in the prior art that the existing unmanned live broadcast system cannot well meet the use requirements of users. The preferred technical solutions in the many technical solutions provided by the present application can produce the many technical effects described below.

[0007] To achieve the above-mentioned purpose, the present application provides the following technical solutions:

[0008] The digital person live broadcast method provided by the present application is applied to the digital person live broadcast system described above and comprises the following steps:

[0009] The local module opens a remote connection protocol, and the remote module acquires first voice information automatically generated by the digital person;

[0010] The remote module starts live broadcast and detects whether a real host is connected;

[0011] If yes, the digital person in the remote module shields the first voice information, acquires second voice information generated by the real host through the local module, plays the second voice information, and resumes playing the first voice information after the second voice information is played completely;

[0012] If no, the digital person in the remote module automatically plays the first voice information.

[0013] Optionally, the digital human in the remote module shields the first voice information, and after obtaining the second voice information generated by the real anchor through the local module, the method further comprises:

[0014] After the remote module receives the second voice information, the remote module creates a virtual device and monitors a temporary audio storage file, and decodes a voice format of the second voice information;

[0015] The remote module creates the temporary audio storage file according to the voice format of the second voice information;

[0016] The remote module writes the decoded second voice information into the temporary audio storage file;

[0017] When the remote module receives a stop live signal, the temporary audio storage file is destroyed.

[0018] Optionally, the remote module creates the temporary audio storage file according to the voice format of the second voice information, comprising:

[0019] Monitoring whether the temporary audio storage file exists;

[0020] If the temporary audio storage file does not exist, creating a channel of the temporary audio storage file according to the voice format;

[0021] If the temporary audio storage file exists, detecting whether the voice format of the second voice information is consistent with a voice format to be currently created;

[0022] If yes, multiplexing a channel of the current temporary audio storage file;

[0023] If no, destroying a channel of the current temporary audio storage file, and re-creating a channel of the temporary audio storage file.

[0024] Optionally, when the local module comprises a software client and a first voice chat client connected in communication, and the remote module comprises a voice chat server, a second voice chat client and a live partner, before the local module starts a remote connection protocol, the method further comprises:

[0025] The live software of the local module is started or the local module starts a remote connection protocol to connect the remote module;

[0026] The remote module obtains a voice chat port of the voice chat server assigned, and detects whether the second voice chat client is opened;

[0027] If yes, the second voice chat client is forcibly closed;

[0028] If not, set the connection IP in the database of the second voice chat client, and then start the second voice chat client to complete the IP address and communication port setting of the voice chat server.

[0029] Optionally, the local module starts a remote connection protocol, the remote module acquires the first voice information automatically generated by the digital person, and after acquiring the second voice information generated by the real anchor through the local module, the method further comprises:

[0030] The remote module fuses the first voice information and the second voice information to generate third voice information, and plays the third voice information.

[0031] Optionally, the remote module fuses the first voice information and the second voice information to generate third voice information, comprising:

[0032] Determine a fusion parameter a according to the first voice information and the second voice information, and set a proportion coefficient of the third voice information as O=0.9+0.1a;

[0033] The calculation formula of the fusion parameter a is:

[0034] a=M / N;

[0035] Wherein, M represents the first voice information, and N represents the second voice information.

[0036] Adjust the first voice information or the second voice information according to the fusion parameter a and the proportion coefficient O of the third voice information to obtain the third voice information.

[0037] A digital person live broadcast system for executing the digital person live broadcast method described above, comprising a local module and a remote module, the local module is in communication connection with the remote module through a remote connection protocol, the local module is used to acquire second voice information generated by a real anchor and transmit the second voice information to the remote module; the remote module is used to play first voice information automatically generated or the second voice information received.

[0038] Optionally, the local module comprises a local microphone and a software client in communication connection, the local microphone is used to acquire the second voice information generated by the real anchor and transmit the second voice information to the software client; the software client is used to transmit the second voice information received to the remote module.

[0039] Optionally, the remote module comprises a software server, a virtual device and a live companion, the software server is in communication connection with the virtual device, the software server is connected with the software client through a communication port, and the virtual device is in communication connection with the live companion; the virtual device is used for receiving the second voice information transmitted by the software server and the first voice information generated by a digital person in the software server, and forwarding the first voice information or the second voice information to the live companion through voice tuning software.

[0040] Optionally, the local module comprises a software client and a first voice chat client in communication connection, the remote module comprises a voice chat server, a second voice chat client and a live companion, and the voice chat server is in communication connection with the live companion through the second voice chat client.

[0041] The first voice chat client and the second voice chat client are connected on the voice chat server, so that the local module and the remote module perform voice communication.

[0042] The above technical solutions of the present application have the following advantages or beneficial effects:

[0043] The present application processes the input and output of sound by transmitting the voice generated by the local microphone to the live companion, flexibly controls the sound, realizes the coexistence of the first voice information automatically generated by the digital person and the second voice information externally input, avoids the live account being controlled, and meets the user demand. BRIEF DESCRIPTION OF DRAWINGS

[0044] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor, and the drawings are as follows:

[0045] Figure 1 is a first structure schematic diagram of the first embodiment of the present application;

[0046] Figure 2 is a second structure schematic diagram of the first embodiment of the present application;

[0047] Figure 3 is a flow chart of the second embodiment of the present application.

[0048] Fig. 1: 1, local module; 11, local microphone; 12, software client; 13, first voice chat client; 2, remote module; 21, software server; 22, virtual device; 23, live companion; 24, voice chat server; 25, second voice chat client. DETAILED DESCRIPTION

[0049] In order to make the objects, technical solutions and advantages of the present application clearer, the various exemplary embodiments to be described below will be described with reference to the corresponding drawings, which form part of the exemplary embodiments, and which describe various exemplary embodiments that can be used to implement the present application. The same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present disclosure. It should be understood that they are merely examples of procedures, methods and apparatuses, etc. consistent with some aspects of the present disclosure as detailed in the appended claims, and other implementations can be used, or modifications can be made to the implementations listed herein, without departing from the scope and spirit of the present application.

[0050] In the description of the present application, it should be understood that the terms "center", "longitudinal", "transverse", etc. indicate the orientation or positional relationship based on the drawings shown, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the elements referred to must have a particular orientation, be constructed and operated in a particular orientation. The terms "first", "second", etc. are only for descriptive purposes and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. The term "a plurality of" means two or more. The terms "connected", "connected" should be interpreted broadly, for example, it can be fixed connection, detachable connection, integral connection, mechanical connection, electrical connection, communication connection, direct connection, indirect connection through intermediate medium, internal communication of two elements or interaction relationship between two elements. The term "and / or" includes any and all combinations of one or more related listed items. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.

[0051] In order to illustrate the technical solutions described in the present application, the following will be described by specific embodiments, only showing the parts related to the embodiments of the present application.

[0052] Embodiment one:

[0053] As Figure 1 and Figure 2As shown, the present application provides a digital person live broadcast system, which comprises a local module 1 and a remote module 2, the local module 1 is in communication connection with the remote module 2 through a remote connection protocol, the local module 1 is used for acquiring second voice information generated by a real person anchor and transmitting the second voice information to the remote module 2, and the remote module 2 is used for playing automatically generated first voice information or received second voice information. Specifically, the local module 1 has the ability to collect the microphone of the client, and sends the second voice information collected by the local module 1 to a server (software server 21 or voice chat server 24 as described below) of the remote module 2 through an open communication end interface on the remote module 2, and the server forwards the automatically generated first voice information or the received second voice information to a virtual device 22, and the voice information in the virtual device 22 is forwarded to a live broadcast partner 23 by a sound mixing element for live broadcast. The remote connection protocol can be RDP protocol.

[0054] The present application processes the input and output of sound by transmitting the voice generated by the local microphone 11 to the live broadcast partner 23, flexibly controls the sound, realizes the coexistence of the automatically generated first voice information of the digital person and the externally input second voice information, avoids the live broadcast account being controlled, and meets the user demand.

[0055] As an optional implementation manner, as shown in the figure, Figure 1 As shown, the local module 1 comprises a local microphone 11 and a software client 12 in communication connection, the local microphone 11 is used for acquiring second voice information generated by a real person anchor and transmitting the second voice information to the software client 12, and the software client 12 is used for transmitting the received second voice information to the remote module 2. Specifically, the second voice information generated by the local microphone 11 is generated by a real person anchor and input to the remote module 2 through the input software client 12.

[0056] As an optional implementation manner, as shown in the figure, Figure 1 As shown, the remote module 2 comprises a software server 21, a virtual device 22 and a live broadcast partner 23, the software server 21 is in communication connection with the virtual device 22, and the software server 21 is connected with the software client 12 through a communication port, the virtual device 22 is in communication connection with the live broadcast partner 23, and the virtual device 22 is used for receiving the second voice information transmitted by the software server 21 and the first voice information generated by the digital person in the software server 21, and forwarding the first voice information or the second voice information to the live broadcast partner 23 through sound mixing software. Specifically, the first voice information generated by the digital person in the software server 21 is automatically generated by the system without external input.

[0057] As an optional implementation manner, as shown in the figure, Figure 2As shown, the local module 1 includes a software client 12 and a first voice chat client 13 connected in communication, the remote module 2 includes a voice chat server 24, a second voice chat client 25 and a live companion 23, the voice chat server 24 is connected in communication with the live companion 23 through the second voice chat client 25, wherein the first voice chat client 13 and the second voice chat client 25 are both connected to the voice chat server 24, so that the local module 1 and the remote module 2 can have voice communication. Specifically, a third-party software is used instead of a local microphone 11, and the third-party software of the embodiment is a voice chat software (Mumble). The voice chat software provides the voice chat server 24, the first voice chat client 13 and the second voice chat client 25. The first voice chat client 13 is deployed on the local module 1, and the voice chat server 24 and the second voice chat client 25 are deployed on the remote module 2. When the first voice chat client 13 and the second voice chat client 25 simultaneously receive a Mumble, voice communication between the local module 1 and the remote module 2 can be realized, so that the second voice information obtained by the local module 1 is transmitted to the remote module 2.

[0058] Embodiment two:

[0059] As shown in the figure, a digital person live broadcast method, comprising: Figure 3

[0060] S10, the local module starts a remote connection protocol, and the remote module obtains first voice information automatically generated by the digital person;

[0061] S20, the remote module starts live broadcast, and detects whether there is a real anchor access;

[0062] S30, if yes, the digital person in the remote module shields the first voice information, and obtains second voice information generated by the real anchor through the local module, plays the second voice information, and restores the playing of the first voice information after the playing of the second voice information is completed;

[0063] S40, if no, the digital person in the remote module automatically plays the first voice information. Specifically, when there is no one on duty, the digital person automatically plays the first voice information to complete the live broadcast. When a real person patrols and accesses, the digital person needs to play the second voice information. At this time, the priority of the second voice information is higher, so the first voice information should be shielded or weakened / muted, and the first voice information should be restored after the second voice information is completed.

[0064] The present application realizes that the first voice information automatically generated by the digital person of the remote module and the second voice information input from outside exist at the same time, and flexibly controls the output of the voice information to meet the different use requirements of users.

[0065] ​As an optional implementation, the digital human in the remote module shields the first voice information, and after obtaining the second voice information generated by the real person anchor through the local module, the remote module further comprises the following steps:

[0066] After the remote module receives the second voice information, the remote module creates a virtual device and monitors a temporary audio storage file, and decodes the voice format of the second voice information;

[0067] The remote module creates a temporary audio storage file according to the voice format of the second voice information;

[0068] The remote module writes the decoded second voice information into the temporary audio storage file;

[0069] When the remote module receives a stop live signal, the temporary audio storage file is destroyed. Specifically, after the local module sends, the software server or the voice chat server of the remote module creates a virtual device and starts monitoring the temporary audio storage file (even if the channel does not exist). The decoded voice format includes the number of channels, the sampling rate, the sampling bit number, etc.

[0070] As an optional implementation, the remote module creates a temporary audio storage file according to the voice format of the second voice information, comprising:

[0071] Monitoring whether the temporary audio storage file exists;

[0072] If the temporary audio storage file does not exist, creating a channel of the temporary audio storage file according to the voice format;

[0073] If the temporary audio storage file exists, detecting whether the voice format of the second voice information is consistent with the voice format to be created currently;

[0074] If yes, multiplexing the channel of the current temporary audio storage file;

[0075] If no, destroying the channel of the current temporary audio storage file, and re-creating the channel of the temporary audio storage file.

[0076] As an optional implementation, when the local module comprises a software client and a first voice chat client connected in communication, and the remote module comprises a voice chat server, a second voice chat client and a live partner, before the local module starts the remote connection protocol, the local module further comprises the following steps:

[0077] The live software of the local module is started or the local module starts the remote connection protocol to connect the remote module;

[0078] The remote module obtains the voice chat port of the voice chat server assigned, and detects whether the second voice chat client is opened;

[0079] If yes, the second voice chat client is forcibly closed;

[0080] If no, the connection IP is set in the database of the second voice chat client, and the second voice chat client is started again to complete the IP address and communication port setting of the voice chat server. Specifically, when the second voice chat server performs the IP:PORT write operation, the second voice chat client must be closed to avoid data consistency problems.

[0081] As an optional implementation, the local module starts the remote connection protocol, the remote module obtains the first voice information automatically generated by the digital person, and after obtaining the second voice information generated by the real anchor through the local module, the method further comprises:

[0082] The remote module fuses the obtained first voice information and second voice information to generate third voice information, and plays the third voice information. Specifically, since the sudden voice of the real anchor is conspicuous, it will affect the live broadcast effect. The second voice information can be taken as the main part, the first voice information can be taken as the auxiliary part, and the third voice information can be generated by fusion and played. The third voice information is more suitable for the rhythm, atmosphere and environment of live broadcast. Using the third voice information for live broadcast can not only avoid the risk control of the live broadcast platform, but also take into account the viewing effect of the live broadcast audience.

[0083] As an optional implementation, the remote module fuses the obtained first voice information and second voice information to generate third voice information, comprising:

[0084] The fusion parameter a is determined according to the first voice information and the second voice information, and the proportion coefficient of the third voice information is set as O=0.9+0.1a;

[0085] The calculation formula of the fusion parameter a is:

[0086] a=M / N;

[0087] Wherein, M represents the first voice information, and N represents the second voice information;

[0088] The first voice information or the second voice information is adjusted according to the fusion parameter a and the proportion coefficient O of the third voice information to obtain the third voice information. Specifically, the first voice information or the second voice information is adjusted according to the fusion parameter a and the proportion coefficient O of the third voice information, comprising: when a>1 and O>1, the second voice information is increased to N*O; when a<1 and O<1, the first voice information is reduced to M*O. Further, the Kalman-GOCNN algorithm combining Kalman filter and genetic optimization convolutional neural network (GO-CNN) is used to fuse the first voice information and the second voice information, which can obtain smooth fused third voice information output.

[0089] The above description is merely that of the preferred embodiments of the application, and various changes or modifications can be suggested to one skilled in the art without departing from the spirit and scope of the application. In addition, various changes or modifications can be suggested to one skilled in the art without departing from the spirit and scope of the application, and the scope of the application is not limited to the specific embodiments disclosed herein.

Claims

1. A method for live streaming digital humans, characterized in that, include: The local module initiates a remote connection protocol, and the remote module obtains the first voice information automatically generated by the digital human. The remote module starts live streaming and detects whether a real broadcaster is connected; If so, the digital human in the remote module blocks the first voice information and obtains the second voice information generated by the real anchor through the local module, plays the second voice information until the second voice information is played, and then resumes playing the first voice information. If not, the digital human in the remote module will automatically play the first voice information; The local module initiates a remote connection protocol, and the remote module acquires the first voice information automatically generated by the digital human. After acquiring the second voice information generated by the real anchor through the local module, the system further includes: The remote module fuses the first and second voice information to generate third voice information, and then plays the third voice information. The remote module fuses the acquired first and second voice information to generate third voice information, including: The fusion parameter a is determined based on the first voice information and the second voice information, and the scaling factor of the third voice information is set to O = 0.9 + 0.1a; The formula for calculating the fusion parameter a is: a = M / N; Where M represents the first voice information and N represents the second voice information; The first or second speech information is adjusted according to the fusion parameter a and the scaling factor O of the third speech information to obtain the third speech information.

2. The digital human live streaming method according to claim 1, characterized in that, The digital human in the remote module masks the first voice information, and after obtaining the second voice information generated by the real anchor through the local module, it further includes: After receiving the second voice information, the remote module creates a virtual device and monitors the temporary audio storage file, while simultaneously decoding the voice format of the second voice information. The remote module creates the temporary audio storage file according to the voice format of the second voice information; The remote module writes the decoded second voice information into the temporary audio storage file; When the remote module receives a stop live broadcast signal, it destroys the temporary audio storage file.

3. The digital human live streaming method according to claim 2, characterized in that, The remote module creates the temporary audio storage file according to the voice format of the second voice information, including: Monitor whether the temporary audio storage file exists; If the temporary audio storage file does not exist, a channel for the temporary audio storage file is created according to the voice format; If the temporary audio storage file exists, then it is checked whether the voice format of the second voice information is consistent with the voice format to be created. If so, then reuse the channel of the temporary audio storage file described above; If not, the channel for the current temporary audio storage file is destroyed, and the channel for the temporary audio storage file is recreated.

4. The digital human live streaming method according to claim 1, characterized in that, When the local module includes a software client and a first voice chat client for communication connection, and the remote module includes a voice chat server, a second voice chat client, and a live streaming companion, before the local module initiates the remote connection protocol, it further includes: The local module's live streaming software is started or the local module enables a remote connection protocol to connect to the remote module. The remote module obtains the assigned voice chat port of the voice chat server and detects whether the second voice chat client is open. If so, then forcibly close the second voice chat client; If not, then set the connection IP in the database of the second voice chat client, and then start the second voice chat client to complete the IP address and communication port settings of the voice chat server.

5. A digital human live streaming system, used to execute the digital human live streaming method according to any one of claims 1-4, characterized in that, It includes a local module (1) and a remote module (2). The local module (1) communicates with the remote module (2) through a remote connection protocol. The local module (1) is used to obtain the second voice information generated by the real anchor and transmit it to the remote module (2). The remote module (2) is used to play the automatically generated first voice information or the received second voice information.

6. The digital human live streaming system according to claim 5, characterized in that, The local module (1) includes a local microphone (11) and a software client (12) connected by communication. The local microphone (11) is used to acquire the second voice information generated by the live anchor and transmit it to the software client (12). The software client (12) is used to transmit the received second voice information to the remote module (2).

7. The digital human live streaming system according to claim 6, characterized in that, The remote module (2) includes a software server (21), a virtual device (22), and a live streaming companion (23). The software server (21) is connected to the virtual device (22) and is connected to the software client (12) through a communication port. The virtual device (22) is connected to the live streaming companion (23). The virtual device (22) is used to receive the second voice information transmitted by the software server (21) and the first voice information generated by the digital human in the software server (21), and forward the first voice information or the second voice information to the live streaming companion (23) through the tuning software.

8. The digital human live streaming system according to claim 5, characterized in that, The local module (1) includes a software client (12) and a first voice chat client (13) connected by communication. The remote module (2) includes a voice chat server (24), a second voice chat client (25), and a live streaming companion (23). The voice chat server (24) is connected to the live streaming companion (23) through the second voice chat client (25). The first voice chat client (13) and the second voice chat client (25) are both connected to the voice chat server (24) so ​​that the local module (1) can make a voice call with the remote module (2).

Citation Information

Patent Citations

  • AI reception digital person and real person non-inductive switching method and system and storage medium

    CN118278448A