Media information processing method, apparatus, device and storage medium based on live broadcast

HK40070425BActive Publication Date: 2026-07-17TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Patent Information

Authority / Receiving Office
HK · HK
Patent Type
Patents
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2022-09-20
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

In existing technologies, the output of media recommendation information during live streaming requires operation by the host or assistant, which reduces the smoothness and efficiency of the live stream.

Method used

By linking live content with media information output instructions, the system automatically identifies the anchor's words, actions, or gaze focus to automatically output media recommendation information. This includes technologies such as speech recognition, image recognition, and gaze detection, and adjusts the area and method of the recommendation information.

Benefits of technology

It improved the output efficiency of media recommendation information, reduced the interference of manual operation, and enhanced the smoothness of live broadcasts and the timeliness of information recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The application provides a live broadcast-based media information processing method, device and equipment and a computer readable storage medium. The method comprises the following steps: presenting a content interface of a live broadcast, and outputting live broadcast content of the live broadcast through the content interface; when the live broadcast content is associated with a media information output instruction, outputting media recommendation information through the content interface according to the media information output instruction; wherein the media recommendation information is used for recommending at least one to-be-recommended object. Through the application, the output efficiency of the media recommendation information can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to live streaming technology, and more particularly to a method, apparatus, device, and computer-readable storage medium for processing media information based on live streaming. Background Technology

[0002] In live video applications, promotional activities or advertisements can be inserted during the live stream, such as a streamer promoting a game during a game stream. Related technologies pre-encode and insert promotional content or advertisements into the live stream to output media recommendations. However, this method requires individual operation by the streamer or assistance from a helper to output the media recommendations, reducing the smoothness of the live stream and the efficiency of media recommendation output. Summary of the Invention

[0003] This application provides a media information processing method, apparatus, device, and computer-readable storage medium based on live streaming, which can realize the output of media recommendation information by associating live streaming content with media information output instructions, thereby improving the output efficiency of media recommendation information.

[0004] The technical solution of this application embodiment is implemented as follows:

[0005] This application provides a media information processing method based on live streaming, including:

[0006] Present the live stream content interface and output the live stream content through the content interface;

[0007] When the live stream content is associated with a media information output instruction, media recommendation information is output through the content interface according to the media information output instruction;

[0008] The media recommendation information is used to recommend at least one object to be recommended.

[0009] This application provides a media information processing device based on live streaming, including:

[0010] The presentation module is used to present the content interface of the live stream and output the live stream content through the content interface.

[0011] The output module is used to output media recommendation information through the content interface according to the media information output instruction when the live content is associated with a media information output instruction;

[0012] The media recommendation information is used to recommend at least one object to be recommended.

[0013] In the above scheme, before outputting media recommendation information through the content interface according to the media information output instruction, the device further includes:

[0014] The instruction determination module is used to determine, when the live stream is a video live stream or an audio live stream, and the live stream content contains a target statement, that the live stream content is associated with a media information output instruction, and

[0015] The instruction associated with the target statement is used as the media information output instruction.

[0016] In the above scheme, the instruction determination module is further used for

[0017] When the live stream is a video or audio live stream, and the live stream content contains a target statement, and the number of times the target statement appears reaches a threshold, it is determined that the live stream content is associated with a media information output instruction, and

[0018] The instruction associated with the target statement is used as the media information output instruction.

[0019] In the above scheme, the instruction determination module is further used for

[0020] When the live stream is a video live stream, and the live stream content contains the broadcaster's target action, it is determined that the live stream content is associated with a media information output instruction, and

[0021] The instruction associated with the target action is used as the media information output instruction.

[0022] In the above scheme, the instruction determination module is further used for

[0023] When the live stream is a video live stream, and the live stream content contains the object to be recommended, it is determined that the live stream content is associated with a media information output instruction, and

[0024] The instruction associated with the object to be recommended is used as the media information output instruction.

[0025] In the above scheme, the instruction determination module is further used for

[0026] When the live stream is a video live stream, and the anchor's gaze is focused on a target area in the content interface for a duration that reaches a duration threshold, it is determined that the live stream content is associated with a media information output instruction, and

[0027] The instruction associated with the target region is used as the media information output instruction.

[0028] In the above scheme, the output module is further configured to output the media recommendation information in the area indicated by the media information location instruction in the content interface, based on the media information output instruction and the media information location instruction, when the live content is also associated with a media information location instruction.

[0029] In the above scheme, the output module is further configured to output the media recommendation information through a target area located in the content interface;

[0030] Accordingly, the device also includes:

[0031] The adjustment module is used to adjust at least one of the following parameters of the target area when the live content contains target content that indicates the adjustment of the target area during the process of outputting the media recommendation information through the target area: area size and area position.

[0032] The media recommendation information is output through the adjusted target area.

[0033] In the above scheme, the output module is further configured to output each media recommendation information corresponding to the target recommendation object from the at least two media recommendation information through the content interface when the number of media recommendation information is at least two and the media information output instruction indicates the output of media recommendation information corresponding to the target recommendation object.

[0034] In the above scheme, the output module is further configured to play the media recommendation information using a floating window when the live stream is a video live stream and the media recommendation information is of video type;

[0035] During the playback of the media recommendation information, the live video is played in a silent mode, and the text corresponding to the audio content in the live video is displayed on the content interface.

[0036] In the above scheme, the output module is also used to present the media recommendation information in the content interface through a bubble-style card overlay when the type of the media recommendation information is text;

[0037] When the media recommendation information is of the type of audio or video, the media recommendation information is played through a sub-interface independent of the content interface.

[0038] In the above scheme, before outputting media recommendation information through the content interface, the device further includes:

[0039] The prompt module is used to present prompt information in the content interface, and the prompt information is used to prompt the output of media recommendation information;

[0040] Accordingly, the output module is also configured to output the media recommendation information through the content interface when the live content contains target content for indicating the output of media recommendation information, based on the prompt information.

[0041] This application provides a computer device, including:

[0042] Memory, used to store executable instructions;

[0043] The processor, when executing executable instructions stored in the memory, implements the live-stream-based media information processing method provided in the embodiments of this application.

[0044] This application provides a computer-readable storage medium storing executable instructions, which, when executed by a processor, implement the live-stream-based media information processing method provided in this application.

[0045] The embodiments of this application have the following beneficial effects:

[0046] This application presents a live streaming content interface and outputs the live streaming content through the content interface; when the live streaming content is associated with a media information output instruction, media recommendation information is output through the content interface according to the media information output instruction; in this way, media recommendation information can be output through live streaming content associated with a media information output instruction, thereby improving the output efficiency of media recommendation information. Attached Figure Description

[0047] Figure 1 A schematic diagram of an optional architecture for a live-stream-based media information processing system provided in this application embodiment;

[0048] Figure 2 An optional flowchart illustrating a live-stream-based media information processing method provided in an embodiment of this application;

[0049] Figure 3 This is a schematic diagram of the architecture of the speech recognition system provided in the embodiments of this application;

[0050] Figure 4 A schematic diagram of the text segmentation process provided in the embodiments of this application;

[0051] Figure 5 A schematic diagram of the text segmentation process provided in the embodiments of this application;

[0052] Figure 6 A schematic diagram of the text segmentation process provided in the embodiments of this application;

[0053] Figure 7 A schematic diagram of the text segmentation process provided in the embodiments of this application;

[0054] Figure 8 A schematic diagram of the text segmentation process provided in the embodiments of this application;

[0055] Figures 9A-9C A schematic diagram illustrating the target action provided in the embodiments of this application;

[0056] Figure 10 This is a schematic diagram illustrating the output of media recommendation information provided in an embodiment of this application.

[0057] Figure 11 This is a schematic diagram of line-of-sight detection provided in an embodiment of this application;

[0058] Figure 12 This is a schematic diagram of the output interface for media recommendation information provided in an embodiment of this application;

[0059] Figure 13A A schematic diagram showing the adjustment of the output interface for media recommendation information provided in this application embodiment;

[0060] Figure 13B A schematic diagram showing the adjustment of the output interface for media recommendation information provided in this application embodiment;

[0061] Figure 14 A schematic diagram of the interface for outputting media information provided in an embodiment of this application;

[0062] Figure 15 This is a schematic diagram of the prompt information display interface provided in an embodiment of this application;

[0063] Figure 16 An optional flowchart illustrating a live-stream-based media information processing method provided in an embodiment of this application;

[0064] Figure 17 A schematic diagram of an optional structure of a live-stream-based media information processing apparatus provided in an embodiment of this application;

[0065] Figure 18 This is an optional structural schematic diagram of a computer device provided in an embodiment of this application. Detailed Implementation

[0066] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0067] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0068] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0069] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.

[0070] 1) Client: An application running on a terminal that provides various services, such as video playback clients, instant messaging clients, live streaming clients, and educational clients.

[0071] 2) In response to, used to indicate the conditions or states on which the operation performed depends. When the conditions or states on which it depends are met, one or more operations performed may be performed in real time or with a set delay. Unless otherwise specified, there is no restriction on the order in which the multiple operations are performed.

[0072] See Figure 1 , Figure 1 This is an optional architecture diagram of a live-stream-based media information processing system 100 provided in an embodiment of this application. To support an exemplary application, the terminal includes a first terminal 400 and a second terminal (second terminal 500-1 and second terminal 500-2 are shown as examples). The first terminal is located on the broadcaster's side, and the second terminal is located on the viewer's side. The terminal is connected to the server 200 through a network 300, which can be a wide area network, a local area network, or a combination of both.

[0073] The terminal can be any type of user terminal such as a smartphone, tablet, or laptop, or any combination of two or more of these data processing devices such as a desktop computer, game console, television, or the server 200 can be a single server configured to support various services, a server cluster, or a cloud server, etc.

[0074] In practical applications, the first terminal 400 presents the live content interface and outputs the live content through the content interface; when the live content is associated with a media information output instruction, it sends a request to obtain the media information to the server 200; based on the request, the server 200 obtains media recommendation information for recommending at least one object to be recommended, and sends the obtained media recommendation information to the terminals (including the first terminal and the second terminal); the terminals (including the first terminal and the second terminal) are used to output the media recommendation information through the live content interface.

[0075] Based on the above description of the live-stream-based media information processing system provided in the embodiments of this application, the live-stream-based media information processing method provided in the embodiments of this application will be described next. In actual implementation, the live-stream-based media information processing method provided in the embodiments of this application can be provided by a server or a terminal (such as...). Figure 1 The first terminal in the process can be implemented independently, or it can be implemented by a server and a terminal (such as...). Figure 1 The first terminal in the process) will be implemented in collaboration with the next step. Figure 1 ,by Figure 1 The following is an example of using a first terminal on the broadcaster's side to implement the live-streaming information processing method provided in this application embodiment, and a second terminal on the viewer's side to display the broadcaster's live-streaming content.

[0076] See Figure 2 , Figure 2 This is an optional flowchart illustrating an embodiment of the live-stream-based information processing method provided in this application, which will be combined with... Figure 2 The steps shown are explained.

[0077] Step 101: The first terminal presents the live broadcast content interface and outputs the live broadcast content through the content interface.

[0078] In practical applications, the first terminal is equipped with video playback clients, instant messaging clients, live streaming clients, etc. The anchor can log in to the anchor end of the client to conduct live broadcasts and convey the live broadcast content to the audience. The types of live broadcast content are diverse, such as talent shows, live teaching, live e-commerce, game live broadcasts, sports event live broadcasts, etc.

[0079] Step 102: When the live content is associated with a media information output instruction, output media recommendation information through the content interface according to the media information output instruction.

[0080] The media recommendation information is used to recommend at least one target audience. This means that during a live stream, other media information can also be recommended, such as game advertising during a game stream or recommending relaxing music during a live lesson break. During the live stream, when the streamer wants to convey media recommendation information to the audience, they can trigger a media information output command through the current live stream content. This command will then cause the first terminal to output the media recommendation information through the live stream's content interface.

[0081] Here, the output method varies depending on the type of media recommendation information. For example, if the media recommendation information is text, the output method is presentation, that is, the media recommendation information is presented through the live content interface; if the media recommendation information is audio or video, the output method is playback, that is, the media recommendation information is played through the live content interface.

[0082] For example, in game live streaming, the content consists of the streamer's commentary and gameplay introduction of the game being streamed. During the live stream, if the streamer wants to promote other games, they can trigger a media information output command by including specific live stream statements or actions. After receiving the media information output command, the first terminal plays the corresponding promotional advertisement (i.e., media recommendation information) on the live stream's content interface. Similarly, in live-streamed lessons, the content consists of the streamer's analysis of the course and explanation of test questions. During the live stream, if the streamer wants to recommend entertainment short videos or music, they can trigger a media information output command by including specific live stream statements or actions. After receiving the media information output command, the first terminal plays the corresponding entertainment short videos or music, etc., on the live stream's content interface.

[0083] In some embodiments, before outputting media recommendation information through the content interface according to the media information output instruction, the first terminal can determine that the live content is associated with a media information output instruction in the following ways:

[0084] When the live stream content contains content corresponding to the streamer's target behavior, it is determined that the live stream content is associated with a media information output instruction. In actual implementation, during the streamer's live stream, the first terminal analyzes the streamer's behavior in the live stream content and obtains the analysis results. When the analysis results indicate that the live stream content contains content corresponding to the streamer's target behavior, it is determined that the live stream content is associated with a media information output instruction.

[0085] Here, the target behavior can be a pre-set behavior that outputs related media information, or it can be a behavior that outputs media recommendation information by the first terminal through analysis of the live broadcast content. Generally speaking, the target behavior of the anchor can include at least one of the anchor's target statements or target actions during the live broadcast.

[0086] During the live broadcast, the first terminal performs semantic recognition on the broadcaster's statements in real time to determine whether the statements contain the target statement. Alternatively, it captures the broadcaster's live images in real time through image acquisition devices (such as cameras), performs image recognition on the live images, and determines whether the live images contain the target action based on the recognition results. This determines whether there is content in the live content that corresponds to the target action. When it is determined that there is content in the live content that corresponds to the target action, it obtains the media information output instruction associated with the target action.

[0087] In some embodiments, before outputting media recommendation information through the content interface according to the media information output instruction, the first terminal may determine the media information output instruction in the following manner:

[0088] When the live stream is a video or audio live stream, and the live stream content contains the target statement, it is determined that the live stream content is associated with a media information output instruction, and the instruction associated with the target statement is used as the media information output instruction.

[0089] Here, the broadcaster can trigger media information output commands via voice, that is, pre-set target statements to instruct the output of media recommendation information. In actual implementation, during the broadcast, the first terminal collects the broadcaster's voice content in real time through audio data acquisition devices (such as microphones), performs speech recognition on the voice content to obtain the corresponding text content, and then determines whether a media information output command triggered by the live broadcast statement has been received by judging whether the text content includes the pre-set target statement to instruct the output of media recommendation information. For example, the target statement to instruct the output of media recommendation information can be set as "play an advertisement" or "relax a bit". When the live broadcast content contains "play an advertisement", it is determined that the live broadcast content is associated with a media information output command, and the command associated with the target statement "play an advertisement" is taken as the media information output command, that is, instructing the output of the corresponding advertisement information.

[0090] See Figure 3 , Figure 3 This is a schematic diagram of the architecture of the speech recognition system provided in the embodiments of this application. See also: Figure 3The speech recognition system includes: an acoustic front-end 301, an acoustic model (AM) 302, a decoder 303, a language model (LM) 304, and a dictionary 305. The acoustic front-end is considered the decoding stage of the sound, involving signal processing to digitize analog signals and convert them into a sequence of feature vectors. The AM represents the acoustic features of the speech units to be recognized; it typically refers to the process of establishing a statistical representation for the feature vector sequence calculated from the speech waveform. The AM has a significant impact on the performance of the speech recognition system. The LM represents the grammar of a language, defining the acceptable sequences of words or phrases that can appear in context. The decoder, through the acoustic front-end, AM, and LM, converts the input sequence of speech features into a character sequence.

[0091] In some embodiments, the first terminal may also perform speech recognition on the collected voice content of the anchor through machine learning algorithms to obtain the text content corresponding to the voice content, and match the obtained text content with the target statement set in advance to indicate the output of media recommendation information, so as to determine whether a media information output instruction triggered based on the live broadcast content has been received.

[0092] In practical implementation, it is necessary to perform text segmentation on the text content corresponding to the speech content, and then match the segmented words with the target sentence. The following will explain the implementation algorithms of text segmentation. Text segmentation algorithms include: dictionary-based methods, Hidden Markov Models (HMMs), maximum entropy models, Conditional Random Fields (CRFs), deep learning models (such as Bi-directional Long Short-Term Memory networks (Bi-LSTMs), and unsupervised learning (such as based on cohesion and degrees of freedom).

[0093] The text segmentation process includes: sample classification training, text semantic analysis, text feature processing, model training, and model evaluation. Common training methods for sample classification training include, but are not limited to, the following: Naive Bayes classifier, Support Vector Machine (SVM), Decision Tree classification algorithm, and Ensemble Method (EM). Text semantic analysis mainly involves word segmentation, named entity recognition, and part-of-speech tagging. Text feature processing includes feature dimensionality reduction, using evaluation functions (such as Term Frequency-Inverse Term Frequency (TF-IDF), mutual information methods, expected cross-entropy, statistical methods, and genetic algorithms) and calculating feature vector weights. Model evaluation includes recall, precision, and F-measure.

[0094] See Figure 4 , Figure 4 This is a schematic diagram of the text segmentation process provided in an embodiment of this application. Figure 4 The model shown is a Chinese word segmentation model based on Bi-LSTM-CRF. This module includes a word embedding layer, a two-layer encoding layer, and a CRF layer. In actual implementation, the text content to be segmented is first processed by word embedding to obtain the corresponding word vector sequence. Then, the word vector sequence is input into the Bi-LSTM layer, which consists of two LSTMs: one for the forward input sequence and one for the backward input sequence, enabling the model to consider contextual features simultaneously, such as the contextual feature vector I above. i (Extracted via a forward process) and the feature vector r below i (Extracted through a backward process), the feature vector l from the above text i And the eigenvector r below i The concatenated feature vector C is obtained by concatenating the features. i The concatenated feature vector C i Input the CRF layer to obtain the optimal word segmentation tag sequence, and complete the word segmentation.

[0095] See Figure 5 , Figure 5 This is a schematic diagram of the text segmentation process provided in an embodiment of this application. Figure 5 The diagram shows a convolutional neural network (CNN) word segmentation model, which includes a word vector layer, a convolutional layer, a pooling layer, and a fully connected layer. In practice, each word in the text is first mapped to a word vector space. Assuming the word vector is k-dimensional, mapping n words is equivalent to generating an n*k-dimensional image. Then, convolutional operations are performed through the convolutional layer. After the convolutional operation, pooling operations are performed through the pooling layer. Finally, a fully connected neural network is used for classification prediction to obtain the classification label (i.e., the word segmentation result).

[0096] See Figure 6 , Figure 6 This is a schematic diagram of the text segmentation process provided in an embodiment of this application. Figure 6 The diagram shows a convolutional neural network (CNN) word segmentation model, which includes a word embedding layer, a two-layer encoding layer, and a concatenation layer. In practice, the text content to be segmented is first processed by word embedding to obtain the corresponding word vector sequence. Then, the word vector sequence is input into a Bi-LSTM layer, which consists of two LSTMs: one for the forward input sequence and one for the backward input sequence. This allows the model to consider contextual features simultaneously, such as the preceding feature vector (extracted through the forward process) and the following feature vector (extracted through the backward process). The preceding and following feature vectors are then concatenated to obtain the concatenated features. Next, the concatenated features are averaged, meaning that each word in each sentence votes for the final classification result to obtain the average features. Finally, the average features are classified using the output activation function to obtain the classification label (i.e., the word segmentation result).

[0097] See Figure 7 , Figure 7 This is a schematic diagram of the text segmentation process provided in an embodiment of this application. Figure 7 The model shown is a Seq2Seq-based word segmentation model, which includes an encoding layer, a decoding layer, and an intermediate state vector C connecting the two. In actual implementation, the text content to be segmented is first input into the encoding layer, which encodes the text content into a fixed-size state vector C. The state vector is then passed to the decoding layer, which learns from the state vector C and outputs the classification label (i.e., the word segmentation result).

[0098] See Figure 8 , Figure 8 This is a schematic diagram of the text segmentation process provided in an embodiment of this application. Figure 8 The diagram illustrates a word segmentation model based on a hierarchical attention network. This model includes an input layer, a word encoding layer, a word attention layer, a sentence encoding layer, a sentence attention layer, and an output layer. In practice, the text content to be segmented is first transmitted to the word encoding layer through the input layer. The word encoding layer encodes the text content to obtain the corresponding word vectors. Then, a bidirectional GRU (Gated Recurrent Unit) is used to summarize information from both directions of each word vector to obtain annotations for the corresponding words, merging the context information into the sentence vectors. Next, the word attention layer performs attention processing on the sentence vectors, the sentence encoding layer encodes the attention-processed sentence vectors, and a bidirectional GRU is used to construct document vectors. After the sentence attention layer processes the document vectors, the output layer outputs the word segmentation results.

[0099] It should be noted that the above-mentioned text segmentation algorithm is applicable to other parts of this application related to speech recognition and text segmentation. The above-mentioned text segmentation algorithm will not be described again in other parts involving speech recognition and text segmentation.

[0100] In some embodiments, before outputting media recommendation information through the content interface according to the media information output instruction, the first terminal may also determine the media information output instruction in the following manner:

[0101] When the live stream is a video live stream or an audio live stream, and the live stream content contains a target statement, and the number of times the target statement appears reaches a threshold, it is determined that the live stream content is associated with a media information output instruction, and the instruction associated with the target statement is used as the media information output instruction.

[0102] Here, the frequency of occurrence of target phrases in the live stream content is counted. Only when the frequency reaches a threshold is it determined that the live stream content is associated with a media information output instruction, and the instruction associated with the target phrase is used as the media information output instruction. For example, the target phrase used to instruct the output of media recommendation information is set as "play an advertisement" or "relax." When the live stream content contains "relax" and the frequency of "relax" reaches 2 times, it is determined that the live stream content is associated with a media information output instruction, and the instruction associated with the target phrase "relax" is used as the media information output instruction.

[0103] In some embodiments, before outputting media recommendation information through the content interface according to the media information output instruction, the first terminal may determine the media information output instruction in the following manner:

[0104] When the live stream is a video live stream and the live stream content contains the broadcaster's target action, it is determined that the live stream content is associated with a media information output instruction, and the instruction associated with the target action is used as the media information output instruction.

[0105] In practice, during the live broadcast, the first terminal captures the broadcaster's live image in real time using an image acquisition device (such as a camera), performs motion recognition on the live image, and determines whether the live image contains the broadcaster's actions and whether the broadcaster's actions are the target actions based on the recognition results. Here, actions include gestures, postures, or facial expressions, etc., and the target actions are used to indicate the output of media recommendation information.

[0106] It should be noted that the target action is preset. For example, if the action is a gesture, the target action can be a single-handed gesture or a two-handed gesture; taking a two-handed gesture as an example... Figures 9A-9C A schematic diagram of the target action provided in the embodiments of this application is shown below. Figures 9A-9CThe target gesture can be an OK sign with both hands, a fist sign with both hands, or a V sign with both hands.

[0107] By using the above method to trigger media information output instructions through target actions, the chances of the anchor accidentally triggering the instructions during the live broadcast can be reduced, thus lowering the risk of misoperation.

[0108] In some embodiments, before outputting media recommendation information through the content interface according to the media information output instruction, the first terminal may determine the media information output instruction in the following manner:

[0109] When the live stream is a video live stream and the live stream content contains the object to be recommended, it is determined that the live stream content is associated with a media information output instruction, and the instruction associated with the object to be recommended is used as the media information output instruction.

[0110] In practice, during the live broadcast, the first terminal can perform semantic recognition on the broadcaster's live broadcast statements to determine whether the live broadcast statements contain objects to be recommended. Alternatively, it can use an image acquisition device (such as a camera) to capture the broadcaster's live broadcast images in real time, perform image recognition on the live broadcast images, and determine whether the live broadcast images contain objects to be recommended based on the recognition results, so as to determine whether the live broadcast content is associated with media information output instructions.

[0111] See Figure 10 , Figure 10 This is a schematic diagram illustrating the output of media recommendation information provided in this application embodiment. During the live broadcast of a game, the host takes out a cake and holds the cake 1001, meaning that the live broadcast content contains the "cake" as a recommended object. In practical applications, the live broadcast content containing the "cake" as a recommended object has been pre-set to be associated with media recommendation information for outputting the corresponding "cake advertisement". Therefore, it can be seen that the first terminal can determine the media information output instruction associated with the "cake advertisement" media recommendation information based on the live broadcast content, and use the instruction associated with the "cake" as a recommended object as the media information output instruction to present the media recommendation information 1002 related to the "cake advertisement" in the live broadcast interface. The size or position of the area of ​​the output interface used to output the media recommendation information 1002 can be adjusted.

[0112] In some embodiments, before outputting media recommendation information through the content interface according to the media information output instruction, the first terminal may determine the media information output instruction in the following manner:

[0113] When the live stream is a video live stream, and the anchor's gaze is focused on the target area in the content interface for a duration that reaches the duration threshold, it is determined that the live stream content is associated with a media information output instruction, and the instruction associated with the target area is used as the media information output instruction.

[0114] In practice, when the live stream is a video live stream, gaze detection is performed on the live image in the live stream content to obtain gaze detection results. When the gaze detection results indicate that the anchor's gaze focus is located in the target area of ​​the content interface and the duration reaches the duration threshold, it is determined that the live stream content is associated with media information output instructions, and the instructions associated with the target area are used as media information output instructions.

[0115] See Figure 11 , Figure 11 This is a schematic diagram of gaze detection provided in an embodiment of this application. The first terminal captures the live broadcast image of the anchor in real time through an image acquisition device (such as a camera) and performs gaze detection on the live broadcast image. When the detection result indicates that the anchor 1102's gaze focus 1104 is located in the target area 1103 area in the live broadcast content interface 1101 for a target duration (such as 5 seconds), it is determined that the live broadcast content is associated with a media information output instruction. If the instruction associated with the target area is an advertisement A playback instruction, then the media information output instruction is the advertisement A playback instruction.

[0116] In some embodiments, the first terminal may output media recommendation information through a content interface based on a media information output instruction in the following manner:

[0117] When the live stream content is also associated with media information location instructions, media recommendation information is output in the area indicated by the media information location instructions in the content interface, based on the media information output instructions and the media information location instructions.

[0118] Here, the media information location instruction is used to instruct the output of media recommendation information in the area indicated by the media information location instruction. The media information location instruction can also be triggered by the anchor content.

[0119] See Figure 12 , Figure 12 This is a schematic diagram of the output interface for media recommendation information provided in this application embodiment. During the live broadcast, the broadcaster's live broadcast content includes the target statement "play an advertisement" and the gesture "swipe left with your left hand". The instruction associated with the target statement "play an advertisement" is a media information output instruction, which is used to instruct the output of the media recommendation information "game advertisement". The instruction associated with the gesture "point to the lower left corner" is a media information location instruction, which is used to instruct the output of the media recommendation information in the lower left corner of the content interface. Then, the first terminal plays a "cake advertisement" in the lower left corner of the live broadcast content interface according to the media information output instruction and the media information location instruction.

[0120] In some embodiments, during the process of the first terminal outputting media recommendation information through a target area located in the content interface, the target area for outputting the media recommendation information can also be adjusted in the following ways;

[0121] In the process of outputting media recommendation information through the target area, when the live content contains target content that indicates adjustments to the target area, at least one of the following parameters of the target area is adjusted: area size and area location; and media recommendation information is output through the adjusted target area.

[0122] Here, the target content includes target statements or target actions. In actual implementation, the first terminal can perform semantic recognition on the live broadcast statements of the anchor to determine whether the live broadcast statements contain target statements. Alternatively, it can capture the live broadcast images of the anchor in real time through an image acquisition device (such as a camera), perform image recognition on the live broadcast images, and determine whether the live broadcast images contain target actions based on the recognition results. This will determine whether the live broadcast content is associated with an adjustment instruction to adjust the size or position of the target area. When the live broadcast content is associated with an adjustment instruction, the first terminal adjusts the size or position of the target area according to the adjustment instruction, so as to output media recommendation information through the adjusted target area.

[0123] See Figure 13A , Figure 13A This is a schematic diagram of the output interface adjustment for media recommendation information provided in this application embodiment. During the process of playing cake advertisement information through the target area 1302, when the anchor's live broadcast content contains the target content "play advertisement in the lower left corner" 1301, the position of the target area is adjusted to the lower left corner 1303 of the live broadcast content interface, that is, the "cake advertisement" is played at the lower left corner 1303 of the live broadcast content interface.

[0124] See Figure 13B , Figure 13B This is a schematic diagram of the output interface adjustment for media recommendation information provided in this application embodiment. During the process of playing video advertisements through the target area 1304, when the live broadcast content of the anchor contains the target content of "full-screen play advertisement", the position of the target area is adjusted to cover the content interface of the live broadcast, that is, the video advertisement is played in the area 1305 that covers the content interface of the live broadcast.

[0125] In some embodiments, the first terminal can output media recommendation information through a content interface in the following manner:

[0126] When there are at least two media recommendation messages, and the media information output instruction indicates that the media recommendation messages corresponding to the target recommendation object should be output, the media recommendation messages corresponding to the target recommendation object from the at least two media recommendation messages will be output through the content interface.

[0127] Here, when there are multiple media recommendation messages, they can be sorted based on their type, duration, or target audience (e.g., when the media recommendation message is an advertisement, the target audience is the celebrity endorsing the advertisement) to form a media recommendation message sequence. The first terminal responds to the media information output instruction and outputs the media recommendation messages corresponding to the target audience in the order of the media recommendation messages in the media recommendation message sequence. For example, the media recommendation messages in the media recommendation sequence include: music xx, cake A advertisement, cake B advertisement, car advertisement, game A promotion. If the media information output instruction indicates that the media recommendation messages corresponding to the target audience "cake" should be output, then cake advertisement A and cake B advertisement will be played in sequence.

[0128] In some embodiments, when there are at least two media recommendation messages and the media information output instruction indicates that the media recommendation message corresponding to the target recommendation object should be output, the first terminal, in response to the media information output instruction, can also simultaneously output each media recommendation message corresponding to the target recommendation object from the at least two media recommendation messages through the content interface. For example, assuming that the at least two media recommendation messages include: music xx, cake A advertisement, cake B advertisement, car advertisement, and game A promotion, if the media information output instruction indicates that the media recommendation message corresponding to the target recommendation object "cake" should be output, then the first terminal can simultaneously play cake advertisement A and cake B advertisement through the same content interface; in this way, users can see multiple media recommendation messages at the same time, meeting the user's need to quickly obtain information.

[0129] In some embodiments, when the number of media recommendation information is at least two, the terminal, in response to a media information output instruction, can also output a target number of media recommendation information from the media recommendation sequence in a forward-to-back order, wherein the target number can be set according to the actual situation. For example, the following media recommendation information sequence is obtained according to the type of media information: {media recommendation information 1, media recommendation information 2, ..., media recommendation information 8}. If the terminal is set to output two media recommendation information in response to a media information output instruction, then it can be seen that when a media information output instruction is associated with the live content for the first time, the terminal responds to the media information output instruction and outputs media recommendation information 1 and media recommendation information 2 in sequence. When a media information output instruction is associated with the live content for the second time, the terminal responds to the media information output instruction and outputs media recommendation information 3 and media recommendation information 4 in sequence, and so on. After all the media recommendation information in the media recommendation information sequence has been output, the output will start from the first media recommendation information and cycle through the output.

[0130] In some embodiments, the first terminal can output media recommendation information through a content interface in the following manner:

[0131] When the live stream is a video live stream and the media recommendation information is of video type, the media recommendation information is played in a floating window; during the playback of the media recommendation information, the video live stream is played in a silent mode, and the text corresponding to the audio content in the video live stream is displayed in the content interface.

[0132] Here, when both the live stream and the media recommendation information are video-based, simultaneous playback of both can cause audio interference. To avoid this, the live stream audio is converted to text, and while the media recommendation information is playing, the live video stream is played in a muted mode, while the corresponding text is displayed on the live stream content interface. This allows users to watch the live stream content while viewing the media recommendation information, improving information retrieval efficiency.

[0133] See Figure 14 , Figure 14 This is a schematic diagram of the interface for outputting media information provided in this application embodiment. The broadcaster's live stream is a video live stream and the media recommendation information 1401 is of type video. During the playback of the media recommendation information 1401, the broadcaster's video live stream is played in a silent playback mode, and at the same time, the text 1402 corresponding to the broadcaster's audio is displayed in the content interface of the live stream.

[0134] In some embodiments, the first terminal can output media recommendation information through a content interface in the following manner:

[0135] When the media recommendation information is text-based, it is presented as a bubble-like card overlay in the content interface; when the media recommendation information is audio or video-based, it is played through a sub-interface independent of the content interface.

[0136] The sub-interface can have a certain degree of transparency and is located above the live content interface. Through the sub-interface, users can view the live content presented in the live content interface. The sub-interface can occupy only a part of the live content interface or occupy the entire live content interface. In this way, presenting media recommendation information through a sub-interface with a certain degree of transparency can allow users to see more information and meet their needs for quick information acquisition. At the same time, the position of the sub-interface on the live content interface moves synchronously with the user's swipe operation.

[0137] In some embodiments, before the first terminal outputs media recommendation information through the content interface, it may also present prompt information in the content interface. The prompt information is used to prompt the output of media recommendation information. Accordingly, the first terminal can output media recommendation information through the content interface in the following way: based on the prompt information, when the live content contains target content for indicating the output of media recommendation information, the media recommendation information is output through the content interface.

[0138] Here, when the live broadcast content is associated with a media information output instruction, a prompt message is displayed to prompt the output of media recommendation information. The host can decide whether to further output media recommendation information based on the prompt message. The host can broadcast the corresponding live broadcast content according to the pre-set target content (target statement or target action). For example, when the host decides to output media recommendation information, the target statement "Confirm Play" is issued, or the target action of "both hands OK gesture" is made. When the host decides not to output media recommendation information, the target action of "crossed arms" is made.

[0139] See Figure 15 , Figure 15 This is a schematic diagram of the prompt information display interface provided in the embodiment of this application. When the live broadcast content is associated with a media information output instruction, the prompt information 1501 "Advertisement is about to play, are you sure you want to play?" is displayed in the live broadcast content interface. When the host makes the target action of making the "OK gesture with both hands", media recommendation information is output through the content interface.

[0140] In some embodiments, after the first terminal outputs media recommendation information through the content interface, it can also cancel the output of media recommendation information in the following ways:

[0141] When outputting media recommendation information through the content interface, if the live content is associated with a cancel output command, the output of media recommendation information will be canceled.

[0142] Here, during the process of outputting media recommendation information through the content interface, the broadcaster can trigger a cancellation command for the media information output via voice, i.e., a pre-set target statement to indicate the cancellation of the media recommendation information output. In actual implementation, during the broadcast, the first terminal collects the broadcaster's voice content in real time through an audio data acquisition device (such as a microphone) and performs speech recognition on the voice content to obtain the corresponding text content. Then, by judging whether the text content includes the pre-set target statement to indicate the cancellation of the media recommendation information output, it is determined whether a cancellation command for the media information output triggered by the live broadcast statement has been received. For example, the target statement to indicate the output of media recommendation information is set to "stop playing advertisements," etc. When the live broadcast content contains "stop playing advertisements," it is determined that the live broadcast content is associated with a cancellation command for media information, and the command associated with the target statement "stop playing advertisements" is used as the cancellation command for the media information output, thus canceling the output of the media recommendation information.

[0143] It should be noted that in practical applications, during the live broadcast, media recommendation information is simultaneously sent to a second terminal located on the viewer's side, so that the media recommendation information is played on the second terminal for the viewer to watch; during the process of outputting media recommendation information through the target area, when the first terminal adjusts the size or position of the target area used to output the media recommendation information, the size or position of the target area on the second terminal located on the viewer's side is also adjusted accordingly, and the second terminal outputs the media recommendation information for the viewer to watch through the adjusted target area.

[0144] Through the above method, this application identifies the live broadcast content. When the live broadcast content is associated with a media information output instruction, it outputs media recommendation information through the content interface according to the media information output instruction. At the same time, it can also adjust the size and position of the output area for outputting media recommendation information through the live broadcast content. In this way, by using live broadcast content associated with relevant instructions, the output of media recommendation information and the adjustment of the output area can be realized without the anchor manually operating the live broadcast interface or human assistance, thereby improving the efficiency and accuracy of obtaining media recommendation information.

[0145] The following will describe an exemplary application of the embodiments of this application in a real-world application scenario, and will continue to be combined with... Figure 1 ,by Figure 1 The following is an example of using a first terminal on the broadcaster's side to implement the live-streaming information processing method provided in this application embodiment, and a second terminal on the viewer's side to display the broadcaster's live-streaming content.

[0146] See Figure 16 , Figure 16 This is an optional flowchart illustrating a live-stream-based media information processing method provided in an embodiment of this application. The method includes a video live-stream client installed on a first terminal, and a backend server corresponding to the video live-stream client. Figure 16 The steps shown are explained.

[0147] Step 201: The first terminal sends a preset target statement associated with media information output instructions to the server.

[0148] Step 202: The server stores the target statement associated with the media information output instruction.

[0149] Step 203: The first terminal presents the live broadcast content interface and outputs the live broadcast content through the content interface.

[0150] In practical applications, the first terminal is equipped with a video live streaming client. The host can log in to the host end of the client to conduct live streaming and convey the content to the audience. The types of live streaming content are diverse, such as talent shows, live lectures, game live streaming, and sports event live streaming.

[0151] Step 204: The first terminal sends the live broadcast content to the server in real time.

[0152] Here, during the live broadcast, the first terminal captures the broadcaster's voice content in real time through an audio data acquisition device (such as a microphone) or captures live images through an image data acquisition device (such as a camera), and sends the captured voice content or live images to the server.

[0153] Step 205: The server performs speech recognition on the speech content to obtain the corresponding text content.

[0154] Step 206: The server matches the text content with the pre-set target statements used to indicate recommended media information for output.

[0155] Step 207: When a match is successful, the server sends the media recommendation information corresponding to the target statement to the first terminal.

[0156] Among them, media recommendation information is used to recommend at least one object to be recommended, such as advertisements or display images. That is, during the live broadcast, other media information can also be recommended, such as the streamer promoting game advertisements during game live broadcasts, or recommending relaxing music during breaks in live teaching.

[0157] It should be noted that in practical applications, the server simultaneously sends the media recommendation information corresponding to the target statement to the second terminal located on the viewer's side, so that the media recommendation information can be played on the second terminal for the viewer to watch.

[0158] Step 208: The first terminal plays media recommendation information through the playback window.

[0159] Step 209: The first terminal determines whether the gesture is the target gesture.

[0160] Here, during the playback of media recommendation information, the first terminal collects the broadcaster's gestures. The target gesture is a pre-stored gesture used to trigger adjustments to the size or position of the playback window. The collected broadcaster's gestures are matched with the target gestures. If the broadcaster's gesture matches the target gesture, step 210 is executed; otherwise, step 208 is executed.

[0161] Step 210: The first terminal adjusts the playback window and plays media recommendation information through the adjusted playback window.

[0162] In practical applications, the playback window on the second terminal located on the viewer's side is also adjusted accordingly. The second terminal plays media recommendation information for the viewer to watch through the adjusted playback window.

[0163] By using the above method, pre-configured advertisements or images and other media recommendation information can be displayed through the specific voice content and gestures of the anchor. Compared with the need for manual assistance to display advertisements or images and other media recommendation information, this can make the live broadcast content more seamless and reduce the cost of human intervention.

[0164] The following continues to describe the exemplary structure of the live-stream-based media information processing device 17 provided in this application embodiment as a software module. In some embodiments, see [link to relevant documentation]. Figure 17 , Figure 17 An optional structural diagram of the live-stream-based media information processing apparatus provided in this application embodiment is shown below. Figure 17 As shown, the media information processing apparatus 17 based on live streaming provided in this application embodiment includes:

[0165] The presentation module 171 is used to present the content interface of the live broadcast and output the live broadcast content through the content interface.

[0166] Output module 172 is used to output media recommendation information through the content interface according to the media information output instruction when the live content is associated with a media information output instruction;

[0167] The media recommendation information is used to recommend at least one object to be recommended.

[0168] In some embodiments, before outputting media recommendation information through the content interface according to the media information output instruction, the device further includes:

[0169] The instruction determination module is used to determine, when the live stream is a video live stream or an audio live stream, and the live stream content contains a target statement, that the live stream content is associated with a media information output instruction, and

[0170] The instruction associated with the target statement is used as the media information output instruction.

[0171] In some embodiments, the instruction determining module is further configured to...

[0172] When the live stream is a video or audio live stream, and the live stream content contains a target statement, and the number of times the target statement appears reaches a threshold, it is determined that the live stream content is associated with a media information output instruction, and

[0173] The instruction associated with the target statement is used as the media information output instruction.

[0174] In some embodiments, the instruction determining module is further configured to...

[0175] When the live stream is a video live stream, and the live stream content contains the broadcaster's target action, it is determined that the live stream content is associated with a media information output instruction, and

[0176] The instruction associated with the target action is used as the media information output instruction.

[0177] In some embodiments, the instruction determining module is further configured to...

[0178] When the live stream is a video live stream, and the live stream content contains the object to be recommended, it is determined that the live stream content is associated with a media information output instruction, and

[0179] The instruction associated with the object to be recommended is used as the media information output instruction.

[0180] In some embodiments, the instruction determining module is further configured to...

[0181] When the live stream is a video live stream, and the anchor's gaze is focused on a target area in the content interface for a duration that reaches a duration threshold, it is determined that the live stream content is associated with a media information output instruction, and

[0182] The instruction associated with the target region is used as the media information output instruction.

[0183] In some embodiments, the output module is further configured to output the media recommendation information in the area indicated by the media information location instruction in the content interface, based on the media information output instruction and the media information location instruction, when the live content is also associated with a media information location instruction.

[0184] In some embodiments, the output module is further configured to output the media recommendation information through a target area located in the content interface;

[0185] Accordingly, the device also includes:

[0186] The adjustment module is used to adjust at least one of the following parameters of the target area when the live content contains target content that indicates the adjustment of the target area during the process of outputting the media recommendation information through the target area: area size and area position.

[0187] The media recommendation information is output through the adjusted target area.

[0188] In some embodiments, the output module is further configured to output, through the content interface, each media recommendation information corresponding to the target recommendation object among the at least two media recommendation information when the number of media recommendation information is at least two and the media information output instruction indicates the output of media recommendation information corresponding to the target recommendation object.

[0189] In some embodiments, the output module is further configured to play the media recommendation information using a floating window when the live stream is a video live stream and the media recommendation information is of video type;

[0190] During the playback of the media recommendation information, the live video is played in a silent mode, and the text corresponding to the audio content in the live video is displayed on the content interface.

[0191] In some embodiments, the output module is further configured to present the media recommendation information in the content interface through a bubble-like card overlay when the type of the media recommendation information is text;

[0192] When the media recommendation information is of the type of audio or video, the media recommendation information is played through a sub-interface independent of the content interface.

[0193] In some embodiments, before outputting media recommendation information through the content interface, the device further includes:

[0194] The prompt module is used to present prompt information in the content interface, and the prompt information is used to prompt the output of media recommendation information;

[0195] Accordingly, the output module is also configured to output the media recommendation information through the content interface when the live content contains target content for indicating the output of media recommendation information, based on the prompt information.

[0196] This application provides a computer device, see [link to relevant documentation] Figure 18 , Figure 18 This is an optional structural diagram of the computer device 500 provided in an embodiment of this application. In practical applications, the computer device 500 can be... Figure 1 The terminal or server in the middle, with computer equipment as the main body Figure 1 Taking the terminal shown as an example, a computer device for implementing the live-stream-based media information processing method of this application embodiment will be described. The computer device includes:

[0197] Memory 550 is used to store executable instructions;

[0198] The processor 510 is used to execute executable instructions stored in the memory to implement the video playback method provided in the embodiments of this application.

[0199] Here, processor 510 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0200] The memory 550 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 550 may optionally include one or more storage devices physically located away from the processor 510.

[0201] The memory 550 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 550 described in this application embodiment is intended to include any suitable type of memory.

[0202] In some embodiments, at least one network interface 520 and a user interface 530 may also be included. Various components in the computer device 500 are coupled together via a bus system 540. It is understood that the bus system 540 is used to enable communication between these components. In addition to a data bus, the bus system 540 also includes a power bus, a control bus, and a status signal bus. However, for clarity, all buses are labeled as bus system 540 in Figure 13.

[0203] This application provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the live-stream-based media information processing method described above in this application.

[0204] This application provides a computer-readable storage medium storing executable instructions. When the executable instructions are executed by a processor, the processor will execute the live streaming-based media information processing method provided in this application.

[0205] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.

[0206] In some embodiments, executable instructions may take the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0207] As an example, executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple collaborating files (e.g., a file that stores one or more modules, subroutines, or code sections).

[0208] As an example, executable instructions can be deployed to execute on a single computing device, or on multiple computing devices located in one location, or on multiple computing devices distributed across multiple locations and interconnected via a communication network.

[0209] In summary, the embodiments of the present invention have the following beneficial effects:

[0210] 1. By recognizing the live content, when the live content is associated with a media information output instruction, the system can output media recommendation information through the live content associated with the media information output instruction. When the live content is associated with an instruction to adjust the output area of ​​the output media recommendation information, the system can adjust the output area by associating the live content with the adjustment instruction. The output of media recommendation information and the adjustment of the output area can be achieved without the anchor manually operating the live interface or human assistance, thereby improving the efficiency and accuracy of obtaining media recommendation information.

[0211] 2. By using specific voice content and gestures from the broadcaster, pre-configured advertisements or images and other media recommendation information can be displayed. Compared to requiring manual assistance to display advertisements or images and other media recommendation information, this can make the live broadcast content flow more smoothly and reduce the cost of human intervention.

[0212] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.

Claims

1. A media information processing method based on live streaming, characterized in that, The method includes: Presents a content interface for the streamer to broadcast live, and outputs the live broadcast content broadcast by the streamer through the content interface; When the live broadcast content contains a first target content associated with a media information output instruction, a prompt message is displayed on the content interface according to the media information output instruction. The prompt message is used to prompt the output of media recommendation information recommended by the broadcaster to the audience. Based on the prompt information, when the live broadcast content contains second target content for indicating the output of the media recommendation information, the media recommendation information recommended by the anchor to the audience is output through the content interface; wherein, the media recommendation information is output in the target area of ​​the content interface and is used to recommend at least one object to be recommended; During the process of outputting the media recommendation information through the target area, when the live content contains third target content that indicates the adjustment of the target area, at least one of the following parameters of the target area is adjusted: area size and area position, and the media recommendation information is output through the adjusted target area. The step of outputting media recommendation information recommended by the anchor to the audience through the content interface includes: When the live stream is a video live stream and the media recommendation information is of the video type, a floating window is used to play the media recommendation information recommended by the anchor to the audience; During the playback of the media recommendation information, the live video is played in a silent mode, and the text corresponding to the audio content in the live video is displayed on the content interface.

2. The method as described in claim 1, characterized in that, Before displaying the prompt information in the content interface according to the media information output instruction, the method further includes: When the live stream is a video or audio live stream, and the live stream content contains a target statement, it is determined that the live stream content is associated with a media information output instruction, and The instruction associated with the target statement is used as the media information output instruction.

3. The method as described in claim 1, characterized in that, Before displaying the prompt information in the content interface according to the media information output instruction, the method further includes: When the live stream is a video or audio live stream, and the live stream content contains a target statement, and the number of times the target statement appears reaches a threshold, it is determined that the live stream content is associated with a media information output instruction, and The instruction associated with the target statement is used as the media information output instruction.

4. The method as described in claim 1, characterized in that, Before displaying the prompt information in the content interface according to the media information output instruction, the method further includes: When the live stream is a video live stream, and the live stream content contains the broadcaster's target action, it is determined that the live stream content is associated with a media information output instruction, and The instruction associated with the target action is used as the media information output instruction.

5. The method as described in claim 1, characterized in that, Before displaying the prompt information in the content interface according to the media information output instruction, the method further includes: When the live stream is a video live stream, and the live stream content contains the object to be recommended, it is determined that the live stream content is associated with a media information output instruction, and The instruction associated with the object to be recommended is used as the media information output instruction.

6. The method as described in claim 1, characterized in that, Before displaying the prompt information in the content interface according to the media information output instruction, the method further includes: When the live stream is a video live stream, and the anchor's gaze is focused on a target area in the content interface for a duration that reaches a duration threshold, it is determined that the live stream content is associated with a media information output instruction, and The instruction associated with the target region is used as the media information output instruction.

7. The method as described in claim 1, characterized in that, The step of outputting media recommendation information recommended by the anchor to the audience through the content interface includes: When the live broadcast content is also associated with media information location instructions, the media recommendation information recommended by the anchor to the audience is output in the area indicated by the media information location instructions in the content interface, based on the media information output instructions and the media information location instructions.

8. The method as described in claim 1, characterized in that, The step of outputting media recommendation information recommended by the anchor to the audience through the content interface includes: When the number of media recommendation information recommended by the anchor to the audience is at least two, and the media information output instruction indicates that the media recommendation information corresponding to the target recommendation object should be output, the media recommendation information corresponding to the target recommendation object among the at least two media recommendation information is output through the content interface.

9. The method as described in claim 1, characterized in that, The step of outputting media recommendation information recommended by the anchor to the audience through the content interface includes: When the media recommendation information recommended by the anchor to the audience is text, the media recommendation information is presented in the content interface through a bubble-style card overlay; When the media recommendation information is of the type of audio or video, the media recommendation information is played through a sub-interface independent of the content interface.

10. A media information processing device based on live streaming, characterized in that, The device includes: The presentation module is used to present the content interface for the broadcaster to live stream, and to output the live streaming content of the broadcaster through the content interface; The output module is used to display prompt information in the content interface according to the media information output instruction when the live content contains a first target content associated with a media information output instruction. The prompt information is used to prompt the output of media recommendation information recommended by the anchor to the audience. Based on the prompt information, when the live broadcast content contains second target content for indicating the output of the media recommendation information, the media recommendation information recommended by the anchor to the audience is output through the content interface; wherein, the media recommendation information is output in the target area of ​​the content interface and is used to recommend at least one object to be recommended; An adjustment module is used to adjust at least one of the following parameters of the target area when the live content contains third target content that indicates the adjustment of the target area during the process of outputting the media recommendation information through the target area: area size and area position. The output module is also used to output the media recommendation information through the adjusted target area; The output module is further configured to, when the live stream is a video live stream and the media recommendation information is of video type, use a floating window to play the media recommendation information recommended by the anchor to the audience; during the playback of the media recommendation information, play the video live stream in a silent mode and display the text corresponding to the audio content in the video live stream in the content interface.

11. The apparatus as claimed in claim 10, characterized in that, Before displaying a prompt message in the content interface according to the media information output instruction, the device further includes: The instruction determination module is used to determine that the live broadcast content is associated with a media information output instruction when the live broadcast is a video live broadcast, the anchor's gaze focus in the live broadcast content is located in the target area of ​​the content interface, and the duration reaches a duration threshold, and to use the instruction associated with the target area as the media information output instruction.

12. The apparatus as claimed in claim 10, characterized in that, The output module is further configured to, when the live content is associated with a media information location instruction, output the media recommendation information recommended by the anchor to the audience in the area indicated by the media information location instruction in the content interface, based on the media information output instruction and the media information location instruction.

13. A computer device, characterized in that, include: Memory, used to store executable instructions; A processor, when executing executable instructions stored in the memory, implements the live-stream-based media information processing method according to any one of claims 1 to 9.

14. A computer-readable storage medium, characterized in that, It stores executable instructions for use by a processor to implement the live-stream-based media information processing method as described in any one of claims 1 to 9.

15. A computer program product, characterized in that, The method includes computer instructions, characterized in that, when the computer instructions are executed by a processor, they implement the live-stream-based media information processing method according to any one of claims 1 to 9.