Live broadcasting room recommendation method, device and system and computer equipment

By extracting semantic features from user search information and multimodal information in the live stream, and calculating similarity and matching degree, the problem of mismatch between live stream recommendations and user interests is solved, achieving accurate recommendations and relevance to cover titles, and adapting to changes in user intent.

CN121012944APending Publication Date: 2025-11-25SHANGHAI BILIBILI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511149611.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-15
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

The existing live streaming recommendation mechanism lacks a deep understanding of users' natural language intent, resulting in a low degree of matching between recommended live streaming rooms and users' interests.

Method used

By acquiring user search information and multimodal live streaming information of candidate live streaming rooms, semantic features are extracted, semantic similarity and matching degree are calculated, and target live streaming rooms are selected for recommendation, including text content processing and keyword matching based on audio, video and interactive information.

Benefits of technology

It improves the accuracy of live stream recommendations, ensuring that recommended live streams are more closely matched with user interests, enhances the relevance of cover images and titles, and adjusts recommendations in real time to address semantic drift.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121012944A_ABST
    Figure CN121012944A_ABST
Patent Text Reader

Abstract

The invention provides a live broadcast room recommendation method, device and system and computer equipment, and relates to the technical field of Internet live broadcast. The live broadcast room recommendation method disclosed by the invention comprises the following steps: acquiring search information; acquiring live broadcast information of the candidate live broadcast room, wherein the live broadcast information comprises at least one of live broadcast audio information, live broadcast picture information and live broadcast interaction information; determining a first semantic feature based on the search information, wherein the first semantic feature is used for representing the overall semantics of the search information; determining a second semantic feature based on the live broadcast information, wherein the second semantic feature is used for representing the overall semantics of the live broadcast information; determining the semantic similarity between the first semantic feature and the second semantic feature; determining a matching degree between the search information and each candidate live broadcast room based on the semantic similarity; and selecting a recommended target live broadcast room from the candidate live broadcast rooms based on the matching degree.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present specification relates to the technical field of Internet live broadcast, and particularly relates to a live broadcast room recommendation method, device, system, computer device, computer readable storage medium and computer program product. BACKGROUND

[0002] With the development of the Internet, video live broadcast platforms gradually become the main scene for users to obtain content and interact. Users can search for live broadcast rooms of interest on a video live broadcast platform to watch live broadcast content. At present, the recommendation mechanism of the live broadcast room is to recommend the live broadcast room to the user according to the label, title and the like set by the host, or to recommend the live broadcast room based on the unified hot list maintained by the video live broadcast platform, which lacks in-depth understanding of the natural language intention of the user, resulting in low matching degree of the recommended live broadcast room and the user interest.

[0003] Therefore, some embodiments of the present specification provide a live broadcast room recommendation method, device, system, computer device, computer readable storage medium and computer program product, aiming to improve the matching degree of the recommended live broadcast room and the user interest, and to realize accurate recommendation of the live broadcast room. SUMMARY

[0004] One or more embodiments of the present specification provide a live broadcast room recommendation method, which comprises: obtaining search information; obtaining live broadcast information of a candidate live broadcast room, the live broadcast information comprising at least one of live broadcast audio information, live broadcast picture information and live broadcast interaction information; determining a first semantic feature based on the search information, the first semantic feature being used to represent the overall semantics of the search information; determining a second semantic feature based on the live broadcast information, the second semantic feature being used to represent the overall semantics of the live broadcast information; determining a semantic similarity between the first semantic feature and the second semantic feature; determining a matching degree of the search information and each candidate live broadcast room based on the semantic similarity; and selecting a recommended target live broadcast room in the candidate live broadcast room based on the matching degree.

[0005] According to the method provided by one or more embodiments of the present specification, the second semantic feature is determined based on the live broadcast information, which comprises: extracting live broadcast related text content based on the live broadcast information to obtain live broadcast text information; generating a content abstract based on the live broadcast text information; and generating the second semantic feature based on the content abstract.

[0006] According to the method provided by one or more embodiments of the present specification, the live broadcast text information comprises one or more of audio text information, picture text information or interactive keywords; the live broadcast related text content is extracted based on the live broadcast information to obtain the live broadcast text information, which comprises at least one of the following processing: converting the live broadcast audio information into a text form to obtain the audio text information; extracting text content in the live broadcast picture information and classifying to obtain the picture text information; and extracting keywords based on the live broadcast interaction information to obtain the interactive keywords.

[0007] According to one or more embodiments of this specification, the method extracts and classifies text content from live stream information to obtain screen text information, including: extracting text content and corresponding location information from live stream information; performing clustering and classification processing based on the text content and corresponding location information from live stream information to determine the classification labels of the text content, and generating screen text information.

[0008] The method provided in one or more embodiments of this specification generates a content summary based on live text information, including: extracting live room topic tags based on live text information; and generating a content summary based on the live text information and the live room topic tags.

[0009] According to one or more embodiments of this specification, the method for determining the matching degree between search information and each candidate live stream based on semantic similarity includes: performing keyword matching based on search information and live stream information, and determining keyword coverage based on the keyword matching results; and determining the matching degree based on keyword coverage and semantic similarity.

[0010] According to one or more embodiments of this specification, the method for matching keywords based on search information and live stream information, and determining keyword coverage based on the keyword matching results, includes: extracting search keywords based on search information; matching the search keywords with live stream topic tags to determine the keywords that are successfully matched in the search keywords; and determining the keyword coverage based on the proportion of successfully matched keywords in the search keywords; wherein, the live stream topic tags are determined based on the live stream information.

[0011] According to one or more embodiments of this specification, the method for selecting a recommended target live stream room from candidate live stream rooms based on matching degree includes: calculating the live stream room popularity of each candidate live stream room based on live stream behavior data, wherein the live stream behavior data includes at least one of the following: number of online viewers, frequency of bullet comments, frequency of comments, frequency of likes, frequency of sending gifts, frequency of giving coins, and the speaking ratio related to live stream room topic tags; selecting a recommended target live stream room from candidate live stream rooms based on live stream room popularity and matching degree; wherein, the live stream room topic tags are determined based on live stream information.

[0012] According to one or more embodiments of this specification, the method for selecting a recommended target live stream from candidate live streams based on live stream popularity and matching degree includes: determining the comprehensive score of each candidate live stream based on the weight of live stream popularity and matching degree; sorting each candidate live stream based on the comprehensive score; and selecting the recommended target live stream to be displayed based on the sorting result.

[0013] The method provided according to one or more embodiments of this specification further includes: when the recommended target live room meets at least one of the following conditions, lowering the ranking of the recommended target live room that meets the conditions or canceling the display of the recommended target live room that meets the conditions: the matching degree between the recommended target live room and the search information drops below a first preset threshold; the number of times the matching degree between the recommended target live room and the search information drops within a preset time is greater than or equal to a second preset threshold; the magnitude of the drop in the matching degree between the recommended target live room and the search information within a preset time is greater than or equal to a third preset threshold.

[0014] The method provided according to one or more embodiments of this specification further includes: generating a cover image of a recommended target live stream room based on search information; generating a cover image of a recommended target live stream room based on search information includes: determining a screenshot of the live stream of the recommended target live stream room based on the semantic similarity between the search information and the image text information, and using the screenshot as the cover image of the recommended target live stream room, wherein the image text information is determined based on the live stream image information; or determining a screenshot of the live stream of the recommended target live stream room based on the image-text similarity between the live stream image information and the search information, and using the screenshot as the cover image of the recommended target live stream room; or extracting emotion tags based on the live stream image information, matching the emotion tags with search keywords, determining a screenshot of the live stream of the recommended target live stream room based on the matching results, and using the screenshot as the cover image of the recommended target live stream room, wherein the search keywords are determined based on the search information.

[0015] The method provided according to one or more embodiments of this specification further includes: generating a cover title for a recommended target live stream based on live stream information; generating a cover title for a recommended target live stream based on live stream information includes: extracting a summary sentence based on audio text information and live stream interaction information, wherein the summary sentence is used to characterize the core topics of the audio text information and live stream interaction information, and the audio text information is determined based on live stream audio information; generating a cover title for a recommended target live stream based on the summary sentence.

[0016] One or more embodiments of this specification also provide another method for recommending live streaming rooms, the method comprising: acquiring search information; displaying recommended target live streaming rooms that match the search information; wherein the recommended target live streaming rooms are determined based on the matching degree between the search information and the live streaming information of the candidate live streaming rooms, the matching degree including the semantic similarity between the search information and the live streaming information, the semantic similarity being determined based on a first semantic feature and a second semantic feature, the first semantic feature being determined based on the search information, the second semantic feature being determined based on the live streaming information, the first semantic feature being used to characterize the overall semantics of the search information, the second semantic feature being used to characterize the overall semantics of the live streaming information, and the live streaming information including at least one of live audio information, live video information, and live interactive information.

[0017] The method provided according to one or more embodiments of this specification further includes: displaying at least one of the following controls on the display interface of the recommended target live stream: a share control, a subscription control, and a persistent page control; sharing the ranking list of the recommended target live stream in response to a trigger operation of the share control; adding the ranking list of the recommended target live stream to favorites in response to a trigger operation of the subscription control, and generating a reminder message when the ranking list of the recommended target live stream changes; and setting the display interface of the recommended target live stream as a persistent page of the target page in response to a trigger operation of the persistent page control; wherein the ranking list of the recommended target live stream is determined based on the matching degree between the search information and the live stream information of the candidate live streams.

[0018] This specification also provides a live streaming room recommendation device in one or more embodiments. The device includes: a first acquisition module for acquiring search information; a second acquisition module for acquiring live streaming information of candidate live streaming rooms, the live streaming information including at least one of live audio information, live video information, and live interactive information; a first determination module for determining a first semantic feature based on the search information, the first semantic feature being used to characterize the overall semantics of the search information; a second determination module for determining a second semantic feature based on the live streaming information, the second semantic feature being used to characterize the overall semantics of the live streaming information; a third determination module for determining the semantic similarity between the first semantic feature and the second semantic feature; a fourth determination module for determining the matching degree between the search information and each candidate live streaming room based on the semantic similarity; and a recommendation module for selecting a target live streaming room from the candidate live streaming rooms based on the matching degree.

[0019] This specification also provides a live streaming recommendation system in one or more embodiments. The system includes a server and a viewer. The viewer is used to acquire search information and display recommended target live streaming rooms that match the search information. The server is used to acquire search information sent from the viewer and live streaming information of candidate live streaming rooms, determine a first semantic feature based on the search information, determine a second semantic feature based on the live streaming information, determine the semantic similarity between the first semantic feature and the second semantic feature, determine the matching degree between the search information and each candidate live streaming room based on the semantic similarity, select recommended target live streaming rooms from the candidate live streaming rooms based on the matching degree, and send them to the viewer. The first semantic feature is used to characterize the overall semantics of the search information, and the second semantic feature is used to characterize the overall semantics of the live streaming information. The live streaming information includes at least one of live audio information, live video information, and live interactive information.

[0020] One or more embodiments of this specification also provide a computer device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it is able to implement the live streaming room recommendation method described in some embodiments of this specification.

[0021] This specification also provides a computer-readable storage medium in one or more embodiments, which stores computer instructions that, when executed by a processor, can implement the live streaming room recommendation method described in some embodiments of this specification.

[0022] One or more embodiments of this specification also provide a computer program product, including a computer program that, when at least a portion of the computer program is executed by a processor, can implement the live streaming room recommendation method described in some embodiments of this specification.

[0023] The beneficial effects that the embodiments of this specification may bring include, but are not limited to: determining semantic similarity based on search information and live streaming information, determining the matching degree between search information and each candidate live streaming room based on semantic similarity, and selecting recommended target live streaming rooms from the candidate live streaming rooms based on the matching degree, thereby improving the matching degree between recommended live streaming rooms and user interests and achieving accurate live streaming room recommendations; live streaming information includes live streaming audio information, live streaming video information, or live streaming interactive information, thus forming multimodal live streaming information, and achieving multimodal semantic matching based on multimodal content modeling, further improving the accuracy of matching; determining the matching degree of each candidate live streaming room based on the keyword coverage rate of search information relative to live streaming information, further improving the accuracy of the matching degree between search information and each candidate live streaming room; selecting recommended target live streaming rooms from the candidate live streaming rooms based on live streaming room popularity and matching degree. This system recommends relevant and popular live streams to users. By generating cover images for recommended live streams based on search information, it can use live stream clips that are semantically strongly related to the user's input as cover images, ensuring that the displayed cover images of recommended live streams accurately reflect the topics the user is interested in, making the cover images more aligned with the user's search intent. By generating cover titles for recommended live streams based on live stream information, it can make the cover titles of recommended live streams more closely match the live stream content, enhancing the consistency between the cover titles and the live stream content. By setting preset conditions and lowering the ranking of recommended live streams that meet the conditions or canceling the display of recommended live streams that meet the conditions, semantic drift detection is achieved. When the live stream content deviates from the user's input intent, the recommendation of that live stream can be reduced, thus recommending live streams that match the user's search intent in real time. It should be noted that different embodiments may produce different beneficial effects. In different embodiments, the beneficial effects that may occur can be any one or a combination of the above, or any other possible beneficial effects. Attached Figure Description

[0024] This specification will be further described by way of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. The same numbers in the drawings denote the same structures or steps.

[0025] Figure 1 This is a schematic diagram of the operating environment of a video live streaming platform according to some embodiments of this specification.

[0026] Figure 2 This is an exemplary flowchart of a live streaming room recommendation method according to some embodiments of this specification.

[0027] Figure 3 This is an exemplary flowchart illustrating a method for determining the matching degree of each candidate live streaming room according to some embodiments of this specification.

[0028] Figure 4 This is an exemplary flowchart illustrating a method for selecting a recommended target live streaming room according to some embodiments of this specification.

[0029] Figure 5 This is an exemplary flowchart of a live streaming room cover generation method according to some embodiments of this specification.

[0030] Figure 6 This is an exemplary flowchart illustrating a method for generating a live streaming room cover title according to some embodiments of this specification.

[0031] Figure 7 This is an exemplary flowchart of another live streaming room recommendation method according to some embodiments of this specification.

[0032] Figure 8 This is a schematic diagram of a recommended target live streaming room display interface according to some embodiments of this specification.

[0033] Figure 9 This is an exemplary block diagram of a live streaming room recommendation device according to some embodiments of this specification.

[0034] Figure 10 This is an exemplary block diagram of another live streaming room recommendation device according to some embodiments of this specification.

[0035] Figure 11 This is an exemplary block diagram of a live streaming room recommendation system according to some embodiments of this specification. Detailed Implementation

[0036] To more clearly illustrate the technical solutions of the embodiments in this specification, the embodiments will be described in detail below with reference to the accompanying drawings. Obviously, the content described below are some examples or embodiments of this specification. For those skilled in the art, without creative effort, the technical solutions or means disclosed in this specification can be applied to other scenarios based on this technical content.

[0037] It should be understood that the terms "system," "device," "unit," and / or "module" used in this specification are a method of distinguishing different components, elements, parts, sections, or assemblies at different levels. However, if other terms can achieve the same purpose, they may be replaced by other expressions.

[0038] Unless otherwise specified, the technical terms used to describe components, elements, etc. in this specification are not singular but may include plural. Generally speaking, terms such as "comprising" or "including" only indicate that explicitly identified steps, elements, or components are included, and these steps, elements, and components do not constitute an exclusive list, as the described method or apparatus may also include other steps or components.

[0039] This specification uses flowcharts to illustrate the operational steps performed by the apparatus or system of related embodiments. However, unless otherwise specified, the order in which these steps are described should not be construed as a limitation on the order of execution. Those skilled in the art can adjust the order of these steps based on the knowledge and information conveyed by the embodiments in this specification. Such adjustments include, but are not limited to, reversing the order of steps, merging multiple steps, and splitting a step.

[0040] Live streaming is a real-time audio and video transmission technology based on the internet. Broadcasters record their live activities (such as performances, explanations, games, and event footage) using devices like cameras and microphones, encode them through a live streaming platform, and then transmit them to users in different regions. Users can also search for live streams of interest on the platform. For example, users can enter search information in the platform's search box, and the platform will recommend relevant live streams based on the user's search.

[0041] Figure 1 This is a schematic diagram illustrating the operating environment of a video live streaming platform according to some embodiments of this specification. For example... Figure 1As shown, the video live streaming platform operating environment 100 may include: a server 110, a viewer 120, a broadcaster 130, and a network 140. The server 110, viewer 120, and broadcaster 130 can transmit data via the network 140. Users can input search information on the viewer 120, which then sends the search information to the server 110 via the network 140. The server 110 can obtain live streaming information from the broadcaster 130 via the network 140, determine recommended target live streaming rooms based on the search information and the live streaming information, and return the recommendations to the viewer 120 for viewers to choose from.

[0042] The server 110 can be a high-performance computer device used to store live stream data and receive and parse user search information. For example, the server 110 can parse the search information sent by the viewer 120, determine recommended target live streams matching the search information based on the live stream information, and return them to the viewer 120. In some embodiments, the server 110 can transmit data based on a streaming media transmission protocol, which may include HLS (HTTP Live Streaming), RTMP (Real-Time Messaging Protocol), DASH (Dynamic Adaptive Streaming over HTTP), etc. In some embodiments, the server 110 may include a local server or a cloud server. Depending on different service requirements, a local server corresponding to that region may be deployed in one or more regions. In some embodiments, the server 110 may include a backend processing server, which can analyze and process the live stream information sent by the broadcaster 130 and the search information sent by the viewer 120, determine recommended target live streams matching the search information, and return them to the viewer 120. In some embodiments, server 110 may include a streaming media server, a transcoding server, a storage server, an authentication server, etc. The streaming media server can transmit data based on a streaming media transmission protocol. The transcoding server can be used to transcode live video data, the storage server can be used to cache and store live data, and the authentication server can be used to verify user access permissions. In some embodiments, server 110 may be a single computer device or a computing cluster composed of multiple computer devices, thereby providing more powerful computing power and more efficient response to user service requests.

[0043] The broadcast client 130 can be used to generate live video data in real time and push the live video data. The live video data may include, for example, a video stream generated by capturing live footage from a camera device, or a video stream generated by the broadcaster sharing their screen, application screen, or document screen. In some embodiments, the broadcast client 130 may be an electronic device such as a desktop computer, smartphone, laptop, or tablet computer.

[0044] Viewer terminal 120 can be used to acquire search information input by the user and display live streams matching the search information to the user. In some embodiments, viewer terminal 120 may include, but is not limited to, terminal devices such as desktop computers, smartphones, laptops, VR (Virtual Reality) devices, tablets, smart TVs, and in-vehicle terminals. Viewer terminal 120 may include a display screen and a processor, the display screen being used to present a graphical user interface. For example, viewer terminal 120 can present recommended target live streams sent by server 110 through a graphical user interface. In some embodiments, the display screen may be separate from the human-machine interface device, allowing the user to operate on the graphical user interface through the human-machine interface device. The processor of viewer terminal 120 can receive operation instructions generated by operating on the graphical user interface through the human-machine interface device, and the display screen can be used to present the graphical user interface, for example, presenting response results generated based on operation instructions input through the graphical user interface. In other embodiments, the display screen may be a touch screen, which can receive operation instructions input by the user based on the graphical user interface. For example, the operation instructions generated by the user on the graphical user interface may include the instruction to input search information and search for live rooms in the graphical user interface. The processor of the viewer terminal 120 may be configured to transmit the instruction to the server terminal 110 for processing after receiving the instruction to search for live rooms. The server terminal 110 will send the live rooms that match the search information to the viewer terminal 120 for display.

[0045] Network 140 can be any form of wired or wireless network, or any combination thereof. As an example, network 140 can be one or more combinations of wired networks, fiber optic networks, telecommunications networks, internal networks, the Internet, local area networks (LANs), wide area networks (WANs), wireless local area networks (WLANs), metropolitan area networks (MANs), public switched telephone networks (PSTNs), Bluetooth networks, etc. Network 140 can have multiple access points, through which server 110, viewer 120, and broadcaster 130 can access network 140.

[0046] It should be noted that, Figure 1The schematic diagram of the video live streaming platform operating environment shown is merely an example. The video live streaming platform operating environment described in the embodiments of this specification is intended to more clearly illustrate the technical solutions of the embodiments of this specification and does not constitute a limitation on the technical solutions provided in the embodiments of this specification. For example, Figure 1 The number of server-side 110, viewer-side 120, and broadcaster-side 130 in this specification is merely illustrative and is not intended to limit the scope of patent protection of this application. Depending on the actual situation, any number of server-side 110, viewer-side 120, and broadcaster-side 130 may be included. Those skilled in the art will understand that with the development of live streaming technology and the emergence of new business scenarios, the technical solutions provided in the embodiments of this specification are also applicable to similar technical problems.

[0047] In some related embodiments, the recommendation mechanism for live streaming rooms recommends live streaming rooms to users based on tags, titles, etc. set by the broadcaster, or recommends live streaming rooms based on a unified popularity ranking maintained by the video live streaming platform. However, it lacks a deep understanding of the user's natural language intent. When the user wants to express their interests through free input, the system has difficulty understanding their intent, resulting in a mismatch between the recommended live streaming rooms and the user's interests.

[0048] To this end, some embodiments of this specification propose a live streaming room recommendation method. The method includes obtaining search information and live streaming information of candidate live streaming rooms, determining semantic similarity based on search information and live streaming information, determining the matching degree between search information and each candidate live streaming room based on semantic similarity, and selecting a target live streaming room to recommend based on the matching degree, thereby improving the matching degree between the target live streaming room to recommend and the user's interests, and achieving accurate recommendation of live streaming rooms to the user.

[0049] Figure 2 This is an exemplary flowchart of a live streaming room recommendation method according to some embodiments of this specification. Figure 2 The process 200 shown can be executed by a processing device, for example, by... Figure 1 The server 110 shown executes the process. In some embodiments, process 200 may be implemented by a live streaming recommendation device 900 deployed on a processing device. Figure 2 As shown, in some embodiments, process 200 may include the following steps.

[0050] Step 210: Obtain search information. In some embodiments, step 210 may be implemented by the first acquisition module 910.

[0051] In some embodiments, server 110 can obtain search information input by the user from viewer 120. For example, a user can install and run video live streaming software on viewer 120 and input search information in the video live streaming software. Viewer 120 can send the search information to server 110, and server 110 can then obtain the search information.

[0052] In some embodiments, the search information entered by the user can be natural language, such as "live stream about AI layoffs". In some embodiments, the search information can be text, audio, or image information. When the search information is audio, the server 110 or the viewer 120 can perform ASR (Automatic Speech Recognition) processing on the audio information to convert it into text. For example, an RNN-T (Recurrent Neural Network Transducer) speech recognition model can be used to convert audio information into text. When the search information is image information, the server 110 or the viewer 120 can perform OCR (Optical Character Recognition) processing on the image information to convert it into text. For example, an OCR model can be used to convert image information into text. The OCR model can automatically extract visual features from the image and perform character recognition. For example, the OCR model can be an optical character recognition model based on the Transformer architecture or an end-to-end optical character recognition framework. The server (110) or the viewer (120) can also call a multimodal large model to understand image semantics and convert image information into text information. For example, a multimodal large model could be a general-purpose large language model with visual understanding capabilities.

[0053] Step 220: Obtain the live streaming information of the candidate live streaming rooms. In some embodiments, step 220 can be implemented by the second acquisition module 920.

[0054] In some embodiments, candidate live stream rooms can be all live stream rooms currently broadcasting. For example, when a streamer starts a live stream, the server 110 can identify that live stream room as a candidate live stream room. The server 110 can also filter live stream rooms currently broadcasting to identify candidate live stream rooms. In some embodiments, the server 110 can filter live stream rooms currently broadcasting based on their activity level within a first preset time period to identify candidate live stream rooms. For example, the first preset time period can be the N minutes prior to obtaining the user's input search information. The activity level of a live stream room can be determined based on the number of viewers watching, the streamer's activity level, and audience (or user) interaction. For example, if the average number of viewers watching a live stream room is more than ten within the first preset time period, and viewers are interacting by sending comments, gifts, or likes, then that live stream room can be considered a candidate live stream room; if the streamer does not speak and the video does not change within the first preset time period, that live stream room can be filtered out, thus excluding live stream rooms that are broadcasting but have no content output, reducing the computational load on the server. In other embodiments, candidate live stream rooms can also be identified based on historical data of the live stream rooms. For example, if a live stream room has been broadcast more than twice in the past week, or if the average duration of a single live stream session in the past week is more than 20 minutes, then the live stream room can be considered as a candidate live stream room when it starts broadcasting.

[0055] In some embodiments, live streaming information may include at least one of live audio information, live video information, or live interactive information. Live audio information may be audio signals in the live stream, such as human voices, ambient sounds, and background music in the live environment. Live video information may be visual signals in the live stream, such as images captured by the camera device of the broadcaster 130, or screen images, application images, or document images shared by the broadcaster 130. Live interactive information may be information generated by viewers' interactive operations in the live stream, such as bullet comments, comments sent by viewers in the live stream, and chat information between viewers and the broadcaster in the chat room.

[0056] In some embodiments, live streaming information of candidate live streaming rooms during a second preset time period can be obtained. The second preset time period can be the period from the start of the live stream to the current time point, or it can be the period M minutes before the current time point; there is no limitation on this.

[0057] In some embodiments, after obtaining search information and live streaming information, the matching degree of each candidate live streaming room can be further determined based on the search information and live streaming information. In some embodiments, the matching degree can include the semantic similarity between the search information and the live streaming information. Here, semantics can refer to the meaning or significance of language. Semantic similarity can be the degree of closeness between two language units (e.g., phrases, sentences, short phrases, documents, etc.) at the semantic level. It is not limited to the matching of literal meanings, but rather measures the similarity between two language units by understanding the intent behind natural language and combining contextual logic. The following describes in detail how to determine the matching degree of each candidate live streaming room.

[0058] Step 230: Determine the first semantic feature based on the search information. In some embodiments, step 230 may be implemented by the first determining module 930.

[0059] In some embodiments, the first semantic feature can be used to characterize the overall semantics of the search information. For example, the first semantic feature can be a feature vector, which encodes the search information into vectors in a semantic space using a language model. The language model can capture the contextual semantic information of the input text and map it into semantically rich vectors. For instance, the language model can be a model for generating sentence-level semantic embeddings.

[0060] In some embodiments, search information (when the search information is text) can be input into the language model, or search information (when the search information is audio or image) can be converted into text and then input into the language model. The language model segments the input text information into smaller units (or tokens), maps each unit to a word vector, and then aggregates the word vectors into a feature vector, which is then output as the first semantic feature. For example, if a user inputs the text information "live broadcast about AI layoffs", the server 110 inputs the text information "live broadcast about AI layoffs" into the language model, the language model converts the text information into a feature vector [0.52, 0.31, ...] and outputs it, which can be used as the first semantic feature.

[0061] Step 240: Determine the second semantic feature based on the live broadcast information. In some embodiments, step 240 may be implemented by the second determining module 940.

[0062] In some embodiments, the second semantic feature can be used to characterize the overall semantics of the live broadcast information. For example, the second semantic feature can be a feature vector. In some embodiments, live broadcast-related text content can be extracted first based on the live broadcast information to obtain live broadcast text information, then a content summary can be generated based on the live broadcast text information, and the second semantic feature can be generated based on the content summary. The following describes in detail how the second semantic feature is generated.

[0063] First, let's explain how to extract live-related text content based on live-stream information to obtain live-stream text information. This live-stream information can include at least one of the following: live-stream audio information, live-stream video information, or live-stream interactive information. We can extract audio text information from live-stream audio information, video text information from live-stream video information, and interactive keywords from live-stream interactive information.

[0064] In some embodiments, live audio information can be converted into text to obtain audio-text information. For example, the live audio information can be preprocessed, such as by filtering background noise and non-human voices using an acoustic model. Then, ASR processing is performed on the preprocessed live audio information to convert it into text. The text-based live audio information is then filtered and refined to obtain audio-text information. In some embodiments, ASR processing can be implemented using a speech recognition model, such as an RNN-T.

[0065] In some embodiments, filtering live audio information in text form may include filtering invalid information from the live audio information in text form to obtain intermediate text information. Invalid information may include interjections (e.g., "uh," "um," etc.), filler words (e.g., "then," "that is," etc.), invalid repeated words (e.g., "we" in "we, we are going to talk about..."), and semantically ambiguous chatter or spoken segments (e.g., "everyone, give me a 666," "welcome new viewers"). For example, interjections, filler words, and invalid repeated words in the live audio information in text form can be filtered using regular expressions; invalid repeated words or semantically ambiguous chatter or spoken segments in the live audio information in text form can be filtered using a pre-trained language model. The pre-trained language model may, for example, be a bidirectional encoding pre-trained model based on the Transformer architecture.

[0066] In some embodiments, refining live audio information in text form may include acquiring filtered intermediate text information and extracting effective information from the intermediate text information to obtain audio text information. Effective information may include sentences with high semantic density, technical terms, entity nouns (e.g., "**AI", "large model", "algorithm adjustment", etc.), and language with structural features (e.g., "Today we'll discuss three points", "The reasons are explained below"). Sentences with high semantic density can be statements that carry a rich amount of information within a finite length. For example, named entities in the intermediate text information can be identified using NER (Named Entity Recognition), determining the number of named entities in each sentence, and identifying sentences with a number of named entities greater than a first preset threshold as sentences with high semantic density. Alternatively, the semantic density of a sentence can be determined based on the number of named entities and the sentence length, and sentences with a semantic density greater than a second preset threshold can be identified as sentences with high semantic density. For example, technical terms and entity nouns in the intermediate text information can be identified using a named entity recognition model or a graph-based ranking algorithm; language with structural features in the intermediate text information can be identified using a pre-trained language model (e.g., a model based on the Transformer architecture).

[0067] In some embodiments, text content can be extracted from the live stream footage and categorized to obtain the video text information. For example, text content and corresponding location information can be extracted from the live stream footage, and clustering and classification processes can be performed based on the text content and corresponding location information to determine the classification labels of the text content and generate the video text information.

[0068] In some embodiments, OCR processing can be performed on the live stream information. For example, an OCR model (e.g., an optical character recognition model based on the Transformer architecture) can be used to extract the text content and corresponding location information from the live stream information, obtaining multiple text fragments, the corresponding screen location of each text fragment, and timestamps. Further, each text fragment can be clustered and classified based on its font size and corresponding screen location to determine its classification label. Classification labels can include, for example, titles, paragraphs / key points, screen tags, data values, and annotation content. For example, if the font size of a text fragment is larger than a preset font size and the text fragment is located at the top of the live stream information, the classification label for that text fragment can be determined as "title". In other embodiments, an OCR model (e.g., an end-to-end optical character recognition framework) can also be used to recognize the text content and corresponding location information in the live stream information. The OCR model can perform clustering and classification based on the text content and corresponding location information to determine the classification label for the text content. For example, live stream information can be input into an OCR model, which then outputs different category labels and the corresponding text content.

[0069] In some embodiments, keywords can also be extracted from the text content of the live stream information to obtain live stream keywords. For example, keywords in the text content of the live stream information can be identified using a named entity recognition model to obtain live stream keywords.

[0070] In some embodiments, the text content, the category corresponding to each text fragment, and keywords from the live stream can be used as the screen text information. In other embodiments, structured screen text information can also be generated based on each text fragment, the category corresponding to each text fragment, and keywords from the live stream. For example, if a host in a knowledge-sharing live stream is sharing a presentation, the content of which includes: {Title: Coping Strategies Under the AI ​​Layoff Wave; Key Point 1: Identifying High-Risk Positions; Key Point 2: Skills Transfer and Retraining; Key Point 3: Internal Job Transfer Mechanisms; Speaker: Zhang San (**AIChina)}, then the generated structured screen text information can be: "Livestream ID":"L789", "Display Content": Title: "Coping Strategies Amidst the AI ​​Layoff Wave" Key Points: "Identifying high-risk positions" "Skills transfer and retraining" "Internal job transfer mechanism" Speaker: Zhang San Source: "PPT slides" "Time point":"00:13:47", Live stream keywords: ["AI layoffs", "risky positions", "retraining", "job transfer"].

[0071] In some embodiments, keywords can be extracted based on live interactive information to obtain interactive keywords. Live interactive information may include bullet comments, comments, and chat messages between viewers and the streamer in the chat room. In some embodiments, the live interactive information can first be converted into semantic vectors. For example, each bullet comment, comment, or chat message can be encoded into a semantic vector using a language model (e.g., a model for generating sentence-level semantic embeddings). For example, a bullet comment like “**T-5 is here!” can be encoded into a semantic vector [0.25, -0.48, ..., 0.13] using a language model. Further, clustering algorithms can be used to cluster the semantic vectors, calculating the cosine similarity of each semantic vector to group semantically similar live interactive information. Clustering algorithms can include K-Means (K-Means Clustering), DBSCAN (Density-Based Spatial Clustering of Applications with Noise), and HDBSCAN (Hierarchical DBSCAN). Then, keywords can be extracted from the live interactive information in each group. For example, the c-TF-IDF (Class-based Term Frequency-Inverse Document Frequency) algorithm can be used to extract keywords from the live interactive information in each group to obtain interactive keywords.

[0072] After obtaining the live text information, a content summary can be generated based on it. In some embodiments, topic modeling can be performed on the live text information to obtain topic keywords, and a content summary can be generated based on these keywords. The topic modeling process can be a process of automatically identifying implicit semantic topic structures from the text to be processed and extracting core keywords representing each topic. For example, topic keywords can be extracted from the text to be processed using a topic modeling model.

[0073] In some embodiments, audio text information can be processed using a topic modeling model to obtain topic keywords. Then, a summary model is used to generate summary information corresponding to the audio text information based on these topic keywords. The topic modeling model can capture deep semantic similarity between texts through semantic embedding, identify topics using clustering algorithms, and extract representative keywords using a category-based TF-IDF weighting method. For example, the topic modeling model can be a topic model based on a pre-trained language model and clustering techniques. The model used to generate the summary can be, for example, a generative pre-trained language model based on the Transformer architecture. For instance, an audio text message can be input into the topic modeling model, which performs semantic clustering on the audio text information, identifies similar text clusters, and extracts topic keywords based on the vocabulary within each text cluster. For example, the extracted topic keywords might be ["**AI", "release", "large model", "layoffs", "global"]. Furthermore, the topic keywords are input into the model used to generate the summary, and the instruction "You are a live content summary generator. Please generate summary information based on the topic keywords. The sentence template is 'The host is talking about the topic [topic]', and the content of [topic] is the summary information generated based on the topic keywords" is input into the model used to generate the summary. Then the model used to generate the summary outputs the summary information corresponding to the audio text information: "The host is talking about the topic '**The recent release of a large AI model has led to a global wave of layoffs'".

[0074] In some embodiments, summary information corresponding to the live screen text information can be generated based on keywords in the live screen text information. For example, the live screen keywords can be processed using a topic modeling model to obtain topic keywords corresponding to the screen text information. Further, a model for generating summaries can generate summary information corresponding to the screen text information based on the topic keywords. For example, the topic modeling model can be a topic model based on a pre-trained language model and clustering techniques. The model for generating summaries can be a generative pre-trained language model based on the Transformer architecture. For instance, the live screen keywords can be input into the topic modeling model, which extracts topic keywords corresponding to the screen text information based on the live screen keywords. Further, the topic keywords corresponding to the screen text information are input into the model for generating summaries, and corresponding instructions are input to enable the model for generating summaries to generate summary information corresponding to the screen text information based on the topic keywords.

[0075] In some embodiments, summary information corresponding to live interactive information can be generated based on interactive keywords. For example, interactive keywords can be processed using a topic modeling model to obtain topic keywords corresponding to the live interactive information. Further, a model for generating summaries can then generate summary information corresponding to the live interactive information based on these topic keywords. For example, the topic modeling model can be a topic model based on a pre-trained language model and clustering techniques. The model for generating summaries can be a generative pre-trained language model based on a Transformer architecture. For instance, interactive keywords can be input into the topic modeling model, which extracts topic keywords corresponding to the live interactive information. Further, the topic keywords corresponding to the live interactive information are input into the model for generating summaries, along with corresponding instructions, so that the model for generating summaries generates summary information corresponding to the live interactive information based on the topic keywords.

[0076] In some embodiments, the summary information corresponding to the audio text information, the summary information corresponding to the on-screen text information, and the summary information corresponding to the live interaction information can be concatenated and merged into a paragraph describing the current semantics of the live stream to obtain a content summary. For example, the summary information corresponding to the audio text information is "The host is discussing the topic of '**AI' recent large-scale model release leading to a global wave of layoffs'"; the summary information corresponding to the on-screen text information is "The PPT title is 'AI Reshaping the Employment Structure'"; and the summary information corresponding to the live interaction information is "Viewers are discussing in the chat whether '**T-5 will replace programmers'". Then, these three summary information sentences can be merged into one paragraph: the host is discussing the topic of "**AI's recent large-scale model release leading to a global wave of layoffs", viewers are discussing in the chat whether "**T-5 will replace programmers", and the PPT title is "AI Reshaping the Employment Structure". This paragraph can be used as the content summary.

[0077] In some embodiments, live stream topic tags can be extracted based on live stream text information, and content summaries can be generated based on the live stream text information and live stream topic tags. For a detailed explanation of how to extract live stream topic tags based on live stream text information, please refer to the description in process 300 below. In some embodiments, live stream topic tags can be processed using a topic modeling model to obtain corresponding topic keywords. Further, a model for generating summaries can generate summary information corresponding to the live stream topic tags based on the topic keywords. For example, the topic modeling model can be a topic model based on a pre-trained language model and clustering techniques; the model for generating summaries can be a generative pre-trained language model based on the Transformer architecture. For instance, live stream topic tags can be input into the topic modeling model, which extracts topic keywords based on the live stream topic tags. Further, the topic keywords corresponding to the live stream topic tags are input into the model for generating summaries, and corresponding instructions are input so that the model for generating summaries generates summary information corresponding to the live stream topic tags based on the topic keywords. In some embodiments, the summary information corresponding to the live room topic tags, the summary information corresponding to the audio text information, the summary information corresponding to the screen text information, and the summary information corresponding to the live interaction information can be concatenated and merged into a paragraph describing the current semantics of the live stream to obtain a content summary.

[0078] In some embodiments, the concatenated content summary can be encoded and converted into a second semantic feature. For example, the content summary can be vectorized and encoded using a language model (e.g., a model for generating sentence-level semantic embeddings) to obtain the second semantic feature. For example, the second semantic feature can be a real-time semantic vector: [0.12, -0.43, ..., 0.87].

[0079] Step 250: Determine the semantic similarity between the first semantic feature and the second semantic feature. In some embodiments, step 250 may be implemented by a third determining module 950.

[0080] In some embodiments, semantic similarity can be the cosine similarity between a first semantic feature and a second semantic feature. For example, cosine similarity can be determined by calculating the cosine of the angle between the first and second semantic features. The value of cosine similarity can range from -1 to 1. Here, 1 indicates that the two feature vectors are completely in the same direction and are highly semantically similar; 0 indicates that the two feature vectors are orthogonal and semantically unrelated; and -1 indicates that the two feature vectors are completely opposite and semantically opposite. By determining the semantic similarity between the first and second semantic features, semantic-level matching between search information and live streaming information can be achieved.

[0081] Step 260: Determine the matching degree between the search information and each candidate live room based on semantic similarity. In some embodiments, step 260 can be implemented by the fourth determining module 960.

[0082] In some embodiments, the semantic similarity between the first semantic feature and the second semantic feature can be directly used as the matching degree between the search information and each candidate live stream. In other embodiments, the matching degree between the search information and each candidate live stream can also be determined based on the keyword coverage of the search information relative to the live stream information and the semantic similarity between the first semantic feature and the second semantic feature. For a description of this part, please refer to the description in flowchart 300 below.

[0083] Step 270: Select a recommended target live stream from the candidate live streams based on the matching degree. In some embodiments, step 270 can be implemented by the recommendation module 970.

[0084] In some embodiments, candidate live streams can be sorted based on the matching degree between search information and each candidate live stream. Based on the sorting results, a recommended target live stream is selected from the candidate live streams to be displayed, and a ranking list of recommended target live streams is generated. For example, candidate live streams can be sorted in descending order of matching degree, and the top N candidate live streams can be selected as recommended target live streams. A ranking list of recommended target live streams is then generated based on the sorting results.

[0085] In one or more embodiments of this specification, the matching degree between search information and each candidate live stream is determined based on semantic similarity, and a recommended target live stream is selected from the candidate live streams based on the matching degree. This achieves an evolution from "keyword tag matching" to "deep semantic understanding," solving the problem of inaccurate capture of user interests, improving the matching degree between recommended live streams and user interests, and achieving accurate live stream recommendations. Furthermore, live stream information includes live audio information, live video information, or live interactive information, thus forming multimodal live stream information. Based on multimodal content modeling, multimodal semantic matching is achieved, further improving the accuracy of the matching.

[0086] To further improve the accuracy of the matching between search information and each candidate live stream, the matching degree of each candidate live stream can also be determined based on the keyword coverage of search information relative to live stream information. Figure 3 This is an exemplary flowchart illustrating a method for determining the matching degree of each candidate live streaming room according to some embodiments of this specification. Figure 3 The process 300 shown can be executed by a processing device, for example, it can be executed by... Figure 1The server 110 shown executes this. In some embodiments, process 300 may be a further description of step 260. In some embodiments, process 300 may be implemented by a fourth determining module 960 in a live room recommendation device 900 deployed on a processing device. Figure 3 As shown, in some embodiments, process 300 may include the following steps.

[0087] Step 310: Perform keyword matching based on search information and live streaming information, and determine keyword coverage based on the keyword matching results.

[0088] In some embodiments, search keywords can be extracted based on search information, live room topic tags can be determined based on live room information, search keywords can be matched with live room topic tags, and keyword coverage can be determined based on keyword matching results.

[0089] In some embodiments, for extracting search keywords, search information can be input into a keyword recognition model (e.g., a keyword extraction model based on graph ranking algorithm), and the keyword recognition model can extract keywords with high semantic weights from the search information and output the extracted search keywords.

[0090] In some embodiments, live stream topic tags can be determined based on live stream information. For example, keywords can be extracted from the live stream information, and semantic clustering and topic modeling can be performed on the extracted keywords to generate live stream topic tags. The following describes in detail how to extract live stream topic tags based on live stream text information.

[0091] In some embodiments, keyword extraction can be performed based on audio text information to obtain keywords corresponding to the live audio information. For example, the audio text information can be input into a keyword extraction model (e.g., a model based on word frequency statistics, a model based on graph ranking algorithms, a named entity recognition model, etc.), and the keyword extraction model extracts keywords with high semantic weights from the audio text information. Further, the keywords extracted from the audio text information, the live screen keywords extracted from the live screen information, and the interactive keywords are used as keywords extracted based on the live information. Then, semantic clustering and topic modeling are performed on the keywords extracted based on the live information to generate live room topic tags. For example, the keywords extracted based on the live information can be processed by a topic modeling model to obtain topic keywords, and the obtained topic keywords are used as live room topic tags. The explanation of how to obtain live screen keywords and interactive keywords can be found in step 240 of process 200 above, and will not be repeated here.

[0092] In some embodiments, search keywords can be matched with live stream topic tags to identify successfully matched keywords. The keyword coverage rate is determined based on the proportion of successfully matched keywords among the search keywords. For example, search keywords include "AI", "layoffs", and "tech news", while live stream topic tags include "AI", "**AI", "layoffs", and "big model". The successfully matched keywords are "AI" and "layoffs", and the keyword coverage rate = number of successfully matched keywords / total number of search keywords = 2 / 3 = 66.7%.

[0093] Step 320: Determine the matching degree based on keyword coverage and semantic similarity.

[0094] In some embodiments, the matching degree can be determined based on the weight of keyword coverage and the weight of semantic similarity between search information and live stream information. For example, if the weight of semantic similarity between search information and live stream information is 0.7 and the weight of keyword coverage is 0.3, then the matching degree = semantic similarity × 0.7 + keyword coverage × 0.3. For instance, if the semantic similarity between search information and live stream information is 0.8 and the keyword coverage is 0.6, then the matching degree = 0.80 * 0.7 + 0.60 * 0.3 = 0.74. In some embodiments, a matching degree score can also be generated for each candidate live stream based on the matching degree. For example, multiplying the matching degree by 100 yields a matching degree score ranging from 0 to 100. Furthermore, the candidate live streams can be sorted based on the matching degree scores, and a recommended target live stream can be selected from the candidate live streams according to the sorting results, generating a ranking list of recommended target live streams.

[0095] To select live streams that better meet user needs, the system can also select recommended target live streams from the candidate live streams based on their popularity and relevance, thus recommending live streams that are both relevant and popular to users. Figure 4 This is an exemplary flowchart illustrating a method for selecting a recommended target live streaming room according to some embodiments of this specification. Figure 4 The process 400 shown can be executed by a processing device, for example, by... Figure 1 The server 110 shown executes this. In some embodiments, process 400 may be a further description of step 270. In some embodiments, process 400 may be implemented by a recommendation module 970 in a live streaming recommendation device 900 deployed on a processing device. Figure 4 As shown, in some embodiments, process 400 may include the following steps.

[0096] Step 410: Calculate the live streaming popularity of each candidate live streaming room based on the live streaming behavior data of the candidate live streaming rooms.

[0097] In some embodiments, live streaming behavior data may include the number of viewers, the frequency of bullet comments, the frequency of comments, the frequency of likes, the frequency of sending gifts, the frequency of giving coins, and the ratio of comments related to live streaming topic tags. Bullet comment frequency can be the number of bullet comments per unit time; comment frequency can be the number of comments per unit time; like frequency can be the number of likes per unit time; gift sending frequency can be the number of gifts sent per unit time; coin giving frequency can be the number of coins given per unit time; and the ratio of comments related to live streaming topic tags can be the percentage of bullet comments (or comments) matching the live streaming topic tags out of the total number of bullet comments (or comments). For example, if the total number of bullet comments (or comments) in the last ten seconds is x, and the number of bullet comments (or comments) matching the live streaming topic tags is y, then the ratio of comments related to the live streaming topic tags can be y / x. The calculation method and related explanations for live streaming topic tags can be found in the description in process 300 above, and will not be repeated here. It is understood that in the specific implementation of this specification, the collection, use, or processing of data (such as the number of online viewers, the frequency of bullet comments, the frequency of comments, the frequency of likes, the frequency of sending gifts, the frequency of giving coins, and the speaking ratio related to the live broadcast room's topic tags) requires the permission or consent of the data subject when one or more embodiments of this specification are applied to specific products or technology implementations. Furthermore, the collection, use, or processing of related data must strictly comply with the relevant laws, regulations, and standards of the data source country, implementation country, and other relevant countries and regions. De-identification technology is used to ensure that the final data used is de-identified data that has been securely processed, thereby protecting the rights and interests of the data subject and data security.

[0098] In some embodiments, the live streaming behavior data can be normalized, and the live streaming popularity can be calculated based on the weights corresponding to each live streaming behavior data. Normalization can be achieved by converting data of different orders of magnitude to a uniform scale based on a data upper limit. For example, dividing the data by its upper limit transforms the data into data within the range [0,1]. For instance, if the number of viewers is 1000, the frequency of bullet comments is 1 per second, the frequency of likes is 5 per second, and 10 bullet comments are received within 10 seconds, of which 3 match "AI," the proportion of comments related to the live streaming topic tag "AI" is 0.3. The upper limit for the number of viewers in the live streaming room is 1000, the upper limit for the bullet comment frequency is 10 per second, the upper limit for the like frequency is 20 per second, and the upper limit for the proportion of comments related to the live streaming topic tag is 1. After normalizing the data of each live broadcast behavior, the number of online viewers is 1000 / 10000=0.1, the frequency of bullet comments is 1 / 10=0.1, the frequency of likes is 5 / 20=0.25, and the ratio of comments related to the live broadcast topic tags is 0.3 / 1=0.3. Furthermore, the normalized data of each live stream behavior is multiplied by its corresponding weight to obtain the live stream popularity. For example, live stream popularity = normalized number of online viewers * 0.3 + normalized frequency of bullet comments * 0.2 + normalized frequency of likes * 0.2 + normalized ratio of comments related to live stream topic tags * 0.3 = 0.1 * 0.3 + 0.1 * 0.2 + 0.25 * 0.2 + 0.3 * 0.3 = 0.03 + 0.02 + 0.05 + 0.09 = 0.19. In some embodiments, a live stream popularity score can also be generated based on the live stream popularity. For example, the live stream popularity can be multiplied by 100 to obtain a live stream popularity score with a value ranging from 0 to 100.

[0099] Step 420: Select the recommended target live stream from the candidate live streams based on the live stream's popularity and matching degree.

[0100] In some embodiments, the overall score of each candidate live stream can be determined based on the weights of live stream popularity and matching degree. For example, if the live stream popularity is 0.19 with a corresponding weight of 0.4 and the matching degree is 0.74 with a corresponding weight of 0.6, the overall score is calculated as follows: (matching degree * 0.6 + live stream popularity * 0.4) * 100 = (0.74 * 0.6 + 0.19 * 0.4) * 100 = 52. In other embodiments, the overall score of each candidate live stream can also be determined based on the weights of the matching degree score and the live stream popularity score. For explanations regarding the matching degree score and the live stream popularity score, please refer to the descriptions in steps 320 and 410 above, which will not be repeated here. For example, if the matching degree score is 74 and the live stream popularity score is 19, the overall score is calculated as follows: (matching degree score * 0.6 + live stream popularity score * 0.4) = 74 * 0.6 + 19 * 0.4 = 52. Since the popularity of a live stream reflects its activity level in the current topic, determining the target live stream to recommend based on its popularity and relevance can be done by recommending live streams that better meet the user's needs from two dimensions: the relevance of the live stream to the search information and the activity level of the live stream itself.

[0101] In some embodiments, candidate live streams can be ranked based on their overall scores, and recommended target live streams can be selected for display based on the ranking results, generating a ranking list of recommended target live streams. For example, the top N candidate live streams can be selected as recommended target live streams, and the display order of the recommended target live streams can be determined based on the ranking results, generating a ranking list of recommended target live streams.

[0102] In some embodiments, when a recommended target live stream meets preset conditions, the ranking of the recommended target live stream that meets the conditions can be lowered or the recommended target live stream that meets the conditions can be canceled from display.

[0103] In some embodiments, the preset conditions may include: the matching degree between the recommended target live stream and the search information drops below a first preset threshold. For example, the first preset threshold is 0.3. When the topic discussed by the host in the recommended target live stream changes from "AI layoffs" to "food sharing," the matching degree between the recommended target live stream and the search information entered by the user decreases accordingly. When the matching degree drops below 0.3, the recommended target live stream can be canceled from display.

[0104] In some embodiments, the preset conditions may further include: the number of times the matching degree between the recommended target live room and the search information decreases within a preset time period is greater than or equal to a second preset threshold. For example, the preset time period is 10 minutes, and the second preset threshold is 5. When the matching degree between a recommended target live room and the user-input search information decreases more than 5 times consecutively within 10 minutes, the ranking of the recommended target live room may be downgraded.

[0105] In some embodiments, the preset condition may further include: the decrease in the matching degree between the recommended target live room and the search information within a preset time period is greater than or equal to a third preset threshold. For example, the preset time period is 10 minutes, and the third preset threshold is 0.2. When the matching degree between a recommended target live room and the search information entered by the user decreases from 0.75 to 0.5 within 10 minutes, the ranking of the recommended target live room may be downgraded.

[0106] In some embodiments, the matching degree mentioned in the above preset conditions may be the matching degree determined in process 200 based on the semantic similarity between search information and live broadcast information, or it may be the matching degree determined in process 300 based on keyword coverage and the semantic similarity between search information and live broadcast information.

[0107] In one or more embodiments of this specification, semantic drift detection is achieved by setting preset conditions and lowering the ranking of recommended target live streams that meet the conditions or canceling the display of recommended target live streams that meet the conditions. When a recommended target live stream meets the preset conditions, it indicates that the recommended target live stream has undergone semantic drift and the live stream content deviates from the user's input intent. Therefore, the recommendation of the live stream is reduced. In this way, live streams that match the user's search intent are recommended to the user in real time, ensuring the real-time nature of the recommended content and the continuous relevance of the topic. This solves the pain points of existing live stream rankings being lagging and misleading.

[0108] In some embodiments, the server 110 can dynamically obtain (e.g., periodically obtain) the live streaming information of the candidate live streaming rooms, recalculate the matching degree of each candidate live streaming room, and reselect the recommended target live streaming room based on the updated matching degree of each candidate live streaming room.

[0109] In some embodiments, the server 110 can dynamically obtain (e.g., periodically obtain) the live streaming information of candidate live streaming rooms, recalculate the matching degree and live streaming popularity of each candidate live streaming room, recalculate the comprehensive score of each candidate live streaming room based on the updated matching degree and live streaming popularity of each candidate live streaming room, reselect the recommended target live streaming room from each candidate live streaming room based on the updated comprehensive score of the live streaming room, and update the ranking of the recommended target live streaming room. In this way, the selected recommended target live streaming room and the ranking list of recommended target live streaming rooms are dynamically updated.

[0110] In some embodiments, when presenting the recommended target live room to the user, the cover of the recommended target live room can also be regenerated. In order to make the cover of the recommended target live room more in line with the user's search intent and attract the user to click on the live room, the cover of the recommended target live room can also be generated based on the search information. Figure 5 This is an exemplary flowchart of a live streaming room cover generation method according to some embodiments of this specification.Figure 5 The process 500 shown can be executed by a processing device, for example, by... Figure 1 The server 110 shown executes this. In some embodiments, process 500 can be implemented by a cover generation module 980 in a live streaming recommendation device 900 deployed on a processing device. Figure 5 As shown, in some embodiments, process 500 may include the following steps.

[0111] Step 510: Select the recommended target live stream from the candidate live streams based on the matching degree.

[0112] The explanation of step 510 is similar to that above. For details, please refer to the descriptions in step 270 and process 400 above. It will not be repeated here.

[0113] Step 520: Generate a cover image for the recommended target live stream room based on the search information.

[0114] In some embodiments, a screenshot of the live stream of the recommended target live stream can be determined based on the semantic similarity between search information and screen text information, and this screenshot can be used as the cover of the recommended target live stream. The screen text information is determined based on the live stream screen information. In some embodiments, screen text information can be extracted from the live stream screen information of each video frame of the candidate live stream within a third preset time period. For example, each video frame in the video stream generated by the candidate live stream within the past minute can be obtained, and screen text information can be extracted based on the live stream screen information of each video frame. Further explanation regarding the extraction of screen text information from live stream screen information can be found in step 240 above, and will not be repeated here.

[0115] In some embodiments, a first semantic feature can be determined based on search information, semantic features of each video frame's text information can be determined based on the text information of each video frame, and a screenshot of the live stream of the target live stream room can be determined based on the semantic similarity between the first semantic feature and the semantic features of each video frame's text information. For example, the video frame with the highest semantic similarity can be used as the screenshot of the live stream of the target live stream room. Further, this screenshot can be used as the cover of the target live stream room. For a detailed explanation of determining the first semantic feature based on search information, please refer to the description in step 230 above, which will not be repeated here. In some embodiments, the text information of each video frame can be converted into semantic vectors using a language model (e.g., a model for generating sentence-level semantic embeddings) to obtain the semantic features of each video frame's text information. Further, the cosine similarity between the first semantic feature and the semantic features of each video frame's text information is calculated, and a screenshot of the live stream room is determined in each video frame of the target live stream room based on the calculation results. For example, the video frame with the highest cosine similarity among the video frames of the target live stream room can be used as the screenshot of the live stream room, and this screenshot can be used as the cover of the target live stream room.

[0116] In some embodiments, the live screen screenshots of the recommended target live room can also be determined based on the semantic similarity between search keywords and live screen keywords. The search keywords can be determined based on search information, and the live screen keywords can be based on the text content in the live screen information. Further descriptions of search keywords and live screen keywords can be found in steps 310 and 240 above, and will not be repeated here. In some embodiments, the matching degree of each candidate live room can be converted into a semantic vector using a language model (e.g., a model for generating sentence-level semantic embeddings) to obtain the semantic features of the search keywords. Similarly, the live screen keywords can be converted into semantic vectors using a language model (e.g., a model for generating sentence-level semantic embeddings) to obtain the semantic features of the live screen keywords. Further, the cosine similarity between the semantic features of the search keywords and the semantic features of the live screen keywords is calculated, and live screen screenshots are determined in each video frame of the recommended target live room based on the calculation result. For example, the video frame with the highest cosine similarity among the video frames of the recommended target live room can be used as the live screen screenshot, and this live screen screenshot can be used as the cover of the recommended target live room.

[0117] In some embodiments, the user intent type can be determined based on search information, the live stream intent type can be determined based on the on-screen text information, and the live stream screenshot of the recommended target live stream room can be determined based on the matching degree between the user intent type and the live stream intent type, as well as the semantic similarity between the search information and the on-screen text information. For example, the search information and the on-screen text information can be input into an intent classification model, which outputs corresponding intent type labels to obtain the user intent type and the live stream intent type. For example, the intent type can include information acquisition, learning, entertainment, etc. The intent classification model can be a bidirectional encoding model based on the Transformer architecture. The intent classification model can preprocess the input text, using a pre-trained language model based on the Transformer architecture to convert the text into a semantic vector representation containing contextual information, and then use a classification layer to discriminate the vector and output the corresponding intent category label.

[0118] In some embodiments, it can be determined whether the user intent type matches the live stream intent type, and the live stream screenshots corresponding to the screen text information that does not match the user intent type can be filtered. For example, if the screen text information of a certain video frame has the highest semantic similarity to the search information, but the live stream intent type corresponding to the screen text information of that video frame does not match the user intent type, then the live stream screenshot of that video frame can be filtered.

[0119] In some embodiments, a screenshot of the live stream of the target live stream can be determined based on the image-text similarity between the live stream information and the search information, and this screenshot can be used as the cover image of the target live stream. For example, the image-text similarity between the live stream information and the search information can be determined using a bidirectional multimodal coding (BMC) model. The BMC model can extract visual features from images and semantic features from text, and calculate the similarity between the visual and semantic features. For instance, the video frames corresponding to the live stream information and the search information can be input into the BMC model. The BMC model converts the video frames into image feature vectors using an image encoder and the search information into text feature vectors using a text encoder. Then, it calculates the cosine similarity between the image feature vectors and the text feature vectors, and outputs this cosine similarity as the image-text similarity between the live stream information and the search information. Similarly, the image-text similarity between the search information and the live stream information corresponding to other video frames of the target live stream is calculated using the same method, and a screenshot of the live stream is determined from each video frame of the target live stream based on the calculation results. For example, the video frame with the highest image-text similarity among all video frames of the recommended target live stream can be used as a screenshot of the live stream and then used as the cover of the recommended target live stream.

[0120] In some embodiments, emotion tags can be extracted based on the live stream footage information, matched with search keywords, and a screenshot of the live stream footage of the recommended target live stream room can be determined based on the matching results. This screenshot can then be used as the cover image of the recommended target live stream room. Emotion tags can be, for example, neutral, surprised, or anxious. Search keywords can be determined based on search information. Further details regarding search keywords can be found in step 310 above and will not be repeated here.

[0121] In some embodiments, a face recognition model can be used to identify facial images in each live stream, and an expression recognition model can be used to identify the expressions in each facial image and generate emotion tags. The emotion tags of each facial image are then matched with search keywords. The face recognition model can be, for example, a face detection and alignment model based on a multi-task cascaded convolutional neural network, and the expression recognition model can be, for example, a facial expression classification model based on a convolutional neural network. In some embodiments, a thesaurus corresponding to different emotion tags can be established. For example, the synonyms for the emotion tag "frowning" include angry, furious, irritable, and raging. Search keywords can be compared with the synonyms corresponding to emotion tags; if they match, the match is successful; otherwise, the match fails. Further, when a match is successful, the video frame corresponding to the emotion tag can be used as a screenshot of the live stream, and this screenshot can be used as the cover image of the recommended target live stream. For example, if the search keywords include "angry layoffs," and the emotion tag identified based on the anchor's expression in a certain video frame is "frowning," then that video frame can be used as a screenshot of the live stream. In some embodiments, when a search keyword successfully matches an emotion tag corresponding to multiple video frames, the frame of any one of the successfully matched video frames can be used as a screenshot of the live stream. In other embodiments, the expression recognition model can also output the confidence score corresponding to the emotion tag when outputting the emotion tag. When a search keyword successfully matches an emotion tag corresponding to multiple video frames, the frame of the video frame corresponding to the emotion tag with the highest confidence score among the successfully matched emotion tags can be used as a screenshot of the live stream.

[0122] In one or more embodiments of this specification, by generating a cover image for a recommended target live streaming room based on search information, a live streaming segment that is semantically strongly related to the user's input can be used as the cover image for the live streaming room. This ensures that the displayed cover image of the recommended target live streaming room can truly reflect the topics that the user is interested in, enhances the consistency between the live streaming room cover image and the live streaming content, solves the problems of "cover image fraud" and "loss of trust" in traditional live streaming rooms, and improves user click-through rate and trust.

[0123] In some embodiments, when presenting the recommended target live stream to the user, the cover title of the recommended target live stream can also be regenerated. To make the cover title of the recommended target live stream more relevant to the live stream content, it can also be generated based on the live stream information. Figure 6 This is an exemplary flowchart illustrating a method for generating a live streaming room cover title according to some embodiments of this specification. Figure 6 The process 600 shown can be executed by a processing device, for example, by... Figure 1 The server 110 shown executes this. In some embodiments, process 600 can be implemented by a cover title generation module 990 in a live streaming recommendation device 900 deployed on a processing device. Figure 6As shown, in some embodiments, process 600 may include the following steps.

[0124] Step 610: Select a recommended target live stream from the candidate live streams based on the matching degree.

[0125] The explanation of step 610 is similar to that above. For details, please refer to the descriptions in step 270 and process 400 above. It will not be repeated here.

[0126] Step 620: Generate the cover title of the recommended target live stream based on the live stream information.

[0127] In some embodiments, a summary sentence can be extracted based on the audio text information and live interaction information, and then a cover title for the recommended target live room can be generated based on the summary sentence. The summary sentence can be used to characterize the core topics of the audio text information and live interaction information. The audio text information can be determined based on the live audio information; further explanation regarding the audio text information can be found in step 240 above, and will not be repeated here.

[0128] In some embodiments, audio text information and live interaction information can be merged into an input segment, which is then fed into a summarization generation model. The model extracts summary content based on the input segment and outputs a summary sentence. For example, the summarization generation model can be a language model based on an encoder-decoder architecture. The encoder can bidirectionally understand the contextual semantics of the input text, and the decoder can progressively generate summary sentences in an autoregressive manner. For instance, the audio text information could be "According to foreign media reports, **AI has conducted large-scale layoffs…", the chat room information could be "**Is T-5 coming soon?**", and the bullet screen information could be "Really?", "**AI is really strong lately", and "Layoffs again". The audio text information, chat room information, and bullet screen information can then be merged into an input segment and fed into the summarization generation model, which outputs "**AI's large-scale layoffs predict **T-5's future." Furthermore, style templates can be set, and a recommended cover title for the target live stream can be generated based on the summary sentence and the set style template. For example, the style template could be "The host is explaining…", "The host interprets…", "A heated discussion is underway…", etc. For example, the generated cover title could be "Anchor interprets the AI ​​mass layoffs and predicts the T-5's future."

[0129] In one or more embodiments of this specification, by generating a cover title for the recommended target live streaming room based on live streaming information, it can be ensured that the cover title of the recommended target live streaming room truly reflects the live streaming content, enhance the consistency between the live streaming room cover title and the live streaming content, solve the problems of "cover title fraud" and "loss of trust" in traditional live streaming rooms, and improve user click-through rate and trust.

[0130] In some embodiments, the server 110 can dynamically acquire (e.g., periodically acquire) the live streaming information of candidate live streaming rooms, and regenerate the cover and cover title of the recommended target live streaming room based on the search information and the updated live streaming information. This automatically identifies the live streaming segment images that are strongly related to the user input and the cover title that is consistent with the live streaming content, and dynamically updates the live streaming room cover and cover title to ensure that the display images and text of the recommended target live streaming room can truly reflect the topics that users are interested in and the real-time live streaming content, thereby improving user click-through rates and trust.

[0131] In some embodiments, the recommended target live streaming rooms selected by the server 110 can be presented to the user through the viewer's terminal 120, so that the user can view the search results obtained based on the search information. This specification also provides another method for recommending live streaming rooms. Figure 7 This is an exemplary flowchart of another live streaming room recommendation method according to some embodiments of this specification. Figure 7 The process 700 shown can be executed by a terminal device. For example, it can be executed by... Figure 1 The process is executed by the viewer terminal 120 shown. In some embodiments, process 700 can be implemented by a live streaming room recommendation device 1000 deployed on the terminal device. Figure 7 As shown, in some embodiments, process 700 may include the following steps.

[0132] Step 710: Obtain search information. In some embodiments, step 710 may be implemented by a third acquisition module 1010.

[0133] Users can install and run live video streaming software on the viewer terminal 120, and enter search information within the software. For example, when a user enters search information and performs a search operation in the graphical user interface of the live video streaming software, the viewer terminal 120 responds to the search operation and obtains the search information entered by the user. In some embodiments, the search information can be text information, audio information, image information, etc. Further explanation regarding search information can be found in the description of process 200 above, and will not be repeated here.

[0134] Step 720: Display recommended target live streaming rooms that match the search information. In some embodiments, step 720 can be implemented by the display module 1020.

[0135] In some embodiments, the recommended target live stream room can be determined based on the matching degree between search information and the live stream information of candidate live stream rooms. The matching degree can include the semantic similarity between search information and live stream information. The semantic similarity can be determined based on a first semantic feature and a second semantic feature. The first semantic feature can be determined based on search information, and the second semantic feature can be determined based on live stream information. The first semantic feature can be used to characterize the overall semantics of search information, and the second semantic feature can be used to characterize the overall semantics of live stream information. Live stream information can include at least one of live audio information, live video information, and live interactive information. For example, server 110 can obtain the live stream information of candidate live stream rooms, determine the first semantic feature based on search information, determine the second semantic feature based on live stream information, and determine the semantic similarity between the first semantic feature and the second semantic feature. Based on the semantic similarity, it can determine the matching degree between search information and each candidate live stream room, and select the recommended target live stream room from the candidate live stream rooms based on the matching degree. For an explanation of this part, please refer to the description in processes 200 to 600 above, which will not be repeated here. Further, server 110 can send the selected recommended target live stream room to viewer 120, and viewer 120 displays the recommended target live stream room that matches the search information.

[0136] In some embodiments, a cover image of the recommended target live stream room can also be displayed. This cover image can be generated based on search information. In some embodiments, a screenshot of the live stream of the recommended target live stream room can be determined based on the semantic similarity between the search information and the text information on the screen, and this screenshot can be used as the cover image of the recommended target live stream room. The text information on the screen can be determined based on the live stream information. In other embodiments, a screenshot of the live stream of the recommended target live stream room can be determined based on the image-text similarity between the live stream information and the search information, and this screenshot can be used as the cover image of the recommended target live stream room. In still other embodiments, emotion tags can be extracted based on the live stream information, matched with search keywords based on the emotion tags, and the screenshot of the live stream of the recommended target live stream room can be determined based on the matching results. This screenshot can be used as the cover image of the recommended target live stream room, and the search keywords can be determined based on the search information. For a detailed explanation of this part, please refer to the description in flowchart 500 above, which will not be repeated here.

[0137] In some embodiments, a cover title for the recommended target live stream may also be displayed. This cover title can be generated based on the live stream information. In some embodiments, a summary sentence can be extracted based on the audio text information and live stream interaction information, and the cover title for the recommended target live stream can be generated based on this summary sentence. This summary sentence can be used to characterize the core topics of the audio text information and live stream interaction information, and the audio text information can be determined based on the live stream audio information. For a detailed explanation of this part, please refer to the description in process 600 above; it will not be repeated here.

[0138] In some embodiments, the recommended target live stream rooms, along with their covers and titles, can be displayed according to the sorting results. For example, server 110 selects three recommended target live stream rooms, with corresponding cover titles of: Host A (23,000 viewers) discussing "AI industry trends," Host B (15,000 viewers) analyzing "**Reasons for AI layoffs," and Host C (12,000 viewers) sharing "Experience interviewing with AI companies." The covers and titles of these three recommended target live stream rooms can be displayed on the recommended target live stream room display interface of viewer 120 according to the sorting results.

[0139] In some embodiments, at least one of the following controls may be displayed on the display interface of the recommended target live stream: a share control, a subscription control, and a persistent page control. These controls may include characters, icons, virtual buttons, etc. In some embodiments, in response to a triggering operation of the share control, the viewer 120 may share the ranking list of the recommended target live stream. For example, sharing the ranking list of the recommended target live stream with other users in the video live streaming software. In some embodiments, in response to a triggering operation of the subscription control, the viewer 120 may favorite the ranking list of the recommended target live stream and generate a notification message when the ranking list of the recommended target live stream changes. The notification message may include character information, image information, audio information, vibration alerts, etc. In some embodiments, in response to a triggering operation of the persistent page control, the viewer 120 may set the display interface of the recommended target live stream as a persistent page of the target page. The target page may be the homepage of the video live streaming software that displays the recommended target live stream, or it may be another user-defined page.

[0140] In some embodiments, a ranking list of recommended target live streams can be generated based on the matching degree between search information and the live stream information of candidate live streams. For example, candidate live streams can be ranked based on the matching degree between search information and each candidate live stream, and recommended target live streams to be displayed can be selected from the candidate live streams according to the ranking results, and a ranking list of recommended target live streams can be generated. For more details on this part, please refer to the descriptions in steps 270, 320 and 420 above, which will not be repeated here.

[0141] Figure 8 This is a schematic diagram illustrating a recommended target live streaming room display interface according to some embodiments of this specification. For example... Figure 8 As shown, the user inputs the search query "live stream of someone cooking home-style dishes." Server 110 can send the matched recommended target live streams to viewer 120, and display the matched recommended target live streams on viewer 120's recommended target live stream display interface. The recommended target live stream display interface can also display the cover image and cover title of each recommended target live stream. For example... Figure 8 As shown, the cover titles of the recommended target live streams are as follows: "The chef is making sweet and sour pork, and the discussion is lively", "The host is making hot and sour shredded potatoes, with recipe", "There is a lively discussion about the host making fish", "The host is explaining the sweet and sour pork he just made", and "There is a discussion about the tofu dish the host made".

[0142] In some embodiments, the recommended target live stream display interface may also display the live stream's overall score and corresponding identifier. For example... Figure 8 As shown, the overall score of the recommended target live stream can be displayed near the cover title "Special Chef is making sweet and sour pork, and the discussion is lively" and the corresponding label "Overall matching degree".

[0143] In some embodiments, the ranking list of recommended target live streams can also be shared or subscribed to. For example, the recommended target live stream display interface can show a share control for sharing the ranking list of recommended target live streams and a subscription control for subscribing to the ranking list of recommended target live streams. The share control can be as follows: Figure 8 As shown by the "Share Ranking" icon, the subscription control can be used as follows: Figure 8 The "Subscription Ranking" icon is shown in the image. In response to the "Share Ranking" icon, the ranking list of the recommended target live stream is shared. In response to the "Subscription Ranking" icon, the ranking list of the recommended target live stream is saved to the user's favorites. When the ranking list of the recommended target live stream is updated, a notification message is generated to inform the user that the subscription ranking list has been updated.

[0144] In some embodiments, the recommended target live stream display interface can also be set as a persistent page and displayed on the target page of the video live streaming software (e.g., the homepage of the video live streaming software). When a user opens the video live streaming software, the recommended target live stream display interface can be presented on the target page of the video live streaming software (e.g., the homepage of the video live streaming software). For example, the recommended target live stream display interface can display a persistent page control, which can be as follows: Figure 8 As shown by the "Permanent on Homepage" icon, responding to the triggering of the "Permanent on Homepage" icon sets the recommended target live stream interface as a permanent page and displays it on the homepage of the video live streaming software. When users open the video live streaming software, they can see the recommended target live stream interface on the homepage. By setting the recommended target live stream interface as a permanent page, it can replace the traditional "Live Stream Hot List" recommendation entry point, becoming a new type of "user interest traffic entry point," and can also enhance user stickiness.

[0145] This manual also provides a live streaming room recommendation device. Figure 9 This is an exemplary block diagram of a live streaming room recommendation device according to some embodiments of this specification. In some embodiments, the live streaming room recommendation device 900 may be deployed on server 110. Figure 9 As shown, in some embodiments, the live streaming recommendation device 900 may include a first acquisition module 910, a second acquisition module 920, a first determination module 930, a second determination module 940, a third determination module 950, a fourth determination module 960, and a recommendation module 970.

[0146] The first acquisition module 910 is used to acquire search information.

[0147] The second acquisition module 920 is used to acquire live streaming information of candidate live streaming rooms. The live streaming information includes at least one of live audio information, live video information, and live interactive information.

[0148] The first determining module 930 is used to determine a first semantic feature based on the search information, wherein the first semantic feature is used to characterize the overall semantics of the search information.

[0149] The second determining module 940 is used to determine a second semantic feature based on the live broadcast information, wherein the second semantic feature is used to characterize the overall semantics of the live broadcast information.

[0150] The third determining module 950 is used to determine the semantic similarity between the first semantic feature and the second semantic feature.

[0151] The fourth determination module 960 is used to determine the matching degree of each candidate live room based on the search information and the live information.

[0152] The recommendation module 970 is used to select a target live stream from the candidate live streams based on the matching degree.

[0153] In some optional embodiments, the second determining module 940 can also be used to: extract live-related text content based on live-stream information to obtain live-stream text information; generate a content summary based on the live-stream text information; and generate a second semantic feature based on the content summary.

[0154] In some optional embodiments, the live text information includes one or more of audio text information, video text information, or interactive keywords; the second determining module 940 can also be used to perform at least one of the following processes: converting the live audio information into text form to obtain audio text information; extracting and classifying the text content in the live video information to obtain video text information; and extracting keywords based on the live interactive information to obtain interactive keywords.

[0155] In some optional embodiments, the second determining module 940 can also be used to extract text content and corresponding location information from the live broadcast screen information; perform clustering and classification processing based on the text content and corresponding location information in the live broadcast screen information to determine the classification label of the text content and generate screen text information.

[0156] In some optional embodiments, the second determining module 940 can also be used to extract live room topic tags based on live text information; and generate a content summary based on the live text information and live room topic tags.

[0157] In some optional embodiments, the fourth determining module 960 can also be used to perform keyword matching based on search information and live broadcast information, and determine keyword coverage based on keyword matching results; and determine matching degree based on keyword coverage and semantic similarity.

[0158] In some optional embodiments, the fourth determining module 960 can also be used to: extract search keywords based on search information; perform keyword matching between search keywords and live room topic tags to determine the keywords that are successfully matched in the search keywords; determine the keyword coverage rate based on the proportion of successfully matched keywords in the search keywords; wherein, the live room topic tags are determined based on live room information.

[0159] In some optional embodiments, the recommendation module 970 can also be used to calculate the live stream popularity of each candidate live stream based on the live stream behavior data of the candidate live streams. The live stream behavior data includes at least one of the following: number of online viewers, frequency of bullet comments, frequency of comments, frequency of likes, frequency of sending gifts, frequency of giving coins, and the speaking ratio related to the live stream topic tags; and select a recommended target live stream from the candidate live streams based on the live stream popularity and matching degree; wherein, the live stream topic tags are determined based on the live stream information.

[0160] In some optional embodiments, the recommendation module 970 can also be used to determine the comprehensive score of each candidate live room based on the weight of the live room popularity and matching degree; sort the candidate live rooms based on the comprehensive score of the live room; and select the recommended target live room to be displayed according to the sorting result.

[0161] In some optional embodiments, the recommendation module 970 can also be used to, when the recommended target live room meets at least one of the following conditions, lower the ranking of the recommended target live room that meets the conditions or cancel the display of the recommended target live room that meets the conditions: the matching degree between the recommended target live room and the search information drops below a first preset threshold; the number of times the matching degree between the recommended target live room and the search information drops within a preset time is greater than or equal to a second preset threshold; the magnitude of the drop in the matching degree between the recommended target live room and the search information within a preset time is greater than or equal to a third preset threshold.

[0162] In some optional embodiments, the live streaming room recommendation device 900 may further include a cover generation module 980 for generating a cover for the recommended target live streaming room based on search information.

[0163] In some optional embodiments, the cover generation module 980 can also be used to: determine a screenshot of the live stream of the target live stream room based on the semantic similarity between search information and image text information, and use the screenshot as the cover of the target live stream room, with the image text information determined based on the live stream information; or determine a screenshot of the live stream of the target live stream room based on the image-text similarity between the live stream information and search information, and use the screenshot as the cover of the target live stream room; or extract emotion tags based on the live stream information, match the emotion tags with search keywords, determine a screenshot of the live stream of the target live stream room based on the matching results, and use the screenshot as the cover of the target live stream room, with the search keywords determined based on the search information.

[0164] In some optional embodiments, the live streaming room recommendation device 900 may further include a cover title generation module 990, which is used to generate a cover title for the recommended target live streaming room based on the live streaming information.

[0165] In some optional embodiments, the cover title generation module 990 can also be used to extract summary sentences based on audio text information and live interaction information, the summary sentences being used to characterize the core topics of the audio text information and live interaction information, the audio text information being determined based on the live audio information; and to generate a cover title for the recommended target live room based on the summary sentences.

[0166] This manual also provides another live streaming recommendation device. Figure 10This is an exemplary block diagram of another live streaming recommendation device according to some embodiments of this specification. In some embodiments, the live streaming recommendation device 1000 may be deployed on the viewer's end 120. Figure 10 As shown, in some embodiments, the live streaming recommendation device 1000 may include a third acquisition module 1010 and a display module 1020.

[0167] The third acquisition module 1010 is used to acquire search information.

[0168] The display module 1020 is used to display recommended target live streaming rooms that match the search information. The recommended target live streaming rooms are determined based on the matching degree between the search information and the live streaming information of the candidate live streaming rooms. The matching degree includes the semantic similarity between the search information and the live streaming information. The semantic similarity is determined based on a first semantic feature and a second semantic feature. The first semantic feature is determined based on the search information, and the second semantic feature is determined based on the live streaming information. The first semantic feature is used to characterize the overall semantics of the search information, and the second semantic feature is used to characterize the overall semantics of the live streaming information. The live streaming information includes at least one of live audio information, live video information, and live interactive information.

[0169] In some optional embodiments, the display module 1020 can also be used to display the cover of the recommended target live stream room, which is generated based on search information.

[0170] In some optional embodiments, the display module 1020 can also be used to display the cover title of the recommended target live room, which is generated based on the live room information.

[0171] This manual also provides a live streaming room recommendation system. Figure 11 This is an exemplary block diagram of a live streaming room recommendation system according to some embodiments of this specification. For example... Figure 11 As shown, in some embodiments, the live streaming recommendation system 1100 may include a viewer terminal 1110 and a server terminal 1120. In some embodiments, the viewer terminal 1110 may also be the viewer terminal 120 in the video live streaming platform operating environment 100, and the server terminal 1120 may also be the server terminal 110 in the video live streaming platform operating environment 100.

[0172] Viewer terminal 1110 is used to obtain search information and display recommended target live rooms that match the search information.

[0173] Server 1120 is used to acquire search information and live streaming information of candidate live streaming rooms sent from the viewer's end, and determine a first semantic feature based on the search information, determine a second semantic feature based on the live streaming information, determine the semantic similarity between the first semantic feature and the second semantic feature, and determine the matching degree between the search information and each candidate live streaming room based on the semantic similarity. Based on the matching degree, it selects a recommended target live streaming room from the candidate live streaming rooms and sends it to the viewer's end. The first semantic feature is used to represent the overall semantics of the search information, and the second semantic feature is used to represent the overall semantics of the live streaming information. The live streaming information includes at least one of live audio information, live video information, and live interactive information.

[0174] For more information on each module, please refer to [link / reference]. Figures 2-8 The relevant explanations will not be repeated here. It should be understood that... Figures 9-11 The apparatus, system, and modules illustrated can be implemented in various ways. For example, in some embodiments, they can be implemented by hardware, software, or a combination of both. The hardware portion can be implemented using dedicated logic; the software portion can be stored in memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated-design hardware. Those skilled in the art will understand that the methods, apparatus, and systems described above can be implemented using computer-executable instructions and / or included in the control code of a processor, such as code provided in the memory of a programmable device on a media such as a disk, CD, or DVD-ROM. The apparatus and modules described in this specification can be implemented not only by hardware circuitry such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips or transistors, or programmable hardware devices such as field-programmable gate arrays or programmable logic devices, but also by software, for example, executed by various types of processors, or by a combination of the aforementioned hardware circuitry and software (e.g., firmware).

[0175] It should be noted that the above descriptions of the devices, systems, and modules are for convenience only and should not be construed as limiting this specification to the embodiments described. It is understood that those skilled in the art, after understanding the principle of the device, can arbitrarily combine the various modules without departing from this principle to form sub-devices connected to other modules. Alternatively, some modules can be split to obtain more modules or multiple units under a single module. Such modifications are all within the scope of this specification.

[0176] Some embodiments of this specification also provide a computer device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it can implement this specification. Figures 2-8 The method shown.

[0177] Some embodiments of this specification also provide a computer-readable storage medium storing computer instructions that, when executed by a processor, can implement this specification. Figures 2-8 The method shown.

[0178] Some embodiments of this specification also provide a computer program product, including a computer program that, when at least a portion of the computer program is executed by a processor, can implement this specification. Figures 2-8 The method is illustrated. In some embodiments, the computer program product may refer only to a computer program, which may be carried on a storage medium or a processing device. In other embodiments, the computer program product may also be a storage medium or a processing device containing the aforementioned computer program. The processing device may include one or more processors, and the storage medium.

[0179] In some embodiments, the processor may be a combination of one or more of the following processors: central processing unit (CPU), application-specific integrated circuit (ASIC), application-specific instruction set processor (ASIP), graphics processing unit (GPU), physical processing unit (PPU), digital signal processor (DSP), field-programmable gate array (FPGA), programmable logic device (PLD), programmable logic controller (PLC), reduced instruction set computer (RISC), and microprocessor.

[0180] In some embodiments, the storage medium may include one or more combinations of the following: mass storage, removable storage, volatile read-write memory, and read-only memory (ROM). Exemplary mass storage may include disks, optical disks, solid-state drives, etc. Exemplary removable storage may include flash drives, floppy disks, optical disks, memory cards, compressed hard disks, magnetic tapes, etc. Exemplary volatile read-write memory may include random access memory (RAM). Exemplary RAM may include dynamic random access memory (DRAM), dual data rate synchronous dynamic random access memory (DDRSDRAM), static random access memory (SRAM), silicon controlled retrieval memory (T-RAM), and zero-capacitance memory (Z-RAM), etc. Exemplary read-only memory may include masked read-only memory (MROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), compressed hard disk read-only memory (CD-ROM), and digital multifunction hard disk read-only memory, etc.

[0181] The basic concepts have been described above. It is obvious that the detailed disclosure above is merely illustrative and does not constitute a limitation of this specification. Although not explicitly stated herein, various modifications, improvements, and corrections may be made to this specification by those skilled in the art. Such modifications, improvements, and corrections are taught in this specification and therefore remain within the spirit and scope of the exemplary embodiments described herein.

Claims

1. A method for recommending live streaming rooms, characterized in that, The method includes: Get search information; Obtain live streaming information from candidate live streaming rooms, wherein the live streaming information includes at least one of live audio information, live video information, and live interactive information; A first semantic feature is determined based on the search information, and the first semantic feature is used to characterize the overall semantics of the search information; A second semantic feature is determined based on the live broadcast information, and the second semantic feature is used to characterize the overall semantics of the live broadcast information; Determine the semantic similarity between the first semantic feature and the second semantic feature; The matching degree between the search information and each of the candidate live streaming rooms is determined based on the semantic similarity. Based on the matching degree, a recommended target live stream is selected from the candidate live streams.

2. The method according to claim 1, characterized in that, Determining the second semantic feature based on the live stream information includes: Based on the live stream information, extract the relevant text content to obtain the live stream text information; A content summary is generated based on the live stream text information; A second semantic feature is generated based on the content summary.

3. The method according to claim 2, characterized in that, The live stream text information includes one or more of audio text information, on-screen text information, or interactive keywords; the extraction of live stream-related text content based on the live stream information to obtain the live stream text information includes at least one of the following processes: The live audio information is converted into text format to obtain the audio text information; Extract the text content from the live stream screen information and classify it to obtain the screen text information; Based on the live interactive information, keywords are extracted to obtain the interactive keywords.

4. The method according to claim 3, characterized in that, The step of extracting and classifying the text content from the live stream screen information to obtain screen text information includes: Extract the text content and corresponding location information from the live stream footage; Clustering and classification processes are performed on the text content and corresponding location information in the live broadcast screen to determine the classification labels of the text content and generate the screen text information.

5. The method according to claim 2, characterized in that, The generation of content summaries based on the live text information includes: Extract live stream topic tags based on the live stream text information; The content summary is generated based on the live stream text information and the live stream topic tags.

6. The method according to claim 1, characterized in that, Determining the matching degree between the search information and each of the candidate live streams based on the semantic similarity includes: Keyword matching is performed based on the search information and the live stream information, and keyword coverage is determined based on the keyword matching results; The matching degree is determined based on the keyword coverage and the semantic similarity.

7. The method according to claim 6, characterized in that, The step of performing keyword matching based on the search information and the live stream information, and determining the keyword coverage based on the keyword matching results, includes: Extract search keywords based on the search information; The search keywords are matched with the live stream topic tags to identify the keywords that successfully match the search keywords. The keyword coverage rate is determined based on the proportion of successfully matched keywords in the search keywords; The live stream topic tags are determined based on the live stream information.

8. The method according to claim 1 or 6, characterized in that, The step of selecting a recommended target live stream from the candidate live streams based on the matching degree includes: The live streaming popularity of each candidate live streaming room is calculated based on the live streaming behavior data of the candidate live streaming rooms. The live streaming behavior data includes at least one of the following: number of online viewers, frequency of bullet comments, frequency of comments, frequency of likes, frequency of sending gifts, frequency of giving coins, and the ratio of comments related to the live streaming room's topic tags. Based on the popularity of the live stream and the matching degree, a recommended target live stream is selected from the candidate live streams. The live stream topic tags are determined based on the live stream information.

9. The method according to claim 8, characterized in that, The step of selecting a recommended target live stream from the candidate live streams based on the live stream popularity and the matching degree includes: The overall score of each candidate live stream is determined based on the weight of the live stream popularity and the matching degree. The candidate live streams are ranked based on their overall scores, and the recommended target live stream is selected to be displayed based on the ranking results.

10. The method according to claim 9, characterized in that, Also includes: When the recommended target live stream meets at least one of the following conditions, the ranking of the recommended target live stream meeting the conditions will be lowered or the recommended target live stream meeting the conditions will be removed from display: The matching degree between the recommended target live room and the search information drops below a first preset threshold; The number of times the matching degree between the recommended target live room and the search information decreases within a preset time is greater than or equal to a second preset threshold. The decrease in the matching degree between the recommended target live room and the search information within a preset time period is greater than or equal to a third preset threshold.

11. The method according to claim 1, characterized in that, Also includes: Generate the cover image of the recommended target live stream room based on the search information; The step of generating the cover image for the recommended target live stream room based on the search information includes: A screenshot of the live stream of the recommended target live stream room is determined based on the semantic similarity between search information and on-screen text information, and this screenshot is used as the cover image of the recommended target live stream room. The on-screen text information is determined based on the live stream information; or Based on the image-text similarity between the live stream information and the search information, a screenshot of the live stream of the recommended target live stream is determined, and the screenshot of the live stream is used as the cover of the recommended target live stream. or Emotional tags are extracted based on the live stream information. These emotional tags are then matched with search keywords. Based on the matching results, a screenshot of the live stream of the recommended target live stream is determined and used as the cover image of the recommended target live stream. The search keywords are determined based on the search information.

12. The method according to claim 1, characterized in that, Also includes: Generate the cover title of the recommended target live stream room based on the live stream information; The step of generating the cover title of the recommended target live stream room based on the live stream information includes: A summary sentence is extracted based on the audio text information and the live interactive information. The summary sentence is used to characterize the core topics of the audio text information and the live interactive information. The audio text information is determined based on the live audio information. The cover title of the recommended target live streaming room is generated based on the summary sentence.

13. A method for recommending live streaming rooms, characterized in that, The method includes: Get search information; Display recommended target live streaming rooms that match the search information; The recommended target live stream is determined based on the matching degree between the search information and the live stream information of the candidate live streams. The matching degree includes the semantic similarity between the search information and the live stream information. The semantic similarity is determined based on a first semantic feature and a second semantic feature. The first semantic feature is determined based on the search information, and the second semantic feature is determined based on the live stream information. The first semantic feature is used to characterize the overall semantics of the search information, and the second semantic feature is used to characterize the overall semantics of the live stream information. The live stream information includes at least one of live audio information, live video information, and live interactive information.

14. The method according to claim 13, characterized in that, Also includes: The display interface of the recommended target live streaming room shall display at least one of the following controls: a share control, a subscription control, and a persistent page control; In response to the triggering operation of the sharing control, the ranking list of the recommended target live streaming rooms is shared; In response to the triggering operation of the subscription control, the ranking list of the recommended target live streaming rooms is added to the favorites, and a reminder message is generated when the ranking list of the recommended target live streaming rooms changes; In response to the triggering operation of the persistent page control, the recommended target live room display interface is set as the persistent page of the target page; The ranking list of recommended target live streaming rooms is determined based on the matching degree between the search information and the live streaming information of the candidate live streaming rooms.

15. A live streaming recommendation device, characterized in that, The device includes: The first acquisition module is used to acquire search information; The second acquisition module is used to acquire the live streaming information of the candidate live streaming room, wherein the live streaming information includes at least one of live audio information, live video information and live interactive information. A first determining module is configured to determine a first semantic feature based on the search information, wherein the first semantic feature is used to characterize the overall semantics of the search information; The second determining module is used to determine a second semantic feature based on the live broadcast information, wherein the second semantic feature is used to characterize the overall semantics of the live broadcast information. The third determining module is used to determine the semantic similarity between the first semantic feature and the second semantic feature; The fourth determining module is used to determine the matching degree between the search information and each of the candidate live streaming rooms based on the semantic similarity. The recommendation module is used to select a target live streaming room from the candidate live streaming rooms based on the matching degree.

16. A live streaming recommendation system, characterized in that, The system includes a server and a viewer; The viewer's terminal is used to obtain search information and display recommended target live streaming rooms that match the search information; The server is configured to acquire the search information and live streaming information of the candidate live streaming rooms sent from the viewer's terminal, determine a first semantic feature based on the search information, determine a second semantic feature based on the live streaming information, determine the semantic similarity between the first semantic feature and the second semantic feature, determine the matching degree between the search information and each of the candidate live streaming rooms based on the semantic similarity, select a recommended target live streaming room from the candidate live streaming rooms based on the matching degree, and send it to the viewer's terminal. Wherein, the first semantic feature is used to characterize the overall semantics of the search information, and the second semantic feature is used to characterize the overall semantics of the live information, wherein the live information includes at least one of live audio information, live video information, and live interactive information.

17. A computer device, characterized in that, The computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, it is able to implement the method as described in any one of claims 1 to 14.

18. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a processor, enable the implementation of the method as described in any one of claims 1 to 14.

19. A computer program product, characterized in that, It includes a computer program that, when at least a portion of the computer program is executed by a processor, enables the implementation of the method as described in any one of claims 1 to 14.