Multimedia content recommendation method and apparatus, device and storage medium
By identifying keywords in instant messaging sessions and utilizing machine learning models to obtain multimedia content, the problem of traditional recommendation systems failing to accurately understand user interests is solved, enabling more accurate and personalized multimedia content recommendations.
Patent Information
- Application Number
- PCT/CN2025/079117
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-24
- Filing Date
- 2025-02-25
- Publication Date
- 2025-11-27
AI Technical Summary
Traditional multimedia content recommendation systems cannot accurately understand user interests or respond promptly to newly released video content, resulting in poor recommendation performance.
By identifying keywords in instant messaging sessions, using machine learning models to obtain multimedia content that matches the keywords, and presenting access points and introductory information in the chat window, personalized recommendations are achieved by combining keyword search, video content analysis, and machine learning model evaluation.
It improves the accuracy and personalization of multimedia content recommendations, ensuring that the recommended multimedia content better meets user needs.
Smart Images

Figure CN2025079117_27112025_PF_FP_ABST
Abstract
Description
Method, device and storage medium for multimedia content recommendation
[0001] The present application claims priority to the Chinese patent application No. 202410659467.0, filed on May 24, 2024, entitled “Method, device and storage medium for multimedia content recommendation”, the whole content of which is incorporated herein by reference. TECHNICAL FIELD
[0002] The exemplary implementation of the present disclosure generally relates to the field of computer, in particular to a method, device, equipment and computer readable storage medium for multimedia content recommendation. BACKGROUND
[0003] With the rapid development of computer technology, more and more applications and platforms are designed to provide various services to users. For example, users can publish, browse and view media item content (also referred to as media item, content item, media data, etc.) of various media item types (such as video type, audio type, picture type, etc.) in the application.
[0004] When a user expects to obtain media resources related to a certain topic, he or she needs to actively search in various media platforms, which is cumbersome. Therefore, it is desirable to propose a solution to actively recommend video works that meet the user's expectations to the user. SUMMARY
[0005] In a first aspect of the present disclosure, a method for multimedia content recommendation is provided, comprising: determining a keyword related to multimedia content based on at least one of the following: conversation information of an instant messaging session or user information of a member of the instant messaging session; obtaining at least one target multimedia content matching the keyword and / or introduction information associated with the at least one target multimedia content based on at least the keyword; and presenting an access portal of the at least one target multimedia content and / or the introduction information in a conversation window of the instant messaging session.
[0006] In a second aspect of the present disclosure, an apparatus for multimedia content recommendation is provided, comprising: a keyword determination module configured to determine a keyword related to multimedia content based on at least one of the following: conversation information of an instant messaging session or user information of a member of the instant messaging session; a target multimedia content obtaining module configured to obtain at least one target multimedia content matching the keyword and / or introduction information associated with the at least one target multimedia content based on at least the keyword; and a presentation module configured to present an access portal of the at least one target multimedia content and / or the introduction information in a conversation window of the instant messaging session.
[0007] In a third aspect of the disclosure, an electronic device is provided. The electronic device includes at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions, when executed by the at least one processor, cause the electronic device to perform the method according to the first aspect of the disclosure.
[0008] In a fourth aspect of the disclosure, a computer-readable storage medium is provided, having stored thereon computer-executable instructions that, when executed by a processor, cause the processor to implement the method according to the first aspect of the disclosure.
[0009] According to a fifth aspect of the disclosure, a computer program product is provided, comprising computer-executable instructions tangibly stored in a computer storage medium and including computer-executable instructions that, when executed by a device, cause the device to perform the method according to the first aspect of the disclosure.
[0010] It is to be understood that the details set forth herein do not purport to be essential or limiting of the subject matter presented. Other features, aspects, and advantages of the subject disclosure will become apparent from the following detailed description, which, taken in conjunction with the accompanying drawings, discloses various implementations of the present disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0011] The above and other features, aspects, and advantages of various implementations of the present disclosure will become more apparent from the following detailed description taken in conjunction with the accompanying drawings. In the drawings, like or similar elements are referred to with like or similar reference numerals, in which:
[0012] FIG. 1 shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;
[0013] FIG. 2A shows a flowchart of a process of multimedia content recommendation according to some embodiments of the present disclosure;
[0014] FIG. 2B shows an example user interface according to some embodiments of the present disclosure;
[0015] FIG. 3 shows a block diagram of a multimedia content recommendation system according to some embodiments of the present disclosure;
[0016] FIG. 4 shows a block diagram of multimedia content recommendation interaction according to some embodiments of the present disclosure;
[0017] FIG. 5 shows a flowchart of a process of multimedia content recommendation according to some embodiments of the present disclosure;
[0018] FIG. 6 shows a schematic structural block diagram of an apparatus for information generation according to some embodiments of the present disclosure; and
[0019] FIG. 7 shows a block diagram of an electronic device in which one or more embodiments of the disclosure can be implemented. DETAILED DESCRIPTION
[0020] Embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be interpreted as being limited to the embodiments set forth herein; rather, these embodiments are provided so that the present disclosure can be more thoroughly and completely understood. It should be understood that the drawings and embodiments of the present disclosure are only for illustrative purposes and are not intended to limit the scope of protection of the present disclosure.
[0021] In the description of embodiments of the present disclosure, the term "comprising" and its conjugations should be understood to encompass the meanings of "consisting of" and "consisting essentially of", i.e., "comprising but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions can also be included below.
[0022] In this document, unless explicitly stated otherwise, performing a step "in response to" an event means that the step can be performed immediately after the event, or it can include one or more intermediate steps.
[0023] It can be understood that the data involved in the technical solutions of the present disclosure (including but not limited to the data itself, the obtaining or use of the data) should comply with the requirements of the relevant laws and regulations and the relevant provisions.
[0024] It can be understood that, before using the technical solutions disclosed in the embodiments of the present disclosure, the type of personal information involved in the present disclosure, the scope of use, the use scenario, etc. should be informed to the user and the authorization of the user should be obtained through appropriate means according to the relevant laws and regulations.
[0025] For example, in response to receiving the active request of the user, a prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed will require obtaining and using the personal information of the user, so that the user can voluntarily choose whether to provide the personal information to the software or hardware such as electronic device, application, server or storage medium, etc. that performs the operation of the technical solutions of the present disclosure according to the prompt information.
[0026] As an optional but not limiting implementation manner, in response to receiving the active request of the user, the manner of sending the prompt information to the user can be, for example, the manner of pop-up window, and the prompt information can be presented in the form of text in the pop-up window. In addition, the pop-up window can also carry selection controls for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0027] It can be understood that the above notification and user authorization obtaining process is only illustrative and does not limit the implementation of the present disclosure, and other ways that meet relevant laws and regulations can also be applied to the implementation of the present disclosure.
[0028] As used herein, the term “model” can learn the relationship between the corresponding input and output from the training data, so that after the training is completed, the corresponding output can be generated for a given input. The generation of the model can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes input and provides a corresponding output by using multiple layers of processing units. The neural network model is an example of a model based on deep learning. In this article, “model” can also be referred to as “machine learning model”, “learning model”, “machine learning network” or “learning network”, which are used interchangeably in this article.
[0029] A “neural network” is a machine learning network based on deep learning. The neural network is capable of processing input and providing a corresponding output, which generally includes an input layer and an output layer and one or more hidden layers between the input layer and the output layer. The neural network used in deep learning applications usually includes many hidden layers, thereby increasing the depth of the network. The layers of the neural network are connected in sequence, so that the output of the previous layer is provided as the input of the next layer, where the input layer receives the input of the neural network, and the output of the output layer is the final output of the neural network. Each layer of the neural network includes one or more nodes (also known as processing nodes or neurons), each of which processes input from the previous layer.
[0030] Generally, machine learning can include three stages, namely a training stage, a testing stage and an application stage (also known as an inference stage). In the training stage, a given model can be trained using a large amount of training data, and the parameter values are updated iteratively until the model can obtain consistent inferences from the training data that meet the expected target. Through training, the model can be considered to learn the relationship between input and output (also known as the mapping from input to output) from the training data. The parameter values of the trained model are determined. In the testing stage, the test input is applied to the trained model to test whether the model can provide the correct output, thereby determining the performance of the model. The testing stage can sometimes be integrated into the training stage. In the application or inference stage, the trained model can be used to process the actual model input based on the parameter values obtained by training to determine the corresponding model output.
[0031] As briefly discussed above, with the rapid development of computer technology, more and more applications and platforms are designed to provide various services to users. For example, users can post, browse, and view media item contents (which can also be referred to as media items, content items, media data, etc.) of various media item types (e.g., video types, audio types, picture types, etc.) in an application. With the booming development of social media and video sharing platforms, a large amount of video contents are posted on video platforms every day. As such, a video platform can be regarded as a content-rich video database.
[0032] When a user expects to obtain media resources related to a certain topic, he or she needs to actively search in various media platforms. However, there are numerous video platforms and a large number of videos on the video platforms. Traditional video recommendation systems usually rely on techniques such as collaborative filtering, content analysis, and simple keyword matching. However, the traditional solutions cannot accurately understand video contents and cannot respond to newly posted videos in a timely manner. Therefore, the traditional solutions are difficult to capture the diversified interests of users and cannot meet the needs of users. In view of this, it is desirable to provide a video recommendation solution to actively recommend video works meeting the expectations of users.
[0033] According to embodiments of the present disclosure, a method, apparatus, device and storage medium for multimedia content recommendation are provided. The method comprises: determining a keyword related to multimedia content based on at least one of: conversation information of an instant messaging session or user information of a member of the instant messaging session; obtaining at least one target multimedia content matching the keyword and / or introduction information associated with the at least one target multimedia content based on at least the keyword; and presenting an access portal of the at least one target multimedia content and / or the introduction information in a conversation window of the instant messaging session.
[0034] In this way, high-quality multimedia content works meeting the needs of users can be provided to the users. Furthermore, according to embodiments of the present disclosure, keyword search, video content analysis, and evaluation based on a machine learning model can be combined, thereby realizing more accurate and personalized multimedia content recommendation services.
[0035] In some embodiments of the present disclosure, for ease of understanding, some embodiments of the present disclosure are described by taking a group chat session as an example scenario. However, it should be understood that embodiments of the present disclosure are not limited to the scenario of a group chat session. In fact, embodiments of the present disclosure can be applied to any instant messaging session scenario, including but not limited to single chat, group chat, chat room, etc. In other words, embodiments of the present disclosure are not limited in terms of the specific type of instant messaging session.
[0036] In some embodiments of the present disclosure, for the convenience of understanding, some embodiments of the present disclosure are described by taking videos as an example of multimedia content. However, it should be understood that embodiments of the present disclosure can be applied to any multimedia type, including but not limited to videos, graphic-text webpages, audios, public number articles, and the like.
[0037] It should be understood that the term "utilizing a machine learning model" used herein refers to utilizing one or more machine learning models. Specifically, a single machine learning model can be utilized to complete the multimedia content recommendation operation, or multiple machine learning models can be utilized to cooperatively complete the multimedia content recommendation operation, wherein each machine learning model completes one link in the multimedia content recommendation operation.
[0038] Some example embodiments of the present disclosure will be described below with continuous reference to the accompanying drawings.
[0039] Example environment
[0040] FIG. 1 shows a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. As shown in FIG. 1, the example environment 100 can include a terminal device 110. In the example environment 100, an application 115 is installed in the terminal device 110. A user 140 can interact with the application 115 via the terminal device 110 and / or an attached device of the terminal device 110.
[0041] The application 115 can provide integration of multiple applications or components to the user 140. These applications can be application modules in the application 115. In some embodiments, the application 115 can be downloaded, installed in the terminal device 110 as an application. In some embodiments, the application 115 can also be accessed by other means, such as accessed by a webpage, and the like.
[0042] The application 115 can be any appropriate type of application capable of providing media content, examples of which can include but are not limited to social applications, audio / video applications, media item playing applications, broadcasting applications, and the like, and embodiments of the present disclosure are not limited in this respect.
[0043] In the environment 100 of FIG. 1, if the application 115 is in an active state, the terminal device 110 can present an interactive page via the application 115. The interactive page can be any appropriate type of page, which can support the user 140 to input any appropriate type of data and present any media item type of media item to the user 140. The interactive interface can include various interfaces that can be provided by the application 115, such as a conversation interface for presenting chat content, an editing interface, a publishing interface, and the like.
[0044] In some embodiments, the terminal device 110 can communicate with the server 130 to implement the provisioning of services of the application 115. The terminal device 110 can be any type of mobile terminal, fixed terminal, or portable terminal including a mobile handset, a tablet computer, a laptop computer, a notebook computer, a netbook computer, a smartbook, a media tablet, a palmtop computer, a portable gaming terminal, a VR / AR device, a Personal Communication System (PCS) terminal, a personal navigation device, a Personal Digital Assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an electronic book device, a game device, or any combination thereof, including an accessory or peripheral device thereof or any combination thereof. In some embodiments, the terminal device 110 can also support any type of interface to the user (such as "wearable" circuitry, etc.).
[0045] The server 130 can be a standalone physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks, and basic cloud computing services such as big data and artificial intelligence platforms. The server 130 may, for example, include a computing system / server, such as a mainframe, an edge computing node, a computing device in a cloud environment, and the like. The server 130 can provide background services for the application 115 in the terminal device 110 that supports a virtual scene.
[0046] A communication connection can be established between the server 130 and the terminal device 110. The communication connection can be established by wired or wireless means. The communication connection can include, but is not limited to, a Bluetooth connection, a mobile network connection, a Universal Serial Bus (USB) connection, a Wireless Fidelity (WiFi) connection, and the like, and embodiments of the present disclosure are not limited in this regard. In embodiments of the present disclosure, the server 130 and the terminal device 110 can implement signaling interaction through the communication connection therebetween.
[0047] The machine learning model 120 is deployed in the environment 100. The machine learning model 120 can be deployed in a suitable electronic device. As an example, the machine learning model 120 can also be deployed in other electronic devices different from the server 130, which can invoke the machine learning model 120 to perform corresponding tasks through service calls, for example. In yet another example, the machine learning model 120 may, for example, also be deployed locally to the server 130.
[0048] As will be described in detail below, the terminal device 110 can utilize the machine learning model 120 to recommend target multimedia content to the user 140. As an example, the terminal device 110 may, for example, invoke the machine learning model 120 through the server 130. As another example, the terminal device 110 may, for example, also directly invoke the machine learning model 120.
[0049] In short, the present disclosure is not limited in terms of the deployment location and invocation manner of the machine learning model 120.
[0050] It should be understood that the structure and function of the various elements in the environment 100 are described for illustrative purposes only, without implying any limitation on the scope of the present disclosure.
[0051] Example process
[0052] FIG. 2A illustrates a schematic diagram of an example process 200A of multimedia content recommendation, according to some embodiments of the present disclosure. The process 200A can be implemented at the terminal device 110. For ease of discussion, the process 200A will be described with reference to the environment 100 of FIG. 1. It is noted that the operations performed by the aforementioned terminal device 110 and the operations performed by the terminal device 110 described hereinafter can be specifically performed by a relevant application (e.g., the application 115) installed on the electronic device 110 and / or the server 130.
[0053] At block 210, the terminal device 110 determines keywords related to the multimedia content search based on at least one of: the conversation information of the instant messaging session or the user information of the members of the instant messaging session. It should be understood that the “keywords” discussed in the embodiments herein can refer to a set of keywords, i.e., can include one or more keywords.
[0054] Next, an example process of how to determine the set of keywords will be further described.
[0055] In some embodiments, the instant messaging session is a group chat session. In this case, the terminal device 110 can determine the keywords related to the multimedia content from the group chat conversation information of the group chat users. In some embodiments, the terminal device 110 determines multimedia content features that the group chat users are interested in based on the group chat conversation information, and further determines the keywords related to the multimedia content based on the multimedia content features. As will be discussed below, the determined keywords can be used to search a set of candidate multimedia content.
[0056] In some embodiments, the multimedia content feature can refer to the type of multimedia content, including but not limited to, video type, text type. In other embodiments, the multimedia content feature can refer to the technical branch to which the multimedia content belongs, e.g., machine learning, image processing, etc. In other embodiments, the multimedia content feature can refer to other features of the multimedia content, e.g., whether it belongs to a dialogue category, whether it belongs to an introduction category, or whether it belongs to a press conference or other media type. In short, the multimedia content feature can refer to any feature related to the type, category of the multimedia content. The various embodiments of the present disclosure are not limited in this respect.
[0057] As an example embodiment, the terminal device 110 can monitor the group chat conversation information and perform semantic analysis on the group chat conversation information to determine whether there is a topic of interest to the group chat user. When it is determined that there is a topic of interest to the group chat user, the terminal device 110 can further determine the keyword associated with the topic.
[0058] In some embodiments, the terminal device 110 can filter out the relevant keywords from the group chat conversation information. In this way, it can be ensured that the candidate multimedia content searched out using the keyword is the multimedia content of interest to the group chat user. Alternatively, in some embodiments, the terminal device 110 can further determine the extended keyword that does not appear in the group chat conversation information based on the determined topic. In this way, it can be ensured that more comprehensive candidate multimedia content is searched out, and the recommended multimedia content can have a guiding and directing effect on the user.
[0059] Alternatively or additionally, the terminal device 110 can also determine the set of keywords based on the user information of the members of the instant messaging session. The user information of the members includes but is not limited to the user preferences authorized by the user and obtained by legal means, the personal information of the user (e.g., personal interests, etc.). In some embodiments, the user information can also include negative recommendation items set by the user in advance: e.g., topics of no interest to the user, multimedia content publishers of no interest to the user, multimedia content types of no interest to the user (such as advertisements), multimedia content of no interest to the user (such as sports), etc. In this case, the terminal device 110 can determine the topic of interest to the user based on the user information, and further determine the keyword associated with the topic.
[0060] In block 220, based on the obtained keywords and using the machine learning model 120, the terminal device 110 obtains at least one target multimedia content matching the keywords and introduction information for the at least one target multimedia content. How to use the machine learning model to determine the at least one target multimedia content will be discussed in detail later.
[0061] At block 230, the terminal device 110 presents an access portal of the at least one target multimedia content and introduction information for the at least one target multimedia content.
[0062] In some embodiments, the introduction information for the at least one target multimedia content can include at least one of a content summary of the at least one target multimedia content, or a recommendation reason for the at least one target multimedia content.
[0063] In some embodiments, the terminal device 110 can present the access portal of the at least one target multimedia content and the introduction information for the at least one target multimedia content in the at least one group chat conversation window.
[0064] Next, the presentation of the target multimedia content will be further described with reference to an example user interface 200B shown in FIG. 2B. The example user interface 200B is a group chat conversation interface, which displays group chat conversation information 250, 251 and 252. It should be understood that the group chat conversation information also includes conversation information that is not displayed in the example interface.
[0065] The example user interface 200B further includes recommendation information 260 of the target multimedia content presented in a card form. In the example embodiment of FIG. 2B, the sender of the recommendation information 260 is a digital assistant, which is used to present the output result of the machine learning model 120.
[0066] As shown in FIG. 2B, the recommendation information 260 can include an access portal of the target multimedia content, for example, an access link of the target multimedia content. In some embodiments, the users in the group can view the target multimedia content by clicking the access portal. Alternatively, in some embodiments, the users in the group can view the target multimedia content by clicking the presentation area of the recommendation information 260.
[0067] In the specific embodiment of FIG. 2B, the recommendation information 260 can further include summary information and / or a recommendation reason of the target multimedia content. In some embodiments, the summary information and / or the recommendation reason can be determined based on at least one of image information, audio information of the target multimedia content, or interaction information related to the at least one target multimedia content. Examples of the interaction information include, but are not limited to, evaluation information of the users, introduction information of the multimedia content, etc.
[0068] Further, the recommendation information 260 can also include image information of the target multimedia content, e.g., a thumbnail of the target multimedia content, a cover of the target multimedia content, a first image frame of the target multimedia content, etc. In some embodiments, a user can click the image information of the target multimedia content to view the details of the target multimedia content without jumping, e.g., playing the multimedia content in a card or in a pop-up window.
[0069] Next, how to determine the at least one target multimedia content will be further described.
[0070] In some embodiments, the introduction information is determined by the machine learning model 120 based on at least one of the following: image information of the at least one target multimedia content, audio information, or interaction information related to the at least one target multimedia content.
[0071] In some embodiments, the terminal device 110 obtains the at least one target multimedia content based on the keyword and by using a first machine learning model, and obtains the introduction information associated with the at least one target multimedia content by using a second machine learning model, wherein the second machine learning model is the same as or different from the first machine learning model. In this way, the recommendation process of the multimedia content will be more flexible.
[0072] In some embodiments, the terminal device 110 first searches for a set of candidate multimedia contents matching the keyword group. Next, the terminal device 110 determines a recommendation score of each candidate multimedia content in the set of candidate multimedia contents by using the machine learning model 120. Finally, the terminal device 110 selects the at least one target multimedia content from the set of candidate multimedia contents as the recommended multimedia content based on the recommendation score of each candidate multimedia content in the set of candidate multimedia contents.
[0073] In some embodiments, the terminal device 110 determines the recommendation score of each candidate multimedia content in the set of candidate multimedia contents based on at least one of the following: image information and / or audio information of each candidate multimedia content in the set of candidate multimedia contents, interaction information and / or user feedback information related to each candidate multimedia content in the set of candidate multimedia contents, multimedia content features that members of the instant messaging session are interested in, conversation information of the instant messaging session, or user information of the members of the instant messaging session. In this way, the accuracy of the recommended multimedia content can be improved.
[0074] In some embodiments, the terminal device 110 utilizes the machine learning model 120 and further determines the recommendation score of each candidate multimedia content in the set of candidate multimedia contents based on the interaction information and / or the user feedback information related to each candidate multimedia content. In some example embodiments, the terminal device 110 sorts and / or filters the search results according to the number of comments, the number of likes, the number of forwards, the number of clicks in the recent period of time, whether it belongs to the current hot topic, and the like. In this way, the quality of the recommended multimedia content is improved.
[0075] In some embodiments, the terminal device 110 first converts the audio information of each candidate multimedia content in the set of candidate multimedia contents into text respectively. Next, the terminal device 110 inputs the text of each candidate multimedia content in the set of candidate multimedia contents into the machine learning model 120 to determine the recommendation score of each candidate multimedia content in the set of candidate multimedia contents. In this way, the content of the multimedia content can be further analyzed, making the evaluation of the candidate multimedia content more objective and accurate.
[0076] In some embodiments, the terminal device 110 creates a download task corresponding to the set of candidate multimedia contents and assigns the download task to a download worker node, where the download worker node downloads the set of candidate multimedia contents to a database accessible by the machine learning model 120. In this way, the localization processing of the candidate multimedia content can be realized, thereby improving the performance of the system.
[0077] In some embodiments, the terminal device 110 assigns the set of candidate multimedia contents to a recommendation worker node and causes the recommendation worker node to determine the recommendation score of each multimedia content in the set of candidate multimedia contents by invoking the machine learning model 120.
[0078] In some embodiments, the terminal device 110 creates a first search task corresponding to the keyword to obtain a set of candidate multimedia contents matching the keyword, and at least one target multimedia content is selected from the set of candidate multimedia contents. In addition, the terminal device 110 can also determine another keyword related to the multimedia content search, and create a second search task corresponding to the other keyword to obtain another set of candidate multimedia contents matching the other keyword. In some embodiments, the first search task is assigned to a first search worker node for execution, and the second search task is assigned to a second search worker node for execution. In this way, the terminal device 110 dynamically determines the keyword group to continuously present the recommended multimedia content to the user 140.
[0079] According to some embodiments of the present disclosure, the machine learning model 120 can also be dynamically / periodically updated. In some embodiments, the terminal device 110 determines feedback information of the group chat user for the at least one target multimedia content based on the user behavior data of the group chat user and / or the group chat conversation information, and updates the machine learning model 120 based on the feedback information.
[0080] In some embodiments, the terminal device 110 can detect the number of clicks of the group chat user for the recommended multimedia content. According to a determination that the number of clicks is low, e.g., lower than a preset number, it is determined that the recommended multimedia content does not meet the user expectation. Accordingly, according to a determination that the number of clicks is high, e.g., higher than the preset number, it is determined that the recommended multimedia content meets the user expectation.
[0081] In some embodiments, the terminal device 110 can detect the group chat conversation information. For example, according to a determination that the group chat user does not have any discussion for the recommended multimedia content, it is determined that the recommended multimedia content does not meet the user expectation. Accordingly, according to a determination that the group chat user has further discussion for the recommended multimedia content, it is determined that the recommended multimedia content meets the user expectation.
[0082] In some embodiments, the group chat conversation information can be evaluation information for the recommended behavior / result, e.g., evaluation information for the recommendation quality, the recommendation time, the recommendation frequency, etc. The terminal device 110 can adjust the recommendation logic based on the evaluation information. In this way, the accuracy of the recommendation can be further improved, and the active recommendation behavior of the terminal device 110 will not disturb the group chat user too much.
[0083] For better understanding of the multimedia content recommendation scheme of some embodiments of the present disclosure, further reference is made to FIG. 3 and FIG. 4, wherein FIG. 3 shows a multimedia content recommendation system block diagram 300 according to some embodiments of the present disclosure, and FIG. 4 shows a multimedia content recommendation interaction block diagram 400 according to some embodiments of the present disclosure.
[0084] In the example embodiments of FIG. 3 and FIG. 4, four types of working nodes can be coordinated to complete the recommendation of multimedia content, i.e., a task creation working node, a search working node, a download working node, and a recommendation working node. It should be understood that the division and naming of the above-mentioned working nodes are only exemplary, and in other embodiments, the example processes of the present disclosure can be cooperatively completed and implemented by other working nodes. The present disclosure is not limited in this respect.
[0085] In operation, the task creation worker node can obtain search keywords and create a multimedia content search task. Specifically, the task creation worker node can monitor and analyze the group chat conversation information to determine whether there is a topic of interest to the group chat users. When it is determined that there is a topic of interest to the group chat users, further determine keywords associated with the topic. In some embodiments, the relevant keywords can be filtered from the group chat conversation information. In this way, it is ensured that the candidate multimedia content searched by the keywords is the multimedia content that the group chat users are truly interested in. Alternatively, in some embodiments, further extended keywords can be determined. In this way, it is ensured that more comprehensive candidate multimedia content is searched.
[0086] Further, the task creation worker node can create a search task based on the determined set of search keywords and store the search task in the database. It should be understood that the operation of the task creation worker node is dynamic and continuous. In other words, the task creation worker node continuously detects the group chat conversation information, and once it detects that there is a topic of interest to the group chat users, it determines the keywords related to the topic and creates a search task. As discussed below, in the following process, one or more recommended multimedia contents can be determined for the search task.
[0087] The search worker node continuously polls the multimedia content search tasks in the database to obtain search keywords and performs automatic search based on the set of keywords. Further, the search worker node stores the search results in the database, which include, for example, the download link, name, interaction data, and metadata of the multimedia content, etc.
[0088] The download worker node continuously polls the multimedia content download tasks in the database to obtain multimedia content download links and further automatically downloads the multimedia content through the multimedia content download links. The downloaded multimedia content can be stored in the database. In some embodiments, the download worker node can also extract the audio data in the multimedia content and obtain the multimedia content text content through voice recognition. The recognized text content can also be stored in the database.
[0089] For each multimedia content search task, the recommendation worker node can continuously poll the database to determine whether the download task of the search task is completed. If it is determined that the download task has been completed, the content of the downloaded multimedia content is evaluated. In some embodiments, the recommendation worker node can determine one or more multimedia contents with the highest score and generate recommendation information of the one or more multimedia contents. The recommendation information can be presented to the user, for example, sent to the group chat session of the user. In this way, through the coordinated work of multiple worker nodes, the accuracy of the recommended multimedia content can be improved.
[0090] Further reference is made to FIG. 5, which illustrates a flowchart 500 of a process of multimedia content recommendation according to some embodiments of the present disclosure. For ease of discussion, the process 500 will be described with reference to the environment 100 of FIG. 1. The process 500 can be performed by the relevant application (e.g., the application 115) installed on the terminal device 110 and / or the server 130. For ease of discussion, the process 500 will be described with the terminal device 110 as the performing subject.
[0091] At block 502, the terminal device 110 obtains the group chat conversation information. At block 504, the terminal device 110 extracts keywords from the group chat conversation information. As an example, the terminal device 110 summarizes the topics (e.g., a certain technical direction) that the group users are interested in from the chat content of the chat group members, and determines the relevant keywords accordingly.
[0092] At block 506, the terminal device 110 performs a search according to the determined keywords. In some embodiments, the terminal device 110 can search for relevant multimedia content on a predetermined multimedia content platform. Alternatively, in some embodiments, the terminal device 110 can search for relevant multimedia content on multiple different multimedia content platforms.
[0093] At block 508, the terminal device 110 obtains multimedia content information, including but not limited to, multimedia content access / download links, titles, descriptions of the multimedia content, and metadata related to the multimedia content. In some embodiments, the terminal device 110 sorts and / or filters the search results according to rules such as multimedia content popularity, publishing time, etc. In this way, it is ensured that newly published multimedia content can be responded to in a timely manner, and the quality of the recommended multimedia content is improved.
[0094] According to some embodiments of the present disclosure, further analysis of the candidate multimedia content is needed to improve the quality of the recommended multimedia content. As shown in FIG. 5, at block 510, the terminal device 110 can download the candidate multimedia content. Next, at block 512, the terminal device 110 extracts the audio in the candidate multimedia content. Specifically, the audio stream is extracted from the multimedia content file as the basis for subsequent speech-to-text conversion. At block 514, the terminal device 110 obtains the text information corresponding to the candidate multimedia content by using speech recognition technology.
[0095] At block 516, the terminal device 110 scores the multiple candidate multimedia contents respectively by using a machine learning model. At block 518, the terminal device 110 performs in-depth analysis and evaluation of the multimedia content by using the machine learning model, in combination with the text content of the multimedia content, user feedback, and / or interaction data, etc. Based on the analysis and evaluation results, one or more multimedia contents with the highest scores are determined.
[0096] At block 520, the terminal device 110 summarizes the multimedia content with the highest score of the one or more multimedia contents by using the machine learning model. Additionally, in some embodiments, the terminal device 110 can also generate a recommendation reason for the one or more multimedia contents. For example, the multimedia content is summarized by the machine learning model, and an automatic multimedia content abstract / overview / recommendation reason is generated to help the user quickly understand the multimedia content theme and quality.
[0097] Finally, the terminal device 110 automatically pushes the one or more multimedia contents with the highest score to the user group that is most consistent with the user interest.
[0098] Through the above example process, the terminal device 110 fully considers the group conversation information when recommending multimedia content, so that the recommended multimedia content is more consistent with the user's expectation. Further, the terminal device 110 uses the machine learning model to complete the multimedia content recommendation, thereby improving the quality of the recommended multimedia content.
[0099] Example apparatus, device, and medium
[0100] Embodiments of the present disclosure also provide a corresponding apparatus for implementing the above method or process. FIG. 6 shows a schematic structural block diagram of an apparatus 500 for multimedia content recommendation according to some embodiments of the present disclosure. The apparatus 600 can be implemented as or included in the terminal device 110. Various modules / components in the apparatus 600 can be implemented by hardware, software, firmware, or any combination thereof.
[0101] As shown in FIG. 6, the apparatus 600 includes a keyword determination module 610 configured to determine a keyword related to the multimedia content based on at least one of the following: conversation information of the instant messaging session or user information of a member of the instant messaging session; a target multimedia content acquisition module 620 configured to acquire at least one target multimedia content matching the keyword and / or introduction information associated with the at least one target multimedia content based on at least the keyword; and a presentation module 630 configured to present an access portal of the at least one target multimedia content and / or the introduction information in a conversation window of the instant messaging session.
[0102] In some embodiments, the target multimedia content acquisition module 620 is further configured to search for a set of candidate multimedia contents matching the keyword based on the keyword; determine a recommendation score of each candidate multimedia content in the set of candidate multimedia contents by using a machine learning model; and select the at least one target multimedia content from the set of candidate multimedia contents based on the recommendation score of each candidate multimedia content in the set of candidate multimedia contents.
[0103] In some embodiments, the target multimedia content obtaining module 620 is further configured to determine a recommendation score of each candidate multimedia content in the set of candidate multimedia contents based on at least one of: image information and / or audio information of each candidate multimedia content in the set of candidate multimedia contents, interaction information and / or user feedback information related to each candidate multimedia content in the set of candidate multimedia contents, multimedia content features of interest to members of the instant messaging session, conversation information of the instant messaging session, or user information of members of the instant messaging session.
[0104] In some embodiments, the target multimedia content obtaining module 620 is further configured to: convert audio information of each candidate multimedia content in the set of candidate multimedia contents into text, respectively; and input the text of each candidate multimedia content in the set of candidate multimedia contents into a machine learning model to determine a recommendation score of each candidate multimedia content in the set of candidate multimedia contents.
[0105] In some embodiments, the apparatus 600 further comprises a downloading module configured to: create a downloading task corresponding to the set of candidate multimedia contents; and assign the downloading task to a downloading worker node, the downloading worker node downloading the set of candidate multimedia contents to a database accessible by the machine learning model.
[0106] In some embodiments, the target multimedia content obtaining module 620 is further configured to: assign the set of candidate multimedia contents to a recommendation worker node; and cause the recommendation worker node to determine a recommendation score of each multimedia content in the set of candidate multimedia contents by invoking the machine learning model.
[0107] In some embodiments, the target multimedia content obtaining module 620 is further configured to: create a first search task corresponding to the keyword to obtain a set of candidate multimedia contents matching the keyword, the at least one target multimedia content being selected from the set of candidate multimedia contents. The apparatus 600 further comprises a second obtaining module configured to: determine another keyword related to the multimedia content search; and create a second search task corresponding to the other keyword to obtain another candidate multimedia content matching the other keyword, wherein the first search task is assigned to a first search worker node for execution, and the second search task is assigned to a second search worker node for execution.
[0108] In some embodiments, the introduction information of the at least one target multimedia content comprises at least one of: a content summary of the at least one target multimedia content, or a recommendation reason for the at least one target multimedia content.
[0109] In some embodiments, the keyword determining module 610 is further configured to determine multimedia content features that are of interest to the members of the instant messaging session based on the user information and / or the conversation information of the members of the instant messaging session, and determine the keywords related to the multimedia content based on the multimedia content features.
[0110] In some embodiments, the introduction information is determined by using a machine learning model and based on at least one of: image information, audio information of the at least one target multimedia content, or interaction information related to the at least one target multimedia content.
[0111] In some embodiments, the apparatus 600 further includes an updating module configured to determine feedback information of the members of the instant messaging session for the at least one target multimedia content based on the user behavior data and / or the conversation information of the members of the instant messaging session, and update the machine learning model based at least on the feedback information, wherein the machine learning model is used to obtain the at least one target multimedia content and / or the introduction information associated with the at least one target multimedia content.
[0112] In some embodiments, the instant messaging session is a group chat session.
[0113] In some embodiments, the target multimedia content obtaining module 620 is further configured to obtain the at least one target multimedia content based on the keywords and by using a first machine learning model, and obtain the introduction information associated with the at least one target multimedia content by using a second machine learning model, wherein the second machine learning model is the same as or different from the first machine learning model.
[0114] The units and / or modules included in the apparatus 600 can be implemented by various means, including software, hardware, firmware, or any combination of these. In some embodiments, one or more units and / or modules can be implemented using software and / or firmware, e.g., machine-executable instructions stored on a storage medium. In addition to or alternatively, some or all of the units and / or modules in the apparatus 500 can be implemented at least partially by one or more hardware logic components. As an example and not by way of limitation, example types of hardware logic components that can be used include Field-Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application-Specific Standard Products (ASSPs), System-on-a-Chip (SOCs), Complex Programmable Logic Devices (CPLDs), etc.
[0115] It should be understood that one or more steps in the above methods can be performed by an appropriate electronic device or combination of electronic devices. Such an electronic device or combination of electronic devices may, for example, include the terminal device 110 in FIG. 1.
[0116] FIG. 7 illustrates a block diagram of an electronic device 700 in which one or more embodiments of the disclosure can be implemented. It should be understood that the electronic device 700 illustrated in FIG. 7 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. The electronic device 700 illustrated in FIG. 7 can be used to implement the terminal device 110 of FIG. 1 or the apparatus 500 of FIG. 5.
[0117] As illustrated in FIG. 7, the electronic device 700 is in the form of a general electronic device. Components of the electronic device 700 can include, but are not limited to, one or more processors 710 or processing units, a memory 720, a storage device 730, one or more communication units 740, one or more input devices 750, and one or more output devices 770. The processor 710 can be a real or virtual processor and is capable of performing various processing according to programs stored in the memory 720. In a multi-processor system, multiple processors perform computer-executable instructions in parallel to improve parallel processing capability of the electronic device 700.
[0118] The electronic device 700 typically includes a number of computer storage media. Such media can be any available media that is accessible by the electronic device 700 and includes both volatile and non-volatile media, removable and non-removable media. The memory 720 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 730 can be a removable or non-removable media and can include a machine-readable medium, such as a flash drive, a magnetic disk, or any other medium that can be used to store information and / or data and that can be accessed by the electronic device 700.
[0119] The electronic device 700 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 7, a disk drive and a disk drive interface can be provided for reading from or writing to a removable, non- volatile magnetic disk (e.g., a "hard disk that can be used for storing software and / or data) and an optical disk drive and an optical disk drive interface can be provided for reading from or writing to a removable, non-volatile optical disk (such as a CD-ROM or other optical medium). In these instances, each drive can be connected to the bus (not shown) by one or more data media interfaces. The memory 720 can include a computer program product 725 having one or more program modules configured to carry out the various methods or actions of the various embodiments of the present disclosure.
[0120] The communication unit 740 enables communication through the communication medium with other electronic devices. Additionally, the functionality of the components of the electronic device 700 can be implemented in a single computing cluster or a plurality of computer machines capable of communicating over a communication connection. As such, the electronic device 700 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.
[0121] The input device 750 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device 770 can be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 700 can also communicate with one or more external devices (not shown), such as a storage device, a display device, etc., through the communication unit 740, as needed, with one or more devices that enable a user to interact with the electronic device 700, or with any device (e.g., a network card, a modem, etc.) that enables the electronic device 700 to communicate with one or more other electronic devices. Such communication can be carried out via an input / output (I / O) interface (not shown).
[0122] According to an example implementation of the present disclosure, there is provided a computer-readable storage medium having computer-executable instructions stored thereon, where the computer-executable instructions are executed by a processor to implement the method described above. According to an example implementation of the present disclosure, there is also provided a computer program product tangibly stored on a non-transitory computer-readable medium and comprising computer-executable instructions, where the computer-executable instructions are executed by a processor to implement the method described above.
[0123] Various aspects of the disclosure can be described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, and computer program products according to implementations of the disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0124] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include a non-transitory computer readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable storage medium having instructions stored therein comprises an article of manufacture including instructions which implement aspects of the function / act specified in the flowchart and / or block diagram block or blocks.
[0125] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0126] The flow diagrams and the block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various implementations of the present disclosure. In this regard, each block in the flow diagrams and the block diagrams can represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logic functions (s). In some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in some cases, be executed substantially concurrently, or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flow diagrams, and combinations thereof, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and
[0127] implementations. Numerous modifications and adaptations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The herein disclosed subject matter is to be considered merely illustrative in nature and is not intended to limit the scope of the described implementations as set forth in the appended claims. Rather, the scope of the described implementations is to be understood only as set forth in the appended claims.
Claims
1. A method of multimedia content recommendation, comprising: determining a keyword related to the multimedia content based on at least one of: conversation information of an instant messaging session or user information of a member of the instant messaging session; obtaining at least one target multimedia content matching the keyword and / or introduction information associated with the at least one target multimedia content based on at least the keyword; and presenting an access portal of the at least one target multimedia content and / or the introduction information in a conversation window of the instant messaging session. 2.The method of claim 1, wherein the obtaining at least one target multimedia content matching the keyword comprises: searching a set of candidate multimedia contents matching the keyword based on the keyword; determining a recommendation score of each candidate multimedia content in the set of candidate multimedia contents using a machine learning model; and selecting the at least one target multimedia content from the set of candidate multimedia contents based on the recommendation score of each candidate multimedia content in the set of candidate multimedia contents. 3.The method of claim 2, wherein the determining a recommendation score of each candidate multimedia content in the set of candidate multimedia contents comprises: determining the recommendation score of each candidate multimedia content in the set of candidate multimedia contents based on at least one of: image information and / or audio information of each candidate multimedia content in the set of candidate multimedia contents, interaction information and / or user feedback information related to each candidate multimedia content in the set of candidate multimedia contents, multimedia content features of interest to the member of the instant messaging session, conversation information of the instant messaging session, or user information of the member of the instant messaging session. 4.The method of claim 3, wherein the determining a recommendation score of each candidate multimedia content in the set of candidate multimedia contents based on audio information of each candidate multimedia content in the set of candidate multimedia contents comprises: converting the audio information of each candidate multimedia content in the set of candidate multimedia contents into text, respectively; and inputting the text of each candidate multimedia content in the set of candidate multimedia contents into the machine learning model to determine the recommendation score of each candidate multimedia content in the set of candidate multimedia contents. 5.The method of claim 2, further comprising: creating a download task corresponding to the set of candidate multimedia contents; and allocating the download task to a download worker node, which downloads the set of candidate multimedia contents to a database accessible by the machine learning model. 6.The method of claim 2, wherein the determining a recommendation score of each candidate multimedia content in the set of candidate multimedia contents comprises: allocating the set of candidate multimedia contents to a recommendation worker node; and causing the recommendation worker node to determine the recommendation score of each candidate multimedia content in the set of candidate multimedia contents by invoking the machine learning model. 7.The method of claim 1, wherein obtaining the at least one target multimedia content matching the keyword comprises: creating a first search task corresponding to the keyword to obtain a set of candidate multimedia contents matching the keyword, the at least one target multimedia content being selected from the set of candidate multimedia contents; and wherein the method further comprises: determining another keyword related to a multimedia content search; and creating a second search task corresponding to the another keyword to obtain another candidate multimedia content matching the another keyword, wherein the first search task is assigned to a first search worker node for execution, and the second search task is assigned to a second search worker node for execution. 8.The method of claim 1, wherein the introduction information for the at least one target multimedia content comprises at least one of: a content summary of the at least one target multimedia content, or a recommendation reason for the at least one target multimedia content. 9.The method of claim 1, wherein determining the keyword related to the multimedia content comprises: determining multimedia content features of interest to members of the instant messaging session based on user information of the members of the instant messaging session and / or the conversation information; and determining the keyword related to the multimedia content based on the multimedia content features. 10.The method of claim 1, wherein the introduction information is determined using a machine learning model and based on at least one of: image information, audio information of the at least one target multimedia content, or interaction information related to the at least one target multimedia content. 11.The method of claim 1, further comprising: determining feedback information of the members of the instant messaging session for the at least one target multimedia content based on user behavior data of the members of the instant messaging session and / or the conversation information; and updating a machine learning model used to obtain the at least one target multimedia content and / or the introduction information associated with the at least one target multimedia content based at least on the feedback information. 12.The method of claim 1, wherein the instant messaging session is a group chat session. 13.The method of claim 1, wherein obtaining the at least one target multimedia content and / or the introduction information associated with the at least one target multimedia content comprises: obtaining the at least one target multimedia content based on the keyword and using a first machine learning model; and obtaining the introduction information associated with the at least one target multimedia content using a second machine learning model, wherein the second machine learning model is the same as or different from the first machine learning model. 14.An apparatus for multimedia content recommendation, comprising: a keyword determination module configured to determine a keyword related to a multimedia content search based on at least one of: conversation information of an instant messaging session or user information of members of the instant messaging session. an obtaining module configured to obtain at least one target multimedia content and introduction information associated with the at least one target multimedia content, based on the keyword; and a presenting module configured to present an access portal of the at least one target multimedia content and / or the introduction information.
15. An electronic device, comprising: at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions, when executed by the at least one processor, cause the electronic device to perform the method according to any one of claims 1-13.
16. A computer-readable storage medium having computer-executable instructions stored thereon that are executable by a processor to implement the method according to any one of claims 1-13.
17. A computer program product comprising computer-executable instructions, wherein the computer program product is tangibly stored on a computer storage medium and comprises computer-executable instructions that, when executed by a device, cause the device to perform the method according to any one of claims 1-13.
Citation Information
Patent Citations
Information recommendation method and device based on instant messaging
CN106777016A
Song recommendation method, device and equipment and computer storage medium
CN112069350A
Multimedia recommendation method and device, equipment and storage medium
CN115129900A
Media information recommendation method and device, electronic equipment and storage medium
CN115221397A
Method and device for displaying music content, electronic equipment and medium
CN117370600A