Multimedia resource recommendation method and device, equipment and storage medium

Through a large language model, the object's historical interaction data is processed, the object's interest characteristics are obtained and the resource library is matched, which solves the problem of insufficient interactiveness in multimedia resource recommendations and achieves a higher accuracy and efficiency recommendation effect.

CN120523975APending Publication Date: 2025-08-22TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202410191744.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-21
Publication Date
2025-08-22

AI Technical Summary

Technical Problem

The existing multimedia resource recommendation methods lack interactivity, resulting in low recommendation accuracy and efficiency, and it is impossible to accurately recommend content that the object is interested in.

Method used

By obtaining the historical interactive data sequence of the object, using a large language model to extract feature, search and query the object's interest characteristics, and match the similarity with the preset resource library, multimedia resources that meet the object's interests are recommended.

Benefits of technology

It improves the accuracy and efficiency of multimedia resource recommendations, meets the actual needs of the object, enriches the content scope of the recommendation system, and improves the relevance and diversity of the recommendation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120523975A_ABST
    Figure CN120523975A_ABST
Patent Text Reader

Abstract

The invention provides a multimedia resource recommendation method and device, equipment and a storage medium, which are applied to various scenes such as cloud technology, artificial intelligence, intelligent traffic and auxiliary driving, and the method comprises the following steps: obtaining a historical interaction data sequence of an object; inputting the historical interaction data sequence to a large language model for feature extraction processing to obtain an object interest feature extraction result; performing search query on the object interest feature extraction result to obtain a preset number of query results representing object interests; performing similarity matching on each query result and a preset multimedia resource in a preset resource library, and determining a matched multimedia resource corresponding to each query result from the preset resource library based on a similarity matching result; and performing multimedia resource recommendation processing on the object according to the matched multimedia resource corresponding to each query result. According to the embodiment of the invention, recommendation precision and recommendation efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of computer technology, and specifically relates to a method, apparatus, device, and storage medium for recommending multimedia resources. Background Art

[0002] With the continuous development of internet technology, various multimedia resource applications have emerged, such as short video applications. These applications include a variety of short videos to meet the needs of users. Therefore, recommending multimedia resources to users has become a research hotspot.

[0003] The logic of the recommendation method in the related art is that when a certain multimedia resource (such as a short video) is liked by most objects, it is likely to be liked by other objects as well. Therefore, the multimedia resource or multimedia resources related to the multimedia resource are directly recommended to other objects.

[0004] However, the multimedia resources recommended in related technologies are not necessarily the content that the subject is interested in, and the recommendation methods in related technologies lack interactivity with the subject, resulting in low accuracy and efficiency in multimedia resource recommendation. Summary of the Invention

[0005] In order to solve the above technical problems, the present application provides a method and device for recommending multimedia resources.

[0006] In one aspect, the present application proposes a method for recommending multimedia resources, the method comprising:

[0007] Acquire a historical interaction data sequence of the object; the historical interaction data sequence is used to represent a sequence formed by data generated by the object interacting with multimedia resources at different historical times;

[0008] Inputting the historical interaction data sequence into a large language model for feature extraction processing to obtain an object interest feature extraction result; wherein the object interest feature extraction result is used to characterize the characteristics of the multimedia resources of interest to the object; the large language model is obtained by training an initial large language model based on the sample interaction data sequence and the object interest feature labels corresponding to the sample interaction data sequence;

[0009] Performing a search query on the object interest feature extraction result to obtain a preset number of query results representing the object interest;

[0010] Performing similarity matching on each query result with preset multimedia resources in a preset resource library, and determining a matching multimedia resource corresponding to each query result from the preset resource library based on the similarity matching results;

[0011] Multimedia resource recommendation processing is performed on the object according to the matching multimedia resources corresponding to each query result.

[0012] On the other hand, the present application proposes a multimedia resource recommendation device, the device comprising:

[0013] A historical interaction data sequence acquisition module is used to acquire a historical interaction data sequence of an object; the historical interaction data sequence is used to represent a sequence formed by data generated by the object interacting with multimedia resources at different historical times;

[0014] a feature extraction module configured to input the historical interaction data sequence into a large language model for feature extraction processing to obtain an object interest feature extraction result; wherein the object interest feature extraction result is used to characterize the characteristics of the multimedia resources of interest to the object; and the large language model is obtained by training an initial large language model based on the sample interaction data sequence and the object interest feature labels corresponding to the sample interaction data sequence;

[0015] A search query module, configured to perform a search query on the object interest feature extraction result to obtain a preset number of query results representing the object interest;

[0016] a matching module, configured to perform similarity matching on each query result with preset multimedia resources in a preset resource library, and determine a matching multimedia resource corresponding to each query result from the preset resource library based on the similarity matching results;

[0017] The recommendation module is used to perform multimedia resource recommendation processing on the object according to the matching multimedia resources corresponding to each query result.

[0018] On the other hand, the present application proposes an electronic device for recommending multimedia resources, which includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by the processor to implement the multimedia resource recommendation method as described above.

[0019] On the other hand, the present application proposes a computer-readable storage medium, which stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by a processor to implement the multimedia resource recommendation method as described above.

[0020] On the other hand, the present application proposes a computer program product, including a computer program, which implements the multimedia resource recommendation method as described above when the computer program is executed by a processor.

[0021] The multimedia resource recommendation method and device proposed in the embodiments of the present application use a large language model to extract the features of the implicit high-order behavior load sequence from the historical interaction data sequence of the object to obtain the object interest feature extraction result. The object interest feature extraction result is then searched and queried to obtain a preset number of query results representing the object's interests, that is, multiple queries representing the object's interests. These multiple query results are then matched with the preset multimedia resources in the preset resource library for similarity, so as to match the matching multimedia resources that meet the actual intention of the object, and multimedia resources are recommended to the object based on the matching multimedia resources. Since the historical interaction data sequence is used as the input of the large language model, the large language model is used to extract the features of the implicit high-order behavior load sequence of the object's historical interaction data sequence to better capture the fine-grained changes and intentions in the historical interaction data sequence, thereby learning the object representation of the language space from the historical interaction data sequence, so that the multimedia resources finally matched can fully reflect the interactivity with the object, thereby improving the recommendation accuracy of multimedia resources; at the same time, since the object interest feature extraction results can be searched and queried, the object interests of different resource aspects and granularities can be obtained, the content range recommended by the recommendation system can be enriched, the relevance and diversity of the recommendation results can be improved, and the recommendation accuracy can be further improved; in addition, since multiple queries representing object interests can be matched with the existing preset resource library in a more fine-grained manner, on the basis of improving the recommendation accuracy, the recommendation search timeliness of the existing preset resource library can also be improved, thereby improving the efficiency of multimedia resource recommendation. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present application or the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0023] Figure 1 The figure is a schematic diagram showing an implementation environment of a multimedia resource recommendation method according to an exemplary embodiment.

[0024] Figure 2 This is a flowchart of a method for recommending multimedia resources according to an exemplary embodiment. Figure 1 .

[0025] Figure 3 It is a schematic diagram of a Transform architecture provided according to an embodiment of the present application.

[0026] Figure 4 This is a flowchart of a method for recommending multimedia resources according to an exemplary embodiment. Figure 2 .

[0027] Figure 5 The figure is a schematic diagram showing a feature extraction process for multimodal multimedia resources according to an exemplary embodiment.

[0028] Figure 6 This is a flowchart of a method for recommending multimedia resources according to an exemplary embodiment. Figure 3 .

[0029] Figure 7 This is a flowchart of a method for recommending multimedia resources according to an exemplary embodiment. Figure 4 .

[0030] Figure 8 This is a flowchart of a method for recommending multimedia resources according to an exemplary embodiment. Figure 5 .

[0031] Figure 9 This is a flowchart of a method for recommending multimedia resources according to an exemplary embodiment. Figure 6 .

[0032] Figure 10 This is a flowchart of a method for recommending multimedia resources according to an exemplary embodiment. Figure 7 .

[0033] Figure 11 The figure is a schematic diagram of a system for recommending multimedia resources according to an exemplary embodiment.

[0034] Figure 12 The figure is a block diagram of a device for recommending multimedia resources according to an exemplary embodiment.

[0035] Figure 13 A hardware structure block diagram of a server provided according to an exemplary embodiment. DETAILED DESCRIPTION

[0036] Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive field of computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0037] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, pre-trained models, operating / interaction systems, and mechatronics. Pre-trained models, also known as large models or basic models, can be fine-tuned and widely applied to downstream tasks across various AI disciplines. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0038] Natural language processing (NLP) is an important field in the fields of computer science and artificial intelligence. It studies various theories and methods that enable effective communication between humans and computers using natural language. Natural language processing involves natural language, which is the language people use in daily life, and is closely related to linguistics research; it also involves computer science and mathematics. Pre-training models, an important technology for model training in the field of artificial intelligence, are developed from large language models (Large Language Models) in the field of NLP. After fine-tuning, large language models can be widely used in downstream tasks. Natural language processing technology generally includes text processing, semantic understanding, machine translation, robot question answering, knowledge graphs, and search-related technologies such as ranking, keywords, and recommendations.

[0039] Specifically, the "search query on the object interest feature extraction results to obtain a preset number of query results representing the object interest; similarity matching each query result with the preset multimedia resources in the preset resource library, and determining the matching multimedia resources corresponding to each query result from the preset resource library based on the similarity matching results; and multimedia resource recommendation processing for the object based on the matching multimedia resources corresponding to each query result" in this application involves search-related technologies in NLP.

[0040] Machine learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning. Pretrained models are the latest development in deep learning, integrating these techniques.

[0041] Specifically, the training process of the large language model in the embodiments of the present application involves deep learning technology in machine learning.

[0042] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, storage, display, etc.), and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the historical interaction data sequences involved in this application were all obtained with full authorization.

[0043] First, the technical terms involved in the embodiments of this application are explained:

[0044] Faiss: It is an open source library for clustering and similarity search. It provides efficient similarity search and clustering for dense vectors and supports searching billions of vectors. It is currently the most mature approximate neighbor search library.

[0045] Elasticsearch is a distributed, highly scalable, and highly real-time search and data analysis engine. It easily enables the search, analysis, and exploration of large amounts of data. Leveraging Elasticsearch's horizontal scalability makes data more valuable in production environments.

[0046] Short videos are short videos, a form of internet content dissemination, typically under five minutes long, that are circulated on new media platforms. With the widespread adoption of mobile devices and increasing network speeds, short, fast-paced, high-traffic content is gaining traction among major platforms, fans, and investors.

[0047] Large language models (LLMs) are computer models capable of processing and generating natural language. They represent a significant advancement in artificial intelligence and are expected to transform the field through their learned knowledge. LLMs can predict the next word or sentence by learning from the statistical patterns and semantic information of language data. Their capabilities improve as the input dataset and parameter space continue to expand. LLMs are used in a variety of applications, such as robotics, machine learning, machine translation, speech recognition, and image processing, earning them the name Multimodal Large Language Models (MLLMs).

[0048] Instruction Tuning: Instruction fine-tuning refers to generating individual instructions for each task, fine-tuning them on several full-shot tasks, and then evaluating generalization capabilities on specific tasks (zero shot). Full-shot refers to fine-tuning all parameters in the pre-trained model.

[0049] RLHF: Reinforcement Learning with Human Feedback is an extension of reinforcement learning (RL) that incorporates human feedback into the training process, providing machines with a natural, human-like interactive learning process. In addition to reward signals, RLHF agents receive feedback from humans, learning with a broader perspective and higher efficiency, similar to how humans learn from the expertise of another person. By building a bridge between agents and humans, RLHF allows humans to directly guide machines and allows machines to master decision-making elements that are clearly embedded in human experience. As an effective alignment technique, RLHF can help mitigate the harmful content generated by large language models (LLMs) and improve information integrity to a certain extent.

[0050] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0051] It should be noted that the terms "first", "second", etc. in the description and claims of the embodiments of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0052] It should be noted that, in the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0053] Figure 1 FIG. 1 is a schematic diagram of an implementation environment of a multimedia resource recommendation method according to an exemplary embodiment. Figure 1 As shown, the implementation environment may include at least a terminal 01 and a server 02. The terminal 01 and the server 02 may be directly or indirectly connected via wired or wireless communication, which is not limited in this embodiment of the present application.

[0054] Specifically, the server 02 can obtain the historical interaction data sequence of the object, generate matching multimedia resources corresponding to each query result, and perform multimedia resource recommendation processing on the object based on the matching multimedia resources corresponding to each query result. Optionally, the server 02 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0055] Specifically, the terminal 01 can be used to display recommended multimedia resources. The terminal 01 can include but is not limited to a mobile phone, a computer, an intelligent voice interaction device, a smart home appliance, a vehicle terminal, an aircraft, etc., but is not limited thereto.

[0056] The embodiments of the present invention can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, assisted driving, etc.

[0057] It should be noted that Figure 1 This is just an example. In other scenarios, other implementation environments may also be included.

[0058] Figure 2 This is a flowchart of a method for recommending multimedia resources according to an exemplary embodiment. Figure 1 This method can be used to Figure 1 In the implementation environment of the embodiment. This specification provides the method operation steps as described above in the embodiment or flowchart, but it may include more or fewer operation steps based on routine or non-creative work. The order of steps listed in the embodiment is only one way of executing the steps among many steps, and does not represent the only execution order. When the actual system or server product is executed, it can be executed in sequence or in parallel (for example, in a parallel processor or multi-threaded processing environment) according to the method shown in the embodiment or the accompanying drawings. Specifically, Figure 2 As shown, the method may include:

[0059] S101. Obtain a historical interaction data sequence of an object; the historical interaction data sequence is used to represent a sequence formed by data generated when an object interacts with multimedia resources at different historical times.

[0060] Optionally, the object refers to a user viewing a multimedia resource.

[0061] Optionally, "data generated by interacting with multimedia resources" may include positive interaction data and negative interaction data. For example, positive interaction data may refer to data generated by an object interacting with multimedia resources of interest, such as data generated by an object liking, sharing, or adding a favorite to a multimedia resource of interest. For example, negative interaction data may refer to data generated by an object interacting with multimedia resources of no interest, such as data generated by an object quickly swiping past, giving negative feedback, or reporting a multimedia resource of no interest.

[0062] Since the historical time of interaction between the object and the multimedia resource is different, the historical interaction data sequence can be generated according to the time of interaction between the object and the multimedia resource and the data generated by the interaction, that is, the historical interaction data sequence includes the data generated by the interaction and the time when the data was generated.

[0063] S103. Input the historical interaction data sequence into the large language model for feature extraction processing to obtain the object interest feature extraction result; wherein the object interest feature extraction result is used to characterize the characteristics of the multimedia resources that the object is interested in; the large language model is obtained by training the initial large language model based on the sample interaction data sequence and the object interest feature label corresponding to the sample interaction data sequence.

[0064] Optionally, a large language model can be pre-trained, and the training process of the large language model can be as follows: obtain a sample interaction data sequence of the object, which carries an object interest feature label; input the sample interaction data sequence into the initial large language model to perform feature extraction of the implicit high-order behavior load sequence, so as to better capture the fine-grained changes and intentions in the sample interaction data sequence, thereby learning the object representation of the language space from the sample interaction data sequence, and obtaining a predicted object interest feature extraction result; calculate the difference between the predicted object interest feature extraction result and the object interest feature label; adjust the model parameters of the initial large language model according to the difference until the difference meets the preset conditions or the number of training times meets the preset conditions, and obtain a trained large language model. The trained large language model has the function of performing feature extraction of the implicit high-order behavior load sequence on the input interaction data sequence, so as to better capture the fine-grained changes and intentions in the interaction data sequence, thereby learning the object representation of the language space from the interaction data sequence, and obtaining an object interest feature extraction result.

[0065] After the large language model is trained, historical interaction data sequences can be fed into the large language model to extract features from implicit high-order behavior payload sequences. This allows for better capture of fine-grained changes and intent within the historical interaction data sequences, thereby learning object representations in the language space from the historical interaction data sequences and extracting object interest features. These object interest feature extraction results are used to characterize the characteristics of multimedia resources of interest to the subject.

[0066] In this way, we can fully utilize the powerful natural language processing capabilities of large language models and the rich knowledge they contain, and have a good ability to capture fine-grained changes and intentions in historical interaction data sequences, effectively enrich the content range recommended by the recommendation system, better meet the actual needs of the objects, and increase the recommendation accuracy of the recommendation system.

[0067] The embodiment of the present application does not limit the structure of the large language model. As long as the model uses the generated Transform architecture, it can be classified into this category. For example, the large language model of the embodiment of the present application is based on the Large Language Model Meta Artificial Intelligence (LLaMa) and the General Linear Model (GLM) model as the basis for construction. The large language model includes a multi-layer Transform architecture. The server can perform feature extraction through the Transform architecture in the large language model. Figure 3 This is a schematic diagram of a Transform architecture provided according to an embodiment of the present application. Figure 3 , the Transform architecture includes a multi-head attention mechanism (Multi-Head Self-Attention) and a position-wise feed-forward neural network (Position-wise Feed-Forward Network). Compared with the self-attention mechanism, the multi-head attention mechanism can simultaneously mine the relationship between each item in the sequence and all other items. Item can be the interaction data in the historical interaction data sequence. That is, the multi-head attention mechanism can be used to mine information from different vector subspaces. The position feed-forward neural network adds a layer of feed-forward network after the attention, giving the model nonlinear expression capabilities and can mine the interaction relationship between different dimensions. Both the multi-head attention mechanism and the position feed-forward neural network use residual networks in the output part and perform a normalization layer (Normalization Layer). Large language models stack multiple Transform architectures together to learn more complex and higher-order interaction information.

[0068] S105. Perform a search query on the object interest feature extraction result to obtain a preset number of query results representing the object interest.

[0069] Optionally, to capture the object's interests across different resource dimensions and granularity, enrich the scope of content recommended by the recommendation system, enhance the relevance and diversity of the recommendation results, and further improve recommendation accuracy, a search query can be performed on the object interest feature extraction results to obtain a preset number of query results representing the object's interests. These preset number of query results representing the object's interests can be considered query text that represents the object's true intent.

[0070] S107 . Perform similarity matching on each query result and the preset multimedia resources in the preset resource library, and determine a matching multimedia resource corresponding to each query result from the preset resource library based on the similarity matching results.

[0071] Optionally, the preset resource library can be understood as a current resource recommendation pool, in which various types of multimedia resources are stored.

[0072] After obtaining each query result, a similarity search query may be performed on each query result in the preset resource library to determine a matching multimedia resource corresponding to each query result from the preset resource library according to the similarity matching result.

[0073] S109. Perform multimedia resource recommendation processing on the object according to the matching multimedia resources corresponding to each query result.

[0074] Optionally, after obtaining the matching multimedia resources corresponding to each query result, multimedia resources may be recommended to the object accordingly.

[0075] Since the historical interaction data sequence is used as the input of the large language model, the large language model directly extracts the features of the implicit high-order behavior load sequence of the historical interaction data sequence of the object, so as to better capture the fine-grained changes and intentions in the historical interaction data sequence, thereby learning the object representation of the language space from the historical interaction data sequence. It can be seen that the large language model in the embodiment of the present application does not need to construct a prompt learning template based on the explicit behavior of the display, that is, the large language model in the embodiment of the present application does not depict the shallow relationship based on the prompt template, but depicts the object representation, so that the multimedia resources finally matched can fully reflect the interactivity between the object and improve the recommendation accuracy of the multimedia resources; at the same time, since the object interest feature extraction results can be searched and queried, the object interests of different resource aspects and granularities can be obtained, the content range recommended by the recommendation system can be enriched, the relevance and diversity of the recommendation results can be improved, and the recommendation accuracy can be further improved; in addition, since multiple queries representing the object interest can be matched with the existing preset resource library in a more fine-grained manner, on the basis of improving the recommendation accuracy, the recommendation search timeliness of the existing preset resource library can also be improved, thereby improving the efficiency of multimedia resource recommendation.

[0076] It should be noted that the above step S101 can be implemented in various ways, which are not specifically limited. In one embodiment, Figure 4 This is a flowchart of a method for recommending multimedia resources according to an exemplary embodiment. Figure 2 ,like Figure 4 As shown, in the above step S101, the above acquisition of the historical interaction data sequence of the object may include:

[0077] S1011. Acquire multimodal multimedia resources.

[0078] Optionally, the multimodal multimedia resources may include but are not limited to: video information, audio information, image information, and text information.

[0079] S1013. Perform feature extraction processing on the multimedia resources of each modality to obtain metadata corresponding to the multimedia resources of each modality.

[0080] Optionally, feature extraction can be performed on video information, audio information, image information, and text information to obtain metadata corresponding to each modality of multimedia resources. Metadata refers to data that describes data, mainly including attribute data and related information.

[0081] The image information includes but is not limited to image content information, such as the cover image of the picture. Exemplarily, the image information may also include the title, summary, release time, etc. of the image. Video information includes but is not limited to video content information, such as video content files, etc. Exemplarily, the video information may also include the cover image link, bit rate, file format, video title, video release time, video author information, etc. The audio information mentioned includes but is not limited to audio content information, such as voice, audio in the video stream, music, etc., which is not limited in the specific embodiments of this application.

[0082] The text information includes but is not limited to text content information. For example, text content information includes text title, text name, multi-level classification and multi-level labeling of text title. Exemplarily, text content information may also include text recognition results in video information, automatic speech recognition results in audio information, and text content in image information (such as abstracts and titles of pictures, etc.), etc., which are not limited in the specific embodiments of this application. In other examples, text information may also include text publisher information, text cover image, text release time, text keywords, etc., which are not limited in the specific embodiments of this application.

[0083] Figure 5 FIG. 1 is a schematic diagram showing a feature extraction process for multimodal multimedia resources according to an exemplary embodiment. Figure 5 As shown:

[0084] Optionally, for video information, video content information, key frame information, and video modality type information can be first extracted from the video information. The video modality type information can indicate the modality to which the feature vector of the video content information belongs. The key frame information is information related to the key frames in the video content information. Thus, after extracting the video content information, key frame information, and video modality type information, feature extraction processing is performed on the video content information, key frame information, and video modality type information based on the video feature extraction model, thereby obtaining the feature vector of the video information.

[0085] Exemplarily, the video feature extraction model includes but is not limited to VideoSwinTransformer, etc., which is not specifically limited. In addition, the described video may include but is not limited to short videos, long videos, etc.

[0086] Optionally, for image information, image content information and image modality type information can be first extracted from the image information. The image modality type information indicates the modality to which the feature vector of the image content information belongs. Thus, after extracting the image content information and the image modality type information, feature extraction processing is performed on the image content information and the image modality type information based on the image feature extraction model to obtain a feature vector of the image information.

[0087] For example, the image feature extraction model may include, but is not limited to, a SwinTransformer model or a Vit model, etc., and is not specifically limited in the embodiments of this application. In addition, the image described may include, but is not limited to, the cover image of video content, the cover image of graphic content, or the image of a post in a channel scene, etc., and is not specifically limited in the embodiments of this application.

[0088] Optionally, for audio information, audio content information, first position information, and audio modal type information can be extracted from the audio information. The audio modal type information can be used to indicate the mode to which the feature vector of the audio content information belongs. In addition, the first position information can indicate the position of each frame of audio in the audio content information. After extracting the audio content information, the first position information, and the audio modal type information, feature extraction processing is performed on the audio content information, the first position information, and the audio modal type information based on the audio feature extraction model to obtain a feature vector of the audio information.

[0089] Exemplarily, the audio feature extraction model may include, but is not limited to, the wavlm-base-plus model, etc., and is not specifically limited in the embodiments of this application. In addition, the audio described may include, but is not limited to, music, speech, audio in film and television videos, audio in video tutorials, etc., and is not specifically limited in the embodiments of this application.

[0090] Optionally, for text information, text content information, second position information, and text modality type information can also be extracted from the text information. The text modality type information is used to indicate the modality to which the feature vector of the text content information belongs. In addition, the second position information is used to indicate the position of each text word in the text content information. Thus, feature extraction processing can be performed on the text content information, the second position information, and the text modality type information based on the text feature extraction model to obtain a feature vector of the text information.

[0091] For example, the text feature extraction model described may include but is not limited to the BERT model, etc., which is not limited in the embodiments of the present application. In addition, the text content information described includes the text title and the text name.

[0092] Continue as Figure 5 As shown, after obtaining the feature vector of video information, the feature vector of image information, the feature vector of audio information, and the feature vector of text information, the feature vector of the video information, the feature vector of image information, the feature vector of audio information, and the feature vector of text information can be input into the multimodal content understanding model and system. The multimodal content understanding model and system analyzes the metadata of each feature vector to obtain metadata of multimedia resources in video mode, metadata of multimedia resources in image mode, metadata of multimedia resources in audio mode, and metadata of multimedia resources in text mode. Among them, the metadata of multimedia resources in video mode is used to describe the data of multimedia resources in video mode, mainly including attribute data and related information of multimedia resources in video mode. The metadata of multimedia resources in image mode is used to describe the data of multimedia resources in image mode, mainly including attribute data and related information of multimedia resources in image mode. The metadata of multimedia resources in audio mode is used to describe the data of multimedia resources in audio mode, mainly including attribute data and related information of multimedia resources in audio mode. The metadata of the multimedia resource in text mode is used to describe the data of the multimedia resource in text mode, and mainly includes attribute data and related information of the multimedia resource in text mode.

[0093] In summary, the metadata corresponding to the multimedia resources of each modality obtained in the above manner in the embodiment of the present application may include but is not limited to: title, multi-level classification and multi-level content tags, author, release time, user-marked keywords, etc.

[0094] S1015. Obtain metadata corresponding to multimedia resources that interacted with the object at different historical times from metadata corresponding to multimedia resources of each modality.

[0095] S1017. Generate a historical interaction data sequence according to different historical times and metadata corresponding to multimedia resources that interacted with the object at different historical times.

[0096] Optionally, after obtaining the metadata corresponding to the multimedia resource of each modality, a mapping relationship between the identification information corresponding to the multimedia resource of each modality and the corresponding metadata can be established, wherein the identification information can be an identity document (Id).

[0097] Next, from the metadata corresponding to the identification information of each modal multimedia resource, metadata corresponding to the ID of the multimedia resource that interacted with the object at different historical times is obtained. This metadata is then mapped one-to-one with the corresponding historical time to form a historical interaction data sequence in sequence.

[0098] Therefore, it is possible to rely on multimodal content understanding to generate meta-information of object-related content, such as title, multi-level classification and multi-level content tags, author, release time, and user-tagged keywords, and generate historical interaction data sequences based on this, thereby enriching the scope of historical interaction data sequence generation, and then effectively enriching the content range recommended by the recommendation system, better meeting the actual needs of the object, and increasing the recommendation accuracy of the recommendation system.

[0099] In other implementations, the data generated when the object interacts with the multimedia resource at different historical times and the historical interaction data sequences at different historical times may be directly generated in chronological order.

[0100] It should be noted that the above step S105 can be implemented in a variety of ways, which are not specifically limited. In an optional embodiment, in the above step S105, the above search query on the object interest feature extraction result may include:

[0101] Determines the beam width of the beam search.

[0102] A beam search is performed on the object interest feature extraction results according to the beam width to obtain a candidate output sequence.

[0103] Repeat the operation of performing a beam search on the object interest feature extraction result according to the beam width to obtain a candidate output sequence until an end symbol is obtained or the length of the text in the candidate output sequence is greater than a preset length threshold.

[0104] The text in the candidate output sequence when the end symbol is obtained, or the text in the candidate output sequence when the length of the text is greater than a preset length threshold, is determined as the query result.

[0105] In this embodiment, to capture object interests across different resource dimensions and granularities, enrich the range of content recommended by the recommendation system, enhance the relevance and diversity of recommendation results, and further improve recommendation accuracy, a beam search (BeamSearch) multi-query generation technique can be used to perform a beam search on the object interest feature extraction results, obtaining a preset number of query results representing object interests. Specifically, the object interest feature extraction results output by the large language model can be used as hypothetical "search queries" to query and retrieve the content to be recommended.

[0106] Among them, the core function of the Beam Search algorithm is the decoding process, which decodes the hidden layer output of a sequence obtained by LLM (i.e., the object interest feature extraction result) to obtain the corresponding words.

[0107] The following example illustrates beam search:

[0108] Step 1: Determine the beam size k of Beam Search, assuming k = 2.

[0109] Step 2: First time step: Generate a candidate output sequence corresponding to the object interest feature extraction result (each candidate word in the candidate output sequence has a probability representing the similarity / correlation with the "object interest feature extraction result"), obtain the two words (word 1 and word 2) with the top 2 probabilities from the candidate output sequence, and obtain the candidate output sequence (word 1, word 2).

[0110] Step 3: Second time step: Generate a candidate output sequence starting with word 1, generate a candidate output sequence starting with word 2, and obtain the two words with the top 2 probabilities from the candidate output sequence starting with word 1 and the candidate output sequence starting with word 2 as the candidate output sequence (assuming they are word 1 and word 3, word 2 and word 4).

[0111] Step 4: The third time step: Generate a candidate output sequence starting with word 1 and word 3, and generate a candidate output sequence starting with word 2 and word 4, and obtain the two words with the top 2 probabilities as candidate output sequences (assuming they are word 1, word 3, and word 5, and word 2, word 4, and word 6).

[0112] Step 5: Repeat the above steps until the end symbol is obtained or the length of the text in the candidate output sequence is greater than the preset length threshold.

[0113] Step 6: Determine the text in the candidate output sequence when the end symbol is obtained, or the text in the candidate output sequence when the length of the text is greater than a preset length threshold, as the query result.

[0114] In other embodiments, a greedy search algorithm may be used to search the object interest feature extraction results for search query to obtain a preset number of query results representing the object interest.

[0115] It should be noted that the above step S107 can be implemented in a variety of ways. In one embodiment, Figure 6 This is a flowchart of a method for recommending multimedia resources according to an exemplary embodiment. Figure 3 ,like Figure 6As shown, in the above step S107, each query result is matched with the preset multimedia resources in the preset resource library by similarity, and a matching multimedia resource corresponding to each query result is determined from the preset resource library based on the similarity matching result, which may include:

[0116] S1071. Extract the query feature vector of each query result; and extract the resource feature vector of the preset multimedia resource.

[0117] S1073. Perform vector similarity matching on the query feature vector of each query result and the resource feature vector to obtain a similarity matching result between the query feature vector of each query result and the resource feature vector.

[0118] S1075. Determine from the resource feature vectors a resource feature vector whose similarity matching result with the query feature vector of each query result meets a preset similarity threshold, and obtain the target resource feature vector of each query result.

[0119] S1077. Determine the preset multimedia resource corresponding to the target resource feature vector of each query result as the matching multimedia resource corresponding to each query result.

[0120] In this embodiment, each query result can be matched with a vector similarity of a preset multimedia resource in a preset resource library. The preset resource library may include at least two preset multimedia resources, and a vector similarity matching is performed between the query feature vector of each query result and the resource feature vector of each preset multimedia resource to obtain a similarity matching result between the query feature vector of each query result and the resource feature vector of each preset multimedia resource.

[0121] It should be noted that the vector can refer to either a BERT-based representation or the implicit text vector representation of the content in the PLM pre-trained large-scale language model. BERT is a language representation model, and PLM refers to a large-scale pre-trained language model.

[0122] It should be noted that satisfying a preset similarity threshold may mean being smaller than a preset similarity threshold. Furthermore, it may mean that the distance between the two vectors is smaller than a preset distance threshold.

[0123] For example, a vector retrieval and matching method based on Faiss or Elasticsearch can be used to match vectors and obtain the target resource feature vector. For example, in Faiss's distributed vector retrieval, the Faiss framework can be used to implement a distributed high-dimensional nearest neighbor retrieval platform. Using a K-nearest neighbor algorithm for large-scale vector retrieval (e.g., the HNSW algorithm), the top k vectors whose distance from the query feature vector is less than a preset distance threshold can be efficiently retrieved from tens of millions of vectors in tens of milliseconds, i.e., the top k vectors similar to the query feature vector, to obtain the target resource feature vector.

[0124] After obtaining the target resource feature vector, the preset multimedia resource corresponding to the target resource feature vector can be determined as the matching multimedia resource for each query result. This allows for more fine-grained matching of multiple queries representing the subject's interests with the preset resource library. This not only improves recommendation accuracy but also enhances the timeliness of recommendation searches within the existing preset resource library, thereby increasing the efficiency of multimedia resource recommendations.

[0125] It should be noted that the above step S109 can be implemented in various ways, which are not specifically limited.

[0126] Figure 7 This is a flowchart of a method for recommending multimedia resources according to an exemplary embodiment. Figure 4 ,like Figure 7 As shown, in an optional embodiment, in the above step S109, the multimedia resource recommendation process for the object according to the matching multimedia resources corresponding to each query result may include:

[0127] S1091-1. Sort the matching multimedia resources corresponding to each query result in descending order according to the similarity matching results between the matching multimedia resources corresponding to each query result and each query result to obtain a multimedia resource sequence.

[0128] S1091-3. Determine the first preset number of multimedia resources in the multimedia resource sequence as the first target multimedia resource.

[0129] S1091-5. Recommend the first target multimedia resource to the object.

[0130] In this embodiment, after obtaining the matching multimedia resources corresponding to each query result, multimedia resources can be directly recommended to the object based on the matching multimedia resources corresponding to each query result. Specifically, the matching multimedia resources corresponding to each query result are sorted in descending order based on the similarity matching results between the matching multimedia resources corresponding to each query result and each query result to obtain a multimedia resource sequence, and then a similarity threshold is set. The first preset number of multimedia resources in the multimedia resource sequence that are greater than the similarity threshold are determined as the first target multimedia resources, and the first target multimedia resources are recommended to the object. In this way, a large pre-trained language model can be directly used to model the historical interaction data sequence, and the object representation in the language space can be learned from the historical interaction data sequence. Then, multiple queries representing the object's interests are generated, and finally the first target multimedia resource that meets the object's actual intention is matched, thereby enhancing the performance of the recommendation system. In addition, instead of directly recommending the matching multimedia resources corresponding to each query result to the object, the first preset number of multimedia resources are selected and recommended to the object, ensuring that the multimedia resources recommended to the object are multimedia resources that the object is interested in, further improving the accuracy of multimedia resource recommendation.

[0131] It should be noted that, in addition to sorting the matching multimedia resources corresponding to each query result in descending order, the matching multimedia resources corresponding to each query result can also be sorted in ascending order, and the last preset number of multimedia resources in the ascending sequence are determined as the first target multimedia resources.

[0132] In another optional embodiment, Figure 8 This is a flowchart of a method for recommending multimedia resources according to an exemplary embodiment. Figure 5 ,like Figure 8 As shown, in the above step S109, the multimedia resource recommendation process for the object according to the matching multimedia resources corresponding to each query result may include:

[0133] S1093-1. Perform multimedia resource recommendation processing on the object according to the matching multimedia resources and the recall system corresponding to each query result; wherein the recall system includes recommending multimedia resources to the object based on at least one resource recall strategy.

[0134] Among them, the recall system is an existing recall system in the traditional recommendation system. The traditional recommendation system includes a recall system (multi-way recall, filtering) and a sorting system (coarse sorting, fine sorting and mixed sorting). At least one resource recall strategy is configured in the recall system. The resource recall strategy can be a collaborative filtering algorithm recall, a social relationship recall or a hot content recall, etc. The traditional recommendation system recommends multimedia resources to the object based on at least one resource recall strategy. On the basis of the traditional recommendation system, by using a large pre-trained language model to model the historical interaction data sequence, the object representation in the language space is learned from the historical interaction data sequence, and then multiple queries representing the object's interests are generated, and finally matching multimedia resources that meet the actual intention of the object are matched. The matching multimedia resources are used to recommend the object, that is, the multimedia resource recommendation method of the embodiment of the present application is equivalent to a new resource recall strategy added on the basis of the recall system in the traditional recommendation system, which can serve as a supplement to the traditional recall system, that is, it can integrate the capabilities contained in the large language model with the recall system, so that multimedia resources can be recommended to the object more accurately.

[0135] In an optional embodiment, in step S1093-1, the multimedia resource recommendation process for the object based on the matching multimedia resources corresponding to each query result and the recall system may include:

[0136] S1093-11. Sort the matching multimedia resources corresponding to each query result in descending order according to the similarity matching results between the matching multimedia resources corresponding to each query result and each query result to obtain a multimedia resource sequence; determine the first preset number of multimedia resources in the multimedia resource sequence as the first target multimedia resource.

[0137] S1093-13. Based on at least one resource recall strategy in the recall system, recall the second target multimedia resource from the preset resource library.

[0138] S1093-15. Perform multimedia resource recommendation processing for the object based on the first target multimedia resource and the second target multimedia resource.

[0139] In this embodiment, on the one hand, as described in the above step S1093-11, the server can directly use a large pre-trained language model to model the historical interaction data sequence, learn the object representation in the language space from the historical interaction data sequence, and then generate multiple queries representing the object's interests, and finally match the first target multimedia resource that meets the actual intention of the object. For the details of the above step S1093-11, please refer to the above steps S1091-1 to S1091-5, and no specific limitation is made to this. On the other hand, as described in the above S1093-13, the server can filter out the second target multimedia resource from the preset resource library based on at least one resource recall strategy in the recall system. Then, as described in the above step S1093-15, the server can recommend multimedia resources to the object based on the first target multimedia resource and the second target multimedia resource.

[0140] Because a large pre-trained language model is used to model historical interaction data sequences and at least one resource recall strategy, multimedia resources to be recommended to the object are filtered from the preset resource library, so that multimedia resources can be recommended to the object from multiple aspects, ensuring that the object can view the multimedia resources of interest, thereby improving the recommendation accuracy and the object's viewing experience.

[0141] In one embodiment, in step S1093-15, the multimedia resource recommendation process for the object based on the first target multimedia resource and the second target multimedia resource may include:

[0142] The first target multimedia resource and the second target multimedia resource are filtered to obtain a third target multimedia resource; the third target multimedia resource is sorted to obtain a fourth target multimedia resource, and the fourth target multimedia resource is recommended to the object.

[0143] In this embodiment, the server can directly combine the first target multimedia resource obtained by modeling the historical interaction data sequence through a large pre-trained language model with the existing recall system, as part of the multi-round recall results, to participate in the processing of the recall system itself, such as deduplication, filtering, etc., and then enter the sorting system for sorting processing, so that the capabilities contained in the large language model and the existing recommendation system can be integrated and supplemented in the recall stage, so that the existing basic features of the large number of object dimensions and content dimensions actually used in the recommendation system can also be fully utilized, without the need for additional processing and modeling costs, reducing the resource recommendation cost and improving the recommendation performance of the resource recommendation system. Specifically, the first target multimedia resource and the second target multimedia resource can be fused, and then the fusion result can be deduplicated and filtered to obtain the third target multimedia resource; then the third target multimedia resource is sorted by the sorting system to obtain the fourth target multimedia resource, and the fourth target multimedia resource is recommended to the object.

[0144] In other implementations, the server may also obtain the union of the first target multimedia resource and the second target multimedia resource, perform deduplication and filtering on the union through a recall system, and then perform coarse sorting, fine sorting, and mixed sorting on the union through a sorting system to obtain multimedia resources recommended to the subject. This can improve the coverage of the recommended multimedia resources, allowing the subject to view a richer range of multimedia resources.

[0145] In other embodiments, the server may also obtain the intersection of the first target multimedia resource and the second target multimedia resource, perform deduplication and filtering on the intersection through a recall system, and then perform coarse sorting, fine sorting, and mixed sorting on the sorting system to obtain multimedia resources recommended for the object. In other words, the server recommends to the object the multimedia resources that are repeated in the first and second sets of candidate resources. This method ensures that the recommended multimedia resources not only conform to the recall strategy of the recall system but also conform to the recommendation criteria of the large language model, achieving the goal of combining the recall system with the recommendation of multimedia resources and ensuring the accuracy of the multimedia resources.

[0146] In another embodiment, in step S1093-15, the multimedia resource recommendation process for the object based on the first target multimedia resource and the second target multimedia resource may include:

[0147] The first target multimedia resource and the second target multimedia resource are sorted to obtain a fifth target multimedia resource, and the fifth target multimedia resource is recommended to the object.

[0148] In this embodiment, the server can use a large pre-trained language model to model the historical interaction data sequence to obtain the first target multimedia resource, without going through the recall system, and directly participate in the processing of the next ranking system together with the multimedia resources recalled by the recall system, so that the capabilities contained in the large language model and the recall results of the recall stage of the existing recommendation system can be integrated and supplemented, and the integrated and supplemented results can be directly input into the ranking system for sorting, so that the existing basic features of the large number of object dimensions and content dimensions actually used in the recommendation system can also be fully utilized, without the need for additional processing and modeling costs, reducing the resource recommendation cost and improving the recommendation performance of the resource recommendation system. Specifically, the first target multimedia resource and the second target multimedia resource can be integrated, and then the integrated result can be input into the next ranking system for sorting processing to obtain the fifth target multimedia resource, and the fifth target multimedia resource can be recommended to the object.

[0149] In other implementations, the server may also obtain the intersection or union of the first target multimedia resource and the second target multimedia resource, and input the intersection or union into the next ranking system for ranking to obtain multimedia resources recommended for the object.

[0150] Figure 9 This is a flowchart of a method for recommending multimedia resources according to an exemplary embodiment. Figure 6 ,like Figure 9 As shown, in an optional embodiment, after performing multimedia resource recommendation processing on the object according to the matching multimedia resources corresponding to each query result, the above method may further include:

[0151] S1011. Obtain a current interaction data sequence between the object and the recommended multimedia resource.

[0152] S1013. Determine forward interactive multimedia resources based on the current interactive data sequence; forward interactive multimedia resources are multimedia resources that the object is interested in.

[0153] S1015. Adjust the model parameters of the large language model based on the positive interaction multimedia resources and the historical interaction data sequence to reduce the difference between the object interest feature extraction results corresponding to the historical interaction data sequence and the positive interaction multimedia resources, thereby obtaining an adjusted large language model. The adjusted large language model is used to perform feature extraction processing on the historical interaction data sequence in the next round of recommendation.

[0154] In this embodiment, after multimedia resources are recommended, the subject can view the multimedia resources recommended by the large language model. The subject can then perform positive or negative interactions with the viewed multimedia resources to generate current positive interaction data or current negative interaction data. The server generates a current interaction data sequence based on the current positive interaction data or current negative interaction data and the corresponding interaction time. The server determines the positively interacted multimedia resources with which the subject has positively interacted based on the current interaction data sequence.

[0155] Then the server adjusts the model parameters of the large language model based on the forward interaction multimedia resources and the historical interaction data sequence to obtain an adjusted large language model. Specifically, it can be: inputting the forward interaction multimedia resources and the historical interaction data sequence into the large language model for feature extraction processing, obtaining the object interest feature extraction result corresponding to the historical interaction data sequence, calculating the difference between the object interest feature extraction result and the forward interaction multimedia resource, and continuously adjusting the model parameters of the large language model according to the difference until the difference meets the preset conditions, or the number of training times meets the preset conditions, to obtain the adjusted large language model. The adjusted large language model is used to perform feature extraction processing on the historical interaction data sequence in the next round of recommendation process. The embodiment of the present application adjusts the parameters of the large language model according to the actual interaction data fed back by the object, which can be regarded as an RLHF method.

[0156] Therefore, the parameters of the large language model are adjusted through actual positive interaction multimedia resources and historical interaction data sequences, so that the multimedia resources recommended by the large language model are more in line with the expectations of the object and closer to the resources that the object is interested in, thereby improving the recommendation performance of the large language model and recommending multimedia resources to the object more accurately.

[0157] The following is an overall description of the recommended methods for the above multimedia resources:

[0158] Figure 10 This is a flowchart of a method for recommending multimedia resources according to an exemplary embodiment. Figure 7 ,like Figure 10 As shown, the method for recommending multimedia resources may include:

[0159] 1) Provide metadata (including but not limited to title, multi-level classification and multi-level tags, author, release time, etc.) through the content understanding system.

[0160] 2) Generate a historical interaction data sequence based on different historical times and metadata corresponding to multimedia resources that interacted with the object at different historical times.

[0161] 3) Input the historical interaction data sequence into a large language model, which may include a 12-layer Transform architecture. The server uses the 12-layer Transform architecture to extract features of the implicit high-order behavior payload sequence of the historical interaction data sequence to better capture the fine-grained changes and intentions in the historical interaction data sequence, thereby learning the object representation of the language space from the historical interaction data sequence and obtaining the object interest feature extraction result.

[0162] 4) Perform beam search on the object interest feature extraction results to obtain a preset number of query results representing the object interest.

[0163] 5) Performing similarity matching on each query result with preset multimedia resources in the preset resource library, and determining matching multimedia resources corresponding to each query result from the preset resource library based on the similarity matching results.

[0164] 6) Recommend multimedia resources to the object based on the matching multimedia resources corresponding to each query result and the recall system. Specifically, this may include: sorting the matching multimedia resources corresponding to each query result in descending order based on the similarity matching results between the matching multimedia resources corresponding to each query result and each query result to obtain a multimedia resource sequence; determining the first preset number of multimedia resources in the multimedia resource sequence as first target multimedia resources; recalling second target multimedia resources from a preset resource library based on at least one resource recall strategy in the recall system; and recommending multimedia resources to the object based on the first target multimedia resources and the second target multimedia resources.

[0165] 7) Obtaining a current interaction data sequence between the object and the recommended multimedia resource; determining a positive interaction multimedia resource based on the current interaction data sequence; the positive interaction multimedia resource is a multimedia resource that the object is interested in; adjusting the model parameters of the large language model based on the positive interaction multimedia resource and the historical interaction data sequence to reduce the difference between the object interest feature extraction result corresponding to the historical interaction data sequence and the positive interaction multimedia resource, thereby obtaining an adjusted large language model.

[0166] Figure 11 is a schematic diagram of a system for recommending multimedia resources according to an exemplary embodiment. Figure 11 As shown, the multimedia resource recommendation system may include:

[0167] 1. Professionally Generated Content (PGC), User-generated Content (UGC), and Consumer End

[0168] (1) Content producers of PGC or UGC, Multi-Channel Network (MCN) or Professional User Generated Content (PUGC) provide local or filmed video content or written self-media articles or picture albums through the application programming interface (API) system of the mobile terminal or back-end interface. The authors can choose to actively upload the cover image of the corresponding content. These are the main sources of distributed content.

[0169] (2) Through communication with the upstream and downstream content interface services, first obtain the upload server interface address, and then upload the local file. During the shooting process, the local video content can choose matching music, filter templates and video beautification functions, etc.

[0170] (3) As a content consumption end, it communicates with the content distribution export server to obtain the index information of the corresponding content.

[0171] (4) Content consumers usually browse consumption data through feeds.

[0172] 2. Uplink and Downlink Content Interface Server

[0173] (1) Communicate directly with the content production end. The content submitted from the front end is usually the title of the content, publisher, summary, cover image, release time, keywords set by the user, etc., or the shot video directly enters the server through the server and stores the file in the video content storage service.

[0174] (2) Write metadata of the multimedia resource content, such as video file size, cover image link, bit rate, file format, title, release time, author, etc., into the content database.

[0175] (3) Submit the uploaded files and content metadata to the dispatch center service for subsequent content processing and circulation.

[0176] 3. Content Database

[0177] (1) The core database of content. The metadata of all content published by producers is stored in this business database.

[0178] (2) The dispatch center's content processing mainly includes machine processing and manual review processing.

[0179] 4. Dispatch Center Service

[0180] (1) Responsible for the entire scheduling process of video and graphic content flow, receiving content through the upstream and downstream content interface servers, and then obtaining the content metadata from the content metadata database.

[0181] (2) As the actual dispatch controller of the text and video links, according to the type of content, the content processing business service system is dispatched to process and review the corresponding content of the image content in the link, directly filter and mark the content with corresponding feature tags for downstream recommendation and distribution.

[0182] 5. Content Storage Service

[0183] (1) After obtaining the content index information, the terminal consumer can also directly access the video content storage server to download the corresponding content.

[0184] (2) In addition to being a data source for external services, it also serves as a data source for internal services, allowing the download file system to obtain raw video data for related processing. The paths of internal and external data sources are usually deployed separately to avoid mutual influence.

[0185] 6. Object Dimension Enhanced Recommendation Model

[0186] Convert historical interaction data sequences into text sequences, and fine-tune the pre-trained language model based on the historical interaction data sequences, including positive and negative interaction data. Positive interaction data is worthy of recognition and should be encouraged and supported by the recommendation system; negative interaction data should be avoided.

[0187] 7. Large Language Models

[0188] (1) It is not limited to a fixed large language model. Any model that uses massive and rich Internet basic corpus to construct a large-scale generative Transform architecture can be classified into this category.

[0189] 8. Object-Dimension Enhanced Recommendation Model and Service

[0190] (1) The object dimension enhanced recommendation model constructed above is transformed into a service, that is, recommendation services are added based on this model library.

[0191] (2) It also includes multiple auxiliary service modules that communicate with the dispatch center server to complete additional model generation based on this service, and at the same time complete the subsequent processing steps of the generated results, including Beam Searh multi-query generation results and retrieval matching with existing content in the content pool to finally obtain the required results.

[0192] 9. Historical Interaction Data Sequence Library

[0193] This historical interaction data sequence includes positive behaviors such as clicks, plays, reposts, and likes, as well as negative behaviors such as quick swipes, negative feedback, and reports. This information contains a large amount of actual intentions of the objects. Different interaction data sequences can well reflect the different intentions of the objects, and this basic data is used to train and upgrade large-scale language models.

[0194] Figure 12 is a block diagram of a multimedia resource recommendation device according to an exemplary embodiment. Figure 12 As shown, the multimedia resource recommendation device includes:

[0195] The historical interaction data sequence acquisition module 201 is used to acquire the historical interaction data sequence of the object; the historical interaction data sequence is used to represent the sequence formed by the data generated by the interaction between the object and the multimedia resources at different historical times;

[0196] Feature extraction module 203 is used to input the historical interaction data sequence into a large language model for feature extraction processing to obtain object interest feature extraction results; wherein the object interest feature extraction results are used to characterize the characteristics of the multimedia resources of interest to the object; the large language model is obtained by training an initial large language model based on the sample interaction data sequence and the object interest feature labels corresponding to the sample interaction data sequence;

[0197] A search query module 205 is configured to perform a search query on the object interest feature extraction result to obtain a preset number of query results representing the object interest;

[0198] A matching module 207 is configured to perform similarity matching on each query result with a preset multimedia resource in a preset resource library, and determine a matching multimedia resource corresponding to each query result from the preset resource library based on the similarity matching results;

[0199] The recommendation module 209 is configured to perform multimedia resource recommendation processing on the object according to the matching multimedia resources corresponding to each query result.

[0200] In an optional embodiment, the historical interaction data sequence acquisition module 201 includes:

[0201] A multimedia resource acquisition unit, configured to acquire multimodal multimedia resources;

[0202] A first metadata acquisition unit is configured to perform feature extraction processing on multimedia resources of each modality to obtain metadata corresponding to the multimedia resources of each modality;

[0203] A second metadata acquisition unit is configured to acquire metadata corresponding to multimedia resources that interacted with the object at different historical times from metadata corresponding to multimedia resources of each modality;

[0204] The sequence generating unit is configured to generate the historical interaction data sequence according to different historical times and metadata corresponding to multimedia resources that interacted with the object at different historical times.

[0205] In an optional embodiment, the search query module includes:

[0206] a beam width determination unit, configured to determine a beam width of a beam search;

[0207] a candidate output sequence generating unit, configured to perform a beam search on the object interest feature extraction result according to the beam width to obtain a candidate output sequence;

[0208] a repeating unit, configured to repeat the operation of performing a beam search on the object interest feature extraction result according to the beam width to obtain a candidate output sequence until an end symbol is obtained or the length of the text in the candidate output sequence exceeds a preset length threshold;

[0209] The query result determining unit is configured to determine the text in the candidate output sequence when the end symbol is obtained, or the text in the candidate output sequence when the length of the text is greater than a preset length threshold, as the query result.

[0210] In an optional embodiment, the matching module includes:

[0211] A vector extraction unit, configured to extract a query feature vector of each query result; and extract a resource feature vector of the preset multimedia resource;

[0212] a similarity determination unit, configured to perform vector similarity matching on the query feature vector of each query result and the resource feature vector, to obtain a similarity matching result between the query feature vector of each query result and the resource feature vector;

[0213] a target vector determining unit, configured to determine, from the resource feature vectors, a resource feature vector whose similarity matching result with the query feature vector of each query result satisfies a preset similarity threshold, and obtain a target resource feature vector for each query result;

[0214] The matching multimedia resource generating unit is configured to determine the preset multimedia resource corresponding to the target resource feature vector of each query result as the matching multimedia resource corresponding to each query result.

[0215] In an optional embodiment, the recommendation module includes:

[0216] a sorting unit, configured to sort the matching multimedia resources corresponding to each query result in descending order according to the similarity matching result between the matching multimedia resources corresponding to each query result and each query result, to obtain a multimedia resource sequence;

[0217] A first target multimedia resource determining unit, configured to determine a first preset number of multimedia resources in the multimedia resource sequence as first target multimedia resources;

[0218] The first recommendation subunit is configured to recommend the first target multimedia resource to the object.

[0219] In an optional embodiment, the recommendation module includes:

[0220] The fusion recommendation unit is used to perform multimedia resource recommendation processing on the object according to the matching multimedia resources corresponding to each query result and the recall system; wherein the recall system includes recommending multimedia resources for the object based on at least one resource recall strategy.

[0221] In an optional embodiment, the fusion recommendation unit includes:

[0222] The first recall subunit is configured to sort the matching multimedia resources corresponding to each query result in descending order based on the similarity matching results between the matching multimedia resources corresponding to each query result and each query result, to obtain a multimedia resource sequence; and determine the first preset number of multimedia resources in the multimedia resource sequence as first target multimedia resources;

[0223] a second recall subunit, configured to recall a second target multimedia resource from the preset resource library based on at least one resource recall strategy in the recall system;

[0224] The second recommendation subunit is configured to perform multimedia resource recommendation processing for the object based on the first target multimedia resource and the second target multimedia resource.

[0225] In an optional embodiment, the second recommendation sub-unit is further used to filter the first target multimedia resource and the second target multimedia resource to obtain a third target multimedia resource; sort the third target multimedia resource to obtain a fourth target multimedia resource, and recommend the fourth target multimedia resource to the object; or, sort the first target multimedia resource and the second target multimedia resource to obtain a fifth target multimedia resource, and recommend the fifth target multimedia resource to the object.

[0226] In an optional embodiment, the device further comprises:

[0227] a current interaction data sequence acquisition module, configured to acquire a current interaction data sequence between the object and the recommended multimedia resource;

[0228] a forward interaction multimedia resource determination module, configured to determine a forward interaction multimedia resource according to the current interaction data sequence; the forward interaction multimedia resource being a multimedia resource of interest to the subject;

[0229] A parameter adjustment module is used to adjust the model parameters of the large language model according to the positive interaction multimedia resource and the historical interaction data sequence to reduce the difference between the object interest feature extraction results corresponding to the historical interaction data sequence and the positive interaction multimedia resource, thereby obtaining an adjusted large language model; wherein the adjusted large language model is used to perform feature extraction processing on the historical interaction data sequence in the next round of recommendation process.

[0230] It should be noted that the device embodiment provided in the embodiments of the present application and the above-mentioned method embodiment are based on the same inventive concept.

[0231] An embodiment of the present application also provides an electronic device for recommending multimedia resources, which includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by the processor to implement a method for recommending multimedia resources as provided in any of the above embodiments.

[0232] An embodiment of the present application also provides a computer-readable storage medium, which can be set in a terminal to store at least one instruction or at least one program for implementing a method for recommending multimedia resources in a method embodiment. The at least one instruction or at least one program is loaded and executed by a processor to implement the method for recommending multimedia resources provided in the above method embodiment.

[0233] Optionally, in the embodiments of this specification, the storage medium may be located in at least one of the multiple network servers of the computer network. Optionally, in this embodiment, the storage medium may include, but is not limited to, various media that can store program code, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0234] The memory of the embodiment of this specification can be used to store software programs and modules. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory. The memory may mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, applications required for functions, etc.; the data storage area can store data created according to the use of the device, etc. In addition, the memory may include a high-speed random access memory and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage device. Accordingly, the memory may also include a memory controller to provide the processor with access to the memory.

[0235] The present application also provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the multimedia resource recommendation method provided in the above method embodiment.

[0236] The embodiment of the multimedia resource recommendation method provided in the embodiment of the present application can be executed in a terminal, a computer terminal, a server or a similar computing device. Taking running on a server as an example, Figure 13 FIG. 1 is a hardware structure diagram of a server provided according to an exemplary embodiment. Figure 13 As shown, the server 300 may have relatively large differences due to different configurations or performances, and may include one or more central processing units (CPUs) 310 (the central processing unit 310 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 330 for storing data, and one or more storage media 320 (such as one or more mass storage devices) for storing application programs 323 or data 322. Among them, the memory 330 and the storage medium 320 can be temporary storage or permanent storage. The program stored in the storage medium 320 may include one or more modules, each module may include a series of instruction operations on the server. Furthermore, the central processing unit 310 can be configured to communicate with the storage medium 320 to execute a series of instruction operations in the storage medium 320 on the server 300. The server 300 may also include one or more power supplies 360, one or more wired or wireless network interfaces 350, one or more input and output interfaces 340, and / or one or more operating systems 321, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0237] The input / output interface 340 can be used to receive or send data via a network. Specific examples of the aforementioned network may include a wireless network provided by the communication provider of the server 300. In one embodiment, the input / output interface 340 includes a network interface controller (NIC), which can be connected to other network devices via a base station to communicate with the Internet. In one embodiment, the input / output interface 340 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0238] It can be understood by those skilled in the art that Figure 13 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 13 More or fewer components than shown, or with Figure 13 Different configurations shown.

[0239] It should be noted that the order of the embodiments of the present application described above is for descriptive purposes only and does not represent the superiority or inferiority of the embodiments. The above description is of specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0240] The various embodiments in this specification are described in a progressive manner. Similar portions between the various embodiments can be referenced to each other. Each embodiment focuses on the differences between the other embodiments. In particular, the device and server embodiments are generally similar to the method embodiments, so their descriptions are relatively simple. For relevant portions, refer to the descriptions of the method embodiments.

[0241] Those skilled in the art will understand that all or part of the steps of implementing the above embodiments may be accomplished by hardware, or by programs instructing related hardware to accomplish the steps. The programs may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk, or an optical disk, etc.

[0242] The above are only preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application.

Claims

1. A method for recommending multimedia resources, characterized in that: The method comprises: Acquire a historical interaction data sequence of the object; the historical interaction data sequence is used to represent a sequence formed by data generated by the object interacting with multimedia resources at different historical times; Inputting the historical interaction data sequence into a large language model for feature extraction processing to obtain an object interest feature extraction result; wherein the object interest feature extraction result is used to characterize the characteristics of the multimedia resources of interest to the object; the large language model is obtained by training an initial large language model based on the sample interaction data sequence and the object interest feature labels corresponding to the sample interaction data sequence; Performing a search query on the object interest feature extraction result to obtain a preset number of query results representing the object interest; Performing similarity matching on each query result with preset multimedia resources in a preset resource library, and determining a matching multimedia resource corresponding to each query result from the preset resource library based on the similarity matching results; Multimedia resource recommendation processing is performed on the object according to the matching multimedia resources corresponding to each query result.

2. The method for recommending multimedia resources according to claim 1, wherein: The acquiring of the historical interaction data sequence of the object includes: Acquire multimodal multimedia resources; Perform feature extraction on multimedia resources of each modality to obtain metadata corresponding to the multimedia resources of each modality; Obtaining metadata corresponding to multimedia resources that interacted with the object at different historical times from metadata corresponding to multimedia resources of each modality; The historical interaction data sequence is generated according to different historical times and metadata corresponding to multimedia resources that interacted with the object at different historical times.

3. The method for recommending multimedia resources according to claim 1, wherein: The searching and querying the object interest feature extraction result to obtain a preset number of query results representing the object interest includes: Determine the beam width of the beam search; Performing a beam search on the object interest feature extraction result according to the beam width to obtain a candidate output sequence; Repeating the operation of performing a beam search on the object interest feature extraction result according to the beam width to obtain a candidate output sequence until an end symbol is obtained or the length of the text in the candidate output sequence is greater than a preset length threshold; The text in the candidate output sequence when the end symbol is obtained, or the text in the candidate output sequence when the length of the text is greater than a preset length threshold, is determined as the query result.

4. The method for recommending multimedia resources according to claim 1, wherein: The performing similarity matching on each query result with a preset multimedia resource in a preset resource library, and determining a matching multimedia resource corresponding to each query result from the preset resource library based on the similarity matching results, includes: Extracting a query feature vector of each query result; and extracting a resource feature vector of the preset multimedia resource; Performing vector similarity matching on the query feature vector of each query result and the resource feature vector to obtain a similarity matching result between the query feature vector of each query result and the resource feature vector; Determine from the resource feature vectors a resource feature vector whose similarity matching result with the query feature vector of each query result satisfies a preset similarity threshold, and obtain a target resource feature vector for each query result; The preset multimedia resource corresponding to the target resource feature vector of each query result is determined as the matching multimedia resource corresponding to each query result.

5. The method for recommending multimedia resources according to any one of claims 1 to 4, characterized in that: The performing multimedia resource recommendation processing on the object according to the matching multimedia resources corresponding to each query result includes: According to the similarity matching results between the matching multimedia resources corresponding to each query result and each query result, the matching multimedia resources corresponding to each query result are sorted in descending order to obtain a multimedia resource sequence; Determining a first preset number of multimedia resources in the multimedia resource sequence as first target multimedia resources; The first target multimedia resource is recommended to the object.

6. The method for recommending multimedia resources according to any one of claims 1 to 4, characterized in that: The performing multimedia resource recommendation processing on the object according to the matching multimedia resources corresponding to each query result includes: According to the matching multimedia resources and the recall system corresponding to each query result, multimedia resource recommendation processing is performed on the object; The recall system includes recommending multimedia resources to the object based on at least one resource recall strategy.

7. The method for recommending multimedia resources according to claim 6, characterized in that: The multimedia resource recommendation process for the object according to the matching multimedia resources corresponding to each query result and the recall system includes: sorting the matching multimedia resources corresponding to each query result in descending order based on similarity matching results between the matching multimedia resources corresponding to each query result and each query result to obtain a multimedia resource sequence; determining the first preset number of multimedia resources in the multimedia resource sequence as first target multimedia resources; Recalling a second target multimedia resource from the preset resource library based on at least one resource recall strategy in the recall system; Perform multimedia resource recommendation processing for the object according to the first target multimedia resource and the second target multimedia resource.

8. The method for recommending multimedia resources according to claim 7, characterized in that: The performing multimedia resource recommendation processing for the object according to the first target multimedia resource and the second target multimedia resource includes: Filtering the first target multimedia resource and the second target multimedia resource to obtain a third target multimedia resource; sorting the third target multimedia resource to obtain a fourth target multimedia resource, and recommending the fourth target multimedia resource to the object; or The first target multimedia resource and the second target multimedia resource are sorted to obtain a fifth target multimedia resource, and the fifth target multimedia resource is recommended to the object.

9. The method for recommending multimedia resources according to any one of claims 1 to 4, characterized in that: After performing multimedia resource recommendation processing on the object according to the matching multimedia resources corresponding to each query result, the method further includes: Obtaining a current interaction data sequence between the object and the recommended multimedia resource; Determining a forward interactive multimedia resource according to the current interactive data sequence; the forward interactive multimedia resource is a multimedia resource that the object is interested in; Adjusting model parameters of the large language model according to the forward interactive multimedia resource and the historical interactive data sequence to reduce the difference between the object interest feature extraction result corresponding to the historical interactive data sequence and the forward interactive multimedia resource, thereby obtaining an adjusted large language model; The adjusted large language model is used to perform feature extraction processing on the historical interaction data sequence in the next round of recommendation process.

10. A multimedia resource recommendation device, characterized in that: The device comprises: A historical interaction data sequence acquisition module is used to acquire a historical interaction data sequence of an object; the historical interaction data sequence is used to represent a sequence formed by data generated by the object interacting with multimedia resources at different historical times; a feature extraction module configured to input the historical interaction data sequence into a large language model for feature extraction processing to obtain an object interest feature extraction result; wherein the object interest feature extraction result is used to characterize the characteristics of the multimedia resources of interest to the object; and the large language model is obtained by training an initial large language model based on the sample interaction data sequence and the object interest feature labels corresponding to the sample interaction data sequence; A search query module, configured to perform a search query on the object interest feature extraction result to obtain a preset number of query results representing the object interest; a matching module, configured to perform similarity matching on each query result with preset multimedia resources in a preset resource library, and determine a matching multimedia resource corresponding to each query result from the preset resource library based on the similarity matching results; The recommendation module is used to perform multimedia resource recommendation processing on the object according to the matching multimedia resources corresponding to each query result.

11. An electronic device for recommending multimedia resources, characterized in that: The electronic device includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the processor loads the at least one instruction or the at least one program to execute the multimedia resource recommendation method according to any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by a processor to implement the multimedia resource recommendation method according to any one of claims 1 to 9.

Citation Information

Cited By

  • Resource content understanding and resource recommendation method and device, and electronic equipment

    CN121210759A