Resource processing method, resource processing device, electronic device and storage medium
By generating a recommended resource set based on video data, the problem of users having to exit the page to search for extended content while watching a video is solved, and a solution for uninterrupted access to extended content is implemented, which improves user experience and resource diversity.
Patent Information
- Application Number
- CN202111103710.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-18
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2041-09-18
AI Technical Summary
When users are watching a video, if they want to know the extended content, they need to exit the video page to search. The operation is complicated and interrupts the playback, affecting the user experience and is not conducive to stimulating demand for diversified products.
Through the video data of the target video, the extension object is determined, a recommended resource set is generated, and the recommended resources are obtained in response to user selection operations, simplifying the process of obtaining extended content and avoiding interruption of video playback.
It enables users to obtain extended video content without searching, improves user experience, enhances the diversity of recommended resources, and stimulates demand for diversified products.
Smart Images

Figure CN113849688B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing, and in particular to intelligent search, cloud computing, computer vision, and deep learning technologies. Specifically, it relates to a resource processing method, a resource processing device, an electronic device, and a storage medium. Background Art
[0002] With the development of terminal devices and Internet technologies, terminal devices have become increasingly versatile, and users can access resources through various functions provided by terminal devices.
[0003] The terminal device may have a video playback function so that the user can watch the video on the terminal device to obtain resources. Summary of the Invention
[0004] The present disclosure provides a resource processing method, a resource processing device, an electronic device, and a storage medium.
[0005] According to one aspect of the present disclosure, a resource processing method is provided, comprising: determining at least one extended object based on video data of a target video, wherein each extended object represents text information of extended content related to the target video; determining a recommended resource set based on the at least one extended object, wherein the recommended resource set includes at least one recommended resource associated with the extended object; and displaying the recommended resource set so as to obtain the recommended resource in response to a selection operation on the recommended resource.
[0006] According to another aspect of the present disclosure, a resource processing device is provided, comprising: a first determination module, configured to determine at least one extended object based on video data of a target video, wherein each extended object represents text information of extended content related to the target video; a second determination module, configured to determine a recommended resource set based on the at least one extended object, wherein the recommended resource set includes at least one recommended resource associated with the extended object; and a display module, configured to display the recommended resource set so that the recommended resource can be obtained in response to a selection operation on the recommended resource.
[0007] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described above.
[0008] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method described above.
[0009] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program, which implements the method described above when executed by a processor.
[0010] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0012] Figure 1 The following schematically illustrates an exemplary system architecture of a resource processing method and apparatus according to an embodiment of the present disclosure;
[0013] Figure 2 The following schematically shows a flow chart of a resource processing method according to an embodiment of the present disclosure;
[0014] Figure 3A The following schematically illustrates an example of a resource processing method according to an embodiment of the present disclosure;
[0015] Figure 3B An example schematic diagram of a resource processing method according to another embodiment of the present disclosure is schematically shown;
[0016] Figure 4 An example schematic diagram of a resource recommendation page according to an embodiment of the present disclosure is schematically shown;
[0017] Figure 5 A block diagram schematically shows a resource processing device according to an embodiment of the present disclosure; and
[0018] Figure 6 A block diagram schematically shows an electronic device suitable for implementing a resource processing method according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0019] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0020] With the development of terminal devices and Internet technology, more and more users of terminal devices like to watch videos on terminal devices, such as watching short videos on terminal devices.
[0021] If a user wants to access extended content while watching a target video, they must exit the target video page and search for the extended content. This process is complex and interrupts the playback of the target video, affecting the user experience and hindering their potential demand for diversified products.
[0022] The disclosed embodiments provide a method for generating a set of recommended resources related to extended content of a target video. Specifically, at least one extended object is determined based on the video data of the target video. Each extended object represents textual information related to the extended content of the target video. A set of recommended resources is determined based on the at least one extended object. The recommended resource set includes at least one recommended resource associated with the extended object, and the recommended resource set is displayed so that the recommended resource can be retrieved in response to a selection operation on the recommended resource.
[0023] By determining at least one extended object related to the extended content of the target video based on the video data of the target video, determining a recommended resource set based on the at least one extended object, and displaying the recommended resource set so that in response to the selection operation for the recommended resource, the extended content of the target video can be obtained without the user searching, thereby simplifying the operation of obtaining the extended content. In addition, the extended content related to the target video can be obtained without exiting the target video. Therefore, the playback of the target video will not be interrupted, thereby improving the user experience. The recommended resource set related to the extended content of the target video can be obtained, thereby improving the diversity of the recommendation types and stimulating the user's potential demand for diversified products.
[0024] Figure 1 An exemplary system architecture of a resource processing method and apparatus according to an embodiment of the present disclosure is schematically shown.
[0025] It should be noted that Figure 1 The examples shown are merely examples of system architectures to which the embodiments of the present disclosure may be applied, to help those skilled in the art understand the technical content of the present disclosure. This does not mean that the embodiments of the present disclosure cannot be applied to other devices, systems, environments, or scenarios. For example, in another embodiment, an exemplary system architecture to which the resource processing method and apparatus may be applied may include a terminal device, but the terminal device may implement the resource processing method and apparatus provided by the embodiments of the present disclosure without interacting with a server.
[0026] like Figure 1As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used as a medium for providing communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links.
[0027] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as short video applications, knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software (only as examples).
[0028] The terminal devices 101 , 102 , and 103 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, and desktop computers.
[0029] The server 105 may be any type of server that provides various services, such as a background management server (for example only) that supports the content browsed by users using the terminal devices 101, 102, and 103. The background management server may analyze and process received data such as user requests, and feed back processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.
[0030] Server 105 can be a cloud server, also known as a cloud computing server or cloud host. It is a host product in the cloud computing service system. It solves the management difficulties and poor business scalability of traditional physical hosts and VPS services. Server 105 can also be a server in a distributed system or a server integrated with blockchain.
[0031] For example, server 105 may determine at least one extended object based on the video data of a target video. Each extended object represents textual information related to extended content of the target video. A recommended resource set is determined based on the at least one extended object. The recommended resource set includes at least one recommended resource associated with the extended object, and the recommended resource set is displayed so that the recommended resource can be obtained in response to a selection operation on the recommended resource.
[0032] It should be noted that the content processing method provided in the embodiment of the present disclosure can generally be executed by the server 105. Accordingly, the resource processing device provided in the embodiment of the present disclosure can also be set in the server 105.
[0033] Alternatively, the content processing method provided by the embodiment of the present disclosure may also be generally executed by the terminal devices 101, 102, 103. Accordingly, the resource processing apparatus provided by the embodiment of the present disclosure may generally be provided in the terminal devices 101, 102, 103. The resource processing method provided by the embodiment of the present disclosure may also be performed by the terminal devices 101, 102, 103 that are different from the terminal devices 101, 102, 103 and that are capable of communicating with the terminal devices 101, 102, 103 and / or the server 105. Accordingly, the resource processing apparatus provided by the embodiment of the present disclosure may also be provided in 101, 102, 103 that are different from the terminal devices 101, 102, 103 and / or the server 105.
[0034] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0035] Figure 2 The flowchart of the resource processing method according to the embodiment of the present disclosure is schematically shown.
[0036] like Figure 2 As shown, the method 200 includes operations S210 to S230.
[0037] In operation S210, at least one extended object is determined based on video data of a target video, wherein each extended object represents text information of extended content related to the target video.
[0038] In operation S220, a recommended resource set is determined according to at least one extended object. The recommended resource set includes at least one recommended resource associated with the extended object.
[0039] In operation S230 , a recommended resource set is displayed so that the recommended resource can be obtained in response to a selection operation on the recommended resource.
[0040] According to an embodiment of the present disclosure, the target video may be the video currently being watched by the user. The user may watch the target video through a video playback application installed on a terminal device, or may watch the target video on a web page through a video playback plug-in provided in a browser. The target video type may include short videos, medium videos, or long videos. Extended content related to the target video may refer to content related to the target video but not the target video itself. An extended object may represent text information of the extended content related to the target video, and the text information may include words or phrases. The extended object may refer to an extended phrase.
[0041] For example, based on the video data of the target video, it is determined that the target video is a video of a tabby cat playing with toys. The extended content related to the target video can be content about how a tabby cat can lose weight. The extended objects can include tabby cats and weight loss.
[0042] According to an embodiment of the present disclosure, a recommended resource set may include at least one recommended resource. Each recommended resource may include a recommended graphic resource or a recommended video resource. The graphic resource may include at least one of an encyclopedia resource and a commentary resource. The video resource may include at least one of a short video resource, a medium video resource, or a long video resource. The recommended graphic resource may include at least one of a recommended encyclopedia resource and a commentary resource. The recommended video resource may include a recommended short video resource. The selection operation may include a touch operation or a voice control command operation. The touch operation may include a click operation or a slide operation.
[0043] According to an embodiment of the present disclosure, video data of a target video can be obtained. The video data can be obtained in response to a user's search operation or can be actively pushed, and this is not limited by the embodiment of the present disclosure. After obtaining the video data of the target video, at least one extended object can be determined based on the video data of the target video. For example, at least one core object can be determined based on the video data of the target video, and then at least one extended object can be determined based on the at least one core object. Alternatively, instead of utilizing a core object, a second target object set can be determined directly from at least one third candidate object based on a video vector corresponding to the video data of the target video and a candidate text vector corresponding to each third candidate object in at least one third candidate object. The second target object set includes at least one second target object, and each second target object is determined as each extended object. The core object can represent text information corresponding to the target video.
[0044] According to an embodiment of the present disclosure, after determining at least one extended object, a recommended resource set can be determined based on the at least one extended object. For example, N extended objects can be selected from M extended objects, and for each of the N extended objects, a recommended resource set corresponding to the extended object can be determined from the candidate resource set, and based on the recommended resource set corresponding to each of the N extended objects, a recommended resource set corresponding to the target video can be determined. M is an integer greater than or equal to 1, and N is an integer greater than or equal to 1 and less than or equal to M. Alternatively, T extended objects can be selected from S extended objects, and the T extended objects can be fused to obtain a fused extended object, and a recommended resource set corresponding to the fused extended object can be determined from the candidate resource set based on the fused extended object. The recommended resource set corresponding to the fused extended object is determined as the recommended resource set corresponding to the target video. S is an integer greater than or equal to 1, and T is an integer greater than or equal to 1 and less than or equal to S.
[0045] According to an embodiment of the present disclosure, after determining the recommended resource set, the recommended resource set can be displayed so that the selected recommended resource can be obtained in response to a selection operation for the recommended resource included in the recommended resource set. The recommended resource can be represented by an extended object. Displaying the recommended resource set may include: displaying the recommended resource set on a preset page. The preset page may be a page related to the target video. For example, the preset page may be a page where the target video is located. The preset page may include a first area and a second area. The first area may be an area for displaying the target video. The second area may be all areas or other partial areas of the preset page except the first area. The target video can be displayed in the first area, and the recommended resource set can be displayed in the second area. Thus, the display of the target video can be achieved without affecting the display of the target video when the recommendation of the recommended resource set is achieved. In addition, the above-mentioned display method of the recommended resource set is only an exemplary embodiment, but is not limited thereto. It can also include display methods known in the art, as long as the display of the recommended resource set can be achieved.
[0046] According to an embodiment of the present disclosure, by determining at least one extended object related to the extended content of the target video based on the video data of the target video, determining a recommended resource set based on the at least one extended object, and displaying the recommended resource set so as to respond to a selection operation for the recommended resource, it is possible to obtain the extended content of the target video without the user searching, thereby simplifying the operation of obtaining the extended content. In addition, the extended content related to the target video can be obtained without exiting the target video. Therefore, the playback of the target video will not be interrupted, thereby improving the user experience. The recommended resource set related to the extended content of the target video can be obtained, thereby improving the diversity of the recommendation types, which is conducive to stimulating the user's potential demand for diversified products.
[0047] According to an embodiment of the present disclosure, operation S210 may include the following operations.
[0048] At least one core object is determined based on the video data of the target video. Each core object represents text information corresponding to the target video. At least one extended object is determined based on the at least one core object.
[0049] According to embodiments of the present disclosure, a core object may be used to represent text information corresponding to a target video, and the text information may include characters or words. A core object may refer to a core term. For example, a core object may include "Chang'e Flying to the Moon." Extended objects may include at least one of "rabbit," "moon," "Mid-Autumn Festival," and "Chang'e."
[0050] According to an embodiment of the present disclosure, after obtaining the video data of a target video, the video data may be classified to determine a category label corresponding to the target video, where the number of category labels may include one or more. At least one core object is determined based on the one or more category labels.
[0051] According to an embodiment of the present disclosure, after determining at least one core object, at least one extended object may be determined based on a preset selection strategy, based on a candidate information set and the at least one core object. The preset selection strategy may be a strategy for determining the at least one extended object based on the candidate information set and the at least one core object. The preset selection strategy may include how to determine the content of the at least one extended object based on the candidate information set and the at least one core object.
[0052] According to an embodiment of the present disclosure, determining at least one core object based on video data of a target video may include the following operations.
[0053] At least one category label is determined based on the video data of the target video, and at least one core object is determined based on the at least one category label.
[0054] According to an embodiment of the present disclosure, category labels can be used to characterize categories of video data. The video data of the target video can be processed using a video classification model to determine at least one category label. The video classification model can be a model trained based on a deep learning model. Based on at least one category label, determining at least one core object may include: determining each category label as each core object. Alternatively, at least one target category label can be determined from one or more category labels based on a preset label selection strategy, and each target category label can be determined as each core object. The preset label selection strategy can be a strategy for determining at least one target category label from one or more category labels. The preset selection strategy may include how to determine the content of at least one target category label from one or more category labels.
[0055] According to an embodiment of the present disclosure, determining at least one core object according to at least one category label may include the following operations.
[0056] The at least one category label is processed using data to determine at least one candidate category label. The data processing includes at least one of the following: entity word filtering, semantic disambiguation, and abnormal word filtering. At least one core object is determined based on the at least one candidate category label.
[0057] According to an embodiment of the present disclosure, entity word filtering can be used to filter out category labels that do not belong to entities, and retain category labels that belong to entities. Semantic disambiguation can be used to determine the semantics related to the video data of the target video for category labels with multiple semantics. For example, if the category label is apple, apple is a word with multiple semantics. It may refer to apples among fruits, or it may refer to the brand of smartphones. In this case, semantic disambiguation can be used to process and determine the semantics related to the target video. Abnormal word filtering can be used to filter out category labels that do not conform to public order and good morals, and retain category labels that conform to public order and good morals.
[0058] According to embodiments of the present disclosure, at least one candidate category label can be determined from at least one category label using at least one of entity word filtering, semantic disambiguation, and abnormal word filtering. This improves the accuracy of the core object subsequently determined based on the at least one candidate category label.
[0059] According to an embodiment of the present disclosure, determining at least one core object according to at least one candidate category label may include the following operations.
[0060] The at least one candidate category label is sorted according to the inverse document frequency of each candidate category label. At least one target category label is determined from the at least one candidate category label based on the sorting result. Each target category label is determined as each core object.
[0061] According to an embodiment of the present disclosure, Term Frequency-Inverse Document Frequency (TF-IDF) is a technology used for information retrieval and data mining. The inverse document frequency can characterize the importance of a term. If the number of the inverse document frequency of a term is larger, it can be said that the category distinction ability of the term is better. The inverse document frequency of a term can be determined by taking the logarithm of the ratio of the number of total documents to the number of documents including the term. In an embodiment of the present disclosure, a term can refer to a candidate category label.
[0062] According to an embodiment of the present disclosure, sorting at least one candidate category tag according to the inverse document frequency of each candidate category tag in at least one candidate category tag may include: determining the inverse document frequency of each candidate category tag in at least one candidate category tag under a preset category. Sorting at least one candidate category tag according to the inverse document frequency of each candidate category tag in at least one candidate category tag to obtain a sorting result. Sorting may include sorting in ascending order of inverse document frequency or sorting in descending order of inverse document frequency. Determining at least one target category tag from at least one candidate category tag according to the sorting result may include: if the sorting is in ascending order of inverse document frequency, a preset number of candidate category tags with a lower ranking may be determined as the target category tag. If the sorting is in descending order of inverse document frequency, a preset number of candidate category tags with a higher ranking may be determined as the target category tag. The preset number may be configured according to actual business needs and is not limited here. For example, the preset number may be set to 2.
[0063] According to an embodiment of the present disclosure, determining at least one core object based on at least one candidate category label may include: determining at least one target category label from at least one candidate category label based on a preset inverse document frequency threshold and the inverse document frequency of each candidate category label in at least one candidate category label, and determining each target category label as each core object.
[0064] According to embodiments of the present disclosure, a preset inverse document frequency threshold can be used as one of the bases for determining whether a candidate category label is a target category label. The preset inverse document frequency threshold can be configured based on actual business needs and is not limited here. For example, the preset inverse document frequency threshold can be set to 0.8.
[0065] According to an embodiment of the present disclosure, determining at least one target category label from at least one candidate category label based on a preset inverse document frequency threshold and the inverse document frequency of each candidate category label in at least one candidate category label may include: for each candidate category label in at least one candidate category label, when it is determined that the inverse document frequency of the candidate category label is greater than or equal to the preset inverse document frequency threshold, determining the candidate category label as the target category label.
[0066] According to an embodiment of the present disclosure, determining at least one extended object according to at least one core object may include the following operations.
[0067] Based on a preset selection strategy, at least one extended object is determined according to the candidate information set and at least one core object.
[0068] According to an embodiment of the present disclosure, a preset selection strategy may include how to determine the content of at least one extended object based on a candidate information set and at least one core object. The candidate information set may include information related to determining the core object. For example, the candidate information set may include a candidate object set. Alternatively, the candidate information set may include a candidate association relationship set.
[0069] According to an embodiment of the present disclosure, the candidate information set includes a candidate object set, and the candidate object set includes at least one first candidate object.
[0070] According to an embodiment of the present disclosure, determining at least one extended object based on a preset selection strategy and according to a candidate information set and at least one core object may include the following operations.
[0071] A core text vector corresponding to at least one core object is determined. A first target object set is determined from a candidate object set based on the core text vector and a candidate text vector corresponding to each first candidate object in the at least one first candidate object. The first target object set includes at least one first target object. Each first target object is determined as an extended object.
[0072] According to an embodiment of the present disclosure, each first candidate object has a candidate text vector corresponding to the first candidate object. Vector processing can be performed on at least one core object to obtain a core text vector corresponding to the at least one core object. The first candidate object set can be processed based on a GCF (Graph Collaborative Filtering Fecommendation) model to obtain a candidate text vector corresponding to each first candidate object in the first candidate object set. The GCF model can be trained based on a search query (i.e., query).
[0073] According to an embodiment of the present disclosure, after determining the core text vector, the respective similarity values between at least one verification object and each first candidate object can be determined based on the core text vector and the candidate text vector corresponding to each first candidate object in at least one first candidate object. Based on the respective similarity values between at least one core object and each first candidate object, a first target object set is determined from the first candidate object set. The similarity value characterizes the degree of similarity between two objects. The similarity value can be configured according to actual business needs and is not limited here. For example, the similarity value may include cosine similarity, Pearson correlation coefficient, Euclidean distance, or Jaccard distance.
[0074] According to an embodiment of the present disclosure, determining a first target object set from a candidate object set based on a core text vector and a candidate text vector corresponding to each first candidate object of at least one first candidate object may include the following operations.
[0075] An approximate nearest neighbor algorithm is used to determine a first target object set from the candidate object set according to the core text vector and the candidate text vector corresponding to each first candidate object of the at least one first candidate object.
[0076] According to an embodiment of the present disclosure, an approximate nearest neighbor (ANN) algorithm refers to an approximate algorithm for solving the nearest neighbor search problem. The approximate nearest neighbor algorithm may include a tree-based approximate nearest neighbor algorithm, a hash-based approximate nearest neighbor algorithm, a vector quantization-based approximate nearest neighbor algorithm, or a neighbor graph-based approximate nearest neighbor algorithm. The tree may include a KD-Tree, a Ball-Tree, or an Annoy. The approximate nearest neighbor algorithm for processing the core text vector corresponding to at least one core object and the candidate text vector of each first candidate object can be determined according to actual business needs, and is not limited here.
[0077] According to an embodiment of the present disclosure, the core text vector corresponding to at least one core object and the candidate text vector of each first candidate object are processed using an approximate nearest neighbor algorithm, which can improve determination efficiency.
[0078] According to an embodiment of the present disclosure, determining a first target object set from a candidate object set based on a core text vector and a candidate text vector corresponding to each first candidate object in at least one first candidate object using an approximate nearest neighbor algorithm may include the following operations.
[0079] Based on the core text vector corresponding to at least one verification object and the candidate text vector of each first candidate object in the first candidate object set, a respective first similarity value between the core text vector corresponding to at least one core object and each first candidate object is determined. Utilizing an approximate nearest neighbor algorithm based on a hash, a first target object set is determined from the first candidate object set according to a target object set determination condition. The target object set determination condition includes a core text vector corresponding to at least one core object, a candidate text vector of each first candidate object in the first candidate object set, and a respective first similarity value between the core text vector corresponding to at least one core object and each first candidate object.
[0080] According to an embodiment of the present disclosure, determining the first target object set from the first candidate object set based on the target object set determination condition may include: determining a first initial candidate object set from the first candidate object set based on the target object set determination condition. Determining the first target object set from the first initial candidate object set based on a first similarity value. Determining the first target object set from the first initial candidate object set based on the first similarity value may include: sorting each first similarity value to obtain a sorting result, and determining the first target object set from the first initial candidate object set based on the sorting result. Alternatively, determining the first target object set from the first initial candidate object set based on a first preset similarity threshold and the first similarity value.
[0081] According to an embodiment of the present disclosure, the first preset similarity threshold can be used as one of the bases for determining the first target object set from the first initial candidate object set. The first preset similarity threshold can be configured based on actual business needs and is not limited here. For example, the first preset similarity threshold can be set to 0.6.
[0082] According to an embodiment of the present disclosure, determining a first initial candidate object set from a first candidate object set based on a target object set determination condition may include: using a target hash function to convert a core text vector corresponding to at least one core object into a hash vector corresponding to at least one core object. Using a target hash function to convert the candidate text vector of each first candidate object in the first candidate object set into a hash vector of the first candidate object. The first initial candidate object set may include at least one first initial candidate object, which is a candidate object having a hash vector consistent with a hash vector corresponding to at least one core object.
[0083] According to an embodiment of the present disclosure, determining a first target object set from a candidate object set based on a core text vector and a candidate text vector corresponding to each first candidate object in at least one first candidate object using an approximate nearest neighbor algorithm may include the following operations.
[0084] Using an approximate nearest neighbor algorithm based on vector quantization, at least one cluster center corresponding to the first candidate object set is determined based on the candidate text vector of each first candidate object in the first candidate object set. Based on the core text vector corresponding to at least one core object and the candidate text vector of each cluster center, a second similarity value is determined between the at least one core object and each cluster center. Based on the second similarity values, a first target object set is determined from the first candidate object set.
[0085] According to an embodiment of the present disclosure, a clustering algorithm may be used to process the candidate text vector of each first candidate object in the first candidate object set to obtain at least one cluster center in the first candidate object set. The clustering algorithm may include a K-means clustering algorithm, a K-center clustering algorithm, a CLARA (Clustering Large Application) algorithm, or a fuzzy C-means algorithm.
[0086] According to an embodiment of the present disclosure, determining the first target object set from the first candidate object set based on the second similarity value may include: sorting the second similarity values to obtain a sorting result, and determining the first target object set from the first candidate object set based on the sorting result. Alternatively, the first target object set may be determined from the first candidate object set based on a second preset similarity threshold and the second similarity value.
[0087] According to an embodiment of the present disclosure, determining the first target object set from the first candidate object set based on the second similarity value may include: determining a second initial candidate object set from the first candidate object set based on the second similarity value. Determining a respective third similarity value between at least one core object and each second initial candidate object based on a core text vector corresponding to at least one core object and a candidate text vector of each second initial candidate object in the second initial candidate object set. Determining the first target object set from the second initial candidate object set based on a second preset similarity threshold and the third similarity value.
[0088] According to an embodiment of the present disclosure, the second preset similarity threshold can be used as one of the bases for determining the first target object set from the second initial candidate object set. The second preset similarity threshold can be configured according to actual business needs and is not limited here. For example, the second preset similarity threshold can be set to 0.6.
[0089] According to the second similarity value, the second initial candidate object set is determined from the first candidate object set. The target cluster center can be determined based on the respective second similarity values between at least one core object and each cluster center, and the first candidate object set corresponding to the target cluster center is determined as the second initial candidate object set.
[0090] According to an embodiment of the present disclosure, the candidate information set includes a candidate association relationship set, the candidate association relationship set includes at least one candidate association relationship, and each candidate association relationship represents an association relationship between an index object and a second candidate object.
[0091] According to an embodiment of the present disclosure, determining at least one extended object based on a preset selection strategy and according to a candidate information set and at least one core object may include the following operations.
[0092] For each core object in the at least one core object, an index object matching the core object is searched from the candidate association relationship set, and a second candidate object corresponding to the index object is determined as an extended object corresponding to the core object.
[0093] According to an embodiment of the present disclosure, the index object may be an object related to the second candidate object. The index object consistent with the core object may be searched from the candidate association relationship set in a retrieval manner.
[0094] According to an embodiment of the present disclosure, determining at least one extended object according to video data of a target video may include the following operations.
[0095] A video vector corresponding to the video data of the target video is determined. A second target object set is determined from the at least one third candidate object based on the video vector and the candidate text vector corresponding to each third candidate object in the at least one third candidate object. The second target object set includes at least one second target object. Each second target object is determined as an extended object.
[0096] According to an embodiment of the present disclosure, determining a video vector corresponding to video data of a target video may include: processing the video data of the target video using a multimodal model to obtain a video vector corresponding to the video data of the target video. The multimodal model may include an Ernie-Vil model. The Ernie-Vil model is a knowledge-enhanced visual-language pre-training model. The Ernie-Vil model incorporates scene graph knowledge into the multimodal training process and enhances cross-modal semantic understanding capabilities by learning a joint representation of scene semantics.
[0097] According to an embodiment of the present disclosure, determining a second target object set from at least one third candidate object based on a video vector and a candidate text vector corresponding to each third candidate object in at least one third candidate object may include: using an approximate nearest neighbor algorithm to determine a second target object set from at least one third candidate object based on a video vector and a candidate text vector corresponding to each third candidate object in at least one third candidate object.
[0098] According to an embodiment of the present disclosure, determining a recommended resource set according to at least one extended object may include the following operations.
[0099] For each of the at least one extended object, a recommended resource set corresponding to the extended object is determined from the candidate resource set. A recommended resource set is determined based on the recommended resource sets corresponding to each of the at least one extended object.
[0100] According to an embodiment of the present disclosure, a candidate resource set may include at least one candidate resource. The type of the candidate resource may include a graphic resource or a video resource, that is, the candidate resource may include a candidate graphic resource or a candidate video resource. A recommended resource set corresponding to an extended object may include at least one recommended resource. The type of the recommended resource may include a graphic resource or a video resource, that is, the recommended resource may include a recommended graphic resource or a recommended video resource.
[0101] According to an embodiment of the present disclosure, the candidate resource set includes at least one of the following: a candidate graphic resource subset and a candidate video resource subset.
[0102] The recommended resource set corresponding to the extended object includes at least one of the following: a recommended graphic and text resource subset and a recommended video resource subset corresponding to the extended object.
[0103] According to an embodiment of the present disclosure, the candidate graphic resource subset may include at least one candidate graphic resource, the candidate video resource subset may include at least one candidate video resource, and the candidate graphic resource subset may include at least one candidate graphic resource.
[0104] According to an embodiment of the present disclosure, a recommended graphic and text resource subset corresponding to an extended object may be determined in the following manner.
[0105] According to the extension object and the index object corresponding to each candidate graphic resource in the candidate graphic resource subset, a target graphic resource subset is determined from the candidate graphic resource subset, and the target graphic resource subset is determined as the recommended graphic resource subset corresponding to the extension object.
[0106] According to an embodiment of the present disclosure, the index object corresponding to each candidate graphic resource can represent the candidate graphic resource. There is a set of index objects corresponding to a subset of candidate graphic resources.
[0107] According to an embodiment of the present disclosure, for an extended object, an index object matching the extended object can be searched from index objects corresponding to a subset of candidate graphic resources, and the candidate graphic resource corresponding to the index object matching the extended object can be determined as the target graphic resource.
[0108] According to an embodiment of the present disclosure, a recommended video resource subset corresponding to an extended object may be determined in the following manner.
[0109] Determine an extended text vector corresponding to the extended object. Determine a target video resource subset from the candidate video resource subset based on the extended text vector corresponding to the extended object and the candidate text vector corresponding to each candidate video resource in the candidate video resource subset. Determine the target video resource subset as the recommended video resource subset corresponding to the extended object.
[0110] According to an embodiment of the present disclosure, a similarity value between the extended object and each candidate video resource in the candidate video resource subset can be determined based on the extended text vector corresponding to the extended object and the candidate text vector corresponding to each candidate video resource in the candidate video resource subset. Based on the similarity value between the extended object and each candidate video resource, a target video resource subset is determined from the candidate video resource subset.
[0111] According to an embodiment of the present disclosure, an approximate nearest neighbor algorithm may be used to determine a target video resource subset from the candidate video resource subset based on an extended text vector corresponding to the extended object and a candidate text vector corresponding to each candidate video resource in the candidate video resource subset.
[0112] Reference below Figure 3A 、 Figure 3B and Figure 4 , the resource processing method according to the embodiment of the present disclosure is further explained in combination with specific embodiments.
[0113] Figure 3A An example diagram of a resource processing method according to an embodiment of the present disclosure is schematically shown.
[0114] like Figure 3A As shown, in resource processing method 300A, at least one core object 302 is determined based on video data 301 of a target video. A core text vector 303 corresponding to the at least one core object 302 is determined. A first target object set is determined from a candidate object set based on the core text vector 303 and a candidate text vector 304 corresponding to each first candidate object of at least one first candidate object. The first target object set includes at least one first target object. Each first target object is determined as an extended object 305. Based on the at least one extended object 305, a recommended resource set 306 corresponding to the target video is determined.
[0115] Figure 3B An example schematic diagram of a resource processing method according to another embodiment of the present disclosure is schematically shown.
[0116] like Figure 3B As shown, in resource processing method 300B, based on video data 301 of a target video, a video vector 307 corresponding to video data 301 is determined. Based on video vector 307 and candidate text vectors 308 corresponding to each of at least one third candidate object, a second target object set is determined from the at least one third candidate object. The second target object set includes at least one second target object. Each second target object is determined as an extended object 309. Based on the at least one extended object 309, a recommended resource set 310 corresponding to the target video is determined.
[0117] The target video data 301 may be about "Chang'e flying to the moon". Figure 3A or Figure 3B In the resource processing method shown, it is determined that at least one extension object 305 or at least one extension object 309 may include a rabbit.
[0118] Figure 4 An example schematic diagram of a resource recommendation page according to an embodiment of the present disclosure is schematically shown.
[0119] Figure 4 It can be used Figure 3A or Figure 3B The resource processing method shown obtains a resource recommendation page corresponding to the recommended resource set.
[0120] like Figure 4 As shown, resource recommendation page 400 may include a target video 401 and a recommended resource set 402. Recommended resource set 402 may include encyclopedia resources 4020, commentary resources 4021, and video resources 4022. Encyclopedia resources 4020 may include explanations of the term "rabbit." Commentary resources 4021 may include comments about "rabbit." Video resources 4022 may include videos about "rabbit." Recommended resources may be obtained in response to a selection operation by user 403.
[0121] The above are merely exemplary embodiments, but are not limited thereto. Other resource processing methods known in the art may also be included as long as resource processing can be achieved.
[0122] It should be noted that the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved in the technical solution of this disclosure are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0123] Figure 5 The block diagram schematically shows a resource processing device according to an embodiment of the present disclosure.
[0124] like Figure 5 As shown, the resource processing device 500 may include a first determination module 510 , a second determination module 520 and a display module 530 .
[0125] The first determining module 510 is configured to determine at least one extended object based on the video data of the target video. Each extended object represents text information of extended content related to the target video.
[0126] The second determining module 520 is configured to determine a recommended resource set according to at least one extended object. The recommended resource set includes at least one recommended resource associated with the extended object.
[0127] The display module 530 is configured to display the recommended resource set so as to obtain the recommended resource in response to a selection operation on the recommended resource.
[0128] According to an embodiment of the present disclosure, the first determining module 510 may include a first determining submodule and a second determining submodule.
[0129] The first determining submodule is configured to determine at least one core object based on the video data of the target video. Each core object represents text information corresponding to the target video.
[0130] The second determining submodule is configured to determine at least one extended object according to at least one core object.
[0131] According to an embodiment of the present disclosure, the first determining submodule may include a first determining unit and a second determining unit.
[0132] The first determining unit is configured to determine at least one category label according to the video data of the target video.
[0133] The second determining unit is configured to determine at least one core object according to at least one category label.
[0134] According to an embodiment of the present disclosure, the second determining unit may include a first determining subunit and a second determining subunit.
[0135] The first determination subunit is configured to process at least one category label using data to determine at least one candidate category label, wherein the data processing includes at least one of the following: entity word filtering, semantic disambiguation, and abnormal word filtering.
[0136] The second determining subunit is configured to determine at least one core object according to at least one candidate category label.
[0137] According to an embodiment of the present disclosure, the second determination subunit is configured to: sort at least one candidate category label according to the inverse document frequency of each candidate category label in the at least one candidate category label; determine at least one target category label from the at least one candidate category label based on the sorting result; and determine each target category label as each core object.
[0138] According to an embodiment of the present disclosure, the second determining submodule may include a third determining unit.
[0139] The third determining unit is configured to determine at least one extended object based on a preset selection strategy and according to the candidate information set and at least one core object.
[0140] According to an embodiment of the present disclosure, the candidate information set includes a candidate object set, and the candidate object set includes at least one first candidate object.
[0141] The third determining unit may include a third determining subunit, a fourth determining subunit, and a fifth determining subunit.
[0142] The third determining subunit is configured to determine a core text vector corresponding to at least one core object.
[0143] The fourth determining subunit is configured to determine a first target object set from the candidate object set based on the core text vector and the candidate text vector corresponding to each first candidate object in the at least one first candidate object. The first target object set includes at least one first target object.
[0144] The fifth determining subunit is configured to determine each first target object as an extended object.
[0145] According to an embodiment of the present disclosure, the fourth determining subunit is used to: determine the first target object set from the candidate object set based on the core text vector and the candidate text vector corresponding to each first candidate object of at least one first candidate object using an approximate nearest neighbor algorithm.
[0146] According to an embodiment of the present disclosure, the candidate information set includes a candidate association relationship set, the candidate association relationship set includes at least one candidate association relationship, and each candidate association relationship represents an association relationship between an index object and a second candidate object.
[0147] The third determining unit may include a searching subunit and a sixth determining subunit.
[0148] The search subunit is configured to search for an index object matching the core object from the candidate association relationship set for each core object in the at least one core object.
[0149] The sixth determining subunit is configured to determine the second candidate object corresponding to the index object as the extended object corresponding to the core object.
[0150] According to an embodiment of the present disclosure, the first determining module 510 may include a third determining submodule, a fourth determining submodule, and a fifth determining submodule.
[0151] The third determination submodule is used to determine a video vector corresponding to the video data of the target video.
[0152] The fourth determination submodule is used to determine a second target object set from at least one third candidate object based on the video vector and the candidate text vector corresponding to each third candidate object in the at least one third candidate object, wherein the second target object set includes at least one second target object.
[0153] The fifth determining submodule is configured to determine each second target object as an extended object.
[0154] According to an embodiment of the present disclosure, the second determining module 520 may include a sixth determining submodule and a seventh determining submodule.
[0155] The sixth determining submodule is configured to determine, for each extended object of the at least one extended object, a recommended resource set corresponding to the extended object from the candidate resource set.
[0156] The seventh determining submodule is configured to determine a recommended resource set according to the recommended resource set corresponding to each extended object in the at least one extended object.
[0157] According to an embodiment of the present disclosure, the candidate resource set includes at least one of the following: a candidate graphic resource subset and a candidate video resource subset.
[0158] The recommended resource set corresponding to the extended object includes at least one of the following: a recommended graphic and text resource subset and a recommended video resource subset corresponding to the extended object.
[0159] The recommended graphic resource subset corresponding to the extended object is determined as follows: a target graphic resource subset is determined from the candidate graphic resource subset based on the extended object and the index object corresponding to each candidate graphic resource in the candidate graphic resource subset, and the target graphic resource subset is determined as the recommended graphic resource subset corresponding to the extended object.
[0160] The recommended video resource subset corresponding to the extended object is determined as follows: an extended text vector corresponding to the extended object is determined. A target video resource subset is determined from the candidate video resource subset based on the extended text vector corresponding to the extended object and the candidate text vector corresponding to each candidate video resource in the candidate video resource subset. The target video resource subset is determined as the recommended video resource subset corresponding to the extended object.
[0161] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0162] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described above.
[0163] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute the method described above.
[0164] According to an embodiment of the present disclosure, a computer program product includes a computer program, and when the computer program is executed by a processor, the computer program implements the method described above.
[0165] Figure 6 A block diagram of an electronic device suitable for implementing a resource processing method according to an embodiment of the present disclosure is schematically shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0166] like Figure 6 As shown, the device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. Various programs and data required for the operation of the device 600 can also be stored in the RAM 603. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0167] Various components in device 600 are connected to I / O interface 605, including an input unit 606, such as a keyboard, mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, optical disk, etc.; and a communication unit 609, such as a network card, modem, wireless communication transceiver, etc. The communication unit 609 allows device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0168] The computing unit 601 can be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 601 performs the various methods and processes described above, such as the resource processing method. For example, in some embodiments, the resource processing method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the resource processing method described above can be performed. Alternatively, in other embodiments, the computing unit 601 can be configured to perform the resource processing method by any other appropriate means (e.g., by means of firmware).
[0169] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0170] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0171] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0172] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0173] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0174] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0175] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.
[0176] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A resource processing method, comprising: Determining at least one extended object based on video data of a target video, wherein each of the extended objects represents text information of extended content related to the target video; determining a recommended resource set according to the at least one extended object, wherein the recommended resource set includes at least one recommended resource associated with the extended object; and Displaying the recommended resource set so that the recommended resource can be obtained in response to a selection operation on the recommended resource; Determining at least one extended object based on the video data of the target video includes: determining at least one core object based on the video data of the target video, wherein each core object represents text information corresponding to the target video; and determining the at least one extended object based on the at least one core object; The determining of the at least one extended object according to the at least one core object comprises: determining the at least one extended object according to the candidate information set and the at least one core object based on a preset selection strategy; The candidate information set includes a candidate object set, and the candidate object set includes at least one first candidate object; determining the at least one extended object based on the candidate information set and the at least one core object based on a preset selection strategy includes: determining a core text vector corresponding to the at least one core object; determining a first target object set from the candidate object set based on the core text vector and the candidate text vector corresponding to each first candidate object in the at least one first candidate object, wherein the first target object set includes at least one first target object; and determining each of the first target objects as each extended object.
2. The method according to claim 1, wherein The determining, based on the video data of the target video, at least one core object comprises: Determining at least one category label based on the video data of the target video; and The at least one core object is determined according to the at least one category label.
3. The method according to claim 2, wherein: The determining the at least one core object according to the at least one category label includes: Processing the at least one category label using data to determine at least one candidate category label, wherein the data processing includes at least one of the following: entity word filtering, semantic disambiguation, and abnormal word filtering; and The at least one core object is determined according to the at least one candidate category label.
4. The method according to claim 3, wherein: The determining the at least one core object according to the at least one candidate category label includes: sorting the at least one candidate category label according to the inverse document frequency of each candidate category label in the at least one candidate category label; Determining at least one target category label from the at least one candidate category label according to the ranking result; and Each target category label is determined as each core object.
5. The method according to claim 1, wherein The determining a first target object set from the candidate object set according to the core text vector and the candidate text vector corresponding to each first candidate object of the at least one first candidate object comprises: The first target object set is determined from the candidate object set by using an approximate nearest neighbor algorithm according to the core text vector and the candidate text vector corresponding to each first candidate object of the at least one first candidate object.
6. The method according to claim 1, wherein The candidate information set includes a candidate association relationship set, the candidate association relationship set includes at least one candidate association relationship, and each candidate association relationship represents an association relationship between the index object and the second candidate object; The determining, based on a preset selection strategy and according to the candidate information set and the at least one core object, the at least one extended object includes: For each core object in the at least one core object, searching for an index object matching the core object from the candidate association relationship set; as well as A second candidate object corresponding to the index object is determined as an extended object corresponding to the core object.
7. The method according to claim 1, wherein The determining, based on the video data of the target video, at least one extended object includes: Determining a video vector corresponding to the video data of the target video; determining a second target object set from the at least one third candidate object based on the video vector and the candidate text vector corresponding to each third candidate object in the at least one third candidate object, wherein the second target object set includes at least one second target object; and Each of the second target objects is determined as each of the extended objects.
8. The method according to any one of claims 1 to 7, wherein The determining of the recommended resource set according to the at least one extended object includes: For each of the at least one extended object, determining a recommended resource set corresponding to the extended object from a candidate resource set; and The recommended resource set is determined according to the recommended resource set corresponding to each extended object of the at least one extended object.
9. The method according to claim 8, wherein The candidate resource set includes at least one of the following: a candidate graphic resource subset and a candidate video resource subset; The recommended resource set corresponding to the extended object includes at least one of the following: a recommended graphic resource subset and a recommended video resource subset corresponding to the extended object; Determine the recommended graphic resource subset corresponding to the extension object in the following manner: determining a target graphic resource subset from the candidate graphic resource subset according to the extension object and the index object corresponding to each candidate graphic resource in the candidate graphic resource subset; as well as Determining the target graphic resource subset as a recommended graphic resource subset corresponding to the extended object; Determine the recommended video resource subset corresponding to the extended object in the following manner: determining an extended text vector corresponding to the extended object; determining a target video resource subset from the candidate video resource subset according to the extended text vector corresponding to the extended object and the candidate text vector corresponding to each candidate video resource in the candidate video resource subset; as well as The target video resource subset is determined as a recommended video resource subset corresponding to the extended object.
10. A resource processing device, comprising: A first determining module is configured to determine at least one extended object based on video data of a target video, wherein each of the extended objects represents text information of extended content related to the target video; a second determining module, configured to determine a recommended resource set according to the at least one extended object, wherein the recommended resource set includes at least one recommended resource associated with the extended object; and a display module, configured to display the recommended resource set so as to obtain the recommended resource in response to a selection operation on the recommended resource; The second determination module includes: a first determination submodule for determining at least one core object based on the video data of the target video, wherein each core object represents text information corresponding to the target video; and a second determination submodule for determining the at least one extended object based on the at least one core object; The second determining submodule includes a third determining unit, the third determining unit being configured to determine the at least one extended object based on a preset selection strategy and according to the candidate information set and the at least one core object; The candidate information set includes a candidate object set, and the candidate object set includes at least one first candidate object; the third determination unit includes a third determination subunit, a fourth determination subunit and a fifth determination subunit; the third determination subunit is used to determine a core text vector corresponding to the at least one core object; the fourth determination subunit is used to determine a first target object set from the candidate object set based on the core text vector and the candidate text vector corresponding to each first candidate object in the at least one first candidate object, wherein the first target object set includes at least one first target object; and the fifth determination subunit is used to determine each of the first target objects as each of the extended objects.
11. The device according to claim 10, wherein The first determining submodule includes: A first determining unit is configured to determine at least one category label based on the video data of the target video; and The second determining unit is configured to determine the at least one core object according to the at least one category label.
12. The device according to claim 11, wherein The second determining unit includes: A first determining subunit is configured to process the at least one category label using data to determine at least one candidate category label, wherein the data processing includes at least one of the following: entity word filtering, semantic disambiguation, and abnormal word filtering; and The second determining subunit is configured to determine the at least one core object according to the at least one candidate category label.
13. The device according to claim 12, wherein The second determining subunit is configured to: sorting the at least one candidate category label according to the inverse document frequency of each candidate category label in the at least one candidate category label; Determining at least one target category label from the at least one candidate category label according to the ranking result; as well as Each target category label is determined as each core object.
14. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 9.
15. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to enable the computer to execute the method according to any one of claims 1 to 9.
16. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the method according to any one of claims 1 to 9.