Content recommendation method and device and related products

By identifying content dimensions and generating recommendation information during video data playback, it solves the problem that users find it difficult to accurately search video content, and efficient information recommendation and acquisition are achieved.

CN120256676APending Publication Date: 2025-07-04DOUYIN VISION CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510362196.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

It is difficult for users to summarize accurate search terms through the pictures in the video data, resulting in inefficient information acquisition.

Method used

Through the content understanding model, identify the content dimensions in the video data, and generate recommendation information based on these dimensions, and display them directly in the playback scene to avoid manual search by users.

Benefits of technology

It improves the efficiency of recommending content in video data to users, and users can obtain relevant information without manual search, which significantly improves the efficiency of information acquisition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256676A_ABST
    Figure CN120256676A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a content recommendation method and device and a related product, and the method comprises the steps: responding to a playing instruction for video data, playing the video data, recognizing at least one piece of content in the video data, and determining a target content dimension corresponding to the recognized content; based on each target content dimension, determining a content recommendation dimension corresponding to the video data, and determining target content belonging to the content recommendation dimension in the identified content; and generating recommendation information for recommending the target content, and displaying the recommendation information in a playing scene of the video data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to a content recommendation method, apparatus, and related products. Background Art

[0002] Currently, users can browse video data through a video data playback program. Video data usually contains multiple items of content. For example, video data includes both the clothing styles of the male and female protagonists and the travel destinations visited by the male and female protagonists. Considering that users may have a relatively high degree of interest in a certain category of content in the video data, but based on the images in the video data, it is difficult for users to summarize the accurate search terms corresponding to this category of content, resulting in the difficulty for users to search for this category of content and low information acquisition efficiency. Therefore, how to improve the efficiency of recommending the content in video data to users and improve the information acquisition efficiency has become one of the problems that need to be solved urgently. Summary of the Invention

[0003] Embodiments of the present disclosure provide a content recommendation method, apparatus, and related products, which can improve the efficiency of recommending the content in video data to users without the need for users to manually search for the content in the video data, and improve the information acquisition efficiency.

[0004] In a first aspect, an embodiment of the present disclosure provides a content recommendation method, including: In response to a playback instruction for video data, play the video data, identify at least one item of content included in the video data, and determine a target content dimension corresponding to the identified content; Based on each of the target content dimensions, determine a content recommendation dimension corresponding to the video data, and determine target content subordinate to the content recommendation dimension from the identified content; Generate recommendation information for recommending the target content, and display the recommendation information in the playback scenario of the video data.

[0005] In a second aspect, an embodiment of the present disclosure provides a content recommendation method, including: Obtain video data, identify at least one item of content included in the video data, and determine a target content dimension corresponding to the identified content; Based on each of the target content dimensions, determine a content recommendation dimension corresponding to the video data, and determine target content subordinate to the content recommendation dimension from the identified content; Generate recommendation information for recommending the target content, and display the recommendation information in a content recommendation scenario associated with the video data.

[0006] In a third aspect, an embodiment of the present disclosure provides a content recommendation apparatus, including: A first recognition unit, configured to play the video data in response to a play instruction for the video data, recognize at least one piece of content included in the video data, and determine a target content dimension corresponding to the recognized content; A first determination unit, configured to determine a content recommendation dimension corresponding to the video data based on each of the target content dimensions, and determine target content subordinate to the content recommendation dimension from the recognized content; A first display unit, configured to generate recommendation information for recommending the target content and display the recommendation information in a play scenario of the video data.

[0007] In a fourth aspect, an embodiment of the present disclosure provides a content recommendation device, including: A second recognition unit, configured to obtain video data, recognize at least one piece of content included in the video data, and determine a target content dimension corresponding to the recognized content; A second determination unit, configured to determine a content recommendation dimension corresponding to the video data based on each of the target content dimensions, and determine target content subordinate to the content recommendation dimension from the recognized content; A second display unit, configured to generate recommendation information for recommending the target content and display the recommendation information in a content recommendation scenario associated with the video data.

[0008] In a fifth aspect, an embodiment of the present disclosure provides an electronic device, including: a processor; and a memory configured to store computer-executable instructions, where the computer-executable instructions, when executed, cause the processor to implement the method described in the first aspect or the method described in the second aspect.

[0009] In a sixth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, where the computer-readable storage medium is used to store computer-executable instructions, and the computer-executable instructions, when executed by a processor, implement the method described in the first aspect or the method described in the second aspect.

[0010] In a seventh aspect, an embodiment of the present disclosure provides a computer program product, where the computer program product includes a computer program, and the computer program, when executed by a processor, implements the method described in the first aspect or the method described in the second aspect.

[0011] In one or more embodiments of the present disclosure, it is possible to identify at least one piece of content included in video data, determine the target content dimension corresponding to the identified content, determine the content recommendation dimension corresponding to the video data based on each target content dimension, determine the target content belonging to the content recommendation dimension among the identified content, generate recommendation information for recommending the target content, and display the recommendation information. Thus, without the user manually searching for the target content in the video data, the user can directly and quickly view the recommendation information of the target content, significantly improving the efficiency of recommending the target content in the video data to the user and enhancing the efficiency of the user obtaining relevant information about the target content. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] To more clearly illustrate the technical solutions in one or more embodiments of the present disclosure or in the related art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the related art. Obviously, the accompanying drawings in the following description are only some embodiments recorded in the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. Figure 1 Schematic flowchart of the content recommendation method provided by an embodiment of the present disclosure; Figure 2a Schematic diagram of the image data of the target content provided by an embodiment of the present disclosure; Figure 2b Schematic diagram of the recommendation information of the target content provided by an embodiment of the present disclosure; Figure 3a Schematic diagram of the image data of the target content provided by another embodiment of the present disclosure; Figure 3b Schematic diagram of the recommendation information of the target content provided by another embodiment of the present disclosure; Figure 4a Schematic diagram of the image data of the target content provided by yet another embodiment of the present disclosure; Figure 4b Schematic diagram of the recommendation information of the target content provided by yet another embodiment of the present disclosure; Figure 5a Schematic diagram of displaying a recommendation component in a video data playback scenario provided by an embodiment of the present disclosure; Figure 5b Schematic diagram of displaying recommendation information in a video data playback scenario provided by an embodiment of the present disclosure; Figure 5c Schematic diagram of displaying recommendation information in a video data playback scenario provided by another embodiment of the present disclosure; Figure 6 Schematic flowchart of the content recommendation method provided by another embodiment of the present disclosure; Figure 7Schematic diagram of a content recommendation scenario associated with video data provided by an embodiment of the present disclosure; Figure 8 Schematic diagram of a content recommendation scenario associated with video data provided by another embodiment of the present disclosure; Figure 9 Schematic diagram of a content recommendation scenario associated with video data provided by yet another embodiment of the present disclosure; Figure 10 Schematic diagram of the structure of a content recommendation device provided by an embodiment of the present disclosure; Figure 11 Schematic diagram of the structure of a content recommendation device provided by another embodiment of the present disclosure; Figure 12 Schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners

[0013] In order to enable those skilled in the art to better understand the technical solutions in one or more embodiments of the present disclosure, the following will clearly and completely describe the technical solutions in one or more embodiments of the present disclosure with reference to the accompanying drawings in one or more embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments. Based on one or more embodiments of the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present disclosure.

[0014] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, relevant parties should be informed of the type, scope of use, usage scenarios, etc. of the information involved in the present disclosure and obtain the authorization of the relevant parties in an appropriate manner in accordance with relevant laws and regulations.

[0015] For example, when receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that executes the operations of the technical solutions of the present disclosure according to the prompt message.

[0016] As an optional but non-limiting implementation manner, the way of sending a prompt message to the user in response to receiving an active request from the user can be, for example, in the form of a pop-up window, and the prompt message can be presented in text in the pop-up window. In addition, the pop-up window can also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0017] It should be understood that the above notification and the process of obtaining user authorization are only illustrative and do not limit the implementation manner of the present disclosure. Other manners that comply with relevant laws and regulations can also be applied to the implementation manner of the present disclosure.

[0018] Embodiments of the present disclosure provide a content recommendation method, apparatus, and related products, which can automatically generate corresponding recommendation information for the content included in video data and display it to users, improving the efficiency of recommending the content in video data to users and the efficiency of information acquisition. Among them, the content recommendation method can be applied to a terminal device and implemented by the terminal device. The terminal device includes, but is not limited to, user terminals with video playback capabilities such as mobile phones, computers, tablet computers, laptop computers, in-vehicle computers, etc.

[0019] Figure 1 It is a schematic flowchart of a content recommendation method provided by an embodiment of the present disclosure. As Figure 1 shown, the process includes: Step S102, in response to a play instruction for video data, play the video data, identify at least one piece of content included in the video data, and determine the target content dimension corresponding to the identified content; Step S104, based on each target content dimension, determine the content recommendation dimension corresponding to the video data, and determine the target content subordinate to the content recommendation dimension among the identified content; Step S106, generate recommendation information for recommending the target content, and display the recommendation information in the playback scene of the video data.

[0020] In this embodiment, in response to a play instruction for video data, the video data can be played, at least one piece of content included in the video data can be identified, and the target content dimension corresponding to the identified content can be determined. Based on each target content dimension, the content recommendation dimension corresponding to the video data can be determined, the target content subordinate to the content recommendation dimension can be determined among the identified content, recommendation information for recommending the target content can be generated, and the recommendation information can be displayed in the playback scene of the video data. Therefore, it is not necessary for the user to manually search for the target content in the video data, and the recommendation information of the target content can be directly and quickly browsed, significantly improving the efficiency of recommending the target content in the video data to the user and the efficiency of the user obtaining relevant information of the target content.

[0021] In the above step S102, in response to a play instruction for video data, the video data is played. The terminal device can detect a triggering operation by the user on the play control for the video data, generate a play instruction for the video data based on this triggering operation, and then play the video data in response to this play instruction. The video data includes any one of short plays, short videos, movies, TV series, video data recorded by the user, and video data received by the user. Any video data that can be played can apply the content recommendation method in this embodiment.

[0022] In the above step S102, during the playing of the video data, at least one piece of content included in the video data is also identified, and the target content dimension corresponding to the identified content is determined. Among them, the content included in the video data includes but is not limited to characters, locations, items, occasions, eras, knowledge points, etc. that appear in the video data. Among them, examples of characters can be the male lead, the female lead, etc. Examples of locations can be cities, villages, parks, shopping malls, etc. Examples of items can be a red dress, a dressing table, a Ferris wheel, a computer, a song, an online game, etc. Examples of occasions can be banquet occasions, office occasions, entertainment occasions, etc. Examples of eras can be the 1980s, the 1990s, etc. Examples of knowledge points can be knowledge related to a certain historical figure, historical knowledge related to a certain scenic spot, etc.

[0023] In one embodiment, identifying at least one piece of content included in the video data includes: Through a content understanding model, according to the physical item dimension, virtual item dimension, character dimension, location dimension, occasion dimension, era dimension, and knowledge point dimension, at least one piece of content included in the video data is identified.

[0024] In this embodiment, the content understanding model can be a visual recognition model trained through AI (Artificial Intelligence) technology, which can perform content understanding on video data. Through the content understanding model, at least one piece of content included in the video data can be identified from a preset multiple dimensions. The preset multiple dimensions include: the physical item dimension, virtual item dimension, character dimension, location dimension, occasion dimension, era dimension, and knowledge point dimension.

[0025] Specifically, the physical item dimension is a classification dimension for items with physical forms, including but not limited to the following sub-dimensions: clothing item sub-dimension, decoration item sub-dimension, play item sub-dimension, food item sub-dimension, office item sub-dimension, etc. Among them, the clothing item sub-dimension may include the following: red dress, white shirt, jeans, leather shoes, sportswear, etc. The decoration item sub-dimension may include the following: sofa, coffee table, cabinet, dressing table, curtain, etc. The play item sub-dimension may include the following: tent, kite, yacht, ferris wheel, etc. The food item sub-dimension may include the following: fried chicken, cola, potato chips, barbecue, etc. The office item sub-dimension may include the following: computer, printer, office desk, etc.

[0026] The virtual item dimension is a classification dimension for items without physical forms, including but not limited to the following sub-dimensions: song sub-dimension, online game sub-dimension, application program sub-dimension, coupon sub-dimension, etc. Among them, the song sub-dimension may include the following: Song A, Song B, etc. The online game sub-dimension may include the following: Online Game C, Online Game D, etc. The application program sub-dimension may include the following: chat application program, travel application program, video application program, etc. The coupon sub-dimension may include the following: food coupon, food group-buying coupon, etc.

[0027] The character dimension is a dimension for classifying characters based on various characteristics and attributes of the characters. For example, the character dimension may include content such as male lead and female lead. The location dimension is a dimension for classifying locations based on various characteristics and attributes of the locations. For example, the location dimension may include content such as city, countryside, park, scenic area, etc. The occasion dimension is a dimension for classifying occasions based on various characteristics, functions, and related elements of the occasions. For example, the occasion dimension may include content such as entertainment occasion, office occasion, banquet occasion, etc. The decade dimension is a time-related classification dimension used to represent the decade in which an event occurred and the era background in which the event took place. For example, the decade dimension may include content such as the 1980s and the 1990s. The knowledge point dimension is a classification dimension for knowledge points related to a certain character, a certain event, or a certain item. For example, the knowledge point dimension may include content such as knowledge points related to historical figure L and knowledge points related to scenic spot Q.

[0028] In an example, assume that the content in the identified video data includes: white dress, Dr. Martens boots, laptop, fried chicken, potato chips, ferris wheel, yacht, chat application program, male lead Little A, female lead Little B, XX Park, City J, annual meeting of G Company, office occasion, 2024, 1990, knowledge points related to scenic spot Q, etc.

[0029] It can be seen that through this embodiment, the content understanding model can identify the content included in the video data from multiple dimensions, making the identified content more comprehensive and accurate, thereby enhancing the diversity and accuracy of content recommendation.

[0030] After identifying the content included in the video data, step S102 further determines the target content dimension corresponding to the identified content.

[0031] Continuing with the above example, among the above-identified items of content, the target content dimensions corresponding to the white dress and Dr. Martens boots are the clothing items sub-dimension under the entity item dimension. The target content dimension corresponding to the laptop is the office items sub-dimension under the entity item dimension. The target content dimensions corresponding to fried chicken and potato chips are the food items sub-dimension under the entity item dimension. The target content dimensions corresponding to the Ferris wheel and yacht are the play items sub-dimension under the entity item dimension. The target content dimension corresponding to the chat application is the application sub-dimension under the virtual item dimension. The target content dimensions corresponding to male lead Little A and female lead Little B are the person dimension. The target content dimensions corresponding to XX Park and City J are the location dimension. The target content dimensions corresponding to the annual meeting of Company G and the office occasion are the occasion dimension. The target content dimensions corresponding to 2024 and 1990 are the decade dimension. The target content dimension corresponding to the knowledge points related to scenic spot Q is the knowledge point dimension.

[0032] In another embodiment, it is also possible to first determine, through the content understanding model, at least one target content dimension included in the video data among the entity item dimension, virtual item dimension, person dimension, location dimension, occasion dimension, decade dimension, and knowledge point dimension. Then, through the content understanding model, identify at least one item of content included in the video data under each target content dimension, and determine the corresponding relationship between the identified content and the target content dimension.

[0033] After identifying at least one item of content in the video data and determining the target content dimension corresponding to the identified content, in step S104, based on each target content dimension, determine the content recommendation dimension corresponding to the video data. The content recommendation dimension can be a combination of multiple target content dimensions. Examples of the content recommendation dimension can be dimensions such as the clothing items of a person, the play items of a location, the decoration items of an occasion, the occasions included in a location, and the songs of an occasion.

[0034] In one embodiment, determining the content recommendation dimension corresponding to the video data based on each target content dimension includes: Determine the second association relationship between each target content dimension according to the first association relationship between the items of content included in the video data; the first association relationship is used to represent the associated content among the items of content; the second association relationship is used to represent the associated dimensions among each target content dimension; According to the second association relationship, combine the associated dimensions to obtain the content recommendation dimensions corresponding to the video data.

[0035] In this embodiment, first, determine the first association relationship between the contents in the video data. The first association relationship is used to represent the associated contents among the contents. The associated contents can be two or more contents that appear simultaneously in the video data. For example, in the video data, the male lead is wearing a white shirt. The content of the male lead and the content of the white shirt appear simultaneously in the video data. Therefore, it is determined that the content of the male lead is associated with the content of the white shirt, and there is a first association relationship between the content of the male lead and the content of the white shirt.

[0036] Next, according to the first association relationship, determine the second association relationship between the respective target content dimensions. The second association relationship is used to represent the associated dimensions among the respective target content dimensions. The target content dimensions corresponding to the associated contents are the associated dimensions. There is a second association relationship between the target content dimensions corresponding to the associated contents. For example, in the video data, the female lead appears in City J. The content of the female lead and the content of City J appear simultaneously. The female lead and City J are associated contents. The target content dimension corresponding to the female lead is the character dimension, and the dimension corresponding to City J is the location dimension. It can be determined that the character dimension is associated with the location dimension, and there is a second association relationship between the character dimension and the location dimension.

[0037] Next, according to the second association relationship between the respective target content dimensions, combine the associated target content dimensions to obtain the content recommendation dimensions corresponding to the video data.

[0038] Continuing with the above example, the contents identified in the video data include: white dress, leather shoes, laptop, fried chicken, potato chips, Ferris wheel, yacht, chat application, male lead Little A, female lead Little B, XX Park, City J, G Company annual meeting, office occasion, 2024, 1990, scenic spot Q, historical knowledge related to scenic spot Q, etc. Assume that in the video data, the female lead Little B often appears in the office occasion wearing a white dress. Then, the female lead Little B, the white dress, and the office occasion are associated contents. The target content dimension corresponding to the content of the female lead Little B is the character dimension, the target content dimension corresponding to the white dress is the sub-dimension of dressing items, and the target content dimension corresponding to the office occasion is the occasion dimension. Based on this, it can be determined that the three dimensions of the character dimension, the sub-dimension of dressing items, and the occasion dimension are associated. Further, since the character dimension, the sub-dimension of dressing items, and the occasion dimension are associated, the character dimension, the sub-dimension of dressing items, and the occasion dimension can be combined to obtain the content recommendation dimension corresponding to this video data as the dressing items of the character in the occasion.

[0039] Also assume that the Ferris wheel appears in City J in the video data. Then, City J and the Ferris wheel are related content. The target content dimension corresponding to City J is the location dimension, and the target content dimension corresponding to the Ferris wheel is the sub-dimension of play items. Based on this, it can be determined that the location dimension and the sub-dimension of play items are related. Therefore, the location dimension and the sub-dimension of play items can be combined to obtain the content recommendation dimension corresponding to this video data as the play items at the location.

[0040] Again assume that the heroine Xiaob often eats potato chips in the video data. Then, the heroine Xiaob and the potato chips are related content. The target content dimension corresponding to the heroine Xiaob is the character dimension, and the target content dimension corresponding to the potato chips is the sub-dimension of food. Based on this, it can be determined that the character dimension and the sub-dimension of food are related. Therefore, the character dimension and the sub-dimension of food can be combined to obtain the content recommendation dimension corresponding to this video data as the food eaten by the character.

[0041] It can be seen that through this embodiment, according to the first association relationship between the various contents included in the video data, the second association relationship between the respective target content dimensions is determined, and according to the second association relationship, the related dimensions are combined to obtain the content recommendation dimension corresponding to the video data, which can start from multiple different dimensions, combine multiple related target content dimensions, generate diverse content recommendation dimensions, and thus improve the diversity and accuracy of content recommendation.

[0042] After determining the content recommendation dimension corresponding to the video data, in step S104, the target content belonging to the content recommendation dimension is also determined among the recognized contents.

[0043] Continuing the above example, the above content recommendation dimensions include: the dimension of the clothing and accessories worn by a person in an occasion, the dimension of the play items at a location, and the dimension of the food eaten by a person. Among the above recognized contents, the target contents belonging to the dimension of the clothing and accessories worn by a person in an occasion include: the white dress worn by the heroine Xiaob in the office occasion and the leather shoes worn by the hero Xiaoa at the annual meeting of Company G. The target contents belonging to the dimension of the play items at a location include: the Ferris wheel in City J and the yacht in City J. The target contents belonging to the dimension of the food eaten by a person include: the potato chips eaten by the heroine Xiaob and the fried chicken eaten by the hero Xiaoa.

[0044] After determining the content recommendation dimension corresponding to the video data and determining the target content belonging to the content recommendation dimension among the recognized contents, in step S106, recommendation information for recommending the target content is generated and the recommendation information is displayed in the playback scene of the video data. The recommendation information is used to recommend the target content in the video data to the user. When the user plays the video data, the recommendation information can be displayed in the playback scene of the video data.

[0045] In one embodiment, generating recommendation information for recommending target content includes: Obtaining, through a large language model, image data of the target content in video data, and generating text data that matches the image data and the target content; Generating, through the large language model, recommendation information based on the image data and the text data.

[0046] In this embodiment, when generating recommendation information for recommending target content, the video data and text representing the target content, such as "the potato chips eaten by female lead Xiaob", can be input into the large language model first. First, through the large language model, the image data of the target content in the video data is obtained. The image data can be video frames in the video data. Then, through the large language model, text data that matches the target content and the image data corresponding to the target content is obtained. The large language model can analyze the target content and the image data corresponding to the target content to generate text data for describing the characteristics and style of the target content. Finally, through the large language model, the image data and the text data corresponding to the target content are combined to obtain recommendation information for recommending the target content. The recommendation information can be in the form of pictures and texts, where the pictures and texts include the above-mentioned image data and text data. The recommendation information can also be in the form of a video, where each video frame in the video is the above-mentioned image data, and the subtitles in the video are the above-mentioned text data.

[0047] In one example, the target content is a white dress worn by female lead Xiaob in an office setting. Figure 2a Schematic diagram of the image data of the target content provided by an embodiment of the present disclosure, as Figure 2a shown, the image data is a video frame of the target content, which is a white dress worn by female lead Xiaob in an office setting, in the short drama "Struggle". The text data that matches the image data is, for example: "Design highlights: The puffed sleeves add a sweet atmosphere, and the waist-cinching design highlights the slender waistline, showing the soft curves of women. Fabric texture: Made of high-quality cotton fabric, skin-friendly, breathable, and comfortable to wear. Style adaptation: The pure white creates a fresh and elegant temperament, which not only meets the decency requirements of the office environment but also can easily handle various office occasions. Matching suggestions: Match with a simple metal necklace, low-heel leather shoes, and a delicate handbag to enhance the overall professional and capable feeling." Figure 2b Schematic diagram of the recommendation information of the target content provided by an embodiment of the present disclosure, as Figure 2b shown, the recommendation information is in the form of pictures and texts, and the recommendation information includes the image data corresponding to the target content, which is a white dress worn by female lead Xiaob in an office setting, and the text data that matches the image data. The recommendation information can be used to recommend to users the target content, which is a white dress worn by female lead Xiaob in an office setting.

[0048] In another example, the target content is the potato chips eaten by the female lead, Xiaob. Figure 3a It is a schematic diagram of the image data of the target content provided by another embodiment of the present disclosure. As Figure 3a shown, this image data is a video frame of the target content, which is the potato chips eaten by the female lead, Xiaob, in the short drama "Struggle". The text data matching this image data is, for example: "Taste: The potato chips are thin, crispy and refreshing, making a 'crunch' sound instantly when bitten, bringing an excellent chewing experience. Appearance: Each piece is of uniform size, with clear surface texture, looking very attractive. Delicious: The taste is rich and flavorful, easily satisfying the taste buds whether for watching dramas or relieving cravings during leisure time. Convenient: Easy to pick up and can be enjoyed at any time when placed in a container. It is a delicious snack for leisure moments." Figure 3b It is a schematic diagram of the recommendation information of the target content provided by another embodiment of the present disclosure. As Figure 3b shown, this recommendation information is in the form of pictures and texts, and it includes the image data corresponding to the target content, which is the potato chips eaten by the female lead, Xiaob, and the text data matching this image data. This recommendation information can be used to recommend the target content, which is the potato chips eaten by the female lead, Xiaob, to users.

[0049] In yet another example, the target content is the Ferris wheel in City J. Figure 4a It is a schematic diagram of the image data of the target content provided by another embodiment of the present disclosure. As Figure 4a shown, this image data is a video frame of the target content, which is the Ferris wheel in City J, in the short drama "Struggle". The text data matching this image data is, for example: "Viewing experience: Taking a ride on it, you can overlook the cityscape of City J 360°. Scenes for playing: Whether it is a romantic date for couples or an outing with friends, it is a good check-in item to leave beautiful memories. Time characteristics: It contrasts with the blue sky and white clouds during the day, has a full sense of atmosphere in the afterglow in the evening, and is even more dazzling when lit up at night." Figure 4b It is a schematic diagram of the recommendation information of the target content provided by another embodiment of the present disclosure. As Figure 4b shown, this recommendation information is in the form of pictures and texts, and it includes the image data corresponding to the target content, which is the Ferris wheel in City J, and the text data matching this image data. This recommendation information can be used to recommend the target content, which is the Ferris wheel in City J, to users.

[0050] It can be seen that through this embodiment, by using a large language model, the image data of the target content in the video data is obtained, and the text data matching the image data and the target content is generated. Based on the image data and the text data, recommendation information is generated. The recommendation information can be in the form of pictures and texts or in the form of videos. The target content is recommended to users based on the form of pictures and texts or videos, improving the accuracy of content recommendation.

[0051] After generating recommendation information for recommending target content, in step S106, the recommendation information is displayed in the playback scenario of the video data. For example, when the user plays the video data, the recommendation information for recommending the target content is displayed to the user on the playback page of the video data.

[0052] In one embodiment, displaying the recommendation information in the playback scenario of the video data includes: On the playback page of the video data, a recommendation component for recommending the target content is displayed; In response to a trigger operation on the recommendation component, the recommendation information is displayed.

[0053] In this embodiment, when displaying the recommendation information in the playback scenario of the video data, a recommendation component for recommending the target content can be first displayed on the playback page of the video data. The recommendation component can be an icon with a jump function or a pop-up window function, and the style of the icon is not specifically limited. The recommendation component can be located in the lower right corner of the playback page of the video data, or other edge positions in the playback page of the video data, as long as it does not block the content in the video data. During the playback of the video data, the terminal device can detect a trigger operation by the user on the recommendation component, and display the recommendation information for recommending the target content according to the trigger operation.

[0054] In this embodiment, considering that there are many target contents belonging to the same content recommendation dimension, the target contents belonging to the same content recommendation dimension can be classified. In one example, when the content recommendation dimension is a combination of the entity item dimension and other target content dimensions, the entity item dimension can be kept unchanged, and the target content can be classified according to the contents included in the other target content dimensions. For example, when the content recommendation dimension is the clothing items of a person in an occasion, the target content can be classified according to different people and different occasions to obtain each type of target content. Examples of each type of target content can be: the clothing items of the female lead in the office occasion, the clothing items of the male lead at the annual meeting.

[0055] In another example, when the content recommendation dimension is a combination of the virtual item dimension and other target content dimensions, the virtual item dimension can be kept unchanged, and the target content can be classified according to the contents included in the other target content dimensions. For example, when the content recommendation dimension is the songs for an occasion, the target content can be classified according to different occasions to obtain each type of target content. Examples of each type of target content can be: songs suitable for the office occasion, songs suitable for the annual meeting, etc.

[0056] Furthermore, corresponding recommendation components are generated for each type of target content. If a user triggers such a recommendation component, they can view the recommendation information for that type of target content. For example, the recommendation component is a component with the text "Click to view the outfits of the female lead in the office". If the user triggers the recommendation component, the recommendation information corresponding to each outfit of the female lead in the office can be shown to the user. It can be understood that one type of target content can include multiple pieces of target content. For example, if one type of target content is the tourist attractions in City A, it can include multiple tourist attractions in City A. Another example is that if one type of target content is the same-style spring and summer outfits of the female lead, it can include multiple spring and summer outfits of the female lead.

[0057] Figure 5a Schematic diagram showing a recommendation component in a video data playback scenario provided by an embodiment of the present disclosure. Figure 5b Schematic diagram showing recommendation information in a video data playback scenario provided by an embodiment of the present disclosure. Figure 5c Schematic diagram showing recommendation information in a video data playback scenario provided by another embodiment of the present disclosure. As Figure 5a shown, at the lower right position on the playback page of the short drama "Struggle", a recommendation component for recommending target content is shown. This recommendation component is a component with the text "Click to view the outfits of the female lead in the office". Assume that in Figure 5a the scenario shown, by identifying the video data, it is obtained that the outfits of the female lead in the office include three pieces of target content: the white dress worn by the female lead Xiaob in the office, the black suit worn by the female lead Xiaob in the office, and the gray suit worn by the female lead Xiaob in the office.

[0058] In one case, if the user triggers the recommendation component as Figure 5a shown, then as Figure 5b shown, the recommendation information corresponding to each outfit of the female lead in the office can be shown in the form of a pop-up window. Figure 5b Taking the white dress worn by the female lead Xiaob in the office as an example in

[0059] shown, the user can switch and view among the recommendation information corresponding to each outfit by sliding the screen. For example, by sliding down, the recommendation information corresponding to the white dress worn by the female lead Xiaob in the office, the recommendation information corresponding to the black suit worn by the female lead Xiaob in the office, and the recommendation information corresponding to the gray suit worn by the female lead Xiaob in the office can be switched and displayed. Figure 5b shown, the recommendation information also includes a "Go to purchase" component. The user can trigger the "Go to purchase" component. In response to this trigger operation, it will jump to the purchase page corresponding to the target content and purchase the target content. For example, when the user is viewing the recommendation information corresponding to the white dress worn by the female lead Xiaob in the office, if the user triggers the "Go to purchase" component, they can jump to the purchase page to buy this white dress. AsFigure 5b As shown, the recommended information also includes a "go to watch drama" component. The user can trigger the "go to watch drama" component, and in response to this trigger operation, the playback page of the video data is returned to continue playing the video data.

[0060] In another case, if the user triggers a recommended component such as Figure 5a shown, then information recommendations can be displayed as shown in Figure 5c In the information recommendation page, the recommended information corresponding to each set of outfits of the female protagonist in the office is displayed. Figure 5c Taking the white dress worn by the female protagonist Xiaob in the office as an example in

[0061] As shown in Figure 5c The recommended information also includes a "go to purchase" component and a "go to watch drama" component. The user can trigger the "go to purchase" component, and in response to this trigger operation, jump to the purchase page corresponding to the target content and purchase the target content. For example, when the user is browsing the recommended information corresponding to the white dress worn by the female protagonist Xiaob in the office, if the user triggers the "go to purchase" component, then they can jump to the purchase page to buy this white dress. As shown in Figure 5c The recommended information also includes a "go to watch drama" component. The user can trigger the "go to watch drama" component, and in response to this trigger operation, the playback page of the video data is returned to continue playing the video data.

[0062] It can be seen that through this embodiment, by displaying a recommended component for recommending target content on the playback page of the video data, the user can obtain the corresponding recommended information by clicking on this recommended component, without the user having to manually search for the target content, improving the efficiency of recommending the target content in the video data to the user and improving the efficiency of the user obtaining relevant information about the target content.

[0063] In one embodiment, the video data includes multiple sets of sub-data played in chronological order on the playback page of the video data; displaying recommended information in the playback scenario of the video data includes: On the playback page, between the first sub-data and the second sub-data in the video data, display the recommended information; Or, On the playback page, after the last set of sub-data in the video data, display the recommended information.

[0064] In this embodiment, the video data may include multiple sub-data, and the multiple sub-data may be played in sequence according to the time order on the playback page. Taking the video data as a short drama as an example, the short drama may include multiple episodes, and each episode of the short drama is played in sequence according to the time order on the playback page. Based on this, on the playback page of the video data, recommendation information may be displayed after the first sub-data in the video data finishes playing and before the second sub-data starts playing. Or, on the playback page of the video data, the recommendation information may be displayed after the last episode of the sub-data in the video data finishes playing. The recommendation information may be in the form of text and pictures or in the form of a video.

[0065] Taking the short drama scenario as an example, in one example, the recommendation information may be displayed to the user every other episode or multiple episodes. The recommendation information is determined based on each episode of the short drama that the user has watched and is used to recommend a type of target content in each episode of the short drama that the user has watched. This type of target content may include multiple target contents. For example, it includes 10 sets of travel outfits of the male lead. For example, after each episode or multiple episodes finish playing, the recommendation information corresponding to the target content in this episode or these episodes is displayed in the form of text and pictures on the playback page of the short drama. In another example, after all the episodes in the short drama finish playing, the recommendation information corresponding to the target content in all the short dramas is displayed in the form of text and pictures on the playback page of the short drama.

[0066] It can be seen that through this embodiment, on the playback page, the recommendation information is displayed between the first sub-data and the second sub-data in the video data, or, on the playback page, the recommendation information is displayed after the last episode of the sub-data in the video data, which can automatically display the recommendation information of the target content to the user, without the user manually searching for the target content, improving the efficiency of recommending the target content in the video data to the user and improving the efficiency of the user obtaining the relevant information of the target content.

[0067] Figure 6 It is a schematic flowchart of the content recommendation method provided by another embodiment of the present disclosure. As Figure 6 shown, this process includes: Step S602, obtain video data, identify at least one piece of content included in the video data, and determine the target content dimension corresponding to the identified content; Step S604, based on each target content dimension, determine the content recommendation dimension corresponding to the video data, and determine the target content belonging to the content recommendation dimension among the identified content; Step S606, generate recommendation information for recommending the target content, and display the recommendation information in the content recommendation scenario associated with the video data.

[0068] In this embodiment, video data can be obtained, at least one piece of content included in the video data can be recognized, the target content dimension corresponding to the recognized content can be determined, based on each target content dimension, the content recommendation dimension corresponding to the video data can be determined, the target content belonging to the content recommendation dimension can be determined from the recognized content, recommendation information for recommending the target content can be generated, and the recommendation information can be displayed in the content recommendation scenario associated with the video data. Thus, there is no need for the user to manually search for the target content in the video data, and the recommendation information of the target content can be directly and quickly browsed, significantly improving the efficiency of recommending the target content in the video data to the user and the efficiency of the user obtaining relevant information about the target content.

[0069] In the above step S602, the obtained video data includes any one of short plays, short videos, movies, TV dramas, video data recorded by users, and video data received by users. Any video data that can be played can apply the content recommendation method in this embodiment. After obtaining the video data, the above step S602 further recognizes at least one piece of content included in the video data and determines the target content dimension corresponding to the recognized content. Different from the above step S102, step S602 performs offline recognition on the video data, that is, recognizes at least one piece of content included in the video data when the video data is not being played.

[0070] In one embodiment, recognizing at least one piece of content included in the video data includes: Through a content understanding model, according to the entity item dimension, virtual item dimension, character dimension, location dimension, occasion dimension, era dimension, and knowledge point dimension, at least one piece of content included in the video data is recognized.

[0071] In this embodiment, when recognizing at least one piece of content included in the video data through the content understanding model, the at least one piece of content included in the video data can be recognized from a plurality of preset dimensions. The plurality of preset dimensions include: entity item dimension, virtual item dimension, character dimension, location dimension, occasion dimension, era dimension, and knowledge point dimension.

[0072] The specific implementation manner and related examples of this embodiment can refer to the implementation manner and examples of the above step S102, and will not be elaborated here.

[0073] It can be seen that through this embodiment, the content included in the video data can be recognized from multiple dimensions through the content understanding model, making the recognized content more comprehensive and accurate, thereby improving the diversity and accuracy of content recommendation.

[0074] After identifying at least one piece of content included in the video data and determining the target content dimensions corresponding to the identified content in each preset dimension, step S604 above determines the content recommendation dimension corresponding to the video data based on each target content dimension.

[0075] In one embodiment, determining the content recommendation dimension corresponding to the video data based on each target content dimension includes: Determining a second association relationship between each target content dimension according to a first association relationship between each piece of content included in the video data; the first association relationship is used to represent the associated content among each piece of content; the second association relationship is used to represent the associated dimensions among each target content dimension; Combining the associated dimensions according to the second association relationship to obtain the content recommendation dimension corresponding to the video data.

[0076] In this embodiment, first, a first association relationship between each piece of content in the video data is determined, and the first association relationship is used to represent the associated content among each piece of content. The associated content may be two or more pieces of content that appear simultaneously in the video data. For example, in the video data, the male protagonist is wearing a white shirt. The content of the male protagonist and the content of the white shirt appear simultaneously in the video data. Therefore, it is determined that the content of the male protagonist is associated with the content of the white shirt, and there is a first association relationship between the content of the male protagonist and the content of the white shirt.

[0077] Next, according to the first association relationship, a second association relationship between each target content dimension is determined, where the second association relationship is used to represent the associated dimensions among each target content dimension. The target content dimensions corresponding to the associated content are the associated dimensions. There is a second association relationship between the target content dimensions corresponding to the associated content. For example, in the video data, the female protagonist appears in City J. The content of the female protagonist and the content of City J appear simultaneously. The female protagonist and City J are associated content. The target content dimension corresponding to the female protagonist is the character dimension, and the dimension corresponding to City J is the location dimension. It can be determined that the character dimension is associated with the location dimension, and there is a second association relationship between the character dimension and the location dimension.

[0078] Next, combining each associated target content dimension according to the second association relationship between each target content dimension to obtain the content recommendation dimension corresponding to the video data.

[0079] It can be seen that through this embodiment, according to the first association relationship between the various contents included in the video data, the second association relationship between each target content dimension is determined, and according to the second association relationship, the associated dimensions are combined to obtain the content recommendation dimension corresponding to the video data. It is possible to start from multiple different dimensions, combine multiple associated target content dimensions, generate diverse content recommendation dimensions, and thus improve the diversity and accuracy of content recommendation.

[0080] After determining the content recommendation dimension corresponding to the video data, in step S604 above, the target content belonging to the content recommendation dimension is also determined among the recognized contents. The specific implementation manner and related examples of this embodiment can refer to the implementation manner and examples of step S104 above, and will not be elaborated here.

[0081] After determining the content recommendation dimension corresponding to the video data and determining the target content belonging to the content recommendation dimension among the recognized contents, step S606 above generates recommendation information for recommending the target content.

[0082] In one embodiment, generating recommendation information for recommending the target content includes: Through the large language model, obtaining the image data of the target content in the video data, and generating text data that matches the image data and the target content; Through the large language model, generating recommendation information based on the image data and the text data.

[0083] In this embodiment, when generating recommendation information for recommending the target content, the video data and the text representing the target content, such as "the heroine Xiaob eating potato chips", can be input into the large language model first. First, through the large language model, the image data of the target content in the video data is obtained, and the image data can be a video frame in the video data. Then, through the large language model, the text data that matches the target content and the image data corresponding to the target content is obtained. The large language model can analyze the target content and the image data corresponding to the target content and generate text data for describing the characteristics and style of the target content. Finally, through the large language model, the image data and the text data corresponding to the target content are combined to obtain the recommendation information for recommending the target content. The recommendation information can be in the form of pictures and texts, and the pictures and texts include the above-mentioned image data and text data. The recommendation information can also be in the form of a video, and each video frame in the video is the above-mentioned image data, and the subtitles in the video are the above-mentioned text data.

[0084] The specific implementation manner and examples of this embodiment can refer to the relevant implementation manner and examples corresponding to step S106 above, and will not be elaborated here.

[0085] It can be seen that through this embodiment, the large language model is used to obtain the image data of the target content in the video data, and generate the text data that matches the image data and the target content. Based on the image data and the text data, the recommendation information is generated. The recommendation information can be in the form of pictures and texts or videos, and the target content is recommended to the user based on the form of pictures and texts or videos, which improves the accuracy of content recommendation.

[0086] After generating the recommendation information for recommending the target content, in step S606 above, the recommendation information is also displayed in the content recommendation scenario associated with the video data. The content recommendation scenario associated with the video data can be exemplified as the recommendation page in the application program or web page that plays the video data, or the recommendation page in the service platform associated with the video data.

[0087] In one embodiment, displaying the recommendation information in the content recommendation scenario associated with the video data includes at least one of the following ways: In the first page for recommending the video data, display the first recommendation card for recommending the target content; in response to the trigger operation for the first recommendation card, display the recommendation information; In the second page for recommending the video content in the video data, display the second recommendation card for recommending the target content; in response to the trigger operation for the second recommendation card, display the recommendation information; In the page of the service platform associated with the video data, display the third recommendation card for recommending the target content; in response to the trigger operation for the third recommendation card, display the recommendation information.

[0088] In this embodiment, the recommendation information can be displayed in the content recommendation scenario associated with the video data in one or more of the following ways. The first way: in the first page, display the first recommendation card for recommending the target content, and in response to the trigger operation for the first recommendation card, display the recommendation information. Among them, the first page can be the page in the application program or website that plays the video data, and the first page is used to recommend the video data. For example, the first page is used to recommend various short plays. The terminal device can detect the trigger operation of the user for the first recommendation card in the first page, and based on this trigger operation, display the recommendation information of the target content. For example, according to this trigger operation, display the information recommendation page, and in the information recommendation page, display the recommendation information of the target content.

[0089] The second method: On the second page for recommending the video content in the video data, display the second recommendation card for recommending the target content. In response to a trigger operation on the second recommendation card, display the recommendation information. Here, the second page can be a page in an application or website for playing video data, and the second page is used to recommend the content in the video data. For example, the second page is used to recommend the same style of good products in each short drama. The terminal device can detect the trigger operation of the user on the second recommendation card, and based on this trigger operation, display the recommendation information of the target content. For example, according to this trigger operation, display the information recommendation page, and in the information recommendation page, display the recommendation information of the target content.

[0090] The third method: On the page of the service platform associated with the video data, display the third recommendation card for recommending the target content. In response to a trigger operation on the third recommendation card, display the recommendation information. The service platform can be, for example, an e-commerce platform associated with the video data. The terminal device can detect the trigger operation of the user on the third recommendation card, and based on this trigger operation, display the recommendation information of the target content. For example, according to this trigger operation, display the information recommendation page, and in the information recommendation page, display the recommendation information of the target content.

[0091] Figure 7 Schematic diagram of a content recommendation scenario associated with video data provided by an embodiment of the present disclosure, as Figure 7 shown, the first page is the "Short Drama" page included in the short drama playback application. On this "Short Drama" page, multiple short drama cards for recommending short dramas to users are displayed. Below each short drama card, a first recommendation card corresponding to this short drama is displayed. This first recommendation card is used to recommend the target content included in this short drama to the user. This first recommendation card can be, for example, the recommendation card for the same workplace outfit as the female lead in the short drama "Struggle". If the user triggers the first recommendation card, the terminal device displays the information recommendation page, and in the information recommendation page, displays the same workplace outfit as the female lead in the short drama "Struggle" in the form of pictures, texts or videos. The schematic diagram of the information recommendation page can refer to the schematic diagram of the information recommendation page shown above Figure 5c shown, and no examples will be given here.

[0092] Figure 8 Schematic diagram of a content recommendation scenario associated with video data provided by another embodiment of the present disclosure, as Figure 8As shown, the second page is the "Same as in the Drama" page included in the short drama playback application. On this "Same as in the Drama" page, multiple second recommendation cards are displayed. The second recommendation cards are used to recommend the same good items in the short drama. For example, the second recommendation card can be the recommendation card for the same food as the heroine in the short drama "Struggle". If the user triggers the second recommendation card, the terminal device displays an information recommendation page. In the information recommendation page, the same food as the heroine in the short drama "Struggle" is displayed in the form of pictures, texts, or videos. For the schematic diagram of the information recommendation page, reference can be made to the schematic diagram of the information recommendation page shown above Figure 5c The schematic diagram of the information recommendation page shown here will not be exemplified again.

[0093] Figure 9 This is a schematic diagram of the content recommendation scenario associated with video data provided by another embodiment of the present disclosure. As Figure 9 shown, the service platform is exemplified as an e-commerce platform associated with video data. In the e-commerce platform, a third recommendation card is displayed. For example, it is the recommendation card for the same items for playing in City W corresponding to the short drama "Struggle". If the user triggers the third recommendation card, the terminal device displays an information recommendation page. In the information recommendation page, the same items for playing in City W corresponding to the short drama "Struggle" are displayed in the form of pictures, texts, or videos. For the schematic diagram of the information recommendation page, reference can be made to the schematic diagram of the information recommendation page shown above Figure 5c The schematic diagram of the information recommendation page shown here will not be exemplified again. In Figure 9 the shown scenario, when the information recommendation page recommends multiple target contents, a "Purchase" component can be set for each target content respectively. By triggering the "Purchase" component, it jumps to the purchase page of the corresponding target content for purchase.

[0094] It can be seen that through this embodiment, by displaying the recommendation information in the content recommendation scenario associated with video data, users can directly obtain the recommendation information corresponding to the target content in the content recommendation scenario associated with video data, without the need for users to manually search for the target content in the video data, improving the efficiency of recommending the target content in the video data to users and improving the efficiency of users obtaining the relevant information of the target content.

[0095] Regarding Figure 6 the process in Figure 1 and the same points of the process in Figure 1 , reference can be made to the description for

[0096] Figure 10 This is a schematic diagram of the structure of the content recommendation device provided by an embodiment of the present disclosure. As Figure 10 shown, the device includes: The first recognition unit 1001 is configured to play the video data in response to a play instruction for the video data, recognize at least one piece of content included in the video data, and determine a target content dimension corresponding to the recognized content; The first determination unit 1002 is configured to determine a content recommendation dimension corresponding to the video data based on each of the target content dimensions, and determine target content subordinate to the content recommendation dimension from the recognized content; The first display unit 1003 is configured to generate recommendation information for recommending the target content and display the recommendation information in the play scene of the video data.

[0097] Optionally, the first recognition unit 1001 is specifically configured to: through a content understanding model, recognize at least one piece of content included in the video data according to an entity item dimension, a virtual item dimension, a person dimension, a location dimension, an occasion dimension, a time dimension, and a knowledge point dimension.

[0098] Optionally, the first determination unit 1002 is specifically configured to: determine a second association relationship between the target content dimensions according to a first association relationship between the pieces of content included in the video data; the first association relationship is used to represent the associated content among the pieces of content; the second association relationship is used to represent the associated dimensions among the target content dimensions; and combine the associated dimensions according to the second association relationship to obtain a content recommendation dimension corresponding to the video data.

[0099] Optionally, the first display unit 1003 is specifically configured to: through a large language model, obtain image data of the target content in the video data and generate text data matching the image data and the target content; and generate the recommendation information based on the image data and the text data through the large language model.

[0100] Optionally, the first display unit 1003 is specifically configured to: display a recommendation component for recommending the target content on the play page of the video data; and display the recommendation information in response to a trigger operation on the recommendation component.

[0101] Optionally, the video data includes multiple sets of sub-data played in chronological order on the play page of the video data; the first display unit 1003 is specifically configured to: display the recommendation information between a first sub-data and a second sub-data in the video data on the play page; or display the recommendation information after the last set of sub-data in the video data on the play page.

[0102] In this embodiment, it is capable of responding to a play instruction for video data, playing the video data, identifying at least one piece of content included in the video data, determining a target content dimension corresponding to the identified content, determining a content recommendation dimension corresponding to the video data based on each target content dimension, determining target content subordinate to the content recommendation dimension among the identified content, generating recommendation information for recommending the target content, and displaying the recommendation information in the play scenario of the video data. Thus, without the user manually searching for the target content in the video data, the user can directly and quickly view the recommendation information of the target content, significantly improving the efficiency of recommending the target content in the video data to the user and enhancing the efficiency of the user obtaining relevant information about the target content.

[0103] The content recommendation device in the embodiments of the present disclosure can implement each process of the content recommendation method embodiment shown in step S102 - step S106 above, and achieve the same effects and functions, which will not be repeated here.

[0104] Figure 11 FIG. is a schematic structural diagram of a content recommendation device provided in another embodiment of the present disclosure, as Figure 11 shown, the device includes: A second recognition unit 1101, configured to obtain video data, identify at least one piece of content included in the video data, and determine a target content dimension corresponding to the identified content; A second determination unit 1102, configured to determine a content recommendation dimension corresponding to the video data based on each target content dimension, and determine target content subordinate to the content recommendation dimension among the identified content; A second display unit 1103, configured to generate recommendation information for recommending the target content and display the recommendation information in a content recommendation scenario associated with the video data.

[0105] Optionally, the second recognition unit 1101 is specifically configured to: through a content understanding model, identify at least one piece of content included in the video data according to an entity item dimension, a virtual item dimension, a person dimension, a location dimension, an occasion dimension, a decade dimension, and a knowledge point dimension.

[0106] Optionally, the second determination unit 1102 is specifically configured to: determine a second association relationship between each target content dimension according to a first association relationship between each piece of content included in the video data; the first association relationship is used to represent the associated content among each piece of content; the second association relationship is used to represent the associated dimensions among each target content dimension; and combine the associated dimensions according to the second association relationship to obtain a content recommendation dimension corresponding to the video data.

[0107] Optionally, the second display unit 1103 is specifically configured to: obtain, through a large language model, the image data of the target content in the video data, and generate text data that matches the image data and the target content; and generate the recommendation information based on the image data and the text data through the large language model.

[0108] Optionally, the second display unit 1103 is specifically configured to implement at least one of the following: display a first recommendation card for recommending the target content on a first page for recommending the video data; display the recommendation information in response to a trigger operation on the first recommendation card; display a second recommendation card for recommending the target content on a second page for recommending video content in the video data; display the recommendation information in response to a trigger operation on the second recommendation card; display a third recommendation card for recommending the target content on a page of a service platform associated with the video data; and display the recommendation information in response to a trigger operation on the third recommendation card.

[0109] In this embodiment, it is possible to obtain video data, identify at least one piece of content included in the video data, determine the target content dimension corresponding to the identified content, determine the content recommendation dimension corresponding to the video data based on each target content dimension, determine the target content subordinate to the content recommendation dimension from the identified content, generate recommendation information for recommending the target content, and display the recommendation information in a content recommendation scenario associated with the video data. Thus, there is no need for the user to manually search for the target content in the video data, and the user can directly and quickly view the recommendation information of the target content, significantly improving the efficiency of recommending the target content in the video data to the user and the efficiency of the user obtaining relevant information about the target content.

[0110] The content recommendation device in the embodiments of the present disclosure can implement each process of the content recommendation method embodiment shown in step S602 - step S606 above, and achieve the same effects and functions, which will not be repeated here.

[0111] An embodiment of the present disclosure further provides an electronic device. Figure 12 It is a schematic structural diagram of the electronic device provided by an embodiment of the present disclosure, as Figure 12As shown, electronic devices can vary significantly due to differences in configuration or performance. They can include one or more processors 1201 and a memory 1202. The memory 1202 can store one or more applications or data. Among them, the memory 1202 can be transient storage or persistent storage. The applications stored in the memory 1202 can include one or more modules (not shown in the figure), and each module can include a series of computer-executable instructions in the electronic device. Further, the processor 1201 can be set to communicate with the memory 1202 and execute a series of computer-executable instructions in the memory 1202 on the electronic device. The electronic device can also include one or more power supplies 203, one or more wired or wireless network interfaces 1204, one or more input or output interfaces 1205, one or more keyboards 1206, etc.

[0112] In a specific embodiment, the electronic device includes a processor; and a memory configured to store computer-executable instructions, which when executed cause the processor to implement the following processes: In response to a play instruction for video data, play the video data, identify at least one piece of content included in the video data, and determine the target content dimension corresponding to the identified content; Based on each of the target content dimensions, determine the content recommendation dimension corresponding to the video data, and determine the target content subordinate to the content recommendation dimension among the identified content; Generate recommendation information for recommending the target content, and display the recommendation information in the play scene of the video data.

[0113] In this embodiment, it is possible to respond to a play instruction for video data, play the video data, identify at least one piece of content included in the video data, and determine the target content dimension corresponding to the identified content. Based on each target content dimension, determine the content recommendation dimension corresponding to the video data, determine the target content subordinate to the content recommendation dimension among the identified content, generate recommendation information for recommending the target content, and display the recommendation information in the play scene of the video data. Thus, there is no need for the user to manually search for the target content in the video data, and the recommendation information of the target content can be directly and quickly browsed, significantly improving the efficiency of recommending the target content in the video data to the user and improving the efficiency of the user obtaining relevant information about the target content.

[0114] The electronic device in the embodiments of the present disclosure can implement each process of the content recommendation method embodiment shown in the above steps S102 - step S106, and achieve the same effects and functions, which will not be repeated here.

[0115] In another specific embodiment, the electronic device includes a processor; and a memory configured to store computer-executable instructions, which when executed cause the processor to implement the following processes: Obtain video data, identify at least one piece of content included in the video data, and determine the target content dimension corresponding to the identified content; Based on each of the target content dimensions, determine the content recommendation dimension corresponding to the video data, and determine the target content subordinate to the content recommendation dimension among the identified content; Generate recommendation information for recommending the target content, and display the recommendation information in the content recommendation scenario associated with the video data.

[0116] In this embodiment, it is possible to obtain video data, identify at least one piece of content included in the video data, determine the target content dimension corresponding to the identified content, determine the content recommendation dimension corresponding to the video data based on each target content dimension, determine the target content subordinate to the content recommendation dimension among the identified content, generate recommendation information for recommending the target content, and display the recommendation information in the content recommendation scenario associated with the video data. Thus, there is no need for the user to manually search for the target content in the video data, and the user can directly and quickly view the recommendation information of the target content, significantly improving the efficiency of recommending the target content in the video data to the user and the efficiency of the user obtaining relevant information of the target content.

[0117] The electronic device in the embodiments of the present disclosure can implement each process of the content recommendation method embodiment shown in step S602-step S606 above, and achieve the same effects and functions, which will not be repeated here.

[0118] Another embodiment of the present disclosure also provides a computer-readable storage medium, which is used to store computer-executable instructions, and the computer-executable instructions, when executed by a processor, implement the following processes: In response to a play instruction for video data, play the video data, identify at least one piece of content included in the video data, and determine the target content dimension corresponding to the identified content; Based on each of the target content dimensions, determine the content recommendation dimension corresponding to the video data, and determine the target content subordinate to the content recommendation dimension among the identified content; Generate recommendation information for recommending the target content, and display the recommendation information in the play scenario of the video data.

[0119] In this embodiment, it is capable of responding to a playback instruction for video data, playing the video data, identifying at least one piece of content included in the video data, determining the target content dimension corresponding to the identified content, determining the content recommendation dimension corresponding to the video data based on each target content dimension, determining the target content subordinate to the content recommendation dimension among the identified content, generating recommendation information for recommending the target content, and presenting the recommendation information in the playback scenario of the video data. Thus, without the user manually searching for the target content in the video data, the user can directly and quickly view the recommendation information of the target content, significantly improving the efficiency of recommending the target content in the video data to the user and the efficiency of the user obtaining relevant information about the target content.

[0120] The computer-readable storage medium in the embodiments of the present disclosure can implement each process of the content recommendation method embodiment shown in the above steps S102 - S106, and achieve the same effects and functions, which will not be repeated here.

[0121] In another specific embodiment, the computer-readable storage medium is used to store computer-executable instructions, and when the computer-executable instructions are executed by a processor, the following process is implemented: Obtain video data, identify at least one piece of content included in the video data, and determine the target content dimension corresponding to the identified content; Based on each of the target content dimensions, determine the content recommendation dimension corresponding to the video data, and determine the target content subordinate to the content recommendation dimension among the identified content; Generate recommendation information for recommending the target content, and present the recommendation information in the content recommendation scenario associated with the video data.

[0122] In this embodiment, it is capable of obtaining video data, identifying at least one piece of content included in the video data, determining the target content dimension corresponding to the identified content, determining the content recommendation dimension corresponding to the video data based on each target content dimension, determining the target content subordinate to the content recommendation dimension among the identified content, generating recommendation information for recommending the target content, and presenting the recommendation information in the content recommendation scenario associated with the video data. Thus, without the user manually searching for the target content in the video data, the user can directly and quickly view the recommendation information of the target content, significantly improving the efficiency of recommending the target content in the video data to the user and the efficiency of the user obtaining relevant information about the target content.

[0123] The computer-readable storage medium in the embodiments of the present disclosure can implement each process of the content recommendation method embodiment shown in the above steps S602 - S606, and achieve the same effects and functions, which will not be repeated here.

[0124] Another embodiment of the present disclosure also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the following processes are implemented: In response to a play instruction for video data, play the video data, identify at least one piece of content included in the video data, and determine a target content dimension corresponding to the identified content; Based on each of the target content dimensions, determine a content recommendation dimension corresponding to the video data, and determine target content subordinate to the content recommendation dimension among the identified content; Generate recommendation information for recommending the target content, and display the recommendation information in the play scenario of the video data.

[0125] In this embodiment, it is possible to respond to a play instruction for video data, play the video data, identify at least one piece of content included in the video data, determine a target content dimension corresponding to the identified content, determine a content recommendation dimension corresponding to the video data based on each of the target content dimensions, determine target content subordinate to the content recommendation dimension among the identified content, generate recommendation information for recommending the target content, and display the recommendation information in the play scenario of the video data. Thus, there is no need for the user to manually search for the target content in the video data, and the recommendation information of the target content can be directly and quickly browsed, significantly improving the efficiency of recommending the target content in the video data to the user and the efficiency of the user obtaining relevant information of the target content.

[0126] The computer program product in the embodiment of the present disclosure can implement each process of the content recommendation method embodiment shown in the above steps S102 - S106, and achieve the same effects and functions, which will not be repeated here.

[0127] In another specific embodiment, the computer program product includes a computer program. When the computer program is executed by a processor, the following processes are implemented: Obtain video data, identify at least one piece of content included in the video data, and determine a target content dimension corresponding to the identified content; Based on each of the target content dimensions, determine a content recommendation dimension corresponding to the video data, and determine target content subordinate to the content recommendation dimension among the identified content; Generate recommendation information for recommending the target content, and display the recommendation information in a content recommendation scenario associated with the video data.

[0128] In this embodiment, video data can be obtained, at least one piece of content included in the video data can be recognized, and a target content dimension corresponding to the recognized content can be determined. Based on each target content dimension, a content recommendation dimension corresponding to the video data can be determined. Target content belonging to the content recommendation dimension can be determined from the recognized content, recommendation information for recommending the target content can be generated, and the recommendation information can be displayed in a content recommendation scenario associated with the video data. Thus, without the user manually searching for the target content in the video data, the recommendation information of the target content can be directly and quickly browsed, significantly improving the efficiency of recommending the target content in the video data to the user and the efficiency of the user obtaining relevant information about the target content.

[0129] The computer program product in the embodiments of the present disclosure can implement each process of the content recommendation method embodiment shown in step S602 - step S606 above, and achieve the same effects and functions, which will not be repeated here.

[0130] In each embodiment of the present disclosure, the computer-readable storage medium includes a read-only memory (ROM for short), a random access memory (RAM for short), a magnetic disk, an optical disc, etc.

[0131] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structures of diodes, transistors, switches, etc.) or software improvements (improvements to method flows). However, with the development of technology, many method flow improvements today can be regarded as direct improvements to hardware circuit structures. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented using a hardware entity module. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is an integrated circuit whose logical function is determined by the user programming the device. Designers can program themselves to "integrate" a digital system onto a single PLD, without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compilers used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a hardware description language (HDL), and there is not just one type of HDL, but many types, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply performing a little logical programming on the method flow using the above-mentioned several hardware description languages and programming it into an integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.

[0132] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to make the controller implement the same function in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or structures within the hardware component.

[0133] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0134] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing the embodiments of the present disclosure, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0135] Those skilled in the art should understand that one or more embodiments of the present disclosure can be provided as a method, a system, or a computer program product. Therefore, one or more embodiments of the present disclosure can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, one or more embodiments of the present disclosure can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.

[0136] This disclosure is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the disclosure. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate means for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.

[0137] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.

[0138] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.

[0139] It should also be noted that the term "comprising", "including", or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity, or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, commodity, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, commodity, or device including the said element.

[0140] One or more embodiments of the present disclosure may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of the present disclosure may also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media including storage devices.

[0141] Each embodiment in the present disclosure is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and for the relevant parts, reference can be made to the description of the method embodiment.

[0142] The above description is only for the embodiments of the present disclosure and is not intended to limit the present disclosure. For those skilled in the art, various changes and modifications can be made to the present disclosure. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present disclosure shall be included within the scope of the claims of the present disclosure.

Claims

1. A content recommendation method, characterized in that, Including: In response to a play instruction for video data, play the video data, identify at least one piece of content included in the video data, and determine a target content dimension corresponding to the identified content; Based on each of the target content dimensions, determine a content recommendation dimension corresponding to the video data, and determine target content subordinate to the content recommendation dimension among the identified content; Generate recommendation information for recommending the target content, and display the recommendation information in the play scenario of the video data.

2. The method according to claim 1, wherein The identifying at least one piece of content included in the video data includes: Through a content understanding model, according to entity item dimension, virtual item dimension, person dimension, location dimension, occasion dimension, era dimension, knowledge point dimension, identify at least one piece of content included in the video data.

3. The method according to claim 1, wherein The determining the content recommendation dimension corresponding to the video data based on each of the target content dimensions includes: According to a first association relationship between each piece of content included in the video data, determine a second association relationship between each of the target content dimensions; the first association relationship is used to represent the associated content among each piece of content; the second association relationship is used to represent the associated dimensions among each of the target content dimensions; According to the second association relationship, combine the associated dimensions to obtain the content recommendation dimension corresponding to the video data.

4. The method according to claim 1, wherein The generating the recommendation information for recommending the target content includes: Through a large language model, obtain image data of the target content in the video data, and generate text data matching the image data and the target content; Through the large language model, generate the recommendation information based on the image data and the text data.

5. The method according to claim 1, characterized in that, The displaying the recommendation information in the play scenario of the video data includes: In the play page of the video data, display a recommendation component for recommending the target content; In response to a trigger operation on the recommendation component, display the recommendation information.

6. The method according to claim 1, characterized in that, The video data includes multiple sets of sub-data played in chronological order in the play page of the video data; the displaying the recommendation information in the play scenario of the video data includes: In the play page, display the recommendation information between a first sub-data and a second sub-data in the video data; Or, In the play page, display the recommendation information after the last set of sub-data in the video data.

7. A content recommendation method, characterized in that, Including: Obtain video data, identify at least one piece of content included in the video data, and determine a target content dimension corresponding to the identified content; Based on each of the target content dimensions, determine a content recommendation dimension corresponding to the video data, and determine target content subordinate to the content recommendation dimension among the identified content; Generate recommendation information for recommending the target content, and display the recommendation information in a content recommendation scenario associated with the video data.

8. The method according to claim 7, characterized in that The identifying at least one piece of content included in the video data includes: Through a content understanding model, at least one piece of content included in the video data is identified according to the dimensions of physical items, virtual items, characters, locations, occasions, eras, and knowledge points.

9. The method according to claim 7, wherein Based on each of the target content dimensions, determining the content recommendation dimension corresponding to the video data includes: Determining a second association relationship between each of the target content dimensions according to a first association relationship between each piece of content included in the video data; the first association relationship is used to represent the associated content among each piece of content; the second association relationship is used to represent the associated dimensions among each of the target content dimensions; Combining the associated dimensions according to the second association relationship to obtain the content recommendation dimension corresponding to the video data.

10. The method according to claim 7, wherein Generating recommendation information for recommending the target content includes: Through a large language model, obtaining image data of the target content in the video data, and generating text data that matches the image data and the target content; Through the large language model, generating the recommendation information based on the image data and the text data.

11. The method according to claim 7, wherein Displaying the recommendation information in a content recommendation scenario associated with the video data includes at least one of the following methods: In a first page for recommending the video data, displaying a first recommendation card for recommending the target content; In response to a trigger operation on the first recommendation card, displaying the recommendation information; In a second page for recommending video content in the video data, displaying a second recommendation card for recommending the target content; In response to a trigger operation on the second recommendation card, displaying the recommendation information; In a page of a service platform associated with the video data, displaying a third recommendation card for recommending the target content; In response to a trigger operation on the third recommendation card, displaying the recommendation information.

12. A content recommendation device, characterized in that, Including: A first recognition unit, configured to play the video data in response to a play instruction for the video data, identify at least one piece of content included in the video data, and determine the target content dimension corresponding to the identified content; A first determination unit, configured to determine the content recommendation dimension corresponding to the video data based on each of the target content dimensions, and determine the target content subordinate to the content recommendation dimension among the identified content; A first display unit, configured to generate recommendation information for recommending the target content and display the recommendation information in a playing scenario of the video data.

13. A content recommendation device, characterized in that, Including: A second recognition unit, configured to obtain video data, identify at least one piece of content included in the video data, and determine the target content dimension corresponding to the identified content; A second determination unit, configured to determine the content recommendation dimension corresponding to the video data based on each of the target content dimensions, and determine the target content subordinate to the content recommendation dimension among the identified content; A second display unit, configured to generate recommendation information for recommending the target content and display the recommendation information in a content recommendation scenario associated with the video data.

14. An electronic device, characterized in that, Comprising: A processor; And A memory configured to store computer-executable instructions, which when executed cause the processor to implement the method according to any one of claims 1-6 or the method according to any one of claims 7-11.

15. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store computer-executable instructions, which when executed by a processor implement the method according to any one of claims 1-6 or the method according to any one of claims 7-11.

16. A computer program product, characterized in that, The computer program product includes a computer program, which when executed by a processor implements the method according to any one of claims 1-6 or the method according to any one of claims 7-11.