Method, apparatus, device, and storage medium for processing multimedia content
The method and device enhance user experience by recognizing and displaying recommended objects in multimedia content, addressing limitations in existing video technologies to provide an enriched interaction and consumption path.
Patent Information
- Application Number
- JP2025504038
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-02-14
- Filing Date
- 2024-01-30
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2044-01-30
AI Technical Summary
Existing video-related technologies lack the ability to enrich user experience by providing an effective path for consuming related information due to limitations in display area and interaction functionality.
A method and device that recognize target resource objects in multimedia content, determine corresponding recommended objects based on pre-set relationships, and display them in a card format, allowing users to explore these objects through slide and interaction operations.
Enhances user experience by providing an extended consumption path for related content, enriching interaction functions and improving the viewing experience.
Smart Images

Figure 2025524054000001_ABST
Abstract
Description
Technical Field
[0001] (Reference to Related Applications) This application claims priority based on a Chinese patent application with the invention title "Method, Apparatus, Device, and Storage Medium for Processing Multimedia Content", application number 2023101598223, filed on February 14, 2023, and incorporates the entire content of the application by reference into this application.
[0002] (Technical Field) The present disclosure relates to the field of data processing, and particularly to a method, apparatus, device, and storage medium for processing multimedia content.
Background Art
[0003] With the continuous development of video processing technology, people's needs for video-related functions are becoming increasingly diverse. Therefore, in order to meet the needs of more users and improve the user experience, how to enrich video-related functions has become an urgent technical problem to be solved.
Summary of the Invention
[0004] To solve the above technical problems, the present disclosure provides a method, apparatus, device, and storage medium for processing multimedia content that can improve the user experience by enriching the video-based interaction function.
[0005] In a first aspect, the present disclosure provides a method for processing multimedia content, the method comprising: responding to a pre-set trigger operation acting on a display page of first multimedia content, and recognizing at least one target resource object attached to the first multimedia content, wherein there is a pre-set correspondence between the type of the target resource object and the type of a recommended object; determining, based on the pre-set correspondence, the type of a recommended object corresponding to a first target resource object among the at least one target resource object; determining at least one recommended object based on the first target resource object, wherein the at least one recommended object belongs to the type of the recommended object corresponding to the first target resource object; and displaying the at least one recommended object.
[0006] In an alternative embodiment, the at least one target resource object further includes a second target resource object, and displaying the at least one recommended object as described above includes classifying and displaying the recommended objects respectively determined based on the first target resource object and the second resource object according to the type of the recommended object.
[0007] In an alternative embodiment, classifying and displaying the recommended objects respectively determined based on the first target resource object and the second resource object according to the type of the recommended object as described above includes displaying at least one first recommended object determined based on the first target resource object on a first card, wherein the first recommended object belongs to the type of the recommended object corresponding to the first target resource object; and displaying at least one second recommended object determined based on the second resource object on a second card, wherein the second recommended object belongs to the type of the recommended object corresponding to the second target resource object, and the first card and the second card belong to a card set that is stacked and displayed.
[0008] In an alternative embodiment, the method further includes scrolling and displaying each card in the card set in response to a pre-set slide operation triggered on the card set.
[0009] In an alternative embodiment, the method further includes, in response to a pre-set trigger operation on a target card in the card set, displaying the recommended object on the target card on a recommended object display page, and receiving a pre-set interaction operation on the target recommended object on the recommended object display page.
[0010] In an alternative embodiment, the at least one target resource object includes an item object, the type of the recommended object corresponding to the item object includes an item type, and determining at least one recommended object based on the first target resource object as described above includes determining at least one recommended item having the same or similar characteristics as the item object and belonging to the item type.
[0011] In an alternative embodiment, the at least one target resource object includes BGM, the type of the recommended object corresponding to the BGM includes a music type, and the at least one recommended object determined based on the first target resource object as described above includes performing music recognition on the BGM to obtain a music recognition result, and determining a vocal music piece resource corresponding to the BGM and belonging to the music type based on the music recognition result.
[0012] In an alternative embodiment, the at least one target resource object includes address information, the type of the recommended object corresponding to the address information includes a life service type, and the at least one recommended object determined based on the first target resource object as described above includes determining at least one life service object that is within a preset distance range centered on the address information and belongs to the life service type.
[0013] In an alternative embodiment, the at least one target resource object includes a target face displayed on a video frame screen, the type of the recommended object corresponding to the target face includes a user account type, and the at least one recommended object determined based on the first target resource object as described above includes determining at least one user account that reaches a preset threshold in similarity between the user avatar and the target face based on the target face displayed on the video frame screen and belongs to the user account type.
[0014] In an alternative embodiment, in response to a pre-set trigger operation acting on the display page of the first multimedia content as described above, recognizing at least one target resource object attached to the first multimedia content includes: in response to a pre-set trigger operation acting on the display page of the first multimedia content, displaying a plurality of video keyframe screens within the first multimedia content on a video recognition page in the form of a transition animation; and based on the plurality of video keyframe screens, recognizing at least one target resource object attached to the first multimedia content.
[0015] In a second aspect, the present disclosure provides a processing device for multimedia content, including a recognition module for recognizing at least one target resource object attached to the first multimedia content in response to a pre-set trigger operation acting on the playback page of the first multimedia content, where the recognition module has a pre-set correspondence relationship between the target resource object and the type of the recommended object; a first determination module for determining the type of the recommended object corresponding to the first target resource object among the at least one target resource object based on the pre-set correspondence relationship; a second determination module for determining at least one recommended object based on the first target resource object, where the at least one recommended object belongs to the type of the recommended object corresponding to the first target resource object; and a display module for displaying the at least one recommended object.
[0016] In a third aspect, the present disclosure provides a computer-readable storage medium, which stores instructions that, when executed on a terminal device, cause the terminal device to implement the method.
[0017] In a fourth aspect, the present disclosure provides a multimedia content processing device, comprising a memory, a processor, and a computer program stored in the memory and executable by the processor, and when the processor executes the computer program, the method is realized.
[0018] In a fifth aspect, the present disclosure provides a computer program product, which stores a computer program / instructions that, when executed by a processor, realizes the above-described method.
[0019] The technical solution provided by the embodiments of the present disclosure has at least the following advantages over the prior art.
[0020] The embodiments of the present disclosure provide a method for processing multimedia content. First, in response to a pre-set trigger operation acting on the display page of the first multimedia content, at least one target resource object attached to the first multimedia content is determined, and among them, there is a pre-set object relationship between the target resource object and the type of the recommended object. Next, based on the pre-set correspondence relationship, the type of the recommended object corresponding to the first target resource object among the at least one target resource object is determined, and at least one recommended object is determined based on the first target resource object. Then, the at least one recommended object is displayed. The embodiments of the present disclosure display, to the user, recommended objects related to the target resource object based on the target resource object attached to the multimedia content during the display of the multimedia content. Therefore, the embodiments of the present disclosure provide an extended consumption path for the content attached to the multimedia content to the user, improving the user experience.
Brief Description of the Drawings
[0021] The accompanying drawings are incorporated in and form a part of the specification, showing embodiments of the present disclosure and used, together with the specification, to explain the principles of the present disclosure.
[0022] Hereinafter, to more clearly explain the embodiments of the present disclosure or the technical configurations in the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly described. It will be obvious to those skilled in the art that other accompanying drawings can be obtained based on the accompanying drawings without creative effort.
[0023]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Embodiments for Carrying Out the Invention
[0024] To more clearly understand the above-mentioned objects, features, and advantages of the present disclosure, the embodiments of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features in the embodiments can be combined with each other as long as they do not conflict with each other.
[0025] To facilitate a complete understanding of the present disclosure, a lot of details are described in the following explanations. However, the present disclosure can also be implemented in other ways different from those described in this specification. It will be obvious that the examples in the specification are only a part of the embodiments of the present disclosure, not all of the embodiments.
[0026] With the continuous development of video processing technology, people's needs for video-related functions are becoming increasingly diversified. In order to meet the needs of more users and improve the user experience, how to enrich video-related functions has become an urgent technical issue to be solved.
[0027] Also, for example, the related information attached to videos, such as the items appearing in the video, the places mentioned, the characters appearing, and the BGM used, is also becoming increasingly rich. There is a possibility that further consumption ideas will arise for the related information attached to the video while the user is watching the video.
[0028] In related technologies, by adopting the method of adding function anchors to the video playback page, a path for the user to further consume the related information attached to the video is provided. However, there are limitations in the display area of the video playback page. If function anchors are overly displayed, it may affect the simplification of the video playback page and the user's video viewing experience.
[0029] For this purpose, the embodiments of the present disclosure provide a method for processing multimedia content. First, in response to a pre-set trigger operation acting on the display page of the first multimedia content, at least one target resource object attached to the first multimedia content is recognized. In some embodiments, there is a pre-set correspondence between the type of the target resource object and the type of the recommended object. Then, based on the pre-set correspondence, the type of the recommended object corresponding to the first target resource object among the at least one target resource object is determined, and at least one recommended object is determined based on the first target resource object. And the at least one recommended object is displayed. The embodiments of the present disclosure can display, to the user, recommended objects related to the target resource object based on the target resource object attached to the multimedia content during the display of the multimedia content, so as to provide the user with an extended consumption path for the content attached to the multimedia content and improve the user experience.
[0030] Based on this, the embodiments of the present disclosure provide a method for processing multimedia content. FIG. 1 is a flowchart of the method for processing multimedia content provided by the embodiments of the present disclosure, and the method includes the following steps.
[0031] S101: In response to a pre-set trigger operation acting on the display page of the first multimedia content, recognize at least one target resource object attached to the first multimedia content.
[0032] In some embodiments, there is a pre-set correspondence between the type of the target resource object and the type of the recommended object.
[0033] In some embodiments, the first multimedia content includes any one of video, audio, graphic content, text content, image content, etc. Specifically, the first multimedia content is any one of the multimedia content in any multimedia information stream, such as any one of the recommended videos in the recommended video stream.
[0034] In some embodiments, the preset trigger operations acting on the display page of the first multimedia content include, but are not limited to in the embodiments of the present disclosure, the knuckle double-click operation, long-press operation, double-click operation, etc. acting on the display page of the first multimedia content.
[0035] In some embodiments, the target resource object may be an object object displayed on the video frame screen attached to the first multimedia content, such as a sweatshirt, a stool, a pet dog, etc., or may be the text content (e.g., subtitles) attached to the first multimedia content or the address information appearing in the location anchor (e.g., a certain tourist attraction, a hooded coat, a supermarket, etc.), or may be the BGM played in the first multimedia content, or may be a public figure appearing on the video frame screen attached to the first multimedia content.
[0036] In some embodiments, when receiving a pre-set trigger operation acting on the display page of the first multimedia content, at least one target resource object attached to the first multimedia content can be recognized, for example, recognizing the first multimedia content using text recognition technology (abbreviated as Optical Character Recognition, OCR), speech recognition technology (abbreviated as Automatic Speech Recognition, ASR), face recognition technology, etc., and obtaining the target resource object attached to the first multimedia content. Regarding the specific recognition method, since different target resource objects will be described in subsequent embodiments, it is omitted here.
[0037] In some embodiments, the type of the recommended object is used to recognize the type to which the recommended object determined based on the target resource object belongs. In some embodiments, the type of the recommended object includes item type, music type, life service type, user account type, etc., and the embodiments of the present disclosure are not limited thereto and can be set according to actual needs.
[0038] In some embodiments, there is a pre-set correspondence between the target resource object and the type of the recommended object, and different target resource objects can correspond to the same or different types of recommended objects. For example, when the target resource object is a sweatshirt, the corresponding type of the recommended object is the item type, and when the target resource object is music, the corresponding type of the recommended object is the music type.
[0039] In actual applications, assuming that the first multimedia content includes video, in order to improve the user's viewing experience of the multimedia content, when a preset trigger operation acting on the first multimedia content is received, with a certain expectation for the function of determining the corresponding recommended object based on the target resource object attached to the first multimedia content, it is also possible to display a plurality of video key-frame screens in the first multimedia content in the form of a transition animation on the video recognition page. Thereby, based on the plurality of video key-frame screens, at least one target resource object attached to the first multimedia content is recognized.
[0040] In an alternative embodiment, for example, every 10 frames, the video key-frame screens in the first multimedia content may be cut at a preset frame interval so as to cut a video frame screen for the first multimedia content, and a plurality of video key-frame screens are obtained.
[0041] In another alternative embodiment, for example, every 1 second, the video key-frame screens in the first multimedia content may be cut at a preset time interval so as to cut a video frame screen for the first multimedia content, and a plurality of video key-frame screens are obtained.
[0042] FIG. 2 is a schematic diagram of a video recognition page provided by an embodiment of the present disclosure, in which a plurality of video key-frame images are displayed on the video recognition page according to a preset movement trajectory or a random running trajectory. In some embodiments, the preset movement trajectory includes, for example, the movement from the center position to the edge position of the video recognition page.
[0043] In an alternative embodiment, based on the style of the cards displayed in a stacked manner, a plurality of key video frames in the first multimedia content may be displayed on different cards of the video recognition page, thereby recognizing at least one target resource object attached to the first multimedia content based on the plurality of video key frames. FIG. 3 shows a schematic diagram of another video recognition page provided by an embodiment of the present disclosure. In some embodiments, the plurality of video key frame screens are superimposed and displayed on the video recognition page in the style of cards.
[0044] S102: Determine the type of the recommended object corresponding to the first target resource object among the at least one target resource object based on the pre-set corresponding relationship.
[0045] S103: Determine at least one recommended object based on the first target resource object.
[0046] In some embodiments, the at least one recommended object belongs to the type of the recommended object corresponding to the first target resource object.
[0047] In some embodiments, the first target resource object is any one of the resource objects among the at least one target resource object recognized from the first multimedia content.
[0048] In some embodiments, after recognizing at least one target resource object attached to the first multimedia content, based on any one of the at least one target resource object, it is possible to determine at least one recommended object belonging to the type of the recommended object corresponding to the target resource object.
[0049] In some embodiments, the first target resource object and at least one recommended object have the same or similar characteristics. For example, assuming that the first target resource object is the rusty part of a tune, based on the type of the rusty part and the corresponding recommended object (i.e., the music type), the entire tune corresponding to the rusty part is determined as the recommended object corresponding to the first target resource object.
[0050] S104: Display the at least one recommended object.
[0051] In some embodiments, after determining at least one recommended object based on the target resource object, individual recommended objects are displayed.
[0052] In an alternative embodiment, the at least one target resource object may further include a second target resource object. In some embodiments, the second target resource object and the first target resource object belong to different types of recommended objects.
[0053] Therefore, displaying the recommended object further includes classifying and displaying the recommended object according to the type of the recommended object to which each recommended object belongs, after determining at least one recommended object based on the first target resource object and the second target resource object.
[0054] As an example, assume that the first target resource object is a white sweatshirt corresponding to the item type, and the second target resource object is the refrain part corresponding to the music type. The recommended objects determined based on the white sweatshirt include a white long-sleeved sweatshirt and a white short-sleeved sweatshirt. The recommended object determined based on music fragment A is the entire piece of music B. Then, according to the type of the recommended object, the white long-sleeved body suit, the white short-sleeved body suit, and music B are classified and displayed. That is, the white long-sleeved body suit and the white short-sleeved body suit are displayed together as recommended objects of the item type, and music B is displayed as a recommended object of the music type.
[0055] Note that the above is only an example in which at least one target resource object includes two target resource objects. The embodiments of the present disclosure do not limit the number of target resource objects recognized from the first multimedia content. Since the display method of the recommended objects corresponding to the plurality of target resource objects can be implemented with reference to the above, it is omitted here.
[0056] In the method for processing multimedia content provided by the embodiments of the present disclosure, first, in response to a pre-set trigger operation acting on the display page of the first multimedia content, at least one target resource object attached to the first multimedia content is recognized, among which, there is a pre-set object relationship between the type of the target resource object and the type of the recommended object. And based on the pre-set corresponding relationship, the type of the recommended object corresponding to the first target resource object among the at least one target resource object is determined, and at least one recommended object is determined based on the first target resource object. Then, at least one recommended object is displayed. The embodiments of the present disclosure can display, to the user, recommended objects related to the target resource object based on the target resource object attached to the multimedia content during the display of the multimedia content. Thereby, the embodiments of the present disclosure provide the user with an extended consumption path for the content attached to the multimedia content and improve the user experience.
[0057] In actual application, in order to enrich the interactive function during the display of multimedia content and improve the viewing experience of the multimedia content by the user, it is also possible to display recommended objects of different types of recommended objects in the form of cards. Specifically, in the form of cards, the recommended objects determined based on the first target resource object and the second target resource object are classified and displayed, among which, the first target resource object and the second target resource object respectively correspond to different types of recommended objects.
[0058] Specifically, at least one first recommended object determined based on the first target resource object is displayed on the first card, and among them, the first recommended object belongs to the type of recommended object corresponding to the first target resource object. At least one second recommended object determined based on the second resource object is displayed on the second card, and among them, the second recommended object belongs to the type of recommended object corresponding to the second target resource object.
[0059] FIG. 4 is a schematic diagram of a display page corresponding to a card set provided by an embodiment of the present disclosure. Taking two cards as an example, on the first card 401, recommended objects determined based on a white top (the first target resource object), such as a white long-sleeved top or a white short-sleeved top, are displayed. On the second card 402, recommended objects determined based on a music clip A (the second target resource object), such as music B, are displayed. In some embodiments, the first card and the second card belong to a card set that is stacked and displayed.
[0060] In an alternative embodiment, while the card set is stacked and displayed on the display page corresponding to the card set, a pre-set slide operation is triggered on the card set to facilitate the user to select and display a desired card based on the content displayed on the individual cards. By this operation, individual cards in the card set can be scrolled and displayed.
[0061] In some embodiments, to trigger the scrolling display of individual cards in the card set, the pre-set slide operation triggered on the card set includes an upward slide operation, a downward slide operation, etc. on the card set.
[0062] As shown in FIG. 4, when an upward slide operation on the card set is received, the currently displayed card 401 is switched upward to the next adjacent card, i.e., card 402, and the individual recommended objects of card 402 are displayed on the display page where the card set is located. When a downward slide operation on the card set is received, the currently displayed card 401 is switched downward to the previous adjacent card, i.e., card 403, in order to be fully displayed on the display page where the card set is located.
[0063] In an alternative embodiment, when the user wants to end the display of the card set, it is possible to return to the playback page of the first multimedia content by a pre-set return operation on the card set. In some embodiments, the pre-set return operation on the card set may include, for example, a left slide operation acting on the card set. Further, it is possible to provide a return control 404 on the display page where the card set is located, and by clicking on the return control, return to the display page of the first multimedia content.
[0064] In actual applications, since the number of cards in the card set is finite, when a pre-set swipe operation triggered on the card set is received, the individual cards in the card set can be repeatedly scrolled and displayed.
[0065] In an alternative embodiment, during the display of the card set, it is also possible to display the individual recommended objects on the target card on the recommended object display page by a pre-set trigger operation on the target card in the card set.
[0066] In some embodiments, the pre-set trigger operation on the target card in the card set includes a click operation, a long press operation, etc. on the target card.
[0067] Specifically, upon receiving a pre-set trigger operation for any card in the card set, the card corresponding to the pre-set trigger operation is determined as the target card, and various recommended objects of the target card are displayed on the recommended object display page.
[0068] As shown in FIG. 4, upon receiving a pre-set trigger operation for card 401, card 401 is determined as the target card, and recommended objects on the target card 401, such as white long-sleeved tops 501 and white short-sleeved tops 502, are displayed in FIG. 5. FIG. 5 is a schematic diagram of the recommended object display page provided by the embodiment of the present disclosure.
[0069] In some embodiments, while stacking and displaying the card set, individual cards in the card set can be scrolled and displayed by a pre-set slide operation triggered for the card set. Upon receiving a pre-set trigger operation for the target card in the card set, by displaying each recommended object on the target card on the recommended object display page, the user's extended consumption of content related to the currently displayed multimedia content is promoted, and the user's viewing experience of the multimedia content is improved.
[0070] In actual applications, the target resource object recognized from the first multimedia content includes resource objects of the first resource type. In some embodiments, the first resource type includes, for example, the item type.
[0071] In addition to the above embodiments, the embodiments of the present disclosure provide a specific method for determining at least one recommended object for a resource object of the first resource type. FIG. 6 is a flowchart of another method for processing multimedia content provided by the embodiment of the present disclosure, and this method includes the following steps.
[0072] S601: In response to a pre-set trigger operation acting on the playback page of the first multimedia content, recognize the item object attached to the first multimedia content.
[0073] In some embodiments, the item objects of the item type attached to the recognized first multimedia content include one or more item objects such as a sweatshirt, trousers, a table, a schoolbag, etc.
[0074] In an alternative embodiment, when receiving a pre-set trigger operation acting on the display page of the first multimedia content, first, capture the video frame screen of the first multimedia content, and then, call the object recognition algorithm to select the video frame screen to which the resource object of the first resource type is attached.
[0075] In actual applications, when capturing the video frame screen in the first multimedia content, for example, based on a pre-set frame interval, it is possible to capture the video frame screen in the first multimedia content so as to capture the first multimedia content once every 10 frames. Also, for example, based on a pre-set time interval, it is possible to capture the video frame screen in the first multimedia content so as to capture the first multimedia content once every second. However, the embodiments of the present disclosure do not limit the method of capturing the video frame screen.
[0076] As an example, when receiving a pre-set trigger operation acting on the display page of the first multimedia content, first, the video frame screen of the first multimedia content is cut out, 10 video frame screens are obtained, and object recognition is performed on each of the 10 cut video frame screens. In some embodiments, two video frame screens are attached with resource objects of a first resource type, such as a white sweatshirt or a short dress.
[0077] S602: Based on the pre-set correspondence relationship, determine a commodity object corresponding to the commodity type.
[0078] S603: Based on the commodity object, determine at least one recommended object.
[0079] In some embodiments, the at least one recommended object belongs to the type of recommended object corresponding to the commodity object.
[0080] Based on the commodity object, at least one recommended object having the same or similar characteristics as it is determined. In some embodiments, the at least one recommended object belongs to the commodity type.
[0081] In some embodiments, after recognizing that a commodity object is attached to the first multimedia content, based on each recognized commodity object, it is possible to respectively determine at least one recommended object having the same or similar characteristics as the commodity object. In some embodiments, at least one recommended object belongs to the commodity type.
[0082] In an alternative embodiment, after recognizing that an item object is attached to the first multimedia content, the video frame image to which the item object is attached is sent to an image similarity calculation model, and based on the image similarity calculation model, at least one recommended object corresponding to the item object can be determined.
[0083] In some embodiments, the image similarity calculation model is used to match the item object with the recommended objects in the recommended object library, and at least one recommended object having characteristics similar to or the same as the item object is selected.
[0084] As an example, assuming that the recognized item objects are a sweatshirt and a short dress respectively, the video frame images with the white sweatshirt and the short dress attached are sent to the image similarity calculation model, and based on the image similarity calculation model, the sweatshirt and the short dress are respectively matched with the items in the item library. In some embodiments, the items having characteristics similar to or the same as the sweatshirt include a white long-sleeved sweatshirt and a white short-sleeved sweatshirt, and the items having characteristics similar to or the same as the short dress include a white short dress and a purple short dress.
[0085] S604: Display the at least one recommended object.
[0086] In an alternative embodiment, during the display of at least one recommended object, when a pre-set interaction operation on the target recommended object on the recommended object display page is received, the user may jump from the recommended object display page to the detailed display page of the target recommended object so as to learn more introduction content about the target recommended object based on the detailed display page. In some embodiments, the pre-set interaction operation on the target recommended object includes a click operation on the target recommended object.
[0087] In the method for processing multimedia content provided by an embodiment of the present disclosure, during the display of multimedia content, based on the target resource object attached to the multimedia content, in order to display to the user a recommended object related to the target resource object, the embodiment of the present disclosure provides the user with an extended consumption path for the content attached to the multimedia content, improving the user experience.
[0088] In actual application, the target resource object recognized from the first multimedia content may include BGM. In some embodiments, the types of recommended objects corresponding to BGM include music types and the like.
[0089] In addition to the above embodiments, the embodiments of the present disclosure provide a specific method for determining a recommended object based on BGM, and the method includes the following steps.
[0090] First, in response to a pre-set trigger operation acting on the display page of the first multimedia content, recognize the BGM attached to the first multimedia content. Next, perform music recognition on the BGM to obtain a music recognition result. Then, based on the music recognition result, determine the music resource corresponding to the BGM. And display the at least one music resource.
[0091] In an alternative embodiment, it is possible to perform music recognition on the BGM by calling a music search algorithm based on voice fingerprints, obtain a music recognition result, and then determine the music resource corresponding to the BGM based on the music recognition result.
[0092] For example, assuming that the BGM is the refrain part of a song, call a music search algorithm based on voice fingerprints, recognize, for example, the song information of the refrain part with the song name "Music A", and search for the song with the song name "Music A" from the song library as the music resource corresponding to the BGM.
[0093] In actual applications, after determining the music resources corresponding to the BGM, it is also possible to display the music resources on the recommended object display page. FIG. 7 is a schematic diagram of another recommended object display page provided by the embodiments of the present disclosure. In some embodiments, on the recommended object display page, the music name, music cover, author information, etc. corresponding to the music resources are displayed.
[0094] Furthermore, on the recommended object display page, when a trigger operation for the music playback control is received, a music playback control 701 for playing the music resources based on the recommended object display page is provided.
[0095] In an alternative embodiment, on the recommended object display page, when a trigger operation for a pre-set return control is received, a pre-set return control 702 that can realize the function of ending the recommended object display page may be provided.
[0096] In some embodiments, when at least one target resource object includes BGM, first, music recognition is performed on the BGM to obtain the music recognition result. Then, based on the music recognition result, the music resources corresponding to the BGM are determined and displayed. The embodiments of the present disclosure provide an extended consumption path for the content attached to the video to the user, further improving the user experience.
[0097] In actual applications, the target resource object recognized from the first multimedia content may include address information. In some embodiments, the type of the recommended object corresponding to the address information may include, for example, a lifestyle service type, etc.
[0098] In addition to the above embodiments, the embodiments of the present disclosure provide a method for determining a recommended object based on address information. Specifically, first, in response to a pre-set trigger operation acting on the display page of the first multimedia content, the address information attached to the first multimedia content is recognized. Centering on the address information, at least one life service object belonging to the pre-set distance range and belonging to the life service type is determined, and at least one life service object is displayed.
[0099] In some embodiments, when receiving a pre-set trigger operation acting on the display page of the first multimedia content, a voice recognition algorithm is called, and by recognizing the audio file in the first multimedia content, it is possible to obtain the address information attached to the first multimedia content.
[0100] In some embodiments, the voice recognition algorithm includes an algorithm based on dynamic time warping, a voice recognition algorithm based on a deep learning neural network, and the like.
[0101] In an alternative embodiment, by calling a text recognition algorithm and recognizing the subtitle content of the first multimedia content, it is possible to obtain the address information attached to the first multimedia content.
[0102] For example, assuming that the address information attached to the first multimedia content is "Location ABC", a search for a mall, supermarket, clothing store, tourist attraction, etc. within 1 km from "Location ABC" is performed centering on "Location ABC".
[0103] In an alternative embodiment, when a specific address anchor is attached to the first multimedia content, the position information corresponding to the address anchor may be directly determined as the life service object corresponding to the address anchor.
[0104] In some embodiments, when at least one target resource object includes address information, first, in response to a preset trigger operation acting on the playback page of the first multimedia content, the address information attached to the first multimedia content is recognized, and then, centering on the address information, at least one life service object belonging to the type of life service within a preset distance range is determined and displayed. Thereby, the embodiments of the present disclosure provide an extended consumption path for the content attached to the multimedia content to the user and improve the user experience.
[0105] In actual application, the target resource object recognized from the first multimedia content includes the target face displayed on the video frame screen. In some embodiments, the type of the recommended object corresponding to the target face displayed on the video frame screen includes the user account type.
[0106] In addition to the above embodiments, the embodiments of the present disclosure provide a method for determining a recommended object based on the target face displayed on the video frame screen. Specifically, when the user to whom the target face belongs permits the use of the target face information, first, in response to a preset trigger operation acting on the playback page of the first multimedia content, the target face attached to the first multimedia content is recognized, and then, based on the target face displayed on the video frame screen, at least one user account whose similarity between the user avatar and the target face reaches a preset threshold is determined. In some embodiments, the user account belongs to the user account type.
[0107] In an alternative embodiment, upon receiving a pre-set trigger operation acting on the playback page of the first multimedia content, first, a video frame screen within the first multimedia content is clipped, and then a face recognition algorithm is invoked to recognize each video frame screen within the first multimedia content, and to recognize the video frame screen with a target face marked therein within the first multimedia content.
[0108] In an alternative embodiment, after recognizing the video frame with a target face marked therein within the first multimedia content, if the user to whom the target face belongs permits the use of the target face information, the video frame with the target face marked therein is transmitted to a face matching service terminal, and based on the face matching service terminal, at least one user account whose similarity between the user's avatar and the target face reaches a pre-set threshold is determined.
[0109] In some embodiments, the pre-set threshold is determined based on actual needs, and may be set, for example, to 80%, 85%, 90%, 95%, etc.
[0110] For example, assuming that the face recognition algorithm is invoked and the recognized target faces are Face A and Face B, the video frames including Face A and Face B are transmitted to the face matching server terminal, and based on the face matching server terminal, in the user avatar database, user avatars whose similarity to Face A and Face B reaches the pre-set threshold are respectively searched. In some embodiments, the user account corresponding to the user avatar whose similarity to Face A reaches the pre-set threshold is "Mr. A", and the user account corresponding to the user avatar whose similarity to Face B reaches the pre-set threshold is "Mr. B".
[0111] In an alternative embodiment, assuming that the first multimedia content is the first video, by invoking the text recognition algorithm, the subtitle content of the first video is recognized, the name of the person attached to the first video is obtained, and based on the name of the person, a user nickname whose similarity to the name of the person reaches a pre-set threshold is searched for, and it is also possible to determine the user account corresponding to the searched user nickname as the target resource object. For example, when performing text recognition on the subtitle content of the first video, the name of a person "Flower" is obtained, and then user nicknames such as "Mr. Flower" and "Store Manager Flower" whose similarity to "Flower" reaches a pre-set threshold are searched for.
[0112] In some embodiments, after determining at least one user account whose similarity between the user avatar and the target face reaches a pre-set threshold, if the user to whom the user account belongs permits the display of the user account, the user account is also displayed on the recommended object display page. FIG. 8 is the recommended object display page provided by the embodiment of the present disclosure.
[0113] In an alternative embodiment, at least one user account displayed on the recommended object display page is respectively provided with a pre-set interaction control, and in response to a trigger operation on the pre-set interaction control corresponding to the first user account among the at least one user account, a pre-set interaction relationship between the current user account and the first user account is established.
[0114] In some embodiments, the trigger operation on the pre-set interaction control corresponding to the first user account of the at least one user account may include a click operation, a long-press operation, etc. on the pre-set interaction control, and the embodiments of the present disclosure are not limited thereto. In some embodiments, the first user account may be any one of the at least one user account.
[0115] In some embodiments, the pre - set interaction relationship between the current user account and the first user account may include determining the first user account as a follower of the current user account.
[0116] As shown in FIG. 8, when receiving a trigger operation on the pre - set interaction control 802 corresponding to the first user account 801, by determining the first user account 801 as a follower of the current user account, the function of establishing the pre - set interaction relationship between the current user account and the first user account is realized.
[0117] In some embodiments, when at least one target resource object includes a target face, first, in response to a pre - set trigger operation acting on the playback page of the first multimedia content, recognize the target face attached to the first multimedia content, and then, based on the target face, determine and display at least one user account whose similarity between the user's avatar and the target face reaches a pre - set threshold. Thereby, the embodiments of the present disclosure provide an extended consumption path for the content attached to the multimedia content to the user and improve the user experience.
[0118] Based on the embodiments of the above - mentioned method, the present disclosure further provides a processing device for multimedia content. FIG. 9 is a schematic diagram of the structure of the processing device for multimedia content provided by the embodiments of the present disclosure. The device includes: A recognition module 901 for recognizing at least one target resource object attached to the first multimedia content in response to a pre - set trigger operation acting on the playback page of the first multimedia content, where there is a pre - set corresponding relationship between the target resource object and the type of the recommended object. A first determination module 902 for determining the type of a recommended object corresponding to a first target resource object among the at least one target resource object based on the pre-set corresponding relationship; A second determination module 903 for determining at least one recommended object based on the first target resource object, wherein the at least one recommended object belongs to the type of the recommended object corresponding to the first target resource object; A display module 904 for displaying the at least one recommended object.
[0119] In an alternative embodiment, the display module Comprises a classification display sub-module for classifying and displaying the recommended objects respectively determined based on the first target resource object and the second resource object according to the type of the recommended object.
[0120] In an alternative embodiment, the classification display sub-module A first determination sub-module for displaying at least one first recommended object determined based on the first target resource object on the first card, wherein the first recommended object belongs to the type of the recommended object corresponding to the first target resource object; A second determination sub-module for displaying at least one second recommended object determined based on the second resource object on a second card, wherein the second recommended object belongs to the type of the recommended object corresponding to the second target resource object, and the second card and the second card belong to a card set that is stacked and displayed.
[0121] In an alternative embodiment, the classification display sub-module It further includes a scroll display sub-module for responding to a pre-set slide operation triggered for the card set and scrolling to display individual cards in the card set.
[0122] In an alternative embodiment, the classification display sub-module includes a recommended object display sub-module for responding to a pre-set trigger operation for the target card in the card set and displaying recommended objects on the target card on a recommended object display page, and an acceptance sub-module for accepting a pre-set interaction operation for a target recommended object on the recommended object display page.
[0123] In an alternative embodiment, the at least one target resource object includes an item object, the type of the recommended object corresponding to the item object includes an item type, and the second determination module includes a third determination sub-module for determining at least one recommended item having characteristics similar to or the same as those of the item object and belonging to the item type.
[0124] In an alternative embodiment, the at least one target resource object includes BGM, the type of the recommended object corresponding to the BGM includes a music type, and the second determination module includes a music recognition sub-module for performing music recognition on the BGM and obtaining a music recognition result, and a fourth determination sub-module for determining a music resource corresponding to the BGM and belonging to the music type based on the music recognition result.
[0125] In an alternative embodiment, the at least one target resource object includes address information, the type of the recommended object corresponding to the address information includes a life service type, and the second determination module comprises a fifth determination sub-module for determining at least one life service object that is within a preset distance range centered on the address information and belongs to the life service type.
[0126] In an alternative embodiment, the at least one target resource object includes a target face displayed on a video frame screen, the type of the recommended object corresponding to the target face includes a user account type, and the second determination module comprises a sixth determination sub-module for determining at least one user account that reaches a preset threshold in similarity between a user avatar and the target face based on the target face displayed on the video frame screen and belongs to the user account type.
[0127] In an alternative embodiment, the recognition module comprises a screen display sub-module for responding to a preset trigger operation acting on the display page of the first multimedia content and displaying a plurality of video key frame screens in the first multimedia content on the video recognition page in the form of a transition animation, and a target resource object recognition sub-module for recognizing at least one target resource object attached to the first multimedia content based on the plurality of video key frame images.
[0128] In the multimedia content processing device provided by the embodiments of the present disclosure, first, in response to a preset trigger operation acting on the display page of the first multimedia content, at least one target resource object attached to the first multimedia content is recognized, among which there is a preset object correspondence relationship between the type of the target resource object and the type of the recommended object. Next, based on the preset correspondence relationship, the type of the recommended object corresponding to the first target resource object among the at least one target resource object is determined, and at least one recommended object is determined based on the first target resource object. Next, the at least one recommended object is displayed. The embodiments of the present disclosure can display, to the user, a recommended object related to the target resource object based on the target resource object attached to the multimedia content during the display of the multimedia content. Therefore, the embodiments of the present disclosure provide an extended consumption path for the content attached to the multimedia content to the user, improving the user experience.
[0129] In addition to the above method and apparatus, when executed on a terminal device, the embodiments of the present disclosure further provide a computer-readable storage medium storing instructions for causing the terminal device to implement the multimedia content processing method described in the embodiments of the present disclosure.
[0130] When executed by a processor, the embodiments of the present disclosure further provide a computer program product including a computer program / instructions for implementing the multimedia content processing method described in the embodiments of the present disclosure.
[0131] Furthermore, the embodiments of the present disclosure provide a multimedia content processing device, which includes a processor 1001, a memory 1002, an input device 1003, and an output device 1004 with reference to FIG. 10.
[0132] The number of processors 1001 in the multimedia content processing device is one or more. In FIG. 10, one processor is used as an example. In some embodiments of the present disclosure, the processor 1001, the memory 1002, the input device 1003, and the output device 1004 may be connected via a bus or other means. In FIG. 10, the connection via a bus is taken as an example.
[0133] The memory 1002 is used to store software programs and modules. The processor 1001 executes the software programs and modules stored in the memory 1002 to perform various functions and data processing of the multimedia content processing device. The memory 1002 mainly includes a storage program area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function, and the like. Further, the memory 1002 may include a high-speed random access memory, for example, at least one disk memory device, a non-volatile memory such as a flash memory device, or other volatile solid-state memory devices. The input device 1003 is used for inputting numerical or character information and generating signal inputs related to user settings and function controls of the multimedia content processing device.
[0134] Specifically, in this embodiment, the processor 1001 loads an executable file corresponding to the process of one or more applications into the memory 1002 according to the following instructions, and realizes various functions of the above-mentioned multimedia content processing device by executing the applications stored in the memory 1002 by the processor 1001.
[0135] In this specification, relational terms such as "first" and "second" are merely used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of such an actual relationship or order between those entities or operations. Also, the terms "comprise", "include" or any other variation thereof are intended to cover non-exclusive inclusion, such that a process, method, product or apparatus consisting of a set of elements includes not only those elements but also other elements not explicitly listed, or other elements specific to such a process, method, product or apparatus. Further, unless otherwise limited, an element defined by the expression "comprising one XX" does not exclude the existence of other identical elements in a process, method, product or apparatus that includes the element.
[0136] The above are only specific examples of the present disclosure to enable those skilled in the art to understand or implement the present disclosure. Various modifications to these examples will be apparent to those skilled in the art, and the general principles defined in this specification can be implemented in other examples without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure is not limited to these examples described in this specification, but follows the broadest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for processing multimedia content, comprising: responding to a pre-set trigger operation acting on a display page of a first multimedia content, and recognizing at least one target resource object attached to the first multimedia content, wherein the target resource object has a pre-set correspondence with the type of a recommended object; determining the type of the recommended object corresponding to a first target resource object among the at least one target resource object based on the pre-set correspondence; determining at least one recommended object based on the first target resource object, wherein the at least one recommended object belongs to the type of the recommended object corresponding to the first target resource object; displaying the at least one recommended object.
2. The method according to claim 1, wherein the at least one target resource object further includes a second target resource object, and the displaying of the at least one recommended object includes: classifying and displaying the recommended objects respectively determined based on the first target resource object and the second resource object according to the type of the recommended object.
3. The classifying and displaying of the recommended objects respectively determined based on the first target resource object and the second resource object according to the type of the recommended object includes: displaying at least one first recommended object determined based on the first target resource object on a first card, wherein the first recommended object belongs to the type of the recommended object corresponding to the first target resource object; displaying, on the second card, at least one second recommended object determined based on the second resource object, wherein the second recommended object belongs to a type of recommended object corresponding to the second target resource object, and the first card and the second card belong to a card set that is stacked and displayed, the method according to claim 2.
4. The method further includes responding to a pre-set slide operation triggered for the card set, and scrolling and displaying each card in the card set, the method according to claim 3.
5. The method further includes responding to a pre-set trigger operation for a target card in the card set, and displaying, on a recommended object display page, the recommended object on the target card; and receiving a pre-set interaction operation for a target recommended object on the recommended object display page, the method according to claim 3 or 4.
6. The at least one target resource object includes an item object, the type of the recommended object corresponding to the item object includes an item type, and determining at least one recommended object based on the first target resource object includes determining at least one recommended item having the same or similar features as the item object and belonging to the item type, the method according to claim 1.
7. The at least one target resource object includes BGM, the type of the recommended object corresponding to the BGM includes a music type, and determining at least one recommended object based on the first target resource object includes performing music recognition on the BGM to obtain a music recognition result; and determining a song resource corresponding to the BGM and belonging to the music type based on the music recognition result, the method according to claim 1.
8. The at least one target resource object includes address information, the type of the recommended object corresponding to the address information includes a life service type, and determining at least one recommended object based on the first target resource object includes determining at least one life service object that is within a preset distance range centered on the address information and belongs to the life service type, and is characterized in that, the method according to claim 1.
9. The at least one target resource object includes a target face displayed on a video frame screen, the type of the recommended object corresponding to the target face includes a user account type, and determining at least one recommended object based on the first target resource object includes determining at least one user account that reaches a preset threshold in similarity to the target face of the user avatar based on the target face displayed on the video frame screen and belongs to the user account type, and is characterized in that, the method according to claim 1.
10. Responding to a preset trigger operation acting on the display page of the first multimedia content, recognizing at least one target resource object attached to the first multimedia content Responding to a preset trigger operation acting on the display page of the first multimedia content, and displaying a plurality of video keyframe screens in the first multimedia content on a video recognition page in the form of a transition animation and recognizing at least one target resource object attached to the first multimedia content based on the plurality of video keyframe screens, and is characterized in that, the method according to claim 1.
11. A processing device for multimedia content, A recognition module for recognizing at least one target resource object attached to the first multimedia content in response to a pre-set trigger operation acting on the playback page of the first multimedia content, wherein the target resource object has a pre-set correspondence relationship with the type of the recommended object, and A first determination module for determining the type of the recommended object corresponding to the first target resource object among the at least one target resource object based on the pre-set correspondence relationship, and A second determination module for determining at least one recommended object based on the first target resource object, wherein the at least one recommended object belongs to the type of the recommended object corresponding to the first target resource object, and A display module for displaying the at least one recommended object, and the apparatus is characterized by comprising the same.
12. A computer-readable storage medium, which when executed on a terminal device, stores instructions for causing the terminal device to implement the method according to any one of Claims 1 to 10.
13. A multimedia content processing device including a memory, a processor, and a computer program stored in the memory and executable by the processor, wherein when the processor executes the computer program, the method according to any one of Claims 1 to 10 is implemented.
Citation Information
Patent Citations
Information display method and device, computer equipment and storage medium
CN115599944A
Search support system, search support method, and search support program
WO2010016281A1