Methods, apparatuses, devices, and storage media for processing multimedia content
Patent Information
- Application Number
- JP2025504038
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-02-14
- Filing Date
- 2024-01-30
- Publication Date
- 2026-08-28
- Estimated Expiration
- 2044-01-30
Smart Images

Figure 0007912667000001 
Figure 0007912667000002 
Figure 0007912667000003
Abstract
Description
[Technical Field]
[0001] (Reference to Related Application) This application claims priority based on the Chinese patent application with the filing number 2023101598223, entitled "Multimedia Content Processing Method, Apparatus, Device and Storage Medium", filed on February 14, 2023, and the entire content of said application is incorporated into this application by reference.
[0002] (Technical Field) The present disclosure relates to the field of data processing, and in particular to a multimedia content processing method, apparatus, device and storage medium. [Background Art]
[0003] With the continuous development of video processing technology, people's demands for video-related functions have become increasingly diversified. Therefore, how to enrich video-related functions to meet the needs of more users and improve user experience has become an urgent technical problem to be solved. [Summary of the Invention]
[0004] To solve the above technical problem, the present disclosure provides a multimedia content processing method, apparatus, device and storage medium capable of improving user experience by enriching video-based interaction functions.
[0005] In a first embodiment, the Disclosure provides a method for processing multimedia content, which, in response to a pre-configured trigger operation acting on a display page of first multimedia content, recognizes at least one target resource object attached to the first multimedia content, wherein there is a pre-configured correspondence between the target resource object and the type of recommended object; determines, based on the pre-configured correspondence, the type of recommended object among the at least one target resource object that corresponds to the first target resource object; determines, based on the first target resource object, at least one recommended object, wherein the at least one recommended object belongs to the type of recommended object that corresponds to the first target resource object; and displays the at least one recommended object.
[0006] In a selective embodiment, the at least one target resource object further includes a second target resource object, and displaying the at least one recommended object as described above classifies and displays the recommended objects determined based on the first target resource object and the second resource object, respectively, according to the type of the recommended object.
[0007] In a selective embodiment, classifying and displaying recommended objects determined based on the first target resource object and the second resource object according to the type of recommended object as described above includes displaying on a first card at least one first recommended object determined based on the first target resource object, the first recommended object belonging to the type of recommended object corresponding to the first target resource object, and displaying on a second card at least one second recommended object determined based on the second resource object, the second recommended object belonging to the type of recommended object corresponding to the second target resource object, and the first card and the second card belonging to a stacked set of cards.
[0008] In a selective embodiment, the method further includes scrolling through each card in the card set in response to a pre-set slide operation triggered on the card set.
[0009] In an optional embodiment, the method further includes responding to a pre-configured trigger operation on a target card in the card set by displaying a recommended object on the target card on a recommended object display page, and accepting a pre-configured interactive operation on the target recommended object on the recommended object display page.
[0010] In a selective embodiment, the at least one target resource object includes an item object, the type of the recommendation object corresponding to the item object includes an item type, and determining at least one recommendation object based on the first target resource object as described above includes determining at least one recommendation item that has similar or identical characteristics to the item object and belongs to the item type.
[0011] In a selective embodiment, the at least one target resource object includes background music (BGM), the type of the recommended object corresponding to the BGM includes a music type, and the at least one recommended object determined based on the first target resource object as described above includes performing music recognition on the BGM and obtaining a music recognition result, and determining a song resource that corresponds to the BGM and belongs to the music type based on the music recognition result.
[0012] In a selective embodiment, the at least one target resource object includes address information, the type of recommended object corresponding to the address information includes a life service type, and the at least one recommended object determined based on the first target resource object as described above includes determining at least one life service object that is within a predetermined distance range centered on the address information and belongs to the life service type.
[0013] In a selective embodiment, the at least one target resource object includes a target face displayed on a video frame screen, the type of the recommendation object corresponding to the target face includes a user account type, and the at least one recommendation object determined based on the first target resource object as described above includes determining at least one user account belonging to the user account type, where the similarity between the user avatar and the target face reaches a preset threshold, based on the target face displayed on the video frame screen.
[0014] In a selective embodiment, recognizing at least one target resource object attached to the first multimedia content in response to a pre-configured trigger operation acting on the display page of the first multimedia content as described above includes, in response to a pre-configured trigger operation acting on the display page of the first multimedia content, displaying a plurality of video keyframe screens within the first multimedia content on a video recognition page in the form of a transition animation, and recognizing at least one target resource object attached to the first multimedia content based on the plurality of video keyframe screens.
[0015] In a second embodiment, the Disclosure provides a multimedia content processing device comprising: a recognition module for recognizing at least one target resource object attached to the first multimedia content in response to a pre-configured trigger operation acting on a playback page of the first multimedia content, the recognition module having a pre-configured correspondence between the target resource object and the type of recommended object; a first determination module for determining the type of recommended object corresponding to the first target resource object among the at least one target resource object based on the pre-configured correspondence; a second determination module for determining at least one recommended object based on the first target resource object, the second determination module wherein the at least one recommended object belongs to the type of recommended object corresponding to the first target resource object; and a display module for displaying the at least one recommended object.
[0016] In a third embodiment, the Disclosure provides a computer-readable storage medium which, when executed by a terminal device, stores instructions causing the terminal device to implement the method.
[0017] In a fourth embodiment, the Disclosure provides a multimedia content processing device comprising memory, a processor, and a computer program stored in the memory and executable by the processor, wherein the processor executes the computer program, thereby realizing the method.
[0018] In a fifth embodiment, the Disclosure provides a computer program product which, when executed by a processor, stores computer programs / instructions that implement the methods described above.
[0019] The technical modes provided by the embodiments of this disclosure have at least the following advantages over the prior art.
[0020] Embodiments of the present disclosure provide a method for processing multimedia content, which first determines, in response to a pre-configured trigger operation acting on a display page of first multimedia content, at least one target resource object attached to the first multimedia content, among which there is a pre-configured object relationship between the target resource object and the type of recommended object, then determines, based on the pre-configured correspondence, the type of the recommended object corresponding to the first target resource object among the at least one target resource object, determines the at least one recommended object based on the first target resource object, and then displays the at least one recommended object. Embodiments of the present disclosure display recommended objects associated with target resource objects based on the target resource objects attached to the multimedia content while the multimedia content is being displayed, thereby providing users with an extended consumption path for content attached to multimedia content and improving the user experience. [Brief explanation of the drawing]
[0021] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate embodiments of the present disclosure and are used together with the specification to explain the principle of the present disclosure.
[0022] In the following, in order to more clearly describe the technical forms in the embodiments of the present disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly described. It will be apparent to those skilled in the art that other accompanying drawings can be obtained based on the accompanying drawings without creative efforts.
[0023] [Figure 1] FIG. 1 is a flowchart of a method for processing multimedia content provided by an embodiment of the present disclosure. [Figure 2] FIG. 2 is a schematic diagram of a video recognition page provided by an embodiment of the present disclosure. [Figure 3] FIG. 3 is a schematic diagram of another video recognition page provided by an embodiment of the present disclosure. [Figure 4] FIG. 4 is a schematic diagram of a display page corresponding to a card set provided by an embodiment of the present disclosure. [Figure 5] FIG. 5 is a schematic diagram of a recommended object display page provided by an embodiment of the present disclosure. [Figure 6] FIG. 6 is a flowchart of another method for processing multimedia content provided by an embodiment of the present disclosure. [Figure 7] FIG. 7 is a schematic diagram of another recommended object display page provided by an embodiment of the present disclosure. [Figure 8] FIG. 8 is a schematic diagram of still another recommended object display page provided by an embodiment of the present disclosure. [Figure 9] FIG. 9 is a schematic diagram of the configuration of a multimedia content processing device provided by an embodiment of the present disclosure. [Figure 10] FIG. 10 is a schematic diagram of the configuration of a multimedia content processing device provided by an embodiment of the present disclosure. MODE FOR CARRYING OUT THE INVENTION
[0024] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the embodiments of this disclosure will be further described below. The embodiments and features of the embodiments of this disclosure can be combined with each other, as long as they do not contradict each other.
[0025] Many details are provided below to facilitate a full understanding of this disclosure, but this disclosure can also be implemented in ways other than those described herein. It will be apparent that the examples in the specification are only a selection of, and not all, of, the examples of this disclosure.
[0026] With the continuous advancements in video processing technology, people's needs for video-related functions are becoming increasingly diverse. Meeting the needs of a wider range of users and enhancing the user experience through video-related features has become an urgent technical challenge that needs to be addressed.
[0027] Furthermore, the related information attached to videos is becoming increasingly rich, such as the items featured, the places mentioned, the people involved, and the background music used. While users are watching a video, the related information attached to the video may generate further consumption ideas.
[0028] Related technologies employ a method of adding feature anchors to video playback pages, providing users with a path to further consume related information attached to the video. However, there are limitations to the display area of the video playback page, and displaying too many feature anchors could negatively impact the simplicity of the video playback page and the user's video viewing experience.
[0029] To this end, the embodiments of this disclosure provide a method for processing multimedia content, which first recognizes at least one target resource object attached to the first multimedia content in response to a pre-configured trigger operation acting on the display page of the first multimedia content. In some embodiments, there is a pre-configured correspondence between the target resource object and the type of recommended object. Based on this pre-configured correspondence, the type of recommended object corresponding to the first target resource object among the at least one target resource object is determined, and at least one recommended object is determined based on the first target resource object. Then, the at least one recommended object is displayed. The embodiments of this disclosure provide the user with an extended consumption path for the content attached to the multimedia content and improve the user experience by displaying recommended objects associated with the target resource object attached to the multimedia content to the user, based on the target resource object attached to the multimedia content while the multimedia content is being displayed.
[0030] Based on this, embodiments of the present disclosure provide a method for processing multimedia content, and Figure 1 is a flowchart of the method for processing multimedia content provided by embodiments of the present disclosure, the method comprising the following steps.
[0031] S101: In response to a pre-configured trigger operation acting on the display page of the first multimedia content, recognize at least one target resource object attached to the first multimedia content.
[0032] In some embodiments, there is a pre-configured correspondence between the target resource object and the type of the recommended object.
[0033] In some embodiments, the first multimedia content includes one of the following: video, audio, graphic content, text content, image content, etc. Specifically, the first multimedia content is one of the multimedia contents in any multimedia information stream, such as one of the recommended videos in a recommended video stream.
[0034] In some embodiments, pre-configured trigger operations acting on the display page of the first multimedia content include, but are not limited to, knuckle double-click operations, long-press operations, double-click operations, etc., acting on the display page of the first multimedia content.
[0035] In some embodiments, the target resource object may be an object displayed on a video frame screen attached to the first multimedia content, such as a sweatshirt, stool, or pet dog; it may be an address information appearing in text content (e.g., subtitles) or location anchors (e.g., a tourist spot, food court, supermarket, etc.) attached to the first multimedia content; or it may be background music played in the first multimedia content, or a public figure appearing on a video frame screen attached to the first multimedia content.
[0036] In some embodiments, upon accepting a pre-configured trigger operation acting on the display page of the first multimedia content, it is possible to recognize at least one target resource object attached to the first multimedia content and obtain the target resource object attached to the first multimedia content using, for example, optical character recognition (OCR), automatic speech recognition (ASR), or facial recognition technology. Specific recognition methods will be described in subsequent embodiments, where different target resource objects will be explained, so they are omitted here.
[0037] In some embodiments, the type of recommended object is used to identify the type to which the recommended object, determined based on the target resource object, belongs. In some embodiments, the type of recommended object includes, but is not limited to, item types, music types, life service types, user account types, etc., and can be set according to actual needs.
[0038] In some embodiments, there is a pre-configured correspondence between the target resource object and the type of the recommended object, and different target resource objects can correspond to the same or different types of recommended objects. For example, if the target resource object is a sweatshirt, its corresponding recommended object type is an item type, and if the target resource object is music, its corresponding recommended object type is a music type.
[0039] In practical applications, assuming that the first multimedia content includes video, and in order to improve the user's viewing experience of the multimedia content, it is also possible to display multiple video keyframe screens within the first multimedia content in the form of a transition animation on the video recognition page when a pre-configured trigger operation acting on the first multimedia content is accepted, so that the user has a certain expectation of being able to determine the corresponding recommended object based on the target resource object attached to the first multimedia content. This allows for the recognition of at least one target resource object attached to the first multimedia content based on the multiple video keyframe screens.
[0040] In a selective embodiment, for example, video keyframes within the first multimedia content may be extracted at a predetermined frame interval, such as every 10 frames, and multiple video keyframes may be obtained.
[0041] In another optional embodiment, for example, video keyframes within the first multimedia content may be extracted at predetermined time intervals, such as every second, resulting in multiple video keyframes.
[0042] Figure 2 is a schematic diagram of a video recognition page provided by an embodiment of the present disclosure, in which multiple video keyframe images are displayed on the video recognition page according to a pre-set motion trajectory or a random travel trajectory. In some embodiments, the pre-set motion trajectory includes movement from the center position to the edge position of the video recognition page.
[0043] In a selective embodiment, based on the style of cards displayed in a stacked manner, multiple key video frames within the first multimedia content may be displayed on different cards of the video recognition page, thereby recognizing at least one target resource object attached to the first multimedia content based on the multiple video keyframes. Figure 3 shows a schematic diagram of another video recognition page provided by an embodiment of the present disclosure. In some embodiments, multiple video keyframe screens are displayed overlaid on the video recognition page in a card style.
[0044] S102: Based on the pre-configured correspondence, determine the type of recommended object corresponding to the first target resource object among the at least one target resource object.
[0045] S103: Based on the first target resource object, determine at least one recommended object.
[0046] In some embodiments, the at least one recommended object belongs to the type of recommended object corresponding to the first target resource object.
[0047] In some embodiments, the first target resource object is one of the resource objects among the at least one target resource object recognized from the first multimedia content.
[0048] In some embodiments, after recognizing at least one target resource object attached to the first multimedia content, it is possible to determine at least one recommended object belonging to the recommended object type corresponding to the target resource object, based on any one of the at least one target resource objects.
[0049] In some embodiments, the first target resource object and at least one recommended object have similar or identical characteristics. For example, assuming the first target resource object is the chorus of a song, the entire song corresponding to the chorus is determined to be the recommended object corresponding to the first target resource object, based on the chorus and the type of the corresponding recommended object (i.e., the music type).
[0050] S104: Display the at least one recommended object.
[0051] In some embodiments, after determining at least one recommended object based on the target resource object, the individual recommended objects are displayed.
[0052] In selective embodiments, at least one target resource object may further include a second target resource object. In some embodiments, the second target resource object and the first target resource object belong to different recommended object types.
[0053] Therefore, displaying recommended objects further includes determining at least one recommended object based on the first target resource object and the second target resource object, and then classifying and displaying the recommended objects according to the type of recommended object to which each recommended object belongs.
[0054] As an example, suppose the first target resource object is a white sweatshirt corresponding to the item type, and the second target resource object is the chorus corresponding to the music type. Based on the white sweatshirt, the recommended objects determined include a white long-sleeved sweatshirt and a white short-sleeved sweatshirt. Based on music fragment A, the recommended object is the entirety of song B. Depending on the type of recommended object, the white long-sleeved bodysuit, the white short-sleeved bodysuit, and song B are classified and displayed; that is, the white long-sleeved bodysuit and the white short-sleeved bodysuit are displayed together as recommended objects for the item type, and song B is displayed as a recommended object for the music type.
[0055] The above is merely an example where at least one target resource object includes two target resource objects. The embodiments of this disclosure do not limit the number of target resource objects recognized from the first multimedia content. The methods for displaying the recommended objects corresponding to multiple target resource objects can be implemented by referring to the above and are therefore omitted here.
[0056] In the multimedia content processing method provided by the embodiments of this disclosure, first, in response to a pre-configured trigger operation acting on the display page of the first multimedia content, at least one target resource object attached to the first multimedia content is recognized, and a pre-configured object relationship exists between the target resource object and the type of recommended object. Then, based on the pre-configured correspondence relationship, the type of recommended object corresponding to the first target resource object among the at least one target resource object is determined, and at least one recommended object is determined based on the first target resource object. Then, at least one recommended object is displayed. The embodiments of this disclosure can display recommended objects associated with target resource objects to the user based on the target resource object attached to the multimedia content while the multimedia content is being displayed. As a result, the embodiments of this disclosure provide the user with an extended consumption path for content attached to the multimedia content and improve the user experience.
[0057] In practical applications, to enhance the interactive functionality during the display of multimedia content and improve the user's viewing experience of multimedia content, it is also possible to display recommended objects of different types in a card style. Specifically, recommended objects determined based on a first target resource object and a second target resource object are classified and displayed in a card style, with the first and second target resource objects corresponding to different types of recommended objects.
[0058] Specifically, at least one first recommended object determined based on the first target resource object is displayed on the first card, and of these, the first recommended object belongs to the recommended object type corresponding to the first target resource object. At least one second recommended object determined based on the second resource object is displayed on the second card, and of these, the second recommended object belongs to the recommended object type corresponding to the second target resource object.
[0059] Figure 4 is a schematic diagram of a display page corresponding to a card set provided in an embodiment of the present disclosure, where two cards are used as an example. The first card 401 displays a recommended object determined based on a white top (first target resource object), such as a white long-sleeved top or a white short-sleeved top, and the second card 402 displays a recommended object determined based on a music fragment A (second target resource object), such as music B. In some embodiments, the first and second cards belong to a card set that is displayed stacked on top of each other.
[0060] In a selective embodiment, while the card set is displayed stacked on the display page corresponding to the card set, it is possible to scroll through the individual cards in the card set by a pre-configured slide operation triggered on the card set, making it easier for the user to select and display the desired card based on the content displayed on each individual card.
[0061] In some embodiments, a pre-configured slide operation triggered on a card set to trigger the scrolling display of individual cards in the card set includes an up slide operation, a down slide operation, etc., on the card set.
[0062] As shown in Figure 4, when an upward slide operation is accepted for the card set, the currently displayed card 401 is switched upward to the next adjacent card, i.e., card 402, and the individual recommended objects of card 402 are displayed on the display page where the card set is located. When a downward slide operation is accepted for the card set, the currently displayed card 401 is switched downward to the previous adjacent card, i.e., card 403, in order to fully display the card set on the display page where it is located.
[0063] In an optional embodiment, if the user wishes to exit the display of the card set, a pre-configured return operation on the card set allows them to return to the playback page of the first multimedia content. In some embodiments, the pre-configured return operation on the card set may include, for example, a left slide operation acting on the card set, and a return control 404 may be provided on the display page where the card set is located, and clicking this return control returns the user to the display page of the first multimedia content.
[0064] In practical applications, since the number of cards in a card set is finite, accepting a pre-configured swipe gesture triggered on the card set allows for repeated scrolling of individual cards within the card set.
[0065] In a selective embodiment, while a card set is being displayed, a pre-configured trigger operation on a target card in the card set can also be used to display individual recommended objects on the recommended object display page.
[0066] In some embodiments, pre-configured trigger operations on a target card in a card set include click operations, long press operations, etc., on the target card.
[0067] Specifically, when a pre-configured trigger operation is accepted for any card in the card set, the card corresponding to the pre-configured trigger operation is identified as the target card, and various recommended objects for that target card are displayed on the recommended object display page.
[0068] As shown in Figure 4, when a pre-configured trigger operation is accepted for card 401, card 401 is identified as the target card, and recommended objects on the target card 401, such as a white long-sleeved top 501 or a white short-sleeved top 502, are displayed in Figure 5. Figure 5 is a schematic diagram of the recommended object display page provided by an embodiment of the present disclosure.
[0069] In some embodiments, while stacking card sets, pre-configured slide operations triggered on the card sets allow for scrolling through individual cards within the set. Accepting a pre-configured trigger operation on a target card in the card set displays each recommended object on the target card on the recommended object display page, thereby promoting extended user consumption of content related to the currently displayed multimedia content and improving the user's viewing experience of the multimedia content.
[0070] In practical applications, the target resource object recognized from the first multimedia content includes a resource object of the first resource type. In some embodiments, the first resource type includes, for example, an item type.
[0071] In addition to the embodiments described above, embodiments of this disclosure provide a specific method for determining at least one recommended object for a resource object of a first resource type. Figure 6 is a flowchart of another multimedia content processing method provided by embodiments of this disclosure, which includes the following steps.
[0072] S601: In response to a pre-configured trigger operation acting on the playback page of the first multimedia content, recognize an item object attached to the first multimedia content.
[0073] In some embodiments, the item objects of the item type attached to the recognized first multimedia content include one or more item objects, such as a sweatshirt, trousers, a table, a school bag, etc.
[0074] In a selective embodiment, upon accepting a pre-configured trigger operation acting on the display page of the first multimedia content, the system first extracts a video frame from the first multimedia content, then invokes a subject recognition algorithm to select the video frame to which a resource object of the first resource type is attached.
[0075] In practical applications, when extracting video frames from a first multimedia content, it is possible to extract video frames from the first multimedia content based on a preset frame interval, for example, once every 10 frames. Alternatively, it is also possible to extract video frames from the first multimedia content based on a preset time interval, for example, once every 1 second. However, the embodiments of this disclosure are not limited to methods for extracting video frames.
[0076] For example, upon accepting a pre-configured trigger operation that acts on the display page of the first multimedia content, the system first extracts video frame screens from the first multimedia content, obtaining 10 video frame screens, and then performs subject recognition on each of the 10 extracted video frame screens. In some embodiments, two video frame screens are attached to resource objects of a first resource type, such as a white sweatshirt or a short dress.
[0077] S602: Based on the pre-configured correspondence, determine the item object corresponding to the item type.
[0078] S603: Based on the item object, determine at least one recommended object.
[0079] In some embodiments, the at least one recommended object belongs to the type of recommended object to which the item object corresponds.
[0080] Based on the aforementioned item object, at least one recommended object having similar or identical characteristics is determined. In some embodiments, the at least one recommended object belongs to the aforementioned item type.
[0081] In some embodiments, after recognizing that an item object is attached to the first multimedia content, it is possible to determine, based on each recognized item object, at least one recommended object having similar or identical characteristics to that item object. In some embodiments, at least one recommended object belongs to an item type.
[0082] In a selective embodiment, after recognizing that an item object is attached to the first multimedia content, the video frame image to which the item object is attached is sent to an image similarity calculation model, and based on the image similarity calculation model, at least one recommended object corresponding to the item object can be determined.
[0083] In some embodiments, an image similarity calculation model is used to match an item object with a recommended object in a recommended object library, selecting at least one recommended object that has similar or identical characteristics to the item object.
[0084] As an example, assuming that the recognized item objects are a sweatshirt and a short dress, a video frame image featuring a white sweatshirt and a short dress is sent to an image similarity calculation model, which then matches the sweatshirt and short dress with items in the item library. In some embodiments, items with similar or identical features to the sweatshirt include a white long-sleeved sweatshirt and a white short-sleeved sweatshirt, while items with similar or identical features to the short dress include a white short dress and a purple short dress.
[0085] S604: Display at least one of the recommended objects.
[0086] In a selective embodiment, if, while displaying at least one recommended object, the recommended object display page receives a pre-configured interaction with the target recommended object, the user may jump from the recommended object display page to the target recommended object's details page so that the user can learn more about the target recommended object based on the details page. In some embodiments, the pre-configured interaction with the target recommended object includes a click operation on the target recommended object.
[0087] In the multimedia content processing method provided by the embodiments of this disclosure, while the multimedia content is being displayed, the embodiments of this disclosure provide the user with an extended consumption path for content attached to the multimedia content, thereby improving the user experience, by displaying recommended objects related to the target resource object attached to the multimedia content to the user, based on the target resource object attached to the multimedia content.
[0088] In actual application, the target resource object recognized from the first multimedia content may include background music (BGM). In some embodiments, the recommended object type corresponding to BGM includes music type, etc.
[0089] In addition to the embodiments described above, embodiments of this disclosure provide a specific method for determining recommended objects based on background music, the method comprising the following steps:
[0090] First, in response to a pre-configured trigger operation acting on the display page of the first multimedia content, the system recognizes the background music (BGM) accompanying the first multimedia content. Next, music recognition is performed on the BGM, and the music recognition result is obtained. Then, based on the music recognition result, the music resource corresponding to the BGM is determined. Finally, at least one of the music resources is displayed.
[0091] In a selective embodiment, it is possible to perform music recognition on background music (BGM) by calling a music search algorithm based on voice fingerprints, obtain the music recognition result, and then determine the music resource corresponding to the BGM based on the music recognition result.
[0092] For example, assuming that the background music (BGM) is the chorus of a song, a music search algorithm based on voice fingerprints is called to recognize the song information for the chorus, for example, titled "Song A," and the BGM is then used as the corresponding music resource to search the music library for a song titled "Song A."
[0093] In practical applications, after determining the music resource corresponding to the background music (BGM), it is also possible to display that music resource on the recommended object display page. Figure 7 is a schematic diagram of another recommended object display page provided by an embodiment of this disclosure. In some embodiments, the recommended object display page displays the music title, music cover, author information, etc., corresponding to the music resource.
[0094] Furthermore, the recommended object display page includes a music playback control 701 that, when a trigger operation is accepted for the music playback control, plays a music resource based on the recommended object display page.
[0095] In an optional embodiment, the recommended object display page may be provided with a pre-configured return control 702 that enables the function of terminating the recommended object display page when a trigger operation on a pre-configured return control is accepted.
[0096] In some embodiments, if at least one target resource object includes background music (BGM), music recognition is first performed on the BGM, and the music recognition result is obtained. Then, based on the music recognition result, the music resource corresponding to the BGM is determined and displayed. Embodiments of this disclosure provide users with an extended consumption path for content attached to videos, further improving the user experience.
[0097] In actual application, the target resource object recognized from the first multimedia content may include address information. In some embodiments, the type of the recommended object corresponding to the address information may include, for example, a lifestyle service type.
[0098] In addition to the embodiments described above, embodiments of the present disclosure provide a method for determining recommended objects based on address information, specifically, first, in response to a pre-configured trigger operation acting on a display page of first multimedia content, recognize address information attached to the first multimedia content, determine at least one life service object belonging to a life service type that is within a pre-configured distance range centered on the address information, and display at least one life service object.
[0099] In some embodiments, accepting a pre-configured trigger operation that acts on the display page of the first multimedia content makes it possible to invoke a speech recognition algorithm and recognize the audio files within the first multimedia content, thereby obtaining address information associated with the first multimedia content.
[0100] In some embodiments, the speech recognition algorithms include algorithms based on dynamic time regularization, speech recognition algorithms based on deep learning neural networks, and the like.
[0101] In a selective embodiment, it is possible to obtain address information attached to the first multimedia content by calling a text recognition algorithm and recognizing the subtitle content of the first multimedia content.
[0102] For example, assuming the address information attached to the first multimedia content is "Location ABC," a search will be performed centered around "Location ABC," including malls, supermarkets, clothing stores, tourist attractions, etc., within 1km of "Location ABC."
[0103] In an optional embodiment, if the first multimedia content is accompanied by a specific address anchor, the location information corresponding to the address anchor may be directly determined as the life service object corresponding to the address anchor.
[0104] In some embodiments, if at least one target resource object contains address information, first, in response to a pre-configured trigger operation acting on the playback page of the first multimedia content, the address information attached to the first multimedia content is recognized, and then at least one life service object belonging to a life service type located within a pre-configured distance range centered on the address information is determined and displayed. Thus, embodiments of the present disclosure provide users with an extended consumption path to content attached to multimedia content, thereby improving the user experience.
[0105] In actual application, the target resource object recognized from the first multimedia content includes the target face displayed on the video frame screen. In some embodiments, the type of the recommended object corresponding to the target face displayed on the video frame screen includes the user account type.
[0106] In addition to the embodiments described above, embodiments of the present disclosure provide a method for determining a recommended object based on a target face displayed on a video frame screen, specifically, if the user to whom the target face belongs allows the use of target face information, first, in response to a pre-configured trigger operation acting on the playback page of a first multimedia content, the method recognizes the target face attached to the first multimedia content, and then, based on the target face displayed on the video frame screen, determines at least one user account whose similarity between the user avatar and the target face reaches a pre-configured threshold. In some embodiments, the user account belongs to a user account type.
[0107] In a selective embodiment, upon accepting a pre-configured trigger operation acting on the playback page of a first multimedia content, the system first extracts a video frame from the first multimedia content, then invokes a face recognition algorithm to recognize each video frame from the first multimedia content, and finally recognizes the video frame from which the target face in the first multimedia content is located.
[0108] In a selective embodiment, after recognizing a video frame in the first multimedia content that has a target face, if the user to whom the target face belongs authorizes the use of the target face information, the video frame with the target face is sent to a face matching service terminal, and based on the face matching service terminal, at least one user account is identified in which the similarity between the user's avatar and the target face has reached a pre-set threshold.
[0109] In some embodiments, the pre-set thresholds are determined based on actual needs and may be set to, for example, 80%, 85%, 90%, 95%, etc.
[0110] For example, assuming that a face recognition algorithm is invoked and the recognized target faces are face A and face B, a video frame containing face A and face B is sent to a face matching server terminal. Based on the face matching server terminal, the system searches the user avatar database for user avatars whose similarity to face A and face B reaches a pre-set threshold. In some embodiments, the user account corresponding to a user avatar whose similarity to face A reaches a pre-set threshold is "Mr. A," and the user account corresponding to a user avatar whose similarity to face B reaches a pre-set threshold is "Mr. B."
[0111] In a selective embodiment, assuming the first multimedia content is a first video, it is also possible to call a text recognition algorithm to recognize the subtitle content of the first video, obtain the names of people attached to the first video, search for user nicknames whose similarity to the names reaches a pre-set threshold, and identify the user accounts corresponding to the searched user nicknames as the target resource object. For example, if text recognition is performed on the subtitle content of the first video, the name of a person "Flower" may be obtained, and then user nicknames such as "Flower Sensei" and "Flower Manager" may be searched for whose similarity to "Flower" reaches a pre-set threshold.
[0112] In some embodiments, after identifying at least one user account whose similarity between the user avatar and the target face reaches a pre-set threshold, the user account is also displayed on the recommended object display page if the user to which the user account belongs allows the user account to be displayed. Figure 8 shows a recommended object display page provided by an embodiment of this disclosure.
[0113] In a selective embodiment, at least one user account displayed on the recommended object display page each has a pre-configured interaction control that responds to a trigger operation on the pre-configured interaction control corresponding to a first user account among the at least one user account, and establishes a pre-configured interaction relationship between the current user account and the first user account.
[0114] In some embodiments, a trigger operation on a pre-configured interactive control corresponding to a first user account of at least one user account may include, but is not limited to, click operations, long press operations, etc., on the pre-configured interactive control. In some embodiments, the first user account may be any one of at least one user account.
[0115] In some embodiments, a pre-configured interaction relationship between the current user account and a first user account may include determining the first user account as a follower of the current user account.
[0116] As shown in Figure 8, when a trigger operation is accepted for a pre-configured dialogue control 802 corresponding to the first user account 801, the function is realized to establish a pre-configured dialogue relationship between the current user account and the first user account by confirming the first user account 801 as a follower of the current user account.
[0117] In some embodiments, if at least one target resource object includes a target face, first, in response to a pre-configured trigger operation acting on the playback page of the first multimedia content, the system recognizes the target face attached to the first multimedia content, and then, based on the target face, identifies and displays at least one user account whose similarity between the user's avatar and the target face reaches a pre-configured threshold. Thus, embodiments of the disclosure provide users with an extended consumption path to the content attached to the multimedia content, thereby improving the user experience.
[0118] Based on embodiments of the above method, the present disclosure further provides a multimedia content processing device, and Figure 9 is a schematic diagram of the structure of the multimedia content processing device provided by embodiments of the present disclosure, and the device is A recognition module 901 for recognizing at least one target resource object attached to the first multimedia content in response to a pre-configured trigger operation acting on the playback page of the first multimedia content, wherein the recognition module 901 has a pre-configured correspondence between the target resource object and the type of recommended object, Based on the pre-configured correspondence, a first determination module 902 for determining the type of recommended object corresponding to the first target resource object among the at least one target resource object, A second determination module 903 that determines at least one recommended object based on the first target resource object, wherein the at least one recommended object belongs to the type of recommended object corresponding to the first target resource object, and the second determination module 903 The system comprises a display module 904 for displaying at least one of the recommended objects.
[0119] In a selective embodiment, the display module is: The system includes a classification display submodule that classifies and displays the recommended objects determined based on the first target resource object and the second resource object, respectively, according to the type of recommended object.
[0120] In a selective embodiment, the classification display submodule is: A first determination submodule for displaying on the first card at least one first recommended object determined based on the first target resource object, wherein the first recommended object comprises a first determination submodule belonging to a type of recommended object corresponding to the first target resource object, The second card comprises a second determination submodule for displaying at least one second recommended object determined based on the second resource object, wherein the second recommended object belongs to a type of recommended object corresponding to the second target resource object, and the second card and the second card belong to a set of cards that are displayed stacked.
[0121] In a selective embodiment, the classification display submodule is: The system further comprises a scroll display submodule for displaying individual cards in the card set in a scrollable manner in response to a pre-configured slide operation triggered on the card set.
[0122] In a selective embodiment, the classification display submodule is: In response to a pre-configured trigger operation on the target card in the card set, the recommended object display page includes a recommended object display submodule for displaying recommended objects on the target card, The system includes an acceptance submodule for accepting pre-configured interactive operations on a target recommended object on the recommended object display page.
[0123] In a selective embodiment, the at least one target resource object includes an item object, the type of the recommended object corresponding to the item object includes an item type, and the second deterministic module is It comprises a third determination submodule for determining at least one recommended item that has similar or identical characteristics to the aforementioned item object and belongs to the aforementioned item type.
[0124] In a selective embodiment, the at least one target resource object includes BGM, the type of the recommended object corresponding to the BGM includes a music type, and the second deterministic module is A music recognition submodule for performing music recognition on the aforementioned background music and obtaining the music recognition result, The system includes a fourth determination submodule for determining a music resource that corresponds to the background music and belongs to the music type, based on the music recognition result.
[0125] In a selective embodiment, the at least one target resource object includes address information, the type of the recommended object corresponding to the address information includes a life service type, and the second deterministic module is The system includes a fifth determination submodule for determining at least one life service object that is located within a pre-set distance range and belongs to the life service type, centered on the address information.
[0126] In a selective embodiment, the at least one target resource object includes a target face displayed on the video frame screen, the type of the recommended object corresponding to the target face includes a user account type, and the second determinant module is The system includes a sixth determination submodule for determining at least one user account belonging to the user account type, based on the target face displayed on the video frame screen, where the similarity between the user avatar and the target face reaches a pre-set threshold.
[0127] In a selective embodiment, the recognition module is: A screen display submodule for displaying multiple video keyframe screens within the first multimedia content on the video recognition page in the form of a transition animation, in response to a pre-configured trigger operation acting on the display page of the first multimedia content, The system includes a target resource object recognition submodule for recognizing at least one target resource object attached to the first multimedia content based on the plurality of video keyframe images.
[0128] In a multimedia content processing device provided by an embodiment of the present disclosure, first, in response to a pre-configured trigger operation acting on the display page of a first multimedia content, the device recognizes at least one target resource object attached to the first multimedia content, and has a pre-configured object correspondence relationship between the target resource object and the type of recommended object. Next, based on the pre-configured correspondence relationship, the device determines the type of recommended object corresponding to the first target resource object among the at least one target resource object, determines at least one recommended object based on the first target resource object, and then displays the at least one recommended object. Because the embodiment of the present disclosure can display recommended objects associated with target resource objects to the multimedia content to the user while the multimedia content is being displayed, the embodiment of the present disclosure provides the user with an extended consumption path for content attached to the multimedia content, thereby improving the user experience.
[0129] In addition to the methods and apparatus described above, embodiments of the present disclosure, when executed on a terminal device, further provide the terminal device with a computer-readable storage medium storing instructions for implementing the multimedia content processing method described on the embodiments of the present disclosure.
[0130] The embodiments of the present disclosure further provide a computer program product that, when executed by a processor, includes a computer program / instruction that implements the multimedia content processing method described in the embodiments of the present disclosure.
[0131] Furthermore, embodiments of the present disclosure provide a multimedia content processing device, referring to Figure 10, which includes a processor 1001, a memory 1002, an input device 1003, and an output device 1004.
[0132] The number of processors 1001 in a multimedia content processing device may be one or more, and in Figure 10, one processor is used as an example. In some embodiments of this disclosure, the processors 1001, memory 1002, input device 1003, and output device 1004 may be connected via a bus or other means, and in Figure 10, connection via a bus is shown as an example.
[0133] Memory 1002 is used to store software programs and modules, and the processor 1001 executes various functions and data processing of the multimedia content processing device by executing the software programs and modules stored in memory 1002. Memory 1002 mainly includes a program storage area and a data storage area, of which the program storage area can store the operating system, application programs required for at least one function, etc. Furthermore, memory 1002 may also include high-speed random access memory, and may include, for example, at least one disk memory device, a non-volatile memory such as a flash memory device, or another volatile solid-state memory device. Input device 1003 is used for input numerical or character information and is used to generate signal inputs related to user settings and function control of the multimedia content processing device.
[0134] Specifically, in this embodiment, the processor 1001 loads executable files corresponding to the processes of one or more applications into memory 1002 according to the following instructions, and the processor 1001 executes the applications stored in memory 1002, thereby realizing the various functions of the multimedia content processing device described above.
[0135] In this specification, relational terms such as “first” and “second” are used simply to distinguish one entity or operation from another, and do not necessarily require or suggest that such an actual relationship or order exists between those entities or operations. Furthermore, the terms “equipment,” “includes,” or any other variations thereof are intended to cover non-exclusive inclusion, where a process, method, product, or apparatus consisting of a set of elements includes not only those elements but also other elements not explicitly enumerated, or other elements specific to such process, method, product, or apparatus. Unless further limited, an element defined in the expression “includes one XX” does not exclude the presence of other identical elements in a process, method, product, or apparatus that includes that element.
[0136] The above are merely specific examples of the disclosure intended to enable those skilled in the art to understand or implement the disclosure. Various modifications to these examples will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other examples without departing from the spirit or scope of the disclosure. Accordingly, the disclosure is not limited to these examples described herein, but rather follows the broadest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for processing multimedia content, In response to a pre-configured trigger operation acting on the display page of the first multimedia content, multiple video keyframe screens within the first multimedia content are displayed on the video recognition page in the form of a transition animation. Based on the plurality of video keyframe screens, recognize at least one target resource object attached to the first multimedia content, wherein the target resource object has a pre-configured correspondence with the type of recommended object. Based on the aforementioned pre-configured correspondence, the type of recommended object corresponding to the first target resource object among the at least one target resource object is determined, Determining at least one recommended object based on the first target resource object, wherein the at least one recommended object belongs to the type of recommended object corresponding to the first target resource object, The features include displaying at least one of the aforementioned recommended objects, Processing method.
2. The at least one target resource object further includes a second target resource object, and the display of the at least one recommended object is: The system is characterized by classifying and displaying the recommended objects determined based on the first target resource object and the second target resource object, respectively, according to the type of recommended object. The processing method according to claim 1.
3. Classifying and displaying the recommended objects determined based on the first target resource object and the second target resource object, respectively, according to the type of recommended object, is: Displaying at least one first recommended object determined based on the first target resource object on the first card, wherein the first recommended object is one of the determined first recommended objects that belongs to the type of recommended object corresponding to the first target resource object. The second card displays at least one second recommended object determined based on the second target resource object, wherein the second recommended object belongs to a type of recommended object corresponding to the second target resource object, and the first card and the second card display the determined at least one second recommended object which belongs to a stacked card set. The processing method according to claim 2.
4. The processing method according to claim 3, further comprising scrolling and displaying each card in the card set in response to a pre-set slide operation triggered on the card set.
5. In response to a pre-configured trigger operation on a target card in the aforementioned card set, the recommended object on the target card is displayed on the recommended object display page. The further comprising accepting pre-configured interactive operations on the target recommended object on the recommended object display page, The processing method according to claim 3.
6. The at least one target resource object includes an item object, and the type of the recommendation object corresponding to the item object includes an item type, and determining at least one recommendation object based on the first target resource object is: This method is characterized by determining at least one recommended item that has the same or similar characteristics as the aforementioned item object and belongs to the aforementioned item type. The processing method according to claim 1.
7. The at least one target resource object includes BGM, the type of the recommended object corresponding to the BGM includes a music type, and determining at least one recommended object based on the first target resource object is: Perform music recognition on the aforementioned background music and obtain the music recognition result. The method is characterized by including, based on the music recognition result, determining a song resource that corresponds to the background music and belongs to the music type, The processing method according to claim 1.
8. The at least one target resource object includes address information, and the type of recommended object corresponding to the address information includes a life service type, and determining at least one recommended object based on the first target resource object is: This method is characterized by including determining at least one life service object that is located within a pre-defined distance range centered on the address information and belongs to the life service type. The processing method according to claim 1.
9. The at least one target resource object includes a target face displayed on the video frame screen, the type of the recommended object corresponding to the target face includes a user account type, and determining at least one recommended object based on the first target resource object is: The method is characterized by determining at least one user account that belongs to the user account type, based on the target face displayed on the video frame screen, where the similarity of the user avatar to the target face reaches a pre-set threshold. The processing method according to claim 1.
10. A multimedia content processing device, A recognition module for displaying multiple video keyframe screens within the first multimedia content on a video recognition page in the form of a transition animation, in response to a pre-configured trigger operation acting on the playback page of the first multimedia content, and for recognizing at least one target resource object attached to the first multimedia content based on the multiple video keyframe screens, wherein the target resource object has a pre-configured correspondence with the type of recommended object. A first determination module for determining the type of recommended object corresponding to the first target resource object among the at least one target resource object, based on the pre-configured correspondence relationship, A second determination module for determining at least one recommended object based on the first target resource object, wherein the at least one recommended object belongs to a type of recommended object corresponding to the first target resource object, The system is characterized by comprising a display module for displaying at least one of the recommended objects, Processing device.
11. A computer-readable storage medium characterized in that, when executed on a terminal device, the terminal device stores an instruction that enables the processing method described in any one of claims 1 to 9.
12. A multimedia content processing device comprising memory, a processor, and a computer program stored in the memory and executable by the processor, wherein the processor executes the computer program to realize the processing method described in any one of claims 1 to 9.
Citation Information
Patent Citations
Information display method and device, computer equipment and storage medium
CN115599944A
Search support system, search support method, and search support program
WO2010016281A1