Multimedia information interaction method, device and storage medium
By automatically generating and displaying video labeling data and allowing users to upload multimedia information, the problem of single user interaction on short video platforms is solved, and the effect of simplifying operations and improving fun is achieved.
Patent Information
- Application Number
- CN201910188271.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-03-13
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2039-03-13
AI Technical Summary
The existing short video platform has a single user interaction form and cumbersome operation. The possibility of users expressing personalized ideas is limited, which reduces the user experience and fun.
By automatically generating multiple label data related to the initial video and displaying these data on the terminal application interface, users can upload multimedia information associated with the label data, simplifying the operation process and improving the fun of interaction.
It realizes personalized interaction between multimedia information between users, simplifies operation steps, and improves user experience and interactive fun.
Smart Images

Figure CN109977303B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer software applications, and more particularly to a method, device, and storage medium for multimedia information interaction. Background Art
[0002] With the development of mobile internet, video is becoming more mobile, fragmented, and social. For example, users of mobile devices like phones can shoot short videos and post them on social platforms to share them with friends. Typically, on short video platforms, users are limited to interacting with videos through limited channels, such as liking, commenting, and forwarding. These interactions are relatively simple.
[0003] To increase the diversity of user interactions, some short video platforms offer a list of audio and video clips, requiring users to select them from the list before interacting. However, the audio and video clips in the list don't express the user's true thoughts. This limits the user's ability to express their individual thoughts, reduces the fun of user interaction, and ultimately degrades the user experience. Furthermore, selecting audio and video from the list makes the entire process cumbersome. Summary of the Invention
[0004] In order to overcome the problem of fixed audio and video used when users interact with each other using audio and video in the related art, the present application discloses a multimedia information interaction method, device and storage medium, which can automatically generate related annotation data for the initial video, and the user can upload multimedia information associated with the annotation data through the interface of the application on the terminal. The entire operation process is simple and smooth, reducing the operation steps and simplifying the operation process; moreover, the user can personalize and upload the user's opinion information on the annotation data, which increases the fun of user interaction and thus improves the user experience.
[0005] According to a first aspect of an embodiment of the present application, a method for interacting with multimedia information is provided, including:
[0006] Get the initial video;
[0007] generating, based on the features of the initial video, a plurality of annotated data corresponding to the initial video, wherein the annotated data is used to annotate derived features of the initial video, and at least two of the plurality of annotated data are used to annotate the same derived feature;
[0008] The multiple annotated data of the initial video are displayed through an interface of an application on a terminal, wherein the interface is also used to upload multimedia information associated with the annotated data.
[0009] Optionally, generating a plurality of annotation data corresponding to the initial video according to the features of the initial video includes:
[0010] Obtaining derived features of the initial video based on a content understanding result of the initial video;
[0011] Obtaining annotation data corresponding to each of the derived features.
[0012] Optionally, obtaining the annotation data corresponding to each of the derived features includes:
[0013] Matching the annotation data corresponding to the derived features from the database; or,
[0014] Keywords in the derived features are extracted to generate labeled data containing the keywords.
[0015] Optionally, obtaining the derived features of the initial video based on the content understanding result of the initial video includes:
[0016] Acquire description information of the initial video from a result of understanding the content of the initial video, wherein the description information is used to describe the content of the initial video;
[0017] Feature learning is performed on the description information to obtain the derived features.
[0018] Optionally, at least two annotation data used to annotate the same derived feature contain keywords with the same semantics, and the semantics of the at least two annotation data are different.
[0019] Optionally, presenting the plurality of annotated data of the initial video through an interface of an application on a terminal includes at least one of the following:
[0020] During the playback of the initial video, displaying part or all of the multiple annotation data on the interface;
[0021] After the initial video is played, part or all of the multiple annotation data are displayed on the interface;
[0022] Using part or all of the multiple annotation data as labels of the initial video, and displaying them on the interface;
[0023] Part or all of the multiple annotation data are used as the cover of the initial video and displayed on the interface.
[0024] Optionally, after displaying the plurality of annotated data of the initial video through an interface of an application on a terminal, the interaction method further includes:
[0025] Uploaded multimedia information associated with the annotation data is received, wherein the multimedia information is generated based on part or all of the multiple annotation information, and all or part of the uploaded multimedia information is set as the initial video.
[0026] Optionally, the multiple annotation data include N-level annotation data, wherein the N-th level annotation data is subordinate to the (N-1)-th level annotation data, and the N-level annotation data is generated based on the initial video and / or the multimedia information, and N is a natural number.
[0027] Optionally, displaying the multiple annotated data of the initial video through an interface of an application on a terminal includes:
[0028] Displaying the plurality of annotated data according to the subordinate relationship of the N-level annotated data;
[0029] Wherein, while displaying the plurality of annotated data, the initial video and / or multimedia information for generating the annotated data is displayed.
[0030] According to a second aspect of an embodiment of the present application, a multimedia information interaction device is provided, including:
[0031] an initial video acquiring unit, configured to acquire an initial video;
[0032] a labeling data generating unit configured to generate a plurality of labeling data corresponding to the initial video based on the features of the initial video, wherein the labeling data is used to label the derived features of the initial video, and at least two of the plurality of labeling data are used to label the same derived feature;
[0033] The interactive unit is configured to display the multiple annotation data of the initial video through an interface of an application on a terminal, wherein the interface is also used to upload multimedia information associated with the annotation data.
[0034] Optionally, generating a plurality of annotation data corresponding to the initial video according to the features of the initial video includes:
[0035] Obtaining derived features of the initial video based on a content understanding result of the initial video;
[0036] Obtaining annotation data corresponding to each of the derived features.
[0037] Optionally, obtaining the annotation data corresponding to each of the derived features includes:
[0038] Matching the annotation data corresponding to the derived features from the database; or,
[0039] Keywords in the derived features are extracted to generate labeled data containing the keywords.
[0040] Optionally, obtaining the derived features of the initial video based on the content understanding result of the initial video includes:
[0041] Acquire description information of the initial video from a result of understanding the content of the initial video, wherein the description information is used to describe the content of the initial video;
[0042] Feature learning is performed on the description information to obtain the derived features.
[0043] Optionally, at least two annotation data used to annotate the same derived feature contain keywords with the same semantics, and the semantics of the at least two annotation data are different.
[0044] Optionally, presenting the plurality of annotated data of the initial video through an interface of an application on a terminal includes at least one of the following:
[0045] During the playback of the initial video, displaying part or all of the multiple annotation data on the interface;
[0046] After the initial video is played, part or all of the multiple annotation data are displayed on the interface;
[0047] Using part or all of the multiple annotation data as labels of the initial video, and displaying them on the interface;
[0048] Part or all of the multiple annotation data are used as the cover of the initial video and displayed on the interface.
[0049] Optionally, after displaying the plurality of annotated data of the initial video through the interface of the application on the terminal, the interactive device further includes:
[0050] Uploaded multimedia information associated with the annotation data is received, wherein the multimedia information is generated based on part or all of the multiple annotation information, and all or part of the uploaded multimedia information is set as the initial video.
[0051] Optionally, the multiple annotation data include N-level annotation data, wherein the N-th level annotation data is subordinate to the (N-1)-th level annotation data, and the N-level annotation data is generated based on the initial video and / or the multimedia information, and N is a natural number.
[0052] Optionally, displaying the multiple annotated data of the initial video through an interface of an application on a terminal includes:
[0053] Displaying the plurality of annotated data according to the subordinate relationship of the N-level annotated data;
[0054] Wherein, while displaying the plurality of annotated data, the initial video and / or multimedia information for generating the annotated data is displayed.
[0055] According to a third aspect of an embodiment of the present application, there is provided an interactive control device for multimedia information, comprising:
[0056] processor;
[0057] a memory for storing instructions executable by the processor;
[0058] The processor is configured to execute any one of the multimedia information interaction methods described above.
[0059] According to the fourth aspect of the embodiment of the present application, a non-temporary computer-readable storage medium is provided. When the instructions in the storage medium are executed by the processor of the mobile terminal, the mobile terminal is enabled to execute an audio and video interaction method, which includes the multimedia information interaction method described in any one of the above items.
[0060] According to a fifth aspect of an embodiment of the present application, a computer program product is provided, including a computer program product, wherein the computer program includes program instructions, and when the program instructions are executed by a mobile terminal, the mobile terminal executes the steps of the above-mentioned multimedia information interaction method.
[0061] The technical solutions provided by the embodiments of the present application may have the following beneficial effects:
[0062] Obtain an initial video. Based on the features of the initial video, generate multiple annotation data corresponding to the initial video. The annotation data is used to annotate the derived features of the initial video. At least two of the multiple annotation data corresponding to the initial video are used to annotate the same derived feature. Display multiple annotation data of the initial video through the interface of the application on the terminal. The interface of the application on the terminal can also be used to upload multimedia information associated with the annotation data. The multimedia information can be a video expressing the user's views on the annotation data, or it can be an audio expressing the user's views on the annotation data. Annotation data related to the initial video can be automatically generated, and multiple annotation data of the initial video can be displayed through the interface of the application on the terminal. The user can upload multimedia information associated with the annotation data through the interface of the application on the terminal. The entire operation process is simple and smooth, and the operation steps are reduced. Moreover, the user can personalize the uploaded user's views on the annotation data, which increases the fun of user interaction and thus improves the user experience.
[0063] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0065] Figure 1 is a flowchart of a multimedia information interaction method according to an exemplary embodiment;
[0066] Figure 2 is a flowchart of a multimedia information interaction method according to an exemplary embodiment;
[0067] Figure 3 is a block diagram of a multimedia information interaction device according to an exemplary embodiment;
[0068] Figure 4 is a block diagram showing an apparatus for executing a multimedia information interaction method according to an exemplary embodiment;
[0069] Figure 5 The present invention is a block diagram showing an apparatus for executing a multimedia information interaction method according to an exemplary embodiment. DETAILED DESCRIPTION
[0070] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.
[0071] Figure 1 This is a flowchart of a multimedia information interaction method according to an exemplary embodiment. Specifically, it includes the following steps:
[0072] In step S101, an initial video is obtained.
[0073] In this step, an initial video is obtained. It can be understood that the initial video is a video uploaded by the user to an application on the terminal, such as a short video platform. The initial video triggers the user to interact with multimedia information in the application on the terminal.
[0074] In step S102, a plurality of annotation data corresponding to the initial video is generated based on the features of the initial video, wherein the annotation data is used to annotate the derived features of the initial video, and at least two of the plurality of annotation data are used to annotate the same derived feature.
[0075] In this step, based on the features of the initial video, multiple annotated data corresponding to the initial video are generated. The annotated data is used to annotate the derived features of the initial video. At least two of the multiple annotated data corresponding to the initial video are used to annotate the same derived feature.
[0076] Derived features are new features derived from feature learning on the original data. Derived features generally arise from two factors: changes in the original data itself, which introduce features that were not originally present; or when learning features from the original data, the algorithm generates derived features based on relationships between features. Sometimes, derived features better reflect the relationships between multiple features in the original data.
[0077] In step S103, the multiple annotated data of the initial video are displayed through an interface of an application on a terminal, wherein the interface is also used to upload multimedia information associated with the annotated data.
[0078] In this step, the multiple annotated data of the initial video are displayed through the interface of the application on the terminal. The interface of the application on the terminal can also be used to upload multimedia information associated with the annotated data. The multimedia information can be a video expressing the user's views on the annotated data, or an audio expressing the user's views on the annotated data.
[0079] According to an embodiment of the present application, an initial video is obtained. Based on the features of the initial video, a plurality of annotation data corresponding to the initial video are generated. The annotation data is used to annotate the derived features of the initial video. At least two of the aforementioned plurality of annotation data corresponding to the initial video are used to annotate the same derived feature. The plurality of annotation data of the initial video are displayed through the interface of the application on the terminal. The interface of the application on the terminal can also be used to upload multimedia information associated with the annotation data. The multimedia information can be a video expressing the user's views on the annotation data, or it can be an audio expressing the user's views on the annotation data. Annotation data related to the initial video can be automatically generated, and the plurality of annotation data of the initial video can be displayed through the interface of the application on the terminal. The user can upload multimedia information associated with the annotation data through the interface of the application on the terminal. The entire operation process is simple and smooth, and the operation steps are reduced. Moreover, the user can personalize the uploaded user's view information on the annotation data, which increases the fun of user interaction and thus improves the user experience.
[0080] Figure 2 This is a flowchart of a multimedia information interaction method according to an exemplary embodiment, which is a more complete embodiment than the previous embodiment. The multimedia information interaction method includes the following steps:
[0081] In step S201, an initial video is obtained.
[0082] This step is the same as Figure 1 The steps are the same as step S101 in FIG. 1 and will not be described in detail here.
[0083] In step S202, a plurality of annotation data corresponding to the initial video is generated based on the features of the initial video, wherein the annotation data is used to annotate the derived features of the initial video, and at least two of the plurality of annotation data are used to annotate the same derived feature.
[0084] In this step, derived features of the initial video are obtained based on the content understanding results of the initial video. Specifically, description information of the initial video is obtained from the content understanding results. The description information is used to describe the content of the initial video. Feature learning is performed on the description information of the initial video to obtain the corresponding derived features.
[0085] Obtaining annotation data corresponding to each derived feature. The annotation data is used to annotate the derived features of the initial video. The same derived feature is annotated by at least two annotation data. The at least two annotation data used to annotate the same derived feature contain keywords with the same semantics, and the at least two annotation data have different semantics.
[0086] Specifically, the annotated data corresponding to the derived features are matched from the database. It is understandable that the matching here can be a semantic match or a string match. Alternatively, the keywords in the derived features are extracted and the annotated data containing the keywords are generated. For example, the derived feature is "the relationship between inner beauty and outer beauty". The keywords in the derived features are "inner beauty" and "outer beauty". Based on the keywords in the derived features, two annotated data containing the keywords are generated, namely "inner beauty is important" and "outer beauty is important". "Inner beauty is important" and "outer beauty is important" have opposite semantics, and at the same time contain the keyword "important" with the same semantics.
[0087] In step S203, the multiple annotation data of the initial video are displayed through the interface of the application on the terminal, wherein the interface is also used to upload multimedia information associated with the annotation data.
[0088] In this step, multiple annotation data of the initial video are displayed on the interface of the application on the terminal. Some or all of the multiple annotation data may be displayed on the interface of the application on the terminal during the playback of the initial video. Some or all of the multiple annotation data may be displayed on the interface of the application on the terminal after the initial video has finished playing. Some or all of the multiple annotation data may be displayed as tags for the initial video on the interface of the application on the terminal. Some or all of the multiple annotation data may be displayed as the cover of the initial video on the interface of the application on the terminal.
[0089] The terminal application interface is also used to upload multimedia information associated with the annotated data. The multimedia information is generated based on some or all of the annotated data. The uploaded multimedia information, in whole or in part, is set as the initial video, thereby enabling further user interaction with the multimedia information within the terminal application.
[0090] Any user is free to choose to watch the initial video and / or uploaded multimedia information.
[0091] In step S204, uploaded multimedia information associated with the annotation data is received, wherein the multimedia information is generated based on part or all of the multiple annotation information, and all or part of the uploaded multimedia information is set as the initial video.
[0092] In this step, multimedia information associated with the annotation data uploaded through the interface of the application on the terminal is received. The multimedia information is generated based on part or all of the multiple annotation information. The multimedia information can be a video expressing the user's views on part or all of the multiple annotation information, or it can be an audio expressing the user's views on part or all of the multiple annotation information. The view is not limited to the user's support, opposition, partial approval, etc. of part or all of the information in the multiple annotation information. All or part of the uploaded multimedia information is set as the initial video, which can further trigger the user to interact with the multimedia information in the application of the terminal.
[0093] Optionally, the multimedia information uploaded by the user based on the annotated data of the initial video can be set as the initial video. The above steps can be performed on all or part of the initial video. Figure 1 or Figure 2 The embodiment shown.
[0094] Specifically, multimedia information associated with the annotation data uploaded through the interface of the application on the terminal is continuously received. All or part of the uploaded multimedia information is set as the initial video, thereby further triggering the user to interact with the multimedia information in the application of the terminal. With the continuous interaction of multimedia information, N-level annotation data is generated based on the initial video and / or multimedia information, where N is a natural number. The N-level annotation data includes N-level annotation data in the multiple annotation data, wherein the N-th level annotation data is subordinate to the (N-1)-th level annotation data.
[0095] It is easy to understand that the multiple annotated data can be displayed through the interface of the application on the terminal according to the subordinate relationship of the N-level annotated data. In particular, when displaying the multiple annotated data, the initial video and / or multimedia information generated by the annotated data is displayed. It is understandable that according to the subordinate relationship of the N-level annotated data, when displaying the multiple annotated data, the initial video and / or multimedia information generated by the annotated data is displayed in a tree-like format.
[0096] A derived feature of an initial video can be a video topic corresponding to the initial video. This derived feature corresponds to two annotated data items: the first topic and the second topic. The first and second topics reflect the pros and cons of the video topic. For example, a user uploads an initial video V_1 of a woman dancing to music. After the algorithm understands the content of the initial video V_1, a derived feature of the initial video V_1 is derived based on the understanding of the video content: the relationship between the girl's appearance and her special talents. Based on this derived feature, two annotated data items are generated for the initial video V_1. These two annotated data items are the first topic (topic 1): it is more important for a girl to be good-looking, and the second topic (topic 2): it is more important for a girl to have special talents.
[0097] The application interface on the terminal, such as the video interaction interface of a short video platform, displays the initial video V_1 and the corresponding two annotated data to the user, encouraging the user to debate their views on the two annotated data. The user who uploaded the initial video V_1 or other users can respond to the two annotated data by uploading multimedia information such as audio or video. Users can then interact with multimedia information such as audio or video based on the annotated data.
[0098] For the two annotated data points mentioned above: topic 1: It's more important for girls to be good-looking, and topic 2: It's more important for girls to have special talents. Users may upload multimedia content such as videos or audio to express their opinions, arguing that some special talents are important, while others are useless. This allows the annotated data points to be further refined. The annotated data points for topic 2: It's more important for girls to have special talents is further refined to yield topic 3: It's more important for girls to have artistic talents, and topic 4: It's more important for girls to have sports talents. Similarly, topic 3 and / or topic 4 can be further refined to yield one or more subtopics. Similarly, during the continuous interaction of multimedia information, users can freely refine branches within one or more annotated data points from the initial video, generating new annotated data from new perspectives.
[0099] As the annotated data continues to differentiate and multimedia information continues to interact, it can be understood that, according to the subordinate relationship of the N-level annotated data, while displaying the multiple annotated data, the initial video and / or multimedia information generated by the annotated data is displayed in a tree-like format. For example, the initial video V_1 is displayed as the root video. The multimedia information corresponding to the first topic (topic1) and the second topic (topic2) is displayed as a sub-video of the initial video V_1. The multimedia information corresponding to the second-level sub-topic (topic3) and the second-level sub-topic (topic4) is displayed as a sub-video of one or more multimedia information in the multimedia information of the second topic (topic2). Similarly, the multimedia information corresponding to the second-level sub-topic (topic3) and / or the second-level sub-topic (topic4) can be displayed as the parent video of the multimedia information corresponding to one or more second-level sub-topics, wherein, optionally, during the process of displaying the video, the annotated data can be displayed together with the video. By analogy, during the continuous interaction of multimedia information, the initial video and / or multimedia information corresponding to the N-level annotated data is displayed in a tree-like format.
[0100] According to an embodiment of the present application, descriptive information of the initial video is obtained from the content understanding results of the initial video. The descriptive information is used to describe the content of the initial video. Feature learning is performed on the descriptive information of the initial video to obtain corresponding derived features. Annotation data corresponding to each derived feature is obtained. The annotated data is used to annotate the derived features of the initial video. The same derived feature is annotated by at least two annotated data sets. The at least two annotated data sets used to annotate the same derived feature contain semantically identical keywords, and the semantics of the at least two annotated data sets are different. Multiple annotated data sets of the initial video are displayed via an application interface on a terminal. Multimedia information associated with the annotated data is received and uploaded via the application interface on the terminal. The multimedia information is generated based on part or all of the multiple annotated information sets. The multimedia information can be a video expressing a user's opinion on part or all of the multiple annotated information sets, or an audio recording expressing the user's opinion on part or all of the multiple annotated information sets. The opinion is not limited to a user's support, opposition, or partial agreement with part or all of the multiple annotated information sets. All or part of the uploaded multimedia information is set as the initial video, thereby further triggering user interaction with the multimedia information within the terminal application. Thus, users can discuss the labeled data at each level by uploading multimedia information, which further increases the fun of user interaction and improves the user experience.
[0101] Figure 3 FIG. 1 is a block diagram of a multimedia information interaction device according to an exemplary embodiment. Figure 3 As shown, the device 30 includes: an initial video acquisition unit 301, a labeling data generation unit 302 and an interaction unit 303.
[0102] The initial video acquiring unit 301 is configured to acquire an initial video.
[0103] The unit is configured to obtain an initial video. It is understood that the initial video is a video uploaded by a user to an application of the terminal, such as a short video platform. The initial video triggers the user to interact with multimedia information in the application of the terminal.
[0104] The annotation data generation unit 302 is configured to generate multiple annotation data corresponding to the initial video based on the features of the initial video, wherein the annotation data is used to annotate the derived features of the initial video, and at least two of the multiple annotation data are used to annotate the same derived feature.
[0105] The unit is configured to generate a plurality of annotated data corresponding to the initial video based on the features of the initial video. The annotated data is used to annotate the derived features of the initial video. At least two of the plurality of annotated data corresponding to the initial video are used to annotate the same derived feature.
[0106] Derived features are new features derived from feature learning on the original data. Derived features generally arise from two factors: changes in the original data itself, which introduce features that were not originally present; or when learning features from the original data, the algorithm generates derived features based on relationships between features. Sometimes, derived features better reflect the relationships between multiple features in the original data.
[0107] The interaction unit 303 is configured to display the multiple annotation data of the initial video through an interface of an application on the terminal, wherein the interface is also used to upload multimedia information associated with the annotation data.
[0108] The unit is configured to display multiple annotated data of the initial video through the interface of the application on the terminal. The interface of the application on the terminal can also be used to upload multimedia information associated with the annotated data. The multimedia information can be a video expressing the user's views on the annotated data, or an audio expressing the user's views on the annotated data.
[0109] According to an optional embodiment of the present application, the annotation data generation unit 302 is configured to obtain derived features of the initial video based on the content understanding results of the initial video. Specifically, description information of the initial video is obtained from the content understanding results of the initial video. The description information is used to describe the content of the initial video. Feature learning is performed on the description information of the initial video to obtain corresponding derived features.
[0110] Obtaining annotation data corresponding to each derived feature. The annotation data is used to annotate the derived features of the initial video. The same derived feature is annotated by at least two annotation data. The at least two annotation data used to annotate the same derived feature contain keywords with the same semantics, and the at least two annotation data have different semantics.
[0111] Specifically, the annotated data corresponding to the derived features are matched from the database. It is understandable that the matching here can be a semantic match or a string match. Alternatively, the keywords in the derived features are extracted and the annotated data containing the keywords are generated. For example, the derived feature is "the relationship between inner beauty and outer beauty". The keywords in the derived features are "inner beauty" and "outer beauty". Based on the keywords in the derived features, two annotated data containing the keywords are generated, namely "inner beauty is important" and "outer beauty is important". "Inner beauty is important" and "outer beauty is important" have opposite semantics, and at the same time contain the keyword "important" with the same semantics.
[0112] The interaction unit 303 is configured to display multiple annotation data of the initial video through the interface of the application on the terminal. It can be that during the playback of the initial video, some or all of the multiple annotation data are displayed on the interface of the application on the terminal. It can also be that after the initial video is played, some or all of the multiple annotation data are displayed on the interface of the application on the terminal. It can also be that some or all of the multiple annotation data are used as tags of the initial video and displayed on the interface of the application on the terminal. It can also be that some or all of the multiple annotation data are used as the cover of the initial video and displayed on the interface of the application on the terminal.
[0113] The terminal application interface is also used to upload multimedia information associated with the annotated data. The multimedia information is generated based on some or all of the annotated data. The uploaded multimedia information, in whole or in part, is set as the initial video, thereby further initiating user interaction with the multimedia information within the terminal application.
[0114] Any user is free to choose to watch the initial video and / or uploaded multimedia information.
[0115] The interaction unit 303 is configured to receive multimedia information associated with the annotation data uploaded through the interface of the application on the terminal. The multimedia information is generated based on part or all of the multiple annotation information. The multimedia information can be a video expressing the user's views on part or all of the multiple annotation information, or it can be an audio expressing the user's views on part or all of the multiple annotation information. The view is not limited to the user's support, opposition, partial approval, etc. of part or all of the information in the multiple annotation information. All or part of the uploaded multimedia information is set as the initial video, which can further trigger the user to interact with the multimedia information in the application of the terminal.
[0116] Optionally, the multimedia information uploaded by the user based on the annotated data of the initial video can be set as the initial video. The above steps can be performed on all or part of the initial video. Figure 1or Figure 2 The embodiment shown.
[0117] Specifically, multimedia information associated with the annotation data uploaded through the interface of the application on the terminal is continuously received. All or part of the uploaded multimedia information is set as the initial video, thereby further triggering the user to interact with the multimedia information in the application of the terminal. With the continuous interaction of multimedia information, N-level annotation data is generated based on the initial video and / or multimedia information, where N is a natural number. The N-level annotation data includes N-level annotation data in the multiple annotation data, wherein the N-th level annotation data is subordinate to the (N-1)-th level annotation data.
[0118] It is easy to understand that the multiple annotated data can be displayed through the interface of the application on the terminal according to the subordinate relationship of the N-level annotated data. In particular, when displaying the multiple annotated data, the initial video and / or multimedia information generated by the annotated data is displayed. It is understandable that according to the subordinate relationship of the N-level annotated data, when displaying the multiple annotated data, the initial video and / or multimedia information generated by the annotated data is displayed in a tree-like format.
[0119] A derived feature of an initial video can be a video topic corresponding to the initial video. This derived feature corresponds to two annotated data items: the first topic and the second topic. The first and second topics reflect the pros and cons of the video topic. For example, a user uploads an initial video V_1 of a woman dancing to music. After the algorithm understands the content of the initial video V_1, a derived feature of the initial video V_1 is derived based on the understanding of the video content: the relationship between the girl's appearance and her special talents. Based on this derived feature, two annotated data items are generated for the initial video V_1. These two annotated data items are the first topic (topic 1): it is more important for a girl to be good-looking, and the second topic (topic 2): it is more important for a girl to have special talents.
[0120] The application interface on the terminal, such as the video interaction interface of a short video platform, displays the initial video V_1 and the corresponding two annotated data to the user, encouraging the user to debate their views on the two annotated data. The user who uploaded the initial video V_1 or other users can respond to the two annotated data by uploading multimedia information such as audio or video. Users can then interact with multimedia information such as audio or video based on the annotated data.
[0121] For the two annotated data points mentioned above: topic 1: It's more important for girls to be good-looking, and topic 2: It's more important for girls to have special talents. Users may upload multimedia content such as videos or audio to express their opinions, arguing that some special talents are important, while others are useless. This allows the annotated data points to be further refined. The annotated data points for topic 2: It's more important for girls to have special talents is further refined to yield topic 3: It's more important for girls to have artistic talents, and topic 4: It's more important for girls to have sports talents. Similarly, topic 3 and / or topic 4 can be further refined to yield one or more subtopics. Similarly, during the continuous interaction of multimedia information, users can freely refine branches within one or more annotated data points from the initial video, generating new annotated data from new perspectives.
[0122] As the annotated data continues to differentiate and multimedia information continues to interact, it can be understood that, according to the subordinate relationship of the N-level annotated data, while displaying the multiple annotated data, the initial video and / or multimedia information generated by the annotated data is displayed in a tree-like format. For example, the initial video V_1 is displayed as the root video. The multimedia information corresponding to the first topic (topic1) and the second topic (topic2) is displayed as a sub-video of the initial video V_1. The multimedia information corresponding to the second-level sub-topic (topic3) and the second-level sub-topic (topic4) is displayed as a sub-video of one or more multimedia information in the multimedia information of the second topic (topic2). Similarly, the multimedia information corresponding to the second-level sub-topic (topic3) and / or the second-level sub-topic (topic4) can be displayed as the parent video of the multimedia information corresponding to one or more second-level sub-topics, wherein, optionally, during the process of displaying the video, the annotated data can be displayed together with the video. By analogy, during the continuous interaction of multimedia information, the initial video and / or multimedia information corresponding to the N-level annotated data is displayed in a tree-like format.
[0123] Figure 4 1 is a block diagram illustrating an apparatus 1200 for performing an audio and video interaction method according to an exemplary embodiment. For example, apparatus 1200 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0124] Reference Figure 4, the device 1200 may include one or more of the following components: a processing component 1202 , a memory 1204 , a power component 1206 , a multimedia component 1208 , an audio component 1210 , an input / output (I / O) interface 1212 , a sensor component 1214 , and a communication component 1216 .
[0125] The processing component 1202 generally controls the overall operation of the device 1200, such as operations associated with display, phone calls, data communications, camera operation, and recording operations. The processing component 1202 may include one or more processors 1220 to execute instructions to perform all or part of the steps of the above-described method. In addition, the processing component 1202 may include one or more modules to facilitate interaction between the processing component 1202 and other components. For example, the processing component 1202 may include a multimedia module to facilitate interaction between the multimedia component 1208 and the processing component 1202.
[0126] The memory 1204 is configured to store various types of data to support the operations of the device 1200. Examples of such data include instructions for any application or method operating on the device 1200, contact data, phone book data, messages, pictures, videos, etc. The memory 1204 can be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0127] The power supply component 1206 provides power to the various components of the device 1200. The power supply component 1206 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device 1200.
[0128] The multimedia component 1208 includes a screen that provides an output interface between the device 1200 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensor can not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 1208 includes a front camera and / or a rear camera. When the device 1200 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each front camera and rear camera can be a fixed optical lens system or have focal length and optical zoom capabilities.
[0129] The audio component 1210 is configured to output and / or input audio signals. For example, the audio component 1210 includes a microphone (MIC) that is configured to receive external audio signals when the device 1200 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals may be further stored in the memory 1204 or transmitted via the communication component 1216. In some embodiments, the audio component 1210 further includes a speaker for outputting audio signals.
[0130] I / O interface 1212 provides an interface between processing component 1202 and peripheral interface modules, such as a keyboard, click wheel, buttons, etc. These buttons may include but are not limited to: a home button, volume buttons, a start button, and a lock button.
[0131] The sensor assembly 1214 includes one or more sensors for providing various aspects of the status assessment of the device 1200. For example, the sensor assembly 1214 can detect the open / closed state of the device 1200, the relative positioning of components, such as the display and keypad of the device 1200. The sensor assembly 1214 can also detect changes in the position of the device 1200 or a component of the device 1200, the presence or absence of user contact with the device 1200, the orientation or acceleration / deceleration of the device 1200, and changes in the temperature of the device 1200. The sensor assembly 1214 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 1214 can also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 1214 can also include an accelerometer, a gyroscope, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0132] The communication component 1216 is configured to facilitate wired or wireless communication between the device 1200 and other devices. The device 1200 can access a wireless network based on a communication standard, such as WiFi, an operator network (such as 2G, 3G, 4G or 5G), or a combination thereof. In an exemplary embodiment, the communication component 1216 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 1216 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.
[0133] In an exemplary embodiment, the apparatus 1200 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-described methods.
[0134] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 1204 including instructions, which can be executed by the processor 1220 of the apparatus 1200 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0135] In an exemplary embodiment, a computer program product is also provided, including a computer program product, wherein the computer program includes program instructions, and when the program instructions are executed by a mobile terminal, the mobile terminal executes the steps of the above-mentioned multimedia information interaction method: obtaining an initial video; generating multiple annotation data corresponding to the initial video based on the characteristics of the initial video, wherein the annotation data is used to annotate the derived features of the initial video, and at least two of the multiple annotation data are used to annotate the same derived feature; displaying the multiple annotation data of the initial video through the interface of the application on the terminal, wherein the interface is also used to upload multimedia information associated with the annotation data.
[0136] Figure 5 1 is a block diagram of a device 1300 for executing an audio and video interactive method according to an exemplary embodiment. For example, the device 1300 can be provided as a server. Figure 5The apparatus 1300 includes a processing component 1322, which further includes one or more processors, and memory resources represented by a memory 1332 for storing instructions executable by the processing component 1322, such as an application. The application stored in the memory 1332 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 1322 is configured to execute the instructions to perform the above-described information list display method.
[0137] The device 1300 may also include a power supply component 1326 configured to perform power management of the device 1300, a wired or wireless network interface 1350 configured to connect the device 1300 to a network, and an input / output (I / O) interface 1358. The device 1300 may operate based on an operating system stored in the memory 1332, such as Windows Server™, MacOS X™, Unix™, Linux™, FreeBSD™, or the like.
[0138] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, and the true scope and spirit of the present application are indicated by the following claims.
[0139] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.
Claims
1. A multimedia information interaction method, characterized in that: include: Get the initial video; A step of generating annotated data: generating a plurality of annotated data corresponding to the initial video based on the features of the initial video, wherein the annotated data is used to annotate derived features of the initial video, and at least two of the plurality of annotated data are used to annotate the same derived feature, wherein a derived feature refers to a new feature obtained by performing feature learning on the original data of the initial video; The multiple annotation data of the initial video are displayed through the interface of the application on the terminal, wherein the interface is also used to upload multimedia information associated with the annotation data, and the multimedia information serves as the initial video in the next annotation data generation step. The multimedia information is used to express the user's views on the information in the annotation data corresponding to the initial video in the current annotation data generation step. After executing multiple annotation data generation steps, multiple annotation data are generated, and the multiple annotation data include N levels of annotation data, wherein the Nth level annotation data is subordinate to the (N-1)th level annotation data, and the Nth level annotation data is generated based on the initial video and / or the multimedia information, and N is a natural number.
2. The interactive method according to claim 1, characterized in that The step of generating a plurality of annotation data corresponding to the initial video according to the features of the initial video includes: Obtaining derived features of the initial video based on a content understanding result of the initial video; Obtaining annotation data corresponding to each of the derived features.
3. The interactive method according to claim 2, characterized in that: The obtaining of the annotation data corresponding to each of the derived features includes: Matching the annotation data corresponding to the derived features from the database; or, Keywords in the derived features are extracted to generate labeled data containing the keywords.
4. The interactive method according to claim 2, characterized in that: The obtaining of the derived features of the initial video based on the content understanding result of the initial video includes: Acquire description information of the initial video from a result of understanding the content of the initial video, wherein the description information is used to describe the content of the initial video; Feature learning is performed on the description information to obtain the derived features.
5. The interactive method according to any one of claims 1 to 4, characterized in that: At least two annotation data used to annotate the same derived feature contain keywords with the same semantics, and the semantics of the at least two annotation data are different.
6. The interactive method according to claim 1, characterized in that: The displaying of the plurality of annotated data of the initial video through an interface of an application on a terminal includes at least one of the following: During the playback of the initial video, displaying part or all of the multiple annotation data on the interface; After the initial video is played, part or all of the multiple annotation data are displayed on the interface; Using part or all of the multiple annotation data as labels of the initial video, and displaying them on the interface; Part or all of the multiple annotation data are used as the cover of the initial video and displayed on the interface.
7. The interactive method according to claim 1, characterized in that: After displaying the plurality of annotated data of the initial video through the interface of the application on the terminal, the interaction method further includes: Uploaded multimedia information associated with the annotation data is received, wherein the multimedia information is generated based on part or all of the information in the plurality of annotation data, and all or part of the uploaded multimedia information is set as the initial video.
8. The interactive method according to claim 1, characterized in that: The displaying of the plurality of annotated data of the initial video through an interface of an application on a terminal includes: Displaying the plurality of annotated data according to the subordinate relationship of the N-level annotated data; Wherein, while displaying the plurality of annotated data, the initial video and / or multimedia information for generating the annotated data is displayed.
9. A multimedia information interactive device, characterized in that: include: an initial video acquiring unit, configured to acquire an initial video; a labeling data generating unit configured to perform a labeling data generating step: generating a plurality of labeling data corresponding to the initial video based on the features of the initial video, wherein the labeling data is used to label derived features of the initial video, and at least two of the plurality of labeling data are used to label the same derived feature, wherein the derived feature refers to a new feature obtained by feature learning the original data of the initial video; An interactive unit is configured to display the multiple annotation data of the initial video through an interface of an application on a terminal, wherein the interface is also used to upload multimedia information associated with the annotation data, and the multimedia information serves as the initial video in the next annotation data generation step. The multimedia information is used to express the user's views on the information in the annotation data corresponding to the initial video in the current annotation data generation step. After executing multiple annotation data generation steps, multiple annotation data are generated, and the multiple annotation data include N levels of annotation data, wherein the Nth level annotation data is subordinate to the (N-1)th level annotation data, and the Nth level annotation data is generated based on the initial video and / or the multimedia information, and N is a natural number.
10. The interactive device according to claim 9, characterized in that The step of generating a plurality of annotation data corresponding to the initial video according to the features of the initial video includes: Obtaining derived features of the initial video based on a content understanding result of the initial video; Obtaining annotation data corresponding to each of the derived features.
11. The interactive device according to claim 10, characterized in that: The obtaining of the annotation data corresponding to each of the derived features includes: Matching the annotation data corresponding to the derived features from the database; or, Keywords in the derived features are extracted to generate labeled data containing the keywords.
12. The interactive device according to claim 10, characterized in that The obtaining of the derived features of the initial video based on the content understanding result of the initial video includes: Acquire description information of the initial video from a result of understanding the content of the initial video, wherein the description information is used to describe the content of the initial video; Feature learning is performed on the description information to obtain the derived features.
13. The interactive device according to any one of claims 9 to 12, characterized in that: At least two annotation data used to annotate the same derived feature contain keywords with the same semantics, and the semantics of the at least two annotation data are different.
14. The interactive device according to claim 9, characterized in that The displaying of the plurality of annotated data of the initial video through an interface of an application on a terminal includes at least one of the following: During the playback of the initial video, displaying part or all of the multiple annotation data on the interface; After the initial video is played, part or all of the multiple annotation data are displayed on the interface; Using part or all of the multiple annotation data as labels of the initial video, and displaying them on the interface; Part or all of the multiple annotation data are used as the cover of the initial video and displayed on the interface.
15. The interactive device according to claim 9, characterized in that After displaying the plurality of annotated data of the initial video through the interface of the application on the terminal, the interactive device further includes: Uploaded multimedia information associated with the annotation data is received, wherein the multimedia information is generated based on part or all of the information in the plurality of annotation data, and all or part of the uploaded multimedia information is set as the initial video.
16. The interactive device according to claim 9, characterized in that The displaying of the plurality of annotated data of the initial video through an interface of an application on a terminal includes: Displaying the plurality of annotated data according to the subordinate relationship of the N-level annotated data; Wherein, while displaying the plurality of annotated data, the initial video and / or multimedia information for generating the annotated data is displayed.
17. An interactive control device for multimedia information, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to execute the multimedia information interaction method according to any one of claims 1 to 8.
18. A non-temporary computer-readable storage medium, which, when the instructions in the storage medium are executed by a processor of a mobile terminal, enables the mobile terminal to execute an audio and video interaction method, the method including the multimedia information interaction method described in any one of claims 1 to 8.
Citation Information
Patent Citations
Method, system and device for acquiring review information during watching programs
CN102780921A
Video interaction method and device, and readable medium
CN108769814A
Bullet screen displaying method, playing device and control terminal
CN108933964A
Multi-modal collaborative web-based video annotation system
US20130145269A1
Video interaction method and device
CN108174247A