Multimedia resource marking method, multimedia resource display method and multimedia resource marking device

By automatically adding voice tags by inputting voice clips in the multimedia resource editing interface, the tedious problem of users manually inputting text is solved, the efficiency of tagging and publishing is improved, and the efficiency of human-computer interaction is enhanced.

CN120751205APending Publication Date: 2025-10-03BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510808265.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

When marking multimedia resources, users need to manually input text, which is cumbersome and affects marking efficiency.

Method used

This paper provides a multimedia resource tagging method. By inputting a voice clip in the resource editing interface, voice tags are automatically added, and the voice clip playback or recognized text display is triggered in the display interface, simplifying the tagging process.

Benefits of technology

It improves the efficiency of marking and publishing multimedia resources, enhances the efficiency of human-computer interaction, and simplifies the operation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120751205A_ABST
    Figure CN120751205A_ABST
Patent Text Reader

Abstract

The invention provides a multimedia resource marking method and device and a multimedia resource display method and device, and belongs to the technical field of multimedia. The method comprises the following steps: displaying multimedia resources in a resource editing interface; in response to a voice marking operation on the multimedia resource, based on an input voice segment, adding a voice mark in the multimedia, the voice mark being used for playing the voice segment after triggering; and in response to a resource publishing operation, publishing the multimedia resource containing the voice mark. According to the technical scheme, the user does not need to manually input the text, compared with text marking, the voice marking operation is simpler, the marking efficiency of the multimedia resources can be improved, the publishing efficiency of the multimedia resources can be improved, and the man-machine interaction efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of multimedia technology, and in particular to a method for marking multimedia resources, and a method and device for displaying multimedia resources. Background Art

[0002] With the development of multimedia technology, more and more users are accustomed to posting multimedia resources such as captured images or videos to social platforms to interact with other users.

[0003] Before publishing a multimedia resource, users can add text to mark it, expressing their own opinions or highlighting key points in the resource. After the multimedia resource is published, other users can see the marked text when viewing the multimedia resource, allowing them to understand the publisher's opinions or key points.

[0004] In the above technical solution, when marking multimedia resources, users need to manually input text, which is cumbersome and affects the marking efficiency of multimedia resources. Summary of the Invention

[0005] The present disclosure provides a multimedia resource marking method, a multimedia resource display method and a device, which can improve the efficiency of multimedia resource marking, multimedia resource publishing efficiency and human-computer interaction efficiency. The technical solution of the present disclosure is as follows:

[0006] According to one aspect of an embodiment of the present disclosure, a method for marking multimedia resources is provided, the method comprising:

[0007] In the resource editing interface, multimedia resources are displayed;

[0008] In response to a voice tagging operation on the multimedia resource, a voice tag is added to the multimedia based on the input voice segment, wherein the voice tag is used to trigger playback of the voice segment or display of text recognized based on the voice segment;

[0009] In response to a resource publishing operation, the multimedia resource including the voice mark is published.

[0010] According to another aspect of an embodiment of the present disclosure, a method for displaying multimedia resources is provided, the method further comprising:

[0011] In the resource display interface, multimedia resources and at least one voice tag of the multimedia resources are displayed, each voice tag is used to play the corresponding voice segment or display the text obtained by recognizing the voice segment after being triggered;

[0012] For any voice mark among the at least one voice mark, in response to a triggering operation on the voice mark, a voice segment corresponding to the voice mark is played or a text obtained by recognizing the voice segment is displayed.

[0013] According to another aspect of an embodiment of the present disclosure, a device for marking multimedia resources is provided, the device comprising:

[0014] A display unit is configured to execute in the resource editing interface and display the multimedia resources;

[0015] The display unit is further configured to execute a voice tagging operation on the multimedia resource, and add a voice tag to the multimedia based on the input voice segment, wherein the voice tag is used to trigger the playback of the voice segment or the display of text obtained by recognizing the voice segment;

[0016] The publishing unit is configured to publish the multimedia resource containing the voice mark in response to a resource publishing operation.

[0017] In some embodiments, the display unit is configured to display a voice input control in response to a mark-adding operation on any position in the multimedia resource, wherein the voice input control is used to obtain a voice segment;

[0018] In response to a triggering operation on the voice input control, the voice mark is displayed at the position based on the input voice segment.

[0019] In some embodiments, the display unit is further configured to perform:

[0020] In response to a mark adding operation on any position in the multimedia resource, displaying an information input box at the position, the information input box being used to obtain a mark for the multimedia resource;

[0021] In response to a triggering operation on the voice input control, displaying the voice mark in the information input box based on the input voice segment;

[0022] In response to a mark confirmation operation in the information input box, the information input box is closed and the voice mark is displayed at the position.

[0023] In some embodiments, the display unit is further configured to perform:

[0024] In response to a text input operation on the information input box, displaying the input first text in the information input box;

[0025] In response to a mark confirmation operation in the information input box, the information input box is closed, and the first text and the voice mark are displayed at the position.

[0026] In some embodiments, the display unit is further configured to execute a mark confirmation operation in response to the information input box, and display a second text at the position, where the second text is obtained based on the recognition of the input voice segment.

[0027] In some embodiments, the display unit is further configured to perform:

[0028] In the resource editing interface, multiple tone options are displayed;

[0029] When any one of the plurality of timbre options is selected, playing the input voice segment based on the target timbre indicated by the audio option;

[0030] In response to a mark confirmation operation in the information input box, the timbre of the voice segment is adjusted to the target timbre.

[0031] In some embodiments, the display unit is further configured to perform:

[0032] In response to an object marking operation on the multimedia resource, displaying an object identifier of a target object indicated by the object marking operation in the information input box;

[0033] In response to a mark confirmation operation in the information input box, an object identifier of the target object is displayed at the position.

[0034] In some embodiments, the display unit is further configured to execute: in response to a resource publishing operation, display the multimedia resource including the voice mark in a conversation interface between the operating object and the target object.

[0035] In some embodiments, the display unit is further configured to: when the target object comments on the multimedia resource, display the comment message of the target object at the top of the comment interface of the multimedia resource.

[0036] According to another aspect of an embodiment of the present disclosure, a device for displaying multimedia resources is provided, the device further comprising:

[0037] A display unit is configured to execute in a resource display interface, displaying a multimedia resource and at least one voice tag of the multimedia resource, each voice tag being used to play a corresponding voice segment or display text obtained by recognizing the voice segment after being triggered;

[0038] The playback unit is configured to execute, for any voice mark among the at least one voice mark, in response to a triggering operation on the voice mark, play the voice segment corresponding to the voice mark or display text obtained by recognizing the voice segment.

[0039] In some embodiments, the playback unit is configured to perform any of the following:

[0040] For any voice mark among the at least one voice mark, in response to a triggering operation on the voice mark, pausing the playing of the multimedia resource and playing the voice segment corresponding to the voice mark;

[0041] In response to a triggering operation on the voice mark, continue playing the multimedia resource at a first volume and play the voice segment corresponding to the voice mark at a second volume, where the first volume is lower than the second volume;

[0042] In response to a triggering operation on the voice mark, the volume of the multimedia resource is lowered, and the voice segment corresponding to the voice mark is played.

[0043] In some embodiments, the multimedia resource has multiple voice tags;

[0044] The playing unit is further configured to, after playing a voice segment corresponding to a voice mark, continue to play a voice segment corresponding to a next voice mark in the order of the multiple voice marks.

[0045] In some embodiments, the display unit is configured to perform:

[0046] Playing the multimedia resource in the resource display interface;

[0047] In response to a pause operation on the multimedia resource, at least one voice mark of the multimedia resource is displayed.

[0048] In some embodiments, the display unit is configured to perform:

[0049] Playing the multimedia resource in the resource display interface;

[0050] In response to the multimedia resource being played to a target time, at least one voice mark of the multimedia resource is displayed, where the target time is the playing time when the at least one voice mark is added to the multimedia resource.

[0051] In some embodiments, the display unit is configured to perform:

[0052] In the resource display interface, display the multimedia resource and a progress bar of the multimedia resource;

[0053] At least one voice mark of the multimedia resource is displayed on the progress bar of the multimedia resource.

[0054] In some embodiments, the playback unit is further configured to execute, in the case of pausing the playback of the voice segment corresponding to any voice mark, in response to the re-triggering operation of the voice mark, continuing to play the voice segment corresponding to the voice mark according to the playback progress achieved last time.

[0055] According to another aspect of an embodiment of the present disclosure, an electronic device is provided, the electronic device including:

[0056] one or more processors;

[0057] a memory for storing program codes executable by the processor;

[0058] The processor is configured to execute the program code to implement the above-mentioned multimedia resource marking method or multimedia resource display method.

[0059] According to another aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided. When the program code in the computer-readable storage medium is executed by a processor of an electronic device, the electronic device can execute the above-mentioned multimedia resource marking method or multimedia resource display method.

[0060] According to another aspect of an embodiment of the present disclosure, a computer program product is provided, including a computer program / instruction, which implements the above-mentioned multimedia resource marking method or multimedia resource display method when executed by a processor.

[0061] The disclosed embodiments provide a novel method for marking multimedia resources. When marking multimedia resources, a voice clip can be directly input in the resource editing interface to voice-mark the multimedia resources without the need for the user to manually input text. Compared with text marking, voice marking is simpler and can improve the marking efficiency of multimedia resources, thereby helping to improve the publishing efficiency of multimedia resources, that is, it can improve the efficiency of human-computer interaction.

[0062] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] The accompanying drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.

[0064] Figure 1 The figure is a schematic diagram of an implementation environment according to an exemplary embodiment.

[0065] Figure 2 The figure is a flowchart of a method for marking multimedia resources according to an exemplary embodiment.

[0066] Figure 3 The figure is a flowchart of another multimedia resource marking method according to an exemplary embodiment.

[0067] Figure 4 The figure is a schematic diagram showing a voice input control according to an exemplary embodiment.

[0068] Figure 5 The figure is a schematic diagram showing a method of adding a speech mark according to an exemplary embodiment.

[0069] Figure 6 The figure is a schematic diagram showing a method of adding a text tag and a voice tag according to an exemplary embodiment.

[0070] Figure 7 is a schematic diagram showing a timbre option according to an exemplary embodiment.

[0071] Figure 8 The figure is a schematic diagram showing a marking object according to an exemplary embodiment.

[0072] Figure 9 The figure is a schematic diagram showing a method of publishing multimedia resources according to an exemplary embodiment.

[0073] Figure 10 The figure is a schematic diagram showing a comment message according to an exemplary embodiment.

[0074] Figure 11 The figure is a flowchart of a method for displaying multimedia resources according to an exemplary embodiment.

[0075] Figure 12 The figure is a schematic diagram showing a method of playing a voice mark according to an exemplary embodiment.

[0076] Figure 13 The figure is a block diagram of a multimedia resource marking device according to an exemplary embodiment.

[0077] Figure 14 The figure is a block diagram of a device for displaying multimedia resources according to an exemplary embodiment.

[0078] Figure 15 It is a block diagram of a terminal according to an exemplary embodiment. DETAILED DESCRIPTION

[0079] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0080] It should be noted that the terms "first," "second," and the like in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the numbers used in this manner are interchangeable where appropriate so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure as detailed in the appended claims.

[0081] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, storage, and display, etc.), and signals involved in this disclosure are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the multimedia resources and voice tags (voice clips) involved in this disclosure are obtained with full authorization.

[0082] Figure 1 FIG. 1 is a schematic diagram of an implementation environment according to an exemplary embodiment. Taking the electronic device as a terminal as an example, see Figure 1 The implementation environment specifically includes: a first terminal 101, a server 102 and a second terminal 103.

[0083] The first terminal 101 is at least one of a smart phone, a smart watch, a desktop computer, a laptop computer, an MP3 player, an MP4 player, and a laptop computer. An application is installed and running on the first terminal 101. The application can be a multimedia application, a social application, a game application, a browser, etc., which is not limited in the embodiment of the present disclosure. The operating object (user) can log in to the application through the first terminal 101 to obtain the services provided by the application. Before publishing the multimedia resources, the operating object can add voice tags to the multimedia resources through the first terminal 101, and then publish the multimedia resources containing the voice tags. Among them, the first terminal 101 can be connected to the server 102 via a wireless network or a wired network. The server 102 is a platform server for the above-mentioned application and can store the multimedia resources after the multimedia resources are published. Then, other terminals (such as the second terminal) can log in to the above-mentioned application to obtain the services provided by the application, that is, display the multimedia resources stored in the server 102.

[0084] The first terminal 101 generally refers to one of multiple terminals. This embodiment uses the first terminal 101 as an example. Those skilled in the art will appreciate that the number of terminals may be greater or lesser. For example, there may be a few terminals, or dozens, hundreds, or even more terminals. This embodiment does not limit the number or device type of terminals.

[0085] The server 102 is at least one of a single server, multiple servers, a cloud computing platform, and a virtualization center. The server 102 can be connected to the first terminal 101 and other terminals via a wireless network or a wired network. The server 102 can store multimedia resources published by the first terminal 101 and can push the multimedia resources to other terminals (such as the second terminal). In some embodiments, the number of the above-mentioned servers can be more or less, and the embodiments of the present disclosure are not limited to this. Of course, the server 102 also includes other functional servers to provide more comprehensive and diversified services.

[0086] The second terminal 103 is at least one of a smart phone, a smart watch, a desktop computer, a laptop computer, an MP3 player, an MP4 player, and a portable laptop computer. An application is installed and running on the second terminal 101. The application can be a multimedia application, a social application, a game application, a browser, etc., which is not limited in the embodiment of the present disclosure. The operating object (user) can log in to the above application through the second terminal 101 to obtain the services provided by the application. The second terminal 103 can be connected to the server 102 through a wireless network or a wired network, display the multimedia resources containing voice tags pushed by the server 102, display the multimedia resources containing voice tags to the operating object, and play the voice clip corresponding to the voice tag, so that the operating object can know the tag added by the publisher of the multimedia resource.

[0087] Figure 2 is a flowchart of a method for marking multimedia resources according to an exemplary embodiment. Figure 2 The multimedia resource marking method is applied to the first terminal, comprising the following steps:

[0088] In step 201, the first terminal displays multimedia resources in a resource editing interface.

[0089] In the embodiments of the present disclosure, multimedia resources may be videos, images, and text, etc., which are not limited in the embodiments of the present disclosure. The first terminal displays a resource editing interface. In response to a resource upload operation in the resource editing interface, the first terminal displays the multimedia resource in the resource editing interface so that various editing operations can be performed on the multimedia resource later. The editing operation may be adding a tag (voice tag or text tag), adding a filter, adding a sticker, adding a special effect, etc. to the multimedia resource, which are not limited in the embodiments of the present disclosure.

[0090] In step 202, in response to a voice tagging operation on a multimedia resource, the first terminal adds a voice tag to the multimedia based on the input voice segment, and the voice tag is used to trigger the playback of the voice segment or the display of text recognized based on the voice segment.

[0091] In an embodiment of the present disclosure, when a voice tag operation is triggered on a multimedia resource, the first terminal obtains a voice segment input by the operating object (user), generates a voice tag based on the input voice segment, and displays the voice tag in the multimedia resource. The voice tag is used to indicate the input voice segment, that is, to indicate the specific content of the tag added to the multimedia resource. The voice segment is equivalent to additional information of the multimedia resource, and can be used to express the views of the publisher of the multimedia resource, the interpretation of the multimedia resource, key prompts or supplementary explanations, etc., which is not limited by the embodiment of the present disclosure.

[0092] The first terminal can add any number of voice tags to a multimedia resource. Each voice tag corresponds to a voice segment. For any voice tag, in response to a triggering operation on the voice tag, the first terminal can play the voice segment corresponding to the voice tag or display the text recognized by the voice segment. In other words, by triggering a voice tag, the user can preview or view the specific content or detailed information of the voice tag.

[0093] In step 203, in response to the resource publishing operation, the first terminal publishes the multimedia resource including the voice tag.

[0094] In an embodiment of the present disclosure, when a resource publishing operation is triggered, a first terminal publishes a multimedia resource containing a voice tag. After the multimedia resource is published, other terminals (such as a second terminal) can obtain and display the multimedia resource. While displaying the multimedia resource, the other terminals can also display at least one voice tag added to the multimedia resource and learn the specific content of the voice tag by playing the voice clip corresponding to the voice tag.

[0095] The disclosed embodiments provide a novel method for marking multimedia resources. When marking multimedia resources, a voice clip can be directly input in the resource editing interface to voice-mark the multimedia resources without the need for the user to manually input text. Compared with text marking, voice marking is simpler and can improve the marking efficiency of multimedia resources, thereby helping to improve the publishing efficiency of multimedia resources, that is, it can improve the efficiency of human-computer interaction.

[0096] In some embodiments, in response to a voice tagging operation on a multimedia resource, adding a voice tag to the multimedia based on an input voice segment includes:

[0097] In response to a mark-adding operation at any position in the multimedia resource, displaying a voice input control, the voice input control being used to obtain a voice segment;

[0098] In response to a triggering operation on the voice input control, a voice mark is displayed at a location based on the input voice segment.

[0099] The solution provided by the embodiment of the present disclosure can add voice tags at any location in multimedia resources through a voice input control with one click. It is not only simple to operate and can improve the tagging efficiency of multimedia resources, but also the operating object can add tags at any location in the multimedia resources according to its own needs, which meets the requirements of object tagging, has good flexibility and strong operability.

[0100] In some embodiments, the method further comprises:

[0101] In response to a mark adding operation on any location in the multimedia resource, an information input box is displayed at the location, the information input box being used to obtain the mark for the multimedia resource;

[0102] In response to a triggering operation on a voice input control, a voice mark is displayed at a location based on the input voice segment, including:

[0103] In response to a triggering operation on the voice input control, displaying a voice mark in the information input box based on the input voice segment;

[0104] In response to a mark confirmation operation in the information input box, the information input box is closed and a voice mark is displayed at the location.

[0105] The solution provided by the embodiment of the present disclosure is that, during the process of adding a voice mark at any location in a multimedia resource, a voice clip is input through an information input box and the voice mark is previewed and displayed. The voice mark will only be displayed at that location after the operation object confirms it. This method makes it easy for the operation object to promptly discover the deficiencies of the voice mark and make adjustments before confirmation, which is conducive to improving the efficiency and accuracy of the mark.

[0106] In some embodiments, the method further comprises:

[0107] In response to a text input operation on the information input box, displaying the input first text in the information input box;

[0108] In response to a mark confirmation operation in the information input box, the information input box is closed and a voice mark is displayed at the location, including:

[0109] In response to a mark confirmation operation in the information input box, the information input box is closed, and the first text and voice mark are displayed at the location.

[0110] The solution provided by the embodiments of the present disclosure can simultaneously perform voice marking and text marking when marking any location in a multimedia resource, thereby achieving marking from both visual and auditory perspectives, enriching the marking content, and facilitating improved marking accuracy, and accurately conveying the marking intent of the object.

[0111] In some embodiments, the method further comprises:

[0112] In response to a mark confirmation operation in the information input box, a second text is displayed at the location, where the second text is obtained based on recognition of the input voice segment.

[0113] The solution provided by the embodiments of the present disclosure can automatically identify the voice segments corresponding to the voice tags in multimedia resources and display the text corresponding to the voice segments, thereby realizing tagging from both visual and auditory perspectives, enriching the tagging content, and helping to improve the accuracy of tagging, and can accurately convey the tagging intention of the object; and without the need for users to manually enter text, it can improve tagging efficiency.

[0114] In some embodiments, the method further comprises:

[0115] In the resource editing interface, multiple tone options are displayed;

[0116] When any one of the plurality of timbre options is selected, playing the input voice clip based on the target timbre indicated by the audio option;

[0117] In response to a mark confirmation operation in the information input box, the timbre of the speech segment is adjusted to a target timbre.

[0118] The solution provided by the embodiment of the present disclosure provides the operating object with multiple timbre options for selection. The operating object can adjust the timbre of the voice segment input by the operating object to the timbre specified by the operating object according to its own needs. This not only meets the marking intention of the operating object, but also is simple to operate and can improve operational efficiency.

[0119] In some embodiments, the method further comprises:

[0120] In response to an object marking operation on a multimedia resource, displaying an object identifier of a target object indicated by the object marking operation in an information input box;

[0121] In response to a mark confirmation operation in the information input box, an object identification of the target object is displayed at the location.

[0122] The solution provided by the embodiment of the present disclosure can also add the object identifier of the target object in the tag when voice tagging multimedia resources, so as to remind the target object to view the multimedia resources and tags. It not only provides a novel tag-based interaction method, but also can enhance the interaction rate between objects through tagging, thereby helping to increase the interaction volume of multimedia resources.

[0123] In some embodiments, the method further comprises:

[0124] In response to the resource publishing operation, the multimedia resource including the voice mark is displayed in the conversation interface between the operating object and the target object.

[0125] The solution provided by the embodiment of the present disclosure is that, when there is an object identifier of the target object in the tag, the multimedia resource containing the voice tag can be automatically displayed in the conversation interface between the operating object and the target object after the resource is released. That is, the multimedia resource is directly sent to the target object through the private message channel between the objects to remind the target object to watch the multimedia resource, which is conducive to improving the interaction rate between the objects, and thus helps to increase the interaction volume of the multimedia resource.

[0126] In some embodiments, the method further comprises:

[0127] In the case where the target object comments on the multimedia resource, the comment message of the target object is displayed at the top of the comment interface of the multimedia resource.

[0128] The solution provided by the embodiment of the present disclosure can prioritize displaying the target object's comments on multimedia resources when the target object comments on the multimedia resources, making it easier for the operator (publisher) of the multimedia resources to discover them in a timely manner, thereby improving information transmission efficiency and interaction efficiency.

[0129] above Figure 2 The following is only a basic process of the present disclosure. The solution provided by the present disclosure is further described based on a specific implementation method. Figure 3 FIG. 1 is a flow chart of another method for marking multimedia resources according to an exemplary embodiment. Taking the electronic device as the first terminal as an example, see Figure 3 , the method comprising:

[0130] In step 301, the first terminal displays multimedia resources in a resource editing interface.

[0131] In the disclosed embodiments, the first terminal is the publisher of a multimedia resource. Before publishing the multimedia resource, the first terminal displays the multimedia resource in a resource editing interface to facilitate editing operations. The disclosed embodiments do not limit these editing operations. The resource editing interface includes a tag control, which allows the user to add voice tags to the multimedia resource. The disclosed embodiments do not limit the display style or location of the tag control.

[0132] In step 302, in response to a mark adding operation on any position in the multimedia resource, the first terminal displays a voice input control, where the voice input control is used to obtain a voice segment.

[0133] In an embodiment of the present disclosure, for any location in a multimedia resource, in response to a triggering operation of a mark control in the resource editing interface, the first terminal displays a voice input control, allowing the user to input the voice segment to be marked through the voice input control. The embodiment of the present disclosure does not limit the display location and display style of the voice input control.

[0134] In some embodiments, the voice input control is a first voice input control. The first voice input control is used to directly collect voice segments of the operation object. Accordingly, in response to the first voice input control, the first terminal collects the voice segments of the operation object from the surrounding environment.

[0135] In other embodiments, the voice input control is a second voice input control. The second voice input control is used to obtain a voice clip from a local storage space. The voice clip can be a voice clip directly stored locally by the first terminal. Alternatively, the voice clip can also be audio from a multimedia resource stored locally by the first terminal. That is, when the voice clip is input through the second voice input control, the first terminal can extract audio from other multimedia resources and use the audio as a voice clip for marking, etc. The embodiments of the present disclosure do not limit the method for obtaining the voice clip.

[0136] The first voice input control and the second voice input control can be displayed simultaneously in the resource editing interface for selection by the operation object.

[0137] For example, Figure 4 FIG is a schematic diagram showing a voice input control according to an exemplary embodiment. Figure 4, the resource editing interface displays multimedia resources and a markup control 401. When markup control 401 is triggered, the first terminal displays a first voice input control 402 and a second voice input control 403 in the resource editing interface. The user can directly record a voice clip by clicking on the first voice input control 402. Alternatively, the user can click on the second voice input control 403 to retrieve the voice clip from local storage.

[0138] In step 303, in response to the triggering operation on the voice input control, the first terminal displays a voice mark at the position based on the input voice segment.

[0139] In an embodiment of the present disclosure, when the voice input control in the resource editing interface is triggered, the first terminal generates a voice tag based on the input voice segment, and displays the voice tag at the location where the tag adding operation is triggered. The embodiment of the present disclosure does not limit the display style of the voice tag. The solution provided by the embodiment of the present disclosure can add a voice tag at any location in the multimedia resource through the voice input control with one click. Not only is the operation simple and can improve the tagging efficiency of the multimedia resource, but the operating object can also add a tag at any location in the multimedia resource according to its own needs, which meets the requirements of adding tags to the object, has good flexibility and strong operability.

[0140] In the process of adding a voice tag, in response to a tag adding operation at any position in the multimedia resource, the first terminal can display an information input box at the position, and the information input box is used to obtain a tag for the multimedia resource. Then, the process of the first terminal displaying the voice tag at the position includes: in response to a triggering operation on the voice input control, the first terminal displays the voice tag in the information input box based on the input voice segment; in response to a tag confirmation operation in the information input box, the first terminal closes the information input box and displays the voice tag at the position. Before the tag confirmation operation, in response to the triggering operation of the voice tag in the information input box, the first terminal can play the input voice segment, which is equivalent to a preview play. That is, before confirming the tag, the operating object can listen to whether the input voice segment is appropriate. Only when it is confirmed to be correct will the first terminal close the information input box and display the voice tag in the multimedia resource. The solution provided by the embodiment of the present disclosure is that, during the process of adding a voice mark at any location in a multimedia resource, a voice clip is input through an information input box and the voice mark is previewed and displayed. The voice mark will only be displayed at that location after the operation object confirms it. This method makes it easy for the operation object to promptly discover the deficiencies of the voice mark and make adjustments before confirmation, which is conducive to improving the efficiency and accuracy of the mark.

[0141] For example, Figure 5 FIG. 1 is a schematic diagram showing a method of adding a speech mark according to an exemplary embodiment. Figure 5 When the mark control 501 is triggered, the first terminal displays a first voice input control 502, a second voice input control 503, and an information input box 504 in the resource editing interface. When the first voice input control 502 or the second voice input control 503 is triggered, the first terminal displays a voice tag 505 in the information input box 504 based on the input voice segment. Then, in response to the triggering operation of the completion control 506, the first terminal closes the information input box 504 and adds a voice tag 505 to the multimedia resource.

[0142] In some embodiments, during the process of voice tagging a multimedia resource, the first terminal may also perform text tagging on the multimedia resource. Accordingly, in response to a text input operation on the information input box, the first terminal displays the first text entered in the information input box. Then, in response to a mark confirmation operation in the information input box, the first terminal closes the information input box and displays the first text and voice tag at the location. The first text is the text tag entered by the operation object. The first text may be a keyword, theme / subject, title, etc. in the voice segment corresponding to the voice tag, which is not limited in the embodiment of the present disclosure. The first text and voice tag may be displayed in the multimedia resource in the order in which they were added, or in the multimedia resource in the default order, which is not limited in the embodiment of the present disclosure. The solution provided by the embodiment of the present disclosure can simultaneously perform voice tagging and text tagging when marking any location in the multimedia resource, thereby achieving tagging from both visual and auditory perspectives, enriching the tag content, and facilitating improving the accuracy of the tag, and accurately conveying the tagging intention of the object.

[0143] For example, Figure 6 FIG2 is a schematic diagram showing a method of adding text tags and voice tags according to an exemplary embodiment. Figure 6 , when the mark control 601 is triggered, the first terminal displays the first voice input control 602, the second voice input control 603 and the information input box 604 in the resource editing interface. When the first voice input control 602 or the second voice input control 603 is triggered, the first terminal displays a voice tag 605 in the information input box 604 based on the input voice segment. In response to the text input operation on the information input box 604, the first terminal displays the first text "Listen well~" in the information input box. Then, in response to the triggering operation on the completion control 606, the first terminal closes the information input box 604 and adds the first text "Listen well~" and the voice tag 605 to the multimedia resource.

[0144] In other embodiments, in addition to manual input of the operation object, the text mark can also be obtained based on the recognition of the input voice segment. Accordingly, in response to the mark confirmation operation in the information input box, the first terminal displays the second text at this position. The second text is obtained based on the recognition of the input voice segment. The second text can include all the text in the voice segment, which is equivalent to the subtitle of the voice segment; or, the second text can also include only the keywords in the voice segment; or, the second text can also be obtained by summarizing the voice segment, for example, the second text is the main point or title of the voice segment, etc., which is not limited by the embodiment of the present disclosure. The solution provided by the embodiment of the present disclosure can automatically identify the voice segment corresponding to the voice mark in the multimedia resource, and display the text corresponding to the voice segment, thereby realizing marking from both visual and auditory perspectives, enriching the marking content, and helping to improve the accuracy of the marking, and can accurately convey the marking intention of the object; and there is no need for the user to manually enter text, which can improve marking efficiency.

[0145] With respect to the aforementioned information input box, the user can adjust at least one of the position and size of the information input box. In response to a move operation on the information input box, the first terminal displays the information input box moved to the position indicated by the move operation. In response to a size adjustment operation on the information input box, the first terminal displays the information input box at the size indicated by the size adjustment operation.

[0146] In some embodiments, for the input voice segment, the operating object can also adjust the timbre of the voice segment. Accordingly, the first terminal displays a plurality of timbre options in the resource editing interface. The timbres indicated by different timbre options are different. When any one of the multiple timbre options is selected, the first terminal plays the input voice segment based on the target timbre indicated by the audio option. Then, in response to the mark confirmation operation in the information input box, the first terminal adjusts the timbre of the voice segment to the target timbre. The solution provided by the embodiment of the present disclosure provides the operating object with a plurality of timbre options for selection. The operating object can adjust the timbre of the voice segment input by the operating object to the timbre specified by the operating object according to its own needs. This not only meets the marking intention of the operating object, but also is simple to operate and can improve operating efficiency.

[0147] For example, Figure 7 FIG is a schematic diagram showing a timbre option according to an exemplary embodiment. Figure 7 , the resource display interface displays three timbre options 701. Different timbre options indicate different timbres. When any of the three timbre options 701 is selected, the first terminal can adjust the timbre of the voice segment corresponding to the voice tag 702 to the timbre indicated by the selected timbre option 701.

[0148] In some embodiments, the first terminal is also able to mark any target object in the multimedia resource so that it can interact with the target object through marking. Accordingly, in response to the object marking operation on the multimedia resource, the first terminal displays the object identifier of the target object indicated by the object marking operation in the information input box. In response to the mark confirmation operation in the information input box, the first terminal displays the object identifier of the target object at the position. Among them, in response to the object marking operation on the multimedia resource, the first terminal displays an object list. The object list includes multiple objects. When the target object in the object list is selected, the first terminal displays the object identifier of the target object in the information input box. The solution provided by the embodiment of the present disclosure can also add the object identifier of the target object in the mark when voice marking the multimedia resource, so as to remind the target object to view the multimedia resource and mark, which not only provides a novel tag-based interaction method, but also can enhance the interaction rate between objects through marking, thereby helping to increase the interaction volume of multimedia resources.

[0149] For example, Figure 8 FIG. 1 is a schematic diagram showing a marking object according to an exemplary embodiment. Figure 8 In response to the input operation on the information input box 801, the first terminal displays the marking operation panel 802 in the resource editing interface. The marking operation panel 802 includes a "mark object" option and a "customize" option. When the "mark object" option is selected, the first terminal displays an object list in the marking operation panel 802, and the object list includes multiple objects. When any object in the object list is selected, the first terminal can display the object identifier (such as the object's nickname, avatar, etc.) of the selected object in the information input box 801. When the "customize" option is selected, the first terminal can display a virtual keyboard in the marking operation panel 802. The virtual keyboard supports the operating object to enter any text mark in the information input box 801 according to its own needs.

[0150] In step 304, in response to the resource publishing operation, the first terminal publishes the multimedia resource including the voice tag.

[0151] In an embodiment of the present disclosure, when a resource publishing operation is triggered, a first terminal publishes a multimedia resource containing a voice tag. The multimedia resource may also contain text tags or object tags, etc., which are not limited in the embodiment of the present disclosure. After the multimedia resource is published, other terminals (such as a second terminal) can display the multimedia resource while also displaying the various tags added to the multimedia resource.

[0152] In some embodiments, for multimedia resources tagged with objects, in response to a resource publishing operation, the first terminal can also display the multimedia resource containing the voice tag in the conversation interface between the operating object and the target object. The solution provided by the embodiment of the present disclosure, when the object identifier of the target object is present in the tag, can automatically display the multimedia resource containing the voice tag in the conversation interface between the operating object and the target object after the resource is published. That is, the multimedia resource is directly sent to the target object through a private message channel between the objects to remind the target object to watch the multimedia resource, which is conducive to improving the interaction rate between the objects, thereby helping to increase the interaction volume of the multimedia resource.

[0153] For example, Figure 9 FIG2 is a schematic diagram showing a method for publishing multimedia resources according to an exemplary embodiment. Figure 9 In the case of publishing a multimedia resource tagged with object 1, the first terminal displays the multimedia resource containing the voice tag in the conversation interface between the operating object and object 1 to remind object 1 to watch the multimedia resource.

[0154] In some embodiments, when a target subject comments on a multimedia resource, the first terminal displays the target subject's comment message at the top of the multimedia resource's comment interface. The solution provided by the disclosed embodiments prioritizes displaying the target subject's comment when the target subject comments on a multimedia resource, facilitating timely discovery by the multimedia resource's operator (publisher), thereby improving information transmission and interaction efficiency.

[0155] For example, Figure 10 FIG. 1 is a schematic diagram showing a comment message according to an exemplary embodiment. Figure 10 , the first terminal is at the first position in the comment interface 1001, and displays the comment message of object 1. The comment message of object 1 may include text messages and voice messages.

[0156] The disclosed embodiments provide a novel method for marking multimedia resources. When marking multimedia resources, a voice clip can be directly input in the resource editing interface to voice-mark the multimedia resources without the need for the user to manually input text. Compared with text marking, voice marking is simpler and can improve the marking efficiency of multimedia resources, thereby helping to improve the publishing efficiency of multimedia resources, that is, it can improve the efficiency of human-computer interaction.

[0157] The above description is only from the perspective of the publisher of multimedia resources. In order to more clearly describe the solution provided by the embodiment of the present disclosure, the following description is further from the perspective of the viewer of multimedia resources to reflect the process of interaction between the publisher and viewer of multimedia resources based on voice tags.

[0158] Figure 11 is a flowchart of a method for displaying multimedia resources according to an exemplary embodiment. Figure 11 The method for displaying multimedia resources is applied to the second terminal, comprising the following steps:

[0159] In step 1101, the second terminal displays multimedia resources and at least one voice tag of the multimedia resources in the resource display interface, where each voice tag is used to play a corresponding voice segment or display text recognized based on the voice segment after being triggered.

[0160] In the disclosed embodiment, the second terminal is a display terminal for multimedia resources. The second terminal displays the multimedia resources in a resource display interface. If the multimedia resources include voice tags, the second terminal can also display at least one voice tag included in the multimedia resources. The disclosed embodiment does not limit the number of voice tags.

[0161] The at least one voice tag of a multimedia resource can be displayed throughout the entire display process of the multimedia resource. That is, the at least one voice tag and the multimedia resource can be displayed and turned off simultaneously. Alternatively, the at least one voice tag of a multimedia resource can be displayed when the multimedia resource is paused. Alternatively, the at least one voice tag of a multimedia resource can be displayed at the corresponding playback time. The present embodiment does not limit the display timing of the voice tag.

[0162] In some embodiments, a voice tag can be displayed when a multimedia resource is paused. Accordingly, step 1101 includes: the second terminal plays the multimedia resource in the resource display interface. Then, in response to the pause playback operation on the multimedia resource, the second terminal displays at least one voice tag of the multimedia resource. The solution provided by the embodiment of the present disclosure displays at least one voice tag of the multimedia resource when the multimedia resource is paused, so that the user can directly play the voice segment corresponding to any voice tag later. Since the multimedia resource has been paused, it will not interfere with the playback of the voice segment; and when playing the multimedia resource, since the audio in the multimedia resource is likely to interfere with the playback of the voice tag, the user is generally unlikely to play the multimedia resource and the voice tag at the same time. This solution does not need to display the voice tag when the multimedia resource is playing, which meets the user's intention to a certain extent and saves operating consumption.

[0163] In other embodiments, the voice tag can be displayed at the corresponding playback time. Accordingly, step 1101 includes: the second terminal plays the multimedia resource in the resource display interface. In response to the multimedia resource being played to the target time, the second terminal displays at least one voice tag of the multimedia resource. The target time is the playback time when at least one voice tag is added to the multimedia resource. The solution provided by the embodiment of the present disclosure displays the corresponding voice tag according to the playback time when the voice tag is added, which facilitates the user to play the corresponding voice tag at the corresponding playback time, so that the intention of the publisher of the multimedia resource can be conveyed to the user more clearly and accurately, and the display effect of the multimedia resource can be improved.

[0164] When the target time is reached, the second terminal can automatically pause the multimedia resource and display at least one corresponding voice mark. Alternatively, the second terminal automatically pauses the multimedia resource and automatically plays the displayed voice mark.

[0165] The at least one voice mark may be displayed at the location where the voice mark is added, or may be displayed on a progress bar of a multimedia resource, etc. The embodiment of the present disclosure does not limit the display location of the voice mark.

[0166] In some embodiments, the voice mark is displayed on the progress bar of the multimedia resource. Accordingly, step 1101 includes: the second terminal displays the multimedia resource and the progress bar of the multimedia resource in the resource display interface. Then, the second terminal displays at least one voice mark of the multimedia resource on the progress bar of the multimedia resource. Among them, at least one voice mark can be displayed on the progress bar according to the adding time (i.e., the playing time when the voice mark is added). The second terminal can display all the voice marks in the multimedia resource at one time, or the second terminal can also display the corresponding voice mark at the playing time according to the playing progress of the multimedia resource, etc., which is not limited by the embodiment of the present disclosure. The solution provided by the embodiment of the present disclosure facilitates the user to intuitively view the various voice marks in the multimedia resource by displaying the voice marks added to the multimedia resource on the progress bar of the multimedia resource, thereby adjusting the playing progress of the multimedia resource through the progress bar, so that the user can directly jump to the playing time where the voice mark is located to view the multimedia resource, which is beneficial to improving the information transmission efficiency of the multimedia resource.

[0167] In step 1102, for any voice mark of the at least one voice mark, in response to a triggering operation on the voice mark, the second terminal plays a voice segment corresponding to the voice mark or displays text recognized based on the voice segment.

[0168] In an embodiment of the present disclosure, for any voice tag in a multimedia resource, when the voice tag is triggered, the second terminal plays the voice segment corresponding to the voice tag or displays the text obtained based on the language segment recognition to convey to the user the specific content marked by the publisher of the multimedia resource. Alternatively, the second terminal can also display the text obtained based on the language segment recognition on the basis of playing the voice segment corresponding to the voice tag. The text obtained based on the language segment recognition can be displayed all at once, or each character in the text can be gradually displayed according to the playback progress of the voice segment, etc., which is not limited by the embodiment of the present disclosure.

[0169] For example, Figure 12 FIG. 1 is a schematic diagram showing a method for playing a voice tag according to an exemplary embodiment. Figure 12 The multimedia resource includes a voice tag 1201 and a voice tag 1202. When the voice tag 1201 is triggered, the second terminal plays the voice segment corresponding to the voice tag 1201.

[0170] In the process of playing the voice clip corresponding to any voice mark, the second terminal can pause playing the multimedia resource; or play the voice clip at a high volume and play the multimedia resource at a low volume; or directly lower the volume of the multimedia resource, etc., which is not limited in the embodiment of the present disclosure.

[0171] In some embodiments, the second terminal can pause the playback of multimedia resources. Accordingly, for any voice mark in at least one voice mark, in response to the triggering operation of the voice mark, the second terminal pauses the playback of the multimedia resource and plays the voice segment corresponding to the voice mark. The solution provided by the embodiment of the present disclosure can directly pause the multimedia resource when playing the voice segment corresponding to any voice mark, which can not only avoid the interference of the multimedia resource and ensure that the user can clearly know the content of the voice segment, but also help improve the accuracy of information transmission; and there is no need for the user to manually pause the multimedia resource. Pausing the multimedia resource and playing the voice segment can be achieved with only one operation, which is simple to operate and helps improve the efficiency of human-computer interaction.

[0172] In other embodiments, the second terminal can play the voice clip at a high volume and play the multimedia resource at a low volume. Accordingly, for any voice mark in at least one voice mark, in response to the triggering operation of the voice mark, the second terminal continues to play the multimedia resource at the first volume and plays the voice clip corresponding to the voice mark at the second volume. The first volume is lower than the second volume. The solution provided by the embodiment of the present disclosure plays the voice clip at a high volume and plays the multimedia resource at a low volume, which can not only reduce the impact of the multimedia resource on the playback of the voice clip and ensure that the user can clearly understand the content of the voice clip, but also does not need to pause the multimedia resource, which is conducive to improving the playback efficiency of the multimedia resource.

[0173] In other embodiments, the second terminal can directly lower the volume of the multimedia resource. Accordingly, for any voice tag in at least one voice tag, in response to the triggering operation of the voice tag, the second terminal lowers the volume of the multimedia resource and plays the voice segment corresponding to the voice tag. The solution provided by the embodiment of the present disclosure, by lowering the volume of the multimedia resource, can not only reduce the impact of the multimedia resource on the playback of the voice segment, ensuring that the user can clearly understand the content of the voice segment, but also eliminate the need to pause the multimedia resource, which is conducive to improving the playback efficiency of the multimedia resource.

[0174] In some embodiments, while playing a voice clip corresponding to any voice tag, the second terminal can also display subtitles for the voice clip. The solution provided by the embodiments of the present disclosure presents the tags in the multimedia resource to the user from both visual and auditory perspectives by displaying subtitles for the voice clip during playback, thereby improving the accuracy of information transmission.

[0175] In some embodiments, there are multiple voice tags for multimedia resources. Multiple voice tags can be played in a certain order. Accordingly, after playing the voice segment corresponding to a voice tag, the second terminal continues to play the voice segment corresponding to the next voice tag in the order of the multiple voice tags. Among them, the order can be the order of adding multiple voice tags, the order of corresponding playback moments, the position order of multiple voice tags (such as from top to bottom, from left to right), or customized by the publisher of the multimedia resource, etc., and the embodiments of the present disclosure are not limited to this. The solution provided by the embodiments of the present disclosure can automatically play multiple voice tags in the order of the voice tags, without the user manually triggering each voice tag. Only one voice tag needs to be triggered to achieve multiple voice tags. The operation is simple and can improve the playback efficiency of the voice tags, that is, it can improve the efficiency of human-computer interaction.

[0176] In some embodiments, when playback of a voice segment corresponding to any voice tag is paused, in response to a re-triggering operation on the voice tag, the second terminal resumes playback of the voice segment corresponding to the voice tag according to the playback progress reached last time. The solution provided by the embodiments of the present disclosure allows playback of a voice segment corresponding to any voice tag to be continued at the previous playback progress when the segment is played again, without having to play the segment from the beginning. This improves the playback efficiency of the voice segments and the efficiency of human-computer interaction.

[0177] The embodiments of the present disclosure provide a method for displaying multimedia resources. During the process of displaying multimedia resources, at least one voice tag in the multimedia resources is displayed, and each tag in the multimedia resources is intuitively presented to the user, so that the user can view it in time. When any voice tag is triggered, the voice clip corresponding to the voice tag is played to convey the specific content of the tag to the user, which is conducive to improving the efficiency of information transmission and human-computer interaction. Moreover, compared with text tags, voice tags usually contain more information such as the emotion and tone of the publisher of the multimedia resource, so that the user can more accurately know and understand the information conveyed by the publisher, which is conducive to improving the accuracy of information transmission.

[0178] All the above optional technical solutions can be arbitrarily combined to form optional embodiments of the present disclosure, and will not be described in detail here.

[0179] Figure 13 FIG1 is a block diagram of a multimedia resource marking device according to an exemplary embodiment. Figure 13 The multimedia resource marking device includes: a display unit 1301 and a publishing unit 1302.

[0180] The display unit 1301 is configured to execute in the resource editing interface and display the multimedia resources;

[0181] The display unit 1301 is further configured to execute a voice tagging operation on a multimedia resource, and add a voice tag to the multimedia based on the input voice segment, wherein the voice tag is used to trigger the playback of the voice segment or the display of text obtained by recognizing the voice segment;

[0182] The publishing unit 1302 is configured to execute, in response to a resource publishing operation, publishing a multimedia resource containing a voice tag.

[0183] In some embodiments, the display unit 1301 is configured to display a voice input control in response to a mark adding operation on any position in the multimedia resource, where the voice input control is used to obtain a voice segment;

[0184] In response to a triggering operation on the voice input control, a voice mark is displayed at the above position based on the input voice segment.

[0185] In some embodiments, the display unit 1301 is further configured to perform:

[0186] In response to a mark adding operation on any position in the multimedia resource, an information input box is displayed at the position, the information input box being used to obtain the mark for the multimedia resource;

[0187] In response to a triggering operation on the voice input control, displaying a voice mark in the information input box based on the input voice segment;

[0188] In response to a mark confirmation operation in the information input box, the information input box is closed and the voice mark is displayed at the above position.

[0189] In some embodiments, the display unit 1301 is further configured to perform:

[0190] In response to a text input operation on the information input box, displaying the input first text in the information input box;

[0191] In response to a mark confirmation operation in the information input box, the information input box is closed, and the first text and voice mark are displayed at the above position.

[0192] In some embodiments, the display unit 1301 is further configured to execute a mark confirmation operation in response to the information input box, and display a second text at the above position, where the second text is obtained based on the recognition of the input voice segment.

[0193] In some embodiments, the display unit 1301 is further configured to perform:

[0194] In the resource editing interface, multiple tone options are displayed;

[0195] When any one of the plurality of timbre options is selected, playing the input voice clip based on the target timbre indicated by the audio option;

[0196] In response to a mark confirmation operation in the information input box, the timbre of the speech segment is adjusted to a target timbre.

[0197] In some embodiments, the display unit 1301 is further configured to perform:

[0198] In response to an object marking operation on a multimedia resource, displaying an object identifier of a target object indicated by the object marking operation in an information input box;

[0199] In response to the mark confirmation operation in the information input box, the object identification of the target object is displayed at the above position.

[0200] In some embodiments, the display unit 1301 is further configured to execute: in response to the resource publishing operation, display the multimedia resource containing the voice mark in the conversation interface between the operation object and the target object.

[0201] In some embodiments, the display unit 1301 is further configured to: when a target object comments on a multimedia resource, display the comment message of the target object at the top of the comment interface of the multimedia resource.

[0202] The disclosed embodiments provide a novel device for marking multimedia resources. When marking multimedia resources, a voice clip can be directly input in the resource editing interface to voice-mark the multimedia resources without the need for the user to manually input text. Compared with text marking, voice marking is simpler and can improve the marking efficiency of multimedia resources, thereby helping to improve the publishing efficiency of multimedia resources, that is, it can improve the efficiency of human-computer interaction.

[0203] It should be noted that the multimedia resource marking device provided in the above embodiment only uses the division of the above functional units as an example to illustrate when marking multimedia resources. In actual applications, the above functions can be assigned to different functional units as needed, that is, the internal structure of the electronic device can be divided into different functional units to complete all or part of the functions described above. In addition, the multimedia resource marking device provided in the above embodiment and the multimedia resource marking method embodiment belong to the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0204] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0205] Figure 14 FIG1 is a block diagram of a device for displaying multimedia resources according to an exemplary embodiment. Figure 14 The multimedia resource display device includes: a display unit 1401 and a playback unit 1402.

[0206] The display unit 1401 is configured to execute, in a resource display interface, displaying multimedia resources and at least one voice tag of the multimedia resources, each voice tag being used to play a corresponding voice segment or display text recognized based on the voice segment after being triggered;

[0207] The playing unit 1402 is configured to play a voice segment corresponding to any one of the at least one voice marker or display text recognized based on the voice segment in response to a triggering operation on the voice marker.

[0208] In some embodiments, the playback unit 1402 is configured to perform any of the following:

[0209] For any one of the at least one voice mark, in response to a triggering operation on the voice mark, pausing playback of the multimedia resource and playing a voice segment corresponding to the voice mark;

[0210] In response to a triggering operation on a voice tag, continue playing the multimedia resource at a first volume and play the voice segment corresponding to the voice tag at a second volume, the first volume being lower than the second volume;

[0211] In response to the triggering operation of the voice mark, the volume of the multimedia resource is lowered and the voice segment corresponding to the voice mark is played.

[0212] In some embodiments, there are multiple voice tags for the multimedia resource;

[0213] The playing unit 1402 is further configured to, after playing the voice segment corresponding to one voice mark, continue to play the voice segment corresponding to the next voice mark in the order of the multiple voice marks.

[0214] In some embodiments, the display unit 1401 is configured to perform:

[0215] In the resource display interface, play multimedia resources;

[0216] In response to a pause operation on the multimedia resource, at least one voice mark of the multimedia resource is displayed.

[0217] In some embodiments, the display unit 1401 is configured to perform:

[0218] In the resource display interface, play multimedia resources;

[0219] In response to the multimedia resource being played to a target time, at least one voice mark of the multimedia resource is displayed, where the target time is the playing time when the at least one voice mark is added to the multimedia resource.

[0220] In some embodiments, the display unit 1401 is configured to perform:

[0221] In the resource display interface, multimedia resources and a progress bar of multimedia resources are displayed;

[0222] At least one voice mark of the multimedia resource is displayed on the progress bar of the multimedia resource.

[0223] In some embodiments, the playback unit 1402 is further configured to execute, in the case of pausing the playback of the voice segment corresponding to any voice mark, in response to the re-triggering operation of the voice mark, continuing to play the voice segment corresponding to the voice mark according to the playback progress achieved last time.

[0224] The embodiments of the present disclosure provide a display device for multimedia resources. During the process of displaying multimedia resources, at least one voice tag in the multimedia resources is displayed, and each tag in the multimedia resources is intuitively presented to the user, which is convenient for the user to view in time. When any voice tag is triggered, the voice segment corresponding to the voice tag is played or the text of the voice segment is displayed, so as to intuitively convey the specific content of the tag to the user, thereby improving the efficiency of information transmission and human-computer interaction. Moreover, compared with text tags, voice tags usually contain more information such as the emotion and tone of the publisher of the multimedia resource, so that the user can more accurately know and understand the information conveyed by the publisher, thereby improving the accuracy of information transmission.

[0225] It should be noted that the multimedia resource display device provided in the above embodiment only uses the division of the above functional units as an example to illustrate when displaying multimedia resources. In actual applications, the above functions can be assigned to different functional units as needed, that is, the internal structure of the electronic device can be divided into different functional units to complete all or part of the functions described above. In addition, the multimedia resource display device provided in the above embodiment and the multimedia resource display method embodiment are of the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0226] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0227] When an electronic device is provided as a terminal, Figure 15 FIG1 is a block diagram of a terminal 1500 according to an exemplary embodiment. Figure 15 The following is a block diagram of a terminal 1500 according to an exemplary embodiment of the present disclosure. Terminal 1500 may be a smartphone, tablet computer, MP3 player (Moving Picture Experts Group Audio Layer III), MP4 player (Moving Picture Experts Group Audio Layer IV), laptop computer, or desktop computer. Terminal 1500 may also be referred to as user equipment, portable terminal, laptop terminal, desktop terminal, or other similar names.

[0228] Typically, the terminal 1500 includes a processor 1501 and a memory 1502 .

[0229] The processor 1501 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 1501 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 1501 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1501 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1501 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.

[0230] The memory 1502 may include one or more computer-readable storage media, which may be non-transitory. The memory 1502 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 1502 is used to store at least one computer program, which is used to be executed by the processor 1501 to implement the multimedia resource marking method or multimedia resource display method provided in the method embodiment of the present application.

[0231] In some embodiments, terminal 1500 may optionally include a peripheral device interface 1503 and at least one peripheral device. The processor 1501, memory 1502, and peripheral device interface 1503 may be connected via a bus or signal lines. Each peripheral device may be connected to peripheral device interface 1503 via a bus, signal lines, or circuit boards. Specifically, the peripheral device may include at least one of a radio frequency circuit 1504, a display screen 1505, a camera assembly 1506, an audio circuit 1507, and a power supply 1508.

[0232] The peripheral device interface 1503 can be used to connect at least one I / O (Input / Output)-related peripheral device to the processor 1501 and the memory 1502. In some embodiments, the processor 1501, the memory 1502, and the peripheral device interface 1503 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1501, the memory 1502, and the peripheral device interface 1503 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.

[0233] RF circuit 1504 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. RF circuit 1504 communicates with communication networks and other communication devices via electromagnetic signals. RF circuit 1504 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals. In some embodiments, RF circuit 1504 includes an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, and the like. RF circuit 1504 can communicate with other terminals via at least one wireless communication protocol. Such wireless communication protocols include, but are not limited to, the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, RF circuit 1504 may also include circuitry related to Near Field Communication (NFC), although this application does not limit this.

[0234] The display screen 1505 is used to display a UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 1505 is a touch screen display, the display screen 1505 also has the ability to collect touch signals on the surface or above the surface of the display screen 1505. The touch signal can be input as a control signal to the processor 1501 for processing. In this case, the display screen 1505 can also be used to provide virtual buttons and / or virtual keyboards, also known as soft buttons and / or soft keyboards. In some embodiments, there can be one display screen 1505, which is set on the front panel of the terminal 1500; in other embodiments, there can be at least two display screens 1505, which are respectively set on different surfaces of the terminal 1500 or in a folding design; in other embodiments, the display screen 1505 can be a flexible display screen, which is set on the curved surface or folding surface of the terminal 1500. Even more, the display screen 1505 can be set to a non-rectangular irregular shape, that is, a special-shaped screen. The display screen 1505 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0235] The camera assembly 1506 is used to capture images or videos. In some embodiments, the camera assembly 1506 includes a front camera and a rear camera. Typically, the front camera is provided on the front panel of the terminal, and the rear camera is provided on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth of field camera, a wide-angle camera, and a telephoto camera, so as to realize the fusion of the main camera and the depth of field camera to realize the background blur function, the fusion of the main camera and the wide-angle camera to realize panoramic shooting and VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera assembly 1506 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation at different color temperatures.

[0236] The audio circuit 1507 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, and convert the sound waves into electrical signals that are input into the processor 1501 for processing, or input into the radio frequency circuit 1504 to achieve voice communication. For the purpose of stereo sound collection or noise reduction, there may be multiple microphones, each located in different parts of the terminal 1500. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert electrical signals from the processor 1501 or the radio frequency circuit 1504 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert electrical signals into sound waves audible to humans, but also convert electrical signals into sound waves inaudible to humans for purposes such as ranging. In some embodiments, the audio circuit 1507 may also include a headphone jack.

[0237] Power supply 1508 is used to power various components in terminal 1500. Power supply 1508 can be AC ​​power, DC power, a disposable battery, or a rechargeable battery. When power supply 1508 includes a rechargeable battery, the rechargeable battery can be wired or wirelessly rechargeable. A wired rechargeable battery is charged via a wired line, while a wireless rechargeable battery is charged via a wireless coil. The rechargeable battery can also support fast charging technology.

[0238] Those skilled in the art will understand that Figure 15 The structure shown in the figure does not constitute a limitation on the terminal 1500, and the terminal 1500 may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component arrangement.

[0239] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory 1502 including instructions. The instructions can be executed by the processor 1501 of the terminal 1500 to implement the above-mentioned multimedia resource marking method or multimedia resource display method. Alternatively, the computer-readable storage medium can be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, optical data storage device, etc.

[0240] A computer program product includes a computer program / instruction, which implements the above-mentioned multimedia resource marking method or multimedia resource display method when executed by a processor.

[0241] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.

[0242] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A method for marking multimedia resources, characterized in that: The method comprises: In the resource editing interface, multimedia resources are displayed; In response to a voice tagging operation on the multimedia resource, a voice tag is added to the multimedia based on the input voice segment, wherein the voice tag is used to trigger playback of the voice segment or display of text recognized based on the voice segment; In response to a resource publishing operation, the multimedia resource including the voice mark is published.

2. The multimedia resource marking method according to claim 1, characterized in that: In response to the voice tagging operation on the multimedia resource, adding a voice tag to the multimedia based on the input voice segment includes: In response to a mark adding operation on any position in the multimedia resource, displaying a voice input control, wherein the voice input control is used to obtain a voice segment; In response to a triggering operation on the voice input control, the voice mark is displayed at the position based on the input voice segment.

3. The multimedia resource marking method according to claim 2, characterized in that: The method further comprises: In response to a mark adding operation on any position in the multimedia resource, displaying an information input box at the position, the information input box being used to obtain a mark for the multimedia resource; The step of displaying the voice mark at the location based on the input voice segment in response to the triggering operation on the voice input control comprises: In response to a triggering operation on the voice input control, displaying the voice mark in the information input box based on the input voice segment; In response to a mark confirmation operation in the information input box, the information input box is closed and the voice mark is displayed at the position.

4. The multimedia resource marking method according to claim 3, characterized in that: The method further comprises: In response to a text input operation on the information input box, displaying the input first text in the information input box; The step of closing the information input box and displaying the voice mark at the position in response to the mark confirmation operation in the information input box comprises: In response to a mark confirmation operation in the information input box, the information input box is closed, and the first text and the voice mark are displayed at the position.

5. The multimedia resource marking method according to claim 3, characterized in that: The method further comprises: In response to a mark confirmation operation in the information input box, a second text is displayed at the position, where the second text is obtained based on recognition of the input voice segment.

6. The multimedia resource marking method according to claim 3, characterized in that: The method further comprises: In the resource editing interface, multiple tone options are displayed; When any one of the plurality of timbre options is selected, playing the input voice segment based on the target timbre indicated by the audio option; In response to a mark confirmation operation in the information input box, the timbre of the voice segment is adjusted to the target timbre.

7. The multimedia resource marking method according to claim 3, characterized in that: The method further comprises: In response to an object marking operation on the multimedia resource, displaying an object identifier of a target object indicated by the object marking operation in the information input box; In response to a mark confirmation operation in the information input box, an object identifier of the target object is displayed at the position.

8. The multimedia resource marking method according to claim 7, characterized in that: The method further comprises: In response to a resource publishing operation, the multimedia resource including the voice mark is displayed in a conversation interface between the operating object and the target object.

9. The multimedia resource marking method according to claim 7, characterized in that: The method further comprises: In the case where the target object comments on the multimedia resource, the comment message of the target object is displayed at the top of the comment interface of the multimedia resource.

10. A method for displaying multimedia resources, characterized in that: The method further comprises: In the resource display interface, multimedia resources and at least one voice tag of the multimedia resources are displayed, each voice tag is used to play the corresponding voice segment or display the text obtained by recognizing the voice segment after being triggered; For any voice mark among the at least one voice mark, in response to a triggering operation on the voice mark, a voice segment corresponding to the voice mark is played or a text obtained by recognizing the voice segment is displayed.

11. The multimedia resource marking method according to claim 10, characterized in that: For any of the at least one voice mark, in response to a triggering operation on the voice mark, playing a voice segment corresponding to the voice mark includes any of the following: For any voice mark among the at least one voice mark, in response to a triggering operation on the voice mark, pausing the playing of the multimedia resource and playing the voice segment corresponding to the voice mark; In response to a triggering operation on the voice mark, continue playing the multimedia resource at a first volume and play the voice segment corresponding to the voice mark at a second volume, where the first volume is lower than the second volume; In response to a triggering operation on the voice mark, the volume of the multimedia resource is lowered, and the voice segment corresponding to the voice mark is played.

12. The multimedia resource marking method according to claim 10, characterized in that: There are multiple voice tags for the multimedia resource; The method further comprises: After playing the voice segment corresponding to one voice mark, continue playing the voice segment corresponding to the next voice mark in the order of the multiple voice marks.

13. The multimedia resource marking method according to claim 10, characterized in that: The displaying of the multimedia resource and at least one voice mark of the multimedia resource in the resource display interface includes: Playing the multimedia resource in the resource display interface; In response to a pause operation on the multimedia resource, at least one voice mark of the multimedia resource is displayed.

14. The multimedia resource marking method according to claim 10, characterized in that: The displaying of the multimedia resource and at least one voice mark of the multimedia resource in the resource display interface includes: Playing the multimedia resource in the resource display interface; In response to the multimedia resource being played to a target time, at least one voice mark of the multimedia resource is displayed, where the target time is the playing time when the at least one voice mark is added to the multimedia resource.

15. The multimedia resource marking method according to claim 10, characterized in that: The displaying of the multimedia resource and at least one voice mark of the multimedia resource in the resource display interface includes: In the resource display interface, display the multimedia resource and a progress bar of the multimedia resource; At least one voice mark of the multimedia resource is displayed on the progress bar of the multimedia resource.

16. The multimedia resource marking method according to claim 10, characterized in that: The method further comprises: In the case of pausing the playing of the voice segment corresponding to any voice mark, in response to the re-triggering operation of the voice mark, the playing of the voice segment corresponding to the voice mark is continued according to the playing progress reached last time.

17. A multimedia resource marking device, characterized in that: The device comprises: A display unit is configured to execute in the resource editing interface and display the multimedia resources; The display unit is further configured to execute a voice tagging operation on the multimedia resource, and add a voice tag to the multimedia based on the input voice segment, wherein the voice tag is used to trigger the playback of the voice segment or the display of text obtained by recognizing the voice segment; The publishing unit is configured to publish the multimedia resource containing the voice mark in response to a resource publishing operation.

18. A multimedia resource display device, characterized in that: The device further comprises: A display unit is configured to execute in a resource display interface, displaying a multimedia resource and at least one voice tag of the multimedia resource, each voice tag being used to play a corresponding voice segment or display text obtained based on recognition of the voice segment after being triggered; The playback unit is configured to execute, for any voice mark among the at least one voice mark, in response to a triggering operation on the voice mark, play the voice segment corresponding to the voice mark or display text obtained by recognizing the voice segment.

19. An electronic device, characterized in that: The electronic device comprises: one or more processors; a memory for storing program code executable by the processor; The processor is configured to execute the program code to implement the multimedia resource marking method according to any one of claims 1 to 9 or the multimedia resource display method according to any one of claims 10 to 15.

20. A computer-readable storage medium, characterized in that When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the multimedia resource marking method as described in any one of claims 1 to 9 or the multimedia resource display method as described in any one of claims 10 to 15.

21. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method for marking multimedia resources according to any one of claims 1 to 9 or the method for displaying multimedia resources according to any one of claims 10 to 15 is implemented.

Citation Information

Patent Citations

  • Voice conversion method and device, file generation method and device, broadcasting method and device, voice processing method and device and medium

    CN110970014A

  • Method and device for carrying out voice marking on pictures and videos

    CN113392272A

  • Voice-to-character intelligent marking method and system, and medium

    CN116033206A

  • Voice comment display method and voice comment processing method

    CN117150071A

  • Resource display method and device, terminal and storage medium

    CN118921532A