Presentation method, apparatus, and electronic device
The method and apparatus enhance multimedia conference understanding by converting audio to subtitle information and displaying annotation, addressing language barriers and unclear expressions to improve dialogue efficiency.
Patent Information
- Application Number
- JP2024563048
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-04-29
- Filing Date
- 2023-04-11
- Publication Date
- 2026-01-09
- Estimated Expiration
- 2043-04-11
Smart Images

Figure 0007796904000001 
Figure 0007796904000002 
Figure 0007796904000003
Abstract
Description
[Technical Field]
[0001] [Cross-reference of related applications] This application claims priority to a Chinese patent application filed on April 29, 2022, bearing application number 202210495727.6 and titled "Presentation method, apparatus, and electronic device," the entire contents of which are incorporated herein by reference. [Technical field] The present invention relates to the field of computer technology, and in particular to a presentation method, apparatus, and electronic device. [Background technology]
[0002] With the development of the Internet, users increasingly use the functions of terminal devices to make their work and lives more convenient. For example, users can use terminal devices to hold online multimedia conferences with other users. Through online multimedia conferences, users can realize long-distance conversations and eliminate the need for users to gather in one place to hold conferences. Multimedia conferences largely avoid the constraints of location and venue of traditional face-to-face conferences. Summary of the Invention
[0003] This Summary section is provided to introduce in a simplified form the ideas that are described in detail later in the Specific Embodiments section. This Summary section is not intended to identify key features or necessary features of the claimed technical solution, nor is it intended to be used to limit the scope of the claimed technical solution.
[0004] According to a first aspect, an embodiment of the present invention provides a presentation method, the method including: obtaining subtitle information of a multimedia conference, where the subtitle information is obtained based on converting audio information during the multimedia conference; determining annotated content in the subtitle information; obtaining annotation information for the annotated content; displaying the subtitle information, and displaying the annotation information corresponding to the annotated content.
[0005] According to a second aspect, an embodiment of the present invention provides a presentation device, the presentation device including: a first acquisition unit for acquiring subtitle information of a multimedia conference, where the subtitle information is obtained based on converting audio information during the multimedia conference; a determination unit for determining annotated content from the subtitle information; a second acquisition unit for acquiring annotation information of the annotated content; and a display unit for displaying the subtitle information and the annotation information corresponding to the annotated content.
[0006] According to a third aspect, an embodiment of the present invention provides an electronic device, the electronic device comprising one or more processors and a storage device for storing one or more programs, the one or more programs, when executed by the one or more processors, causing the one or more processors to implement the presentation method described in the first aspect.
[0007] According to a fourth aspect, an embodiment of the present invention provides a computer-readable medium having stored thereon a computer program, the program being adapted to, when executed by a processor, perform the steps of the presentation method according to the first aspect.
[0008] The presentation method, apparatus, and electronic device provided by the embodiments of the present invention can obtain subtitle information of a multimedia conference, which can be obtained by converting audio information during a multimedia conference, determine words from the subtitle information to obtain content to be annotated, then obtain annotation information for the content to be annotated, and then display the subtitle information and display annotation information corresponding to the annotation words in the subtitle information, thereby obtaining a new presentation method. According to this presentation method, annotation information can be presented to participants in a multimedia conference, and the annotation information can be used to improve the speed at which the participants understand the multimedia conference, avoiding interruptions to the progress of the conference or misunderstandings caused by not being able to understand the meaning of other participants, thereby improving the accuracy and efficiency of dialogue in the multimedia conference. [Brief explanation of the drawings]
[0009] The above-mentioned and other features, advantages, and aspects of each embodiment of the present invention will become more apparent by reference to the following specific embodiments in conjunction with the drawings. Identical or similar reference numerals refer to identical or similar elements throughout the drawings. It should be understood that the drawings are schematic and that the drawings and elements are not necessarily drawn to scale. [Figure 1] FIG. 1 is a flowchart illustrating an embodiment of a presentation method according to the present invention. [Figure 2] 1 is a diagram showing a schematic diagram of an application scene of a presentation method according to the present invention; [Figure 3] 1 is a diagram showing a schematic diagram of an application scene of a presentation method according to the present invention; [Figure 4] 1 is a diagram showing a schematic diagram of an application scene of a presentation method according to the present invention; [Figure 5] 1 is a diagram showing a schematic diagram of the structure of an embodiment of a presentation device according to the present invention; [Figure 6] FIG. 1 illustrates an exemplary system architecture to which the presented method of one embodiment of the present invention may be applied. [Figure 7]1 shows a schematic diagram of the basic structure of an electronic device provided by an embodiment of the present invention; DETAILED DESCRIPTION OF THE INVENTION
[0010] Hereinafter, the embodiments of the present invention will be described in more detail with reference to the drawings. Although the drawings show specific embodiments of the present invention, it should be understood that the present invention can be realized in various forms and should not be construed as being limited to the embodiments described herein, but rather these embodiments are provided for a more thorough and complete understanding of the present invention. It should be understood that the drawings and embodiments of the present invention are for illustrative purposes and are not used to limit the protection scope of the present invention.
[0011] It should be understood that the steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel, and that method embodiments may include additional steps and / or omit illustrated steps, and the scope of the present invention is not limited in this respect.
[0012] As used herein, the term "comprising" and variations thereof are open-ended inclusions, including, but not limited to, the term "based on" means "based at least in part on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one other embodiment," and the term "some embodiments" means "at least some embodiments." Relevant definitions of other terms are provided below.
[0013] It should be noted that the concepts of "first", "second", etc. referred to in the present invention are merely used to distinguish between different devices, modules, or units, and are not intended to limit the order or interdependence of functions performed by these devices, modules, or units.
[0014] It should be noted that the modifications "one" and "multiple" referred to in the present invention are exemplary rather than limiting, and that those skilled in the art should understand "one or more" unless the context clearly indicates otherwise.
[0015] The names of messages or information exchanged between devices in the embodiments of the present invention are used for descriptive purposes only and are not intended to limit the scope of these messages or information.
[0016] 1, a flow chart of an embodiment of the method for displaying information according to the present invention is shown. The method for displaying information shown in FIG. 1 includes the following steps:
[0017] In step 101, the closed caption information of the multimedia conference is obtained.
[0018] Here, the subtitle information is obtained based on converting audio information during a multimedia conference.
[0019] In this embodiment, the multimedia conference may be an online conference using a multimedia method, the multimedia may include but is not limited to audio and / or video, and the multimedia conference interface may be a related interface for the multimedia conference.
[0020] In this embodiment, the application for holding a multimedia conference may be any type of application, and is not limited thereto, for example, the application may be an instant video conference type application, a communication type application, a video playback type application, an email type application, etc.
[0021] In this embodiment, the audio information of participants in a multimedia conference can be converted into subtitle information. For example, participants in a multimedia conference may include Party A, Party B, and Party C. When a participant speaks, the participant's speech can be converted into subtitle information.
[0022] In this embodiment, the subtitle information may be presented in real time or with a lag relative to the audio time. Optionally, the subtitle information may be presented in a multimedia stream presentation interface, i.e., in the real-time multimedia stream of the conference participants. Optionally, the interface for presenting the subtitle information may be presented alongside the multimedia stream presentation interface.
[0023] In some embodiments, the multimedia conference is a real-time multimedia conference currently in progress, and during the real-time multimedia conference, closed caption information and presentation annotation information are presented.
[0024] In some embodiments, the multimedia conference is a completed multimedia conference. Optionally, the closed caption information may be presented after the completion of the multimedia conference.
[0025] In step 102, the content to be annotated from the subtitle information is determined.
[0026] Here, predefined criteria are used to determine the content to be annotated from the subtitle information.
[0027] In this embodiment, the annotated content may be the content to which the annotation is added, and alternatively, the annotated content may be a word, i.e., the annotated content may be referred to as annotated word.
[0028] Here, the annotated content may be a word that is difficult for the intended conference participants to understand. The annotated content may be determined based on a predetermined condition. The predetermined condition may be set according to an actual situation and is not limited thereto.
[0029] In step 103, annotation information of the content to be annotated is obtained.
[0030] In this embodiment, annotation information of the annotated content can be obtained, which may be information that interprets the annotated content.
[0031] Alternatively, the annotation information may be obtained from a predetermined database. Alternatively, the annotation information may be obtained from the Internet.
[0032] In step 104, the subtitle information is displayed, and the annotation information corresponding to the annotated content is displayed.
[0033] In this embodiment, subtitle information can be displayed, and the displayed subtitle information may include the annotated content, that is, the annotated content or the annotation information of the annotated content may be displayed simultaneously with the display of the subtitle information.
[0034] Optionally, annotation information may be displayed corresponding to the displayed annotated content, for example, adjacent to the display location of the annotated content.
[0035] The presentation method provided by this embodiment can obtain subtitle information for a multimedia conference, which can be obtained by converting audio information during the multimedia conference, determine words from the subtitle information to obtain annotated content, obtain annotation information for the annotated content, and then display the subtitle information and display annotation information corresponding to the annotation words in the subtitle information, thereby obtaining a new presentation method. It should be noted that this presentation method can provide annotation information to participants in a multimedia conference, improve the speed at which participants understand the multimedia conference using the annotation information, and avoid interruptions to the progress of the conference or misunderstandings due to inability to understand the meaning of other participants, thereby improving the accuracy and efficiency of dialogue in the multimedia conference.
[0036] In some embodiments, step 102 may include: determining a user group thesaurus corresponding to a user group to which a conference participant of the multimedia conference belongs; and determining annotated content of the subtitle information based on the user group thesaurus.
[0037] Here, the user group thesaurus includes phrases and phrase interpretations. The method of setting phrases in the user group thesaurus is not limited here.
[0038] Alternatively, a user group thesaurus corresponding to a user group can be determined based on the user group to which all participants in a multimedia conference belong. For example, if participants A, B, and C belong to the same user group, the user group thesaurus for that user group can be used to determine the content to be annotated.
[0039] Alternatively, a user group thesaurus corresponding to a user group may be determined based on the user group to which some conference participants in a multimedia conference belong. For example, if conference participants A and B belong to a first user group and conference participant User C belongs to a second user group, the user group thesaurus corresponding to the first user group may be selectively used to determine the annotated content. Alternatively, the user group thesaurus corresponding to the first user group and the user group thesaurus corresponding to the second user group may be selectively used to determine the annotated content, or the user group thesaurus corresponding to the second user group may be selectively used to determine the annotated content.
[0040] Alternatively, for annotated content determined using the user group thesaurus corresponding to the first user group, annotation information of the annotated content may be presented to conference participants belonging to the first user group or to conference participants belonging to the second user group. For annotated content determined using the user group thesaurus corresponding to the first user group and the user group thesaurus corresponding to the second user group, annotation information of the annotated content may be presented to conference participants belonging to the first user group and to conference participants belonging to the second user group. For annotated content determined using the user group thesaurus corresponding to the second user group, annotation information of the annotated content may be presented to conference participants belonging to the second user group or to conference participants belonging to the second user group.
[0041] It should be explained that by determining a user group thesaurus using the user group to which the conference participants belong and determining the content to be annotated based on the user group thesaurus, conference participants can be presented with phrases that may have specific meanings to the company, helping users accurately understand the meanings of expressions used by other conference participants and improving the efficiency of multimedia conference interactions.
[0042] As an example, referring to Figure 2, Figure 2 shows a scene in which annotation information is presented for a company phrase in subtitle information. In Figure 2, a multimedia stream presentation area 201 can present a multimedia stream of a multimedia conference (e.g., real-time video of conference participants, shared content, etc.). A subtitle presentation area 202 can present a subtitle corresponding to the voice of user A, saying, "The bean value in the calculation result is relatively reasonable." Here, "bean value" may be a phrase in a user group thesaurus corresponding to the user group, and annotation information "growth rate" can be presented corresponding to the "bean value."
[0043] Optionally, determining the annotated content of the subtitle information based on the user group thesaurus may include selecting, from among words of the subtitle information, words that can be hit or included in a phrase in the user group thesaurus as the annotated content.
[0044] In some embodiments, determining content to be annotated from the subtitle information based on the user group thesaurus may include selecting, from among words of the subtitle information, words that match phrases or phrases included in the user group thesaurus that have predefined characteristics, as the content to be annotated.
[0045] Here, a phrase in a user group thesaurus may have several attributes, which may include, but are not limited to, at least one of: phrase length, phrase language, whether the phrase is an abbreviated phrase, the number of times the phrase appears in a given document set, the frequency with which the phrase appears in a given document set, and whether the phrase appears in a given dictionary set.
[0046] Here, the predefined features may be predefined features that indicate words that may be difficult for the conference participants to understand.
[0047] By way of example, the predefined characteristics may include, but are not limited to, at least one of: the phrase length is greater than a predetermined length threshold; the phrase language is a predetermined language; the phrase is an abbreviated phrase; the phrase appears in a predetermined document set less than or equal to a predetermined number threshold; the phrase appears in a predetermined document set less than or equal to a predetermined frequency threshold; and the phrase does not appear in a predetermined dictionary set.
[0048] It should be noted that by using phrases with predefined characteristics from the user group thesaurus as the criterion for determining the content to be annotated, the scope of the content to be annotated can be effectively narrowed, and it is possible to avoid a situation where key points are obscured by annotating all of the large number of words in the subtitle information.
[0049] In some embodiments, the predefined characteristics include the phrase being an abbreviated phrase, and the annotation information includes a full name of the phrase. Step 104 includes displaying the full name of the phrase corresponding to the annotated content.
[0050] Here, the phrase can be an abbreviated phrase, which indicates that the phrase is an abbreviation of a phrase with the same meaning. Usually, in communication, people sometimes use a simpler word to represent a more complex word in order to improve communication efficiency. However, for people who are not proficient in using simpler words (e.g., abbreviations), the meaning of the simpler word may be confused.
[0051] For example, the abbreviation for Internet Data Center is IDC. When "IDC" appears in the subtitle information and a phrase having predefined characteristics is included in the user group thesaurus as "IDC" (the phrase is an abbreviated phrase), "IDC" in the subtitle information can be determined as a phrase to be annotated, and the full phrase name of the phrase to be annotated may include Internet Data Center and / or Internet Data Center.
[0052] As an example, referring to Figure 3, Figure 3 shows a scene in which annotation information is presented corresponding to an abbreviated phrase among corporate phrases in subtitle information. In Figure 3, multimedia streams of a multimedia conference (e.g., real-time video of conference participants, shared content, etc.) can be presented in multimedia stream presentation area 301. Subtitles corresponding to the voice of user B, such as "IDCs are generally installed in remote locations," can be presented in subtitle presentation area 302. Here, "IDC" may be an abbreviated phrase in a user group thesaurus corresponding to a user group, and in this case, annotation information "Internet Data Center" can be presented corresponding to "IDC."
[0053] It should be explained that by using hits or inclusions in abbreviated phrases in the user group thesaurus as the basis for determining the content to be annotated, the efficiency of dialogue can be improved. Specifically, conference participants who use abbreviated phrases do not need to worry about other conference participants not understanding what they want to say (which requires additional time to think), and conference participants who receive abbreviated phrases but do not know their meaning can immediately find out the meaning of the abbreviated phrase by referring to the annotation information, thereby avoiding misunderstandings due to lack of understanding, thereby improving the efficiency of dialogue in multimedia conferences.
[0054] In some embodiments, step 102 may include determining a first language of the multimedia conference and determining annotated content from words in a second language of the closed caption information.
[0055] Here, the first language may be a base language of the multimedia conference, and the base language may be a language that is primarily used during the multimedia conference.
[0056] In some embodiments, the first language may be determined based on at least one of, but not limited to, a geographic region in which the participants of the multimedia conference are located, a native language of the participants of the multimedia conference, and a predetermined language of the multimedia conference.
[0057] In some embodiments, the first language can be set by the conference participant.
[0058] In some embodiments, the first language may be determined by recognizing the speech of the conferee.
[0059] In some embodiments, the first language can be determined based on the conference participant, for example, based on the geographic region in which the conference participant is located, and can also be determined based on the language the conference participant uses in the application.
[0060] wherein the second language is a language different from the first language; optionally, the second language is a language other than the first language; optionally, the second language is a designated language different from the first language.
[0061] For example, the first language of a multimedia conference may be determined as Chinese, and a language other than Chinese may be determined as a second language, for example, the second language may include English. When Chinese and English appear during the multimedia conference, the annotated content may be selected from among the English words.
[0062] It is necessary to explain that when a language other than the first language appears during a multimedia conference, it may hinder the participants' understanding of the multimedia conference, and that by setting words in the second language as content to be annotated, the participants' full understanding of the meanings expressed by the words in the second language can be improved, thereby improving the efficiency of the dialogue in the multimedia conference.
[0063] In some embodiments, determining the content to be annotated from among the second language words in the subtitle information includes selecting, from among the second language words in the subtitle information, second language words that match or are included in a predetermined second language thesaurus as the content to be annotated.
[0064] Here, the second language thesaurus includes second language phrases and corresponding first language definitions.
[0065] Here, the form of the predetermined second language thesaurus can be set according to the actual situation, and is not limited here.
[0066] In some embodiments, a predetermined second language thesaurus may include phrases whose frequency of use is less than a predetermined frequency threshold.
[0067] Some common words in the second language have a difficulty in comprehension that does not hinder the conference participants. By selecting some words from the words in the second language and setting them in advance as a second language thesaurus, and using words that are hit by or included in the thesaurus as content to be annotated, it is possible to reduce the number of times that second language words in the subtitle information are determined as content to be annotated, thereby enabling more accurate selection of words that are difficult for conference participants to understand, avoiding large-scale annotation of second language words in the subtitle information, and reducing interference with conference participants.
[0068] In some embodiments, the first language interpretation includes at least two sub-interpretations, and step 104 may include selecting, based on contextual information in the subtitle information in which the annotated content is located, a sub-interpretation from the first language interpretation corresponding to the annotated content that is related or correlated to the contextual information, and displaying the selected sub-interpretation corresponding to the annotated content.
[0069] For example, the second language may include English, and the annotated content (English word) selected from the subtitle information may have a Chinese interpretation, and the Chinese interpretation may include at least two sub-interpretations. In this case, based on context information in the subtitle information where the annotated content (English word) is located, a sub-interpretation related or correlated to the context information may be selected from the Chinese interpretation of the annotated content.
[0070] In some embodiments, the context may include first language words and / or second language words, and a sub-explanation related to or correlated with the context is selected based on the meaning expressed by the first language words and / or second language words.
[0071] It should be noted that by recognizing and processing contextual information, the selected sub-interpretation can be adapted to the context, improving the accuracy of the displayed sub-interpretation, thereby improving the accuracy of the conference participant's understanding of the second language words, and further improving the accuracy of the conference participant's understanding of the multimedia conference, thereby improving the accuracy and efficiency of the multimedia conference information interaction.
[0072] In some embodiments, step 102 may include selecting, from among the words of the closed caption information, words that hit or are included in a predetermined rare word (rarely used characters) thesaurus as content to be annotated.
[0073] In some implementations, the language to which the rare words of the rare word thesaurus belong may include words of any language, for example, first language words and non-first language words.
[0074] It is necessary to explain that a rare word thesaurus is set, the rare word thesaurus is used to select the content to be annotated, and annotation information of the selected rare word from the subtitle information is displayed, so that even if other participants in the multimedia conference use rare words to express themselves, the participants can accurately understand them using the annotation information of the rare words, thereby improving the participants' understanding of the multimedia conference and further improving the efficiency of the dialogue in the multimedia conference.
[0075] In some embodiments, step 104 described above may include displaying the annotated content of the displayed closed caption information in association with the corresponding annotation information.
[0076] Here, the specific implementation of the associated display can be set depending on the actual application scene, and is not limited here.
[0077] As an example, annotation information can be displayed above or below the content to be annotated, thereby realizing associated display.
[0078] As an example, annotation information can be displayed in a bubble that points to the content to be annotated, thereby enabling associated display.
[0079] It should be explained that by displaying the annotation information in association with the content to be annotated, the annotation information for the content to be annotated can be clearly presented and the accuracy of the annotation information provided can be improved.
[0080] In some embodiments, displaying the annotated content of the displayed subtitle information in association with the corresponding annotation information includes, but is not limited to, at least one of: there being at least one difference in presentation style between the annotated content and the content of the displayed subtitles other than the annotated content; the annotation information presentation area not overlapping with the subtitle presentation area; the annotation information presentation area being embedded within the subtitle presentation area; presenting the annotation information corresponding to the annotated content in response to a trigger operation on the annotated content; presenting the annotation information in association with the corresponding annotated content in the form of a floating window; and during a real-time multimedia conference, a first presentation duration of the annotated content of the subtitle information is less than or equal to a second presentation duration of the corresponding annotation information.
[0081] Here, there is at least one difference in the presentation style between the annotated content and the other content in the displayed subtitles. For example, the annotated content (small values) in Figure 2 can be presented in a different presentation style from the other content (relative validity in the calculation results). For example, the annotated content can be presented in bold or italics (differentiated display not shown).
[0082] Here, the annotation information presentation area does not overlap with the subtitle presentation area; in other words, the annotation information can be presented in a separate area from the subtitle information.
[0083] Here, the annotation information presentation area is embedded in the subtitle presentation area, in other words, the annotation information can be displayed in the subtitle presentation area, for example, as shown in FIG.
[0084] In this case, annotation information corresponding to the annotated content can be presented in response to a trigger operation on the annotated content, i.e., the annotation information is presented in response to a user's trigger, thereby reducing the interference of the annotation information with the inquiry about the subtitle information of the conference participant.
[0085] Optionally, the above-mentioned step 102 includes, but is not limited to, at least one of determining the annotated content in response to a user's operation, and the electronic device automatically determining the annotated content.
[0086] Here, the annotation information is presented in association with the corresponding annotated content in the form of a floating window, thereby not only presenting the annotation information in association with the annotated content, but also not affecting the presentation of the subtitle information.
[0087] During a real-time multimedia conference, the first presentation duration of the annotated content in the subtitle information is equal to or shorter than the second presentation duration of the corresponding annotation information. In other words, the duration of the annotation information staying on the interface is equal to or longer than the duration of the annotated content on the subtitle information terminal. This allows for sufficient presentation of the annotation information and helps users understand the annotated content.
[0088] In some embodiments, the method further includes, for the same annotated content that appears at least twice during the multimedia conference, determining an interval between a non-first appearance of the annotated content and a previous corresponding annotated content for which annotation information was presented, and displaying annotation information corresponding to the non-first appearance of the annotated content in response to the interval satisfying a predetermined interval condition.
[0089] Here, annotation information can be displayed corresponding to the first-appearing annotation target content.
[0090] For example, the same annotated content may appear multiple times during a multimedia conference. For the same word that appears multiple times, the frequency of the word's appearance can be controlled. For example, the time interval between two annotations for the same word is greater than one minute. For example, the position interval between two annotations for the same word is greater than five lines.
[0091] For example, refer to FIG. 4, which shows an exemplary scenario for controlling the annotation frequency for annotated content. In FIG. 4, subtitle information corresponding to the voices of conference participants can be presented in a subtitle information presentation area 201. The English phrase "cute," which corresponds to the Chinese word "cute," appears in the subtitle information corresponding to the voices of User A, User B, User C, and User D. That is, "cute" appears multiple times during a multimedia conference. The word "cute" spoken by User A appears first, while the words "cute" spoken by User B, User C, and User D appear non-first in the annotated content. In Figure 4, if the subtitle information of user B's audio information contains the annotation target content (cute), it can be determined that the interval between the cute spoken by user B and the cute in the annotation information presented correspondingly last time is 0 lines, and if the interval condition is 1 line or more, it can be determined that the cute spoken by user B does not satisfy the interval condition, and annotation information (i.e., cute) will not be presented corresponding to the cute spoken by user B.
[0092] In Figure 4, if the subtitle information of user C's audio information contains the annotation target content (cute), it can be determined that the interval between the cute spoken by user C and the cute in the annotation information presented correspondingly last time is one line, and if the interval condition is one line or more, it can be determined that the cute spoken by user C satisfies the interval condition, and annotation information (i.e., cute) is presented corresponding to the cute spoken by user C.
[0093] It is necessary to explain that controlling the frequency of presenting annotation information for the same annotated content can reduce the interference caused by frequent presentation of annotation information in situations where a user may temporarily understand the annotated content, and can also help the user understand the annotated content in situations where a conference participant user has not seen the annotation information for the annotated content for a long period of time. This can improve the efficiency of dialogue in multimedia conferences by reducing interference to the user and improving the degree of timely attention to the user.
[0094] In some embodiments, the interval comprises a time interval, the time interval indicating the interval between audio information corresponding to the annotated content.
[0095] In some embodiments, the method further includes determining that a time interval between a non-first appearance of the annotated content and the last corresponding annotated content for which annotation information was presented is greater than a predetermined duration threshold, and the time interval satisfies a predetermined interval condition.
[0096] For example, the same annotated content may appear multiple times during a multimedia conference. For the same word that appears multiple times, the frequency of the word's appearance can be controlled. For example, the time interval between presenting annotation information for the same word twice can be set to be greater than one minute.
[0097] By adjusting the interval at which annotation information for the same annotated content is presented, it is possible to present the annotation information effectively in accordance with the law of human forgetting. For example, if a human's memory is short (5 minutes), by controlling the interval at which the same annotation information is presented to be 5 minutes or more, it is possible to not present the annotation information for the annotation content for 5 minutes when the conference participant user may still remember it, and to present the annotation information again 5 minutes after the annotation information for the annotation content that the conference participant user may have already forgotten.
[0098] In some embodiments, the spacing includes text spacing, which is used to indicate positional spacing between the same annotated content in the closed caption information.
[0099] In some embodiments, the method further includes determining that a text spacing between a non-first occurrence of the annotated content and the last corresponding annotated content for which annotation information was presented is greater than a predetermined text length threshold, and the text spacing satisfies a predetermined spacing condition.
[0100] For example, the same annotated content may appear multiple times during a multimedia conference. For the same word that appears multiple times, the frequency of the word's appearance can be controlled. For example, for the same word, the interval between two annotations can be greater than five lines.
[0101] Alternatively, the text spacing may indicate the spacing between texts. The text spacing may include, but is not limited to, at least one of the size between display positions of texts on an interface and the spacing between statements to which the text belongs. The size between display positions of texts on an interface may be indicated by an absolute size or by the line gap between lines of text in which the texts in the subtitle information are located.
[0102] In some embodiments, the text length threshold comprises a line gap between lines of text located in the closed caption information that is greater than a predetermined line gap.
[0103] It is necessary to explain that by adjusting the text interval for presenting annotation information for the same annotated content, it is possible to present the information effectively in accordance with the presentation format of the subtitle information and the user's viewing format. As an example, if five lines of subtitle information can be presented at a time in the subtitle information presentation area in Figure 2, by controlling the presentation interval of the same annotation information to be five lines or more, if the conference participant user can see five lines of annotation information for the annotated content, the annotation information will not be presented, and if the conference participant user may have already forgotten the annotation information and it is not presented on the current page (i.e., after an interval of five lines), the annotation information will be presented again.
[0104] Further referring to FIG. 5, as an implementation of the method shown in each of the above-mentioned drawings, the present invention provides an embodiment of a presentation apparatus, which corresponds to the embodiment of the method shown in FIG. 1, and which can be specifically applied to various electronic devices.
[0105] 5, the presentation device of this embodiment includes a first acquisition unit 501, a determination unit 502, a second acquisition unit 503, and a display unit 504. Here, the first acquisition unit acquires subtitle information of a multimedia conference, where the subtitle information is obtained based on converting audio information during the multimedia conference, the determination unit determines annotated content from the subtitle information, the second acquisition unit acquires annotation information of the annotated content, and the display unit displays the subtitle information and the annotation information corresponding to the annotated content.
[0106] In this embodiment, the specific processing and technical effects brought about by the first acquisition unit 501, the determination unit 502, the second acquisition unit 503, and the display unit 504 of the presentation device can refer to the relevant descriptions of step 101, step 102, step 103, and step 104 of the corresponding embodiment in Figure 1, respectively, and will not be further described here.
[0107] In some embodiments, determining the annotated content of the subtitle information includes: determining a user group thesaurus corresponding to a user group based on a user group to which a conference participant user of the multimedia conference belongs, the user group thesaurus including phrases and phrase interpretations; and determining the annotated content of the subtitle information based on the user group thesaurus.
[0108] In some embodiments, determining content to be annotated in the subtitle information based on the user group thesaurus includes selecting, from words in the subtitle information, phrases in the user group thesaurus having predefined characteristics, or words of phrases included in the user group thesaurus, or words that match the phrases, as the content to be annotated.
[0109] In some embodiments, the predefined characteristics include the phrase being an abbreviated phrase, the annotation information includes a phrase full name, and displaying the subtitle information and the annotation information corresponding to the annotated content includes displaying the phrase full name corresponding to the annotated content.
[0110] In some embodiments, determining content to be annotated from the subtitle information includes determining a first language of the multimedia conference, where the first language is determined based on at least one of a geographical region in which participants of the multimedia conference are located, a native language of the participants of the multimedia conference, and a predetermined language of the multimedia conference; and determining content to be annotated from second language words of the subtitle information, where the second language is a language other than the first language.
[0111] In some embodiments, determining content to be annotated from second language words in the subtitle information includes selecting, from the second language words in the subtitle information, second language words that hit or are included in a predetermined second language thesaurus as content to be annotated, wherein the second language thesaurus includes second language phrases and corresponding first language interpretations.
[0112] In some embodiments, the first language interpretation includes at least two sub-interpretations, and displaying the subtitle information and the annotation information corresponding to the annotated content includes:
[0113] The method includes selecting a sub-interpretation that is related or correlated to the context information from among the first language interpretations corresponding to the content to be annotated based on context information in the subtitle information in which the content to be annotated is located, and displaying the selected sub-interpretation corresponding to the content to be annotated.
[0114] In some embodiments, determining the annotated content of the subtitle information includes selecting, from among words of the subtitle information, words that hit or are included in a predetermined rare word thesaurus as the annotated content.
[0115] In some embodiments, displaying the subtitle information and displaying annotation information corresponding to the annotated content includes displaying the annotated content of the displayed subtitle information in association with the corresponding annotation information.
[0116] In some embodiments, displaying the annotated content of the displayed subtitle information in association with the corresponding annotation information includes at least one of: there being at least one difference in presentation style between the annotated content and the content of the displayed subtitles other than the annotated content; the annotation information presentation area not overlapping with the subtitle presentation area; the annotation information presentation area being embedded within the subtitle presentation area; presenting the annotation information corresponding to the annotated content in response to a trigger operation on the annotated content; presenting the annotation information in association with the corresponding annotated content in the form of a floating window; and during a real-time multimedia conference, a first presentation duration of the annotated content of the subtitle information is less than or equal to a second presentation duration of the corresponding annotation information.
[0117] In some embodiments, the device further performs the following for the same annotated content that appears at least twice during the multimedia conference: determining an interval between a non-first appearance of the annotated content and a previous corresponding annotated content for which annotation information was presented; and displaying annotation information corresponding to the non-first appearance of the annotated content in response to the interval satisfying a predetermined interval condition.
[0118] In some embodiments, the interval includes a time interval, the time interval indicating the interval between audio information corresponding to the annotated content, and the device further performs determining that the time interval satisfies a predetermined interval condition if the time interval between a non-first appearance of the annotated content and the last corresponding annotated content for which annotation information was presented is greater than a predetermined duration threshold.
[0119] In some embodiments, the intervals include text intervals, which are used to indicate positional intervals between identical annotated content in subtitle information that are adjacent in appearance order, and the device further executes determining that the text interval satisfies a predetermined interval condition if the text interval between a non-first-appearing annotated content and the annotated content for which annotation information was previously presented is greater than a predetermined text length threshold.
[0120] In some embodiments, the text length threshold comprises a line gap between lines of text located in the closed caption information that is greater than a predetermined line gap.
[0121] In some embodiments, the multimedia conference is an ongoing real-time multimedia conference.
[0122] Referring to FIG. 6, FIG. 6 illustrates an exemplary system architecture to which the presented method of one embodiment of the present invention may be applied.
[0123] 6, the system architecture may include terminal devices 601, 602, 603, a network 604, and a server 605. The network 604 is used to provide a medium for a communication link between the terminal devices 601, 602, 603 and the server 605. The network 604 may include various connection types, such as wired, wireless communication links, or fiber optic cables.
[0124] The terminal devices 601, 602, and 603 interact with a server 605 via a network 604 to receive or send messages, etc. Various client applications may be installed on the terminal devices 601, 602, and 603, such as a web browser application, a search-type application, a news information-type application, etc. The client applications of the terminal devices 601, 602, and 603 can receive user instructions and complete corresponding functions according to the user instructions, such as adding corresponding information to information according to the user instructions.
[0125] The terminal devices 601, 602, and 603 may be hardware or software. If the terminal devices 601, 602, and 603 are hardware, they may be various electronic devices that have a display screen and support web browsing, including, but not limited to, smartphones, tablet computers, e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III, a standard audio compression technology for moving pictures), MP4 players (Moving Picture Experts Group Audio Layer IV, a standard audio compression technology for moving pictures), laptop computers, and desktop computers. If the terminal devices 601, 602, and 603 are software, they may be installed on the electronic devices listed above. The terminal devices 601, 602, and 603 may be implemented as multiple software programs or software modules (e.g., software programs or software modules that provide distributed services) or as a single software program or software module. This does not limit the scope of the present disclosure.
[0126] The server 605 may be a server that provides various services, such as receiving information acquisition requests sent by the terminal devices 601, 602, and 603, acquiring presentation information corresponding to the information acquisition requests in various ways based on the information acquisition requests, and transmitting data related to the presentation information to the terminal devices 601, 602, and 603.
[0127] It should be mentioned that the presentation method provided by the embodiment of the present invention may be executed by a terminal device, and correspondingly, the presentation apparatus may be installed in the terminal device 601, 602, 603. It should be noted that the presentation method provided by the embodiment of the present invention may be executed by a server 605, and correspondingly, the presentation apparatus may be installed in the server 605.
[0128] It should be understood that the number of terminal devices, networks, and servers in Figure 6 is merely approximate, and any number of terminal devices, networks, and servers may be included as required for implementation.
[0129] Referring to Figure 7 below, a schematic diagram of the structure of an electronic device (e.g., the terminal device or server in Figure 6) suitable for implementing an embodiment of the present invention is shown. The terminal device in the embodiment of the present invention may include, but is not limited to, mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and in-vehicle terminals (e.g., car navigation terminals), as well as fixed terminals such as digital TVs (televisions) and desktop computers. The electronic device shown in Figure 7 is merely an example and does not impose any limitations on the functionality and scope of use of the embodiment of the present invention.
[0130] 7, electronic device 700 may include a processing unit (e.g., a central processing unit, a graphics processor, etc.) 701 that can perform various appropriate operations and processes based on programs stored in read-only memory (ROM) 702 or programs loaded from storage device 708 into random access memory (RAM) 703. RAM 703 further stores various programs and data necessary for the operation of electronic device 700. The processing unit 701, ROM 702, and RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to bus 704.
[0131] Typically, devices such as input devices 706 including a touchscreen, touchpad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc., output devices 707 including a liquid crystal display (LCD), speaker, vibrator, etc., storage devices 708 including magnetic tape, hard disk, etc., and communication devices 709 may be connected to the I / O interface 705. The communication devices 709 may enable the electronic device 700 to communicate wirelessly or via wires to exchange data with other devices. While the figures show the electronic device 700 with various devices, it should be understood that it need not implement or include all of the illustrated devices. It may alternatively implement or include more or fewer devices.
[0132] In particular, according to embodiments of the present invention, the processes described with reference to the flowcharts above may be implemented as a computer software program. For example, embodiments of the present invention include a computer program product including a computer program carried on a computer-readable medium, the computer program including program code for performing the methods illustrated in the flowcharts. In such embodiments, the computer program may be downloaded and installed from a network via the communication device 709, or may be installed from the storage device 708 or the ROM 702. When the computer program is executed by the processing device 701, the functions described above, which are specific to the methods of the embodiments of the present invention, are performed.
[0133] It should be noted that the above-mentioned computer-readable medium of the present invention may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the above. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of the computer-readable storage medium may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be a tangible medium that contains or stores a program that can be used by or in combination with the program, instruction execution system, apparatus, or device. In addition, in the present invention, the computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier carrying computer-readable program code. Such propagated data signals may take a variety of forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium, other than a computer-readable storage medium, that transmits, propagates, or transmits a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained in the computer-readable medium may be transmitted using any suitable medium, including, but not limited to, wire, fiber optic cable, RF (radio frequency), or any suitable combination of the above.
[0134] In some embodiments, clients and servers may communicate using any now known or later developed network protocol, such as HTTP (HyperText Transfer Protocol), and may be connected to each other via any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), the Internet (e.g., the Internet), and an end-to-end network (e.g., an ad hoc end-to-end network), and any now known or later developed network.
[0135] The computer readable medium described above may be included in the electronic device described above, or may exist separately and not be assembled to said electronic device.
[0136] The computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to obtain subtitle information of a multimedia conference, where the subtitle information is obtained based on converting audio information during the multimedia conference; determine annotated content from the subtitle information; obtain annotation information for the annotated content; display the subtitle information, and display the annotation information corresponding to the annotated content.
[0137] Computer program code for carrying out the operations of the present invention can be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and the like, as well as conventional procedural programming languages such as "C" or similar programming languages. The program code may run entirely on the user computer, partially on the user computer, as a standalone software package, partially on the user computer and partially on a remote computer, or entirely on a remote computer or server. When referring to a remote computer, the remote computer may be connected to the user computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet Service Provider).
[0138] The flowcharts and block diagrams in the figures illustrate possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. From this perspective, each box in a flowchart or block diagram contains one or more executable instructions for implementing the specified logical function(s) and may represent a module, program segment, or portion of code. It should also be noted that in some alternative implementations, the functions depicted in the boxes may occur in a different order than that depicted in the figures. For example, two boxes shown in succession may actually be executed substantially in parallel, or may be executed in the reverse order, depending on the functionality involved. It should also be noted that each box in the block diagrams and / or flowcharts, and combinations of boxes in the block diagrams and / or flowcharts, may be implemented in a dedicated hardware-based system that performs the specified function(s) or operation(s), or may be implemented as a combination of dedicated hardware and computer instructions.
[0139] The modules described as being involved in the embodiments of the present invention may be implemented by software or hardware, and the names of the units do not constitute limitations on the units themselves in a particular case, for example, the first acquisition unit may be described as a "unit for acquiring subtitle information."
[0140] The functions described herein may be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), etc.
[0141] In the context of the present invention, a machine-readable medium may be a tangible medium that can contain or store a program that can be used by or in combination with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination thereof. More specific examples of machine-readable storage media include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0142] The above description is merely a preferred embodiment of the present invention and merely describes the applied technical principles. Those skilled in the art should understand that the scope of the present invention is not limited to the technical solution formed by a specific combination of the above technical features, but also includes other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the concept of the present invention. For example, it should be understood that the present invention also includes technical solutions formed by substituting the above features with technical features having similar functions disclosed in the present invention (but not limited to these).
[0143] It should be noted that although operations are depicted in a particular order, this should not be understood as requiring that these operations be performed in the particular order or sequence shown. Multitasking and parallel processing may be advantageous in certain environments. Similarly, although the above discussion includes some specific implementation details, these should not be construed as limiting the scope of the invention. Certain features that are described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment may also be implemented in multiple embodiments individually or in any suitable subcombination.
[0144] 18. Although the present subject matter has been described in language specific to structural features and / or logical operations of a method, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or operations described above. Rather, the specific features and operations described above are merely example forms of implementing the claims.
Claims
1. A presentation method performed by an electronic device, comprising: Obtaining subtitle information of a multimedia conference, wherein the subtitle information is obtained based on converting audio information during the multimedia conference; determining annotated content from the subtitle information; acquiring annotation information for the annotation target content; displaying subtitle information and annotation information corresponding to the annotated content; For the same annotated content that occurs at least two times during the multimedia conference, determining an interval between a non-first occurrence of the annotated content and a previous occurrence of the annotated content for which annotation information was presented; and displaying annotation information corresponding to the non-first occurrence of the annotated content in response to the interval satisfying a predetermined interval condition. A presentation method characterized by:
2. Determining the content to be annotated from the subtitle information includes: determining a user group thesaurus corresponding to a user group based on a user group to which a conference participant of the multimedia conference belongs, the user group thesaurus including phrases and phrase interpretations; determining annotated content from the subtitle information based on the user group thesaurus; The presentation method according to claim 1 .
3. determining annotated content from the subtitle information based on the user group thesaurus, selecting, from among the words of the subtitle information, words of phrases included in the user group thesaurus that have predefined characteristics as the annotated content; The presentation method according to claim 2 .
4. the predefined characteristics include the phrase being an abbreviated phrase, and the annotation information includes the phrase's full name; Displaying the subtitle information and the annotation information corresponding to the annotation target content includes: displaying a phrase full name corresponding to the annotated content; The presentation method according to claim 3 .
5. Determining the content to be annotated from the subtitle information includes: determining a first language for the multimedia conference, the first language being determined based on at least one of a geographic region in which conferees of the multimedia conference are located, a native language of the conferees of the multimedia conference, and a predetermined language for the multimedia conference; determining annotated content from words in a second language in the subtitle information, the second language being a language other than the first language; The presentation method according to claim 1 .
6. determining the annotation target content from the second language words in the subtitle information, selecting second language words included in a predetermined second language thesaurus from the second language words of the subtitle information as content to be annotated, the second language thesaurus including second language phrases and corresponding first language interpretations; The presentation method according to claim 5 .
7. the first language interpretation includes at least two sub-interpretations; Displaying the subtitle information and the annotation information corresponding to the annotation target content includes: selecting a sub-interpretation correlated with context information from the first language interpretation corresponding to the annotated content based on context information in the subtitle information in which the annotated content is located; and displaying the selected sub-explanation corresponding to the annotated content. The presentation method according to claim 6 .
8. Determining the content to be annotated from the subtitle information includes: selecting words included in a predetermined rare word thesaurus from the words of the subtitle information as content to be annotated; The presentation method according to claim 1 .
9. Displaying the subtitle information and the annotation information corresponding to the annotation target content includes: and displaying the annotated content of the displayed subtitle information in association with the corresponding annotation information. The presentation method according to claim 1 .
10. The display of the annotation target content in the displayed subtitle information in association with the corresponding annotation information includes: There is at least one difference in the presentation style of both the annotated content and the displayed subtitles other than the annotated content; The annotation information presentation area does not overlap with the subtitle presentation area. The annotation information presentation area is embedded within the subtitle presentation area. presenting annotation information corresponding to the annotated content in response to a trigger operation on the annotated content; presenting the annotation information in the form of a floating window in association with the corresponding annotated content; during the real-time multimedia conference, a first presentation duration of the annotated content of the closed caption information is less than or equal to a second presentation duration of the corresponding annotation information; The presentation method according to claim 9 .
11. the intervals include time intervals, the time intervals indicating intervals between audio information corresponding to the annotated content; The presentation method includes: and determining that a time interval between a non-first appearance of the annotated content and a previous corresponding annotated content for which annotation information was presented is greater than a predetermined duration threshold, the time interval satisfying the predetermined interval condition. The presentation method according to claim 1 .
12. the spacing includes a text spacing, which is used to indicate a positional spacing between the same annotated content in the subtitle information; The presentation method includes: and determining that a text interval between a non-first appearance of the annotated content and a previously corresponding annotated content for which annotation information has been presented is greater than a predetermined text length threshold, so that the text interval satisfies a predetermined interval condition. The presentation method according to claim 1 .
13. the text length threshold includes a line gap between lines of text located in the subtitle information being greater than a predetermined line gap; The presentation method of claim 12 .
14. The multimedia conference is a real-time multimedia conference currently being held. The presentation method according to claim 1 .
15. A presentation device, a first acquiring unit for acquiring subtitle information of a multimedia conference, the subtitle information being obtained based on converting audio information during the multimedia conference; a determination unit for determining an annotation target content from the subtitle information; a second acquisition unit for acquiring annotation information of the annotation target content; a display unit that displays subtitle information and annotation information corresponding to the annotation target content; The presentation device is For the same annotated content that occurs at least two times during the multimedia conference, determining an interval between a non-first occurrence of the annotated content and a previous corresponding occurrence of the annotated content for which annotation information was presented; and displaying annotation information corresponding to the non-first occurrence of the annotated content in response to the interval satisfying a predetermined interval condition. A presentation device characterized by:
16. one or more processors; a storage device for storing one or more programs; The one or more programs, when executed by the one or more processors, cause the one or more processors to implement the method of any one of claims 1 to 14. An electronic device characterized by:
17. A computer-readable medium having a computer program stored thereon, When the computer program is executed by a processor, the method according to any one of claims 1 to 14 is realized.
10. A computer-readable medium comprising:
Citation Information
Patent Citations
Method and device for editing caption of video
CN108650543A
Data processing method and device, electronic equipment and storage medium
CN111161737A
Conference file generation method and device and electronic equipment
CN112084756A
Autonomous collaboration agent for meetings
US20170154264A1
Persisting annotations applied to an electronic hosted whiteboard
US20170262419A1