Interaction methods, devices, electronic devices, storage media, and computer programs

The interaction method and device enhance information acquisition from video and audio by displaying text sentences alongside audio, addressing inefficiencies in existing methods and improving user experience.

JP7861165B2Active Publication Date: 2026-05-18BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2024571844
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-08-15
Filing Date
2023-08-14
Publication Date
2026-05-18
Estimated Expiration
2043-08-14

AI Technical Summary

Technical Problem

Existing methods for obtaining valuable information from video and audio are limited and inefficient.

Method used

An interaction method and device that displays a list of text sentences from video content alongside audio, allowing users to access information through text display operations, including speech recognition and interactive controls.

Benefits of technology

Enriches the methods of information acquisition from video and audio, enhancing user experience by providing alternative means to access information, especially in noisy environments or for those with hearing impairments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007861165000001
    Figure 0007861165000001
  • Figure 0007861165000002
    Figure 0007861165000002
  • Figure 0007861165000003
    Figure 0007861165000003
Patent Text Reader

Abstract

An interaction method, apparatus, electronic device, and storage medium. The method includes receiving a text display operation on a first media content including video content (S101), and in response to the text display operation, displaying a text sentence list of the first media content in a predetermined area, where the text sentence list includes at least two text sentences, and for each of the at least two text sentences, there is a corresponding audio sentence in the target audio data of the first media content (S102).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] [Cross-reference to Related Applications] This application claims the priority of a Chinese patent application with the application number 202210977564.5, which was filed with the China National Intellectual Property Administration on August 15, 2022, and all the contents of the said application are incorporated herein by reference.

[0002] This disclosure relates to the field of computer technology, for example, interaction methods, devices, electronic devices 、 storage media , and computer programs and relates to.

Background Art

[0003] Users can obtain valuable information in video and audio by listening to it. However, the methods for obtaining valuable information in video and audio are relatively single, and the efficiency of information acquisition is low.

Summary of the Invention

Problems to be Solved by the Invention

[0004] This disclosure provides an interaction method, device, electronic device, and storage media that enrich the acquisition form of valuable information in video and audio and improve the acquisition efficiency of valuable information in video and audio.

Means for Solving the Problems

[0005] Embodiments of this disclosure include receiving a text display operation on a first media content including video content, and in response to the text display operation, displaying a list of text sentences of the first media content in a predetermined area, where the list of text sentences includes at least two text sentences, and the at least two text sentences respectively have corresponding audio sentences in the target audio data of the first media content, and provides an interaction method.

[0006] The embodiments of this disclosure are An operation receiving module configured to receive text display operations for a first media content including video content, The interaction device further provides a list display module configured to display a list of text statements of the first media content in a predetermined area in response to the text display operation, wherein the list of text statements includes at least two text statements, and each of the at least two text statements has a corresponding audio statement in the target audio data of the first media content.

[0007] The embodiments of this disclosure are One or more processors, A memory configured to store one or more programs, The present invention further provides an electronic device that, when the one or more programs are executed by the one or more processors, causes the one or more processors to implement the interaction method described in any embodiment of the present disclosure.

[0008] Embodiments of the present disclosure further provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the interaction method described in the embodiment of the present disclosure. [Brief explanation of the drawing]

[0009] [Figure 1] This is a schematic flowchart of the interaction method provided in the embodiments of this disclosure. [Figure 2] This is a schematic diagram illustrating the display of a text list provided in the embodiments of this disclosure. [Figure 3] This is a schematic diagram illustrating the display of the position control provided in the embodiments of this disclosure. [Figure 4] This is a schematic diagram of the text display control provided in the embodiments of this disclosure. [Figure 5] This is a schematic diagram of the display of the media content jump control provided in the embodiments of this disclosure. [Figure 6] This is a schematic flowchart of another interaction method provided by the embodiments of this disclosure. [Figure 7] This is a schematic diagram illustrating the display of a first text sentence provided by the embodiments of this disclosure. [Figure 8] This is a schematic diagram of the editing page for the second media content provided in the embodiments of this disclosure. [Figure 9] This is a block diagram of the configuration of an interaction device provided in the embodiments of this disclosure. [Figure 10] This is a schematic diagram of the configuration of an electronic device provided in the embodiments of this disclosure. [Modes for carrying out the invention]

[0010] The embodiments of this disclosure will be described below with reference to the drawings. While the drawings show several embodiments of this disclosure, this disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein; rather, these embodiments are provided for the purpose of understanding this disclosure. The drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0011] The steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. The method embodiments may also include additional steps and / or omit the performance of the indicated steps. The scope of this disclosure is not limited thereto.

[0012] As used herein, the term "comprising" and its variations are open-ended, meaning "including, but not limited to." The term "based on" means "at least partially based on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one another embodiment," and the term "some embodiments" means "at least some embodiments." Related definitions of other terms are given in the following description.

[0013] The concepts such as "first," "second," etc. mentioned in this disclosure are only for distinguishing different devices, modules or units, and are not for limiting the order or interdependence of the functions executed by these devices, modules or units.

[0014] The terms "one" and "a plurality" mentioned in this disclosure are illustrative rather than restrictive, and should be understood as "one or a plurality" unless clearly indicated otherwise in the context.

[0015] The names of the messages or information that interact between multiple devices in the embodiments of this disclosure are only for the purpose of explanation, and are not for limiting the scope of these messages or information.

[0016] Before using the technical solutions disclosed in multiple embodiments of this disclosure, the user should be informed in an appropriate form about the types, scope of use, usage scenarios, etc. of the personal information related to this disclosure in accordance with relevant laws and regulations, and the approval of the user should be obtained.

[0017] For example, when responding to an active request from a user, prompt information is sent to the user to explicitly prompt that the operation requesting execution needs to obtain and use the user's personal information. Thereby, the user can autonomously select whether to provide personal information to software or hardware such as an electronic device, an application, a server, or a storage medium that executes the operation of the technical solution of the present disclosure according to the prompt information.

[0018] As an optional but non-limiting implementation form, in response to receiving an active request from a user, the form of sending prompt information to the user may be, for example, in the form of a pop-up window, and the prompt information may be presented in the form of text in the pop-up window. Further, the pop-up window may be equipped with a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0019] <了 The above-described notification and user approval acquisition procedures are merely schematic and do not limit the implementation forms of the present disclosure. Other forms that comply with relevant laws and regulations can also be applied to the implementation forms of the present disclosure.

[0020] FIG. 1 is a schematic flowchart of an interaction method provided by an embodiment of the present disclosure. The method can be executed by an interaction device, where the device can be implemented by software and / or hardware and can be configured as an electronic device, such as a mobile phone or a tablet computer. The interaction method provided by the embodiment of the present disclosure is applicable to a scenario of checking a text list of a video, for example, a scenario of checking a text list of a video while watching the video. As shown in FIG. 1, the interaction method provided by this embodiment includes the following.

[0021] S101: Receive a text display operation on the first media content including video content.

[0022] Among these, the text display operation may be, for example, an operation that triggers the text display control of one media content, an operation that triggers the media content jump control of another media content (e.g., a third media content) generated based on the text in one media content, or a trigger operation that instructs the display of a text list of media content, such as a gesture operation to instruct the display of a text list of other media content. The first media content may be media content whose text list display is instructed by the text display operation, and such media content may be, for example, video content or graphic text content, and this embodiment is not limited to such types. Optionally, the first media content may be video content that includes audio data other than background music (e.g., human voices), and below, the example will be given that the first media content is video content.

[0023] For example, the system may receive text display operations on the first media content, such as when the first media content is displayed and the user triggers a text display control for the first media content, or when other media content generated based on the text in the first media content is displayed and the user performs a media content jump operation, such as triggering a media content jump control corresponding to a third media content.

[0024] S102: In response to the text display operation, a list of text statements of the first media content is displayed in a predetermined area, the list of text statements includes at least two text statements, and each text statement has a corresponding audio statement in the target audio data of the first media content.

[0025] The predetermined area may be an area for displaying a list of text sentences from the first media content, and the position and size of the predetermined area may be flexibly set as needed. The target audio data may be audio data from the first media content, for example, human voice data other than background music (e.g., valid human voice data) from the audio data of the first media content.

[0026] The text list may be a list for displaying text sentences corresponding to at least one audio sentence in the target audio data of the first media content, and the text list may contain text sentences corresponding to audio sentences in the target audio data, for example, text sentences that have a one-to-one correspondence with audio sentences in the target audio data. If there are at least two audio sentences in the target audio data, the text list may contain at least two text sentences. If there is only one audio sentence in the target audio data, the text list may contain only one text sentence. Each text list may be arranged so that the audio sentences corresponding to the text list are in chronological order in the target audio data.

[0027] The text in the text list may be recognized and input by the distributor of the first media content listening to the target audio data, for example, by speech recognition of the target audio data, or for example, after the distribution of the first media content is successful, speech recognition of the target audio data of the first media content may be performed to obtain the text list of the first media content. The text in the text list may be divided based on predetermined punctuation marks in the text list, for example, the content between two adjacent predetermined punctuation marks in the text list may be treated as a single text. The predetermined punctuation marks may be set as needed and may include, for example, commas, semicolons, periods, etc.

[0028] Whether or not the text list is displayed depends on whether or not subtitles exist for the first media content. In other words, if subtitles exist for the first media content, the text list for the first media content may be displayed. For example, the text list for the first media content may be displayed simultaneously with the first media content and the currently playing subtitles for the first media content. The text list for the first media content may also be displayed even if subtitles do not exist for the first media content. The text list referred to in this embodiment differs from subtitles. When subtitles are displayed, usually only the subtitles corresponding to the currently playing audio sentences are displayed, and not all subtitle content for the video content is displayed. However, the text list displayed in this embodiment may include text corresponding to at least one audio sentence in the target audio data. Therefore, through this text list, the user can check not only the currently playing text, but also text that has already been played and text that has not yet been played.

[0029] For example, when a text display operation is received for the first media content, the text list 20 of the first media content may be displayed in a predetermined area, as shown in Figure 2. For example, the text list 20 of the first media content may be displayed simultaneously with the playback of the first media content. For instance, the first media content may be played in the media content playback area of ​​the media content display page, and the text list 20 of the first media content may be displayed in a predetermined area of ​​the media content display page. Here, the media content playback area and the predetermined area may or may not overlap, and this embodiment is not limited to this.

[0030] Furthermore, when displaying a text list of the first media content in a predetermined area, pre-configured interaction controls for the first media content may also be displayed in the predetermined area. For example, the predetermined area may display a distributor marker for the first media content (e.g., an avatar), a like control, a comment control, a favorite control, or a share control for the first media content. Thus, by triggering the distributor marker for the first media content, the user can check the profile of the distributor of the first media content; by triggering the like control, the user can like the first media content; by triggering the comment control, the user can check the comment information for the first media content or comment on the first media content, instructing the current application to display the comment panel for the first media content; by triggering the favorite control, the user can add the first media content to their favorites list; or by triggering the share control, the user can share the first media content, instructing the current application to display the share panel for the first media content. When a user does not intend to check the text list of the first media content, they may instruct the current application to stop displaying the text list of the first media content by a corresponding trigger operation, which may include, for example, triggering a close control displayed in a given area, clicking the media content playback area, and / or continuing to slide down in the given area when the first text in the text area list is displayed.

[0031] In this embodiment, by displaying a list of text snippets for the video content, users can obtain information from the video (e.g., video audio) not only by watching the video screen or listening to the video audio, but also by checking the list of text snippets for the video content. Therefore, even in situations where it is inconvenient for the user to listen to the video audio, such as when it is difficult to turn on the video audio, when the user is in a noisy environment, or when the user has a hearing impairment, the information contained in the video audio can still be quickly obtained. Compared to conventional technologies in which users obtain information from a video only by watching the video screen or listening to the video audio, this method enriches the ways in which information from the video (e.g., video audio) can be obtained, improves the efficiency of information acquisition from the video, and enhances the user experience.

[0032] In this embodiment, referring to Figure 2 again, when displaying the text list 20 of the first media content, the current text 21 in the text list 20 (shown in Figure 2 as "These are the things we need to seriously consider") can be automatically moved and displayed within a predetermined area. The current text 21 may have a different display configuration from other texts in the text list 20 to make it easier for the user to clearly identify and check the current text, for example, by having a different font, font size and / or color.

[0033] In this case, optionally, the at least two text statements include the current text statement and other text statements other than the current text statement that have a different display mode from the current text statement, wherein the current text statement is the text statement currently being played.

[0034] In this context, the current text may be the text corresponding to the audio text currently being played in the target audio data of the first media content. This current text may change as the first media content is played.

[0035] In this embodiment, if a search keyword 22 is included, the search keyword 22 and the content other than the search keyword 22 in the text list 20 may be displayed in different ways in the text list 20. For example, the search keyword 22 and the content other than the search keyword 22 in the text list 20 may be displayed in different fonts, font sizes, and / or colors. For example, to make it easier for the user to identify the search keyword 22 in the text list 20 and search for it quickly, a search tag (a magnifying glass icon displayed to the upper right of the search keyword "human capital" as shown in Figure 2) may be added to the search keyword 22.

[0036] In this case, optionally, the search-waiting keywords included in the at least two text sentences are displayed in a different manner from other content in the at least two text sentences other than the search-waiting keywords, and the search-waiting keywords are used to trigger the display of search results that match the triggered search-waiting keywords.

[0037] Among these, the keywords awaiting search may be determined based on pre-set decision rules, for example, candidate keywords may be manually marked in advance, and / or candidate keywords may be determined in advance by the search frequency of multiple keywords in the current application or by a candidate keyword determination model, and candidate keywords contained in at least one text sentence of the text sentence list of the first media content will be used as keywords awaiting search in the text sentence list.

[0038] Therefore, when a user triggers a search for a keyword in a list of text sentences, the system may retrieve and display search results that match that keyword. For example, it may retrieve search results that match the keyword from the server and display them on the current page (e.g., a designated area), or it may switch the current page from a media content display page to a search results page and display the search results on the search results page.

[0039] In one embodiment, when the currently playing sentence is changed, the display position and / or display mode of the currently playing sentence within a predetermined area before and after the change may be automatically adjusted. For example, before the currently playing sentence is changed, the currently playing sentence before the change may be displayed in a first display mode at a set position in the predetermined area, and sentences not currently playing in the text sentence list may be displayed in a second display mode. When the currently playing sentence is changed, at least one text sentence in the text sentence list (including the currently playing sentence before the change) may be controlled to move so that the currently playing sentence after the change is moved to the set position in the predetermined area and displayed. The currently playing sentence before the change may be switched from the first display mode to the second display mode, and the currently playing sentence after the change may be switched from the second display mode to the first display mode. In this case, optionally, the current text sentence is displayed at the set position in the predetermined area, and the interaction method provided in this embodiment further includes moving the current text sentence after the change to the set position and displaying it when the current text sentence is changed.

[0040] In one embodiment, after displaying the text list of the first media content in a predetermined area, the system further includes displaying a position control to trigger a switch in the text displayed in the predetermined area in response to a word switching operation performed within the predetermined area, thereby moving the current text to a set position in the predetermined area for display.

[0041] Among these, the word switching operation may be, for example, an operation that triggers a text switching control, or a slide operation that acts within a predetermined area, which is a trigger operation that switches the text displayed in a predetermined area. The position control may be a control for triggering the movement and display of the current text at a set position in the predetermined area, and the position control may be displayed when a word switching operation is received, or when a word switching operation is received and the currently playing text is not displayed within the predetermined area, or it may be set as needed.

[0042] For example, when a word switching operation is received, for instance, when it is detected that the user is sliding vertically within a predetermined area, the text displayed in the predetermined area may be switched. For example, in order to move a text that is not displayed in the predetermined area into the predetermined area and display it, at least one text displayed in the predetermined area may be controlled to move in the direction of the user's slide. Then, as shown in Figure 3, when a word switching operation is received and / or when the currently playing sentence is not displayed in the predetermined area, the position control 30 may be displayed. Therefore, the user may trigger the position control 30 when trying to check the currently playing sentence. Accordingly, when the current application program detects that the user has triggered the position control 30, it may control the movement of multiple texts displayed in the predetermined area so that the currently playing sentence is moved to a set position in the predetermined area and displayed.

[0043] In the above-described embodiment, if the currently playing sentence is changed after a word switching operation, the changed sentence may still be automatically moved to a predetermined location within the set area and displayed.

[0044] For example, to satisfy the user's need to check the text displayed in the designated area after a word switching operation, the currently playing text after the change may not be automatically moved to the designated area's set position for display. In this case, when it is detected that the user has triggered a position control, or when a trigger operation is received to switch the playback progress of the first media content (for example, when a playback operation is received for a single text in a text text list), the system may be restored to automatically move the currently playing text after the change to the designated area's set position for display. That is, when the currently playing text is changed, the currently playing text after the change is automatically moved to the designated area's set position for display.

[0045] In one embodiment, before receiving a text display operation for the first media content, the first media content is further played, and displaying a list of text sentences of the first media content in a predetermined area includes displaying a list of text sentences of the first media content in a predetermined area and adjusting the first media content outside the predetermined area for playback.

[0046] In the embodiments described above, the user may perform a text display operation on the first media content when viewing the first media content.

[0047] As shown in Figure 4, the current application may play the first media content in the media content playback area of ​​the media content display page. Therefore, when a user attempts to check the text list 20 of the first media content, they may perform a text display operation on the first media content, for example, by triggering the text display control 40 of the first media content. Accordingly, when the current application program receives a text display operation on the first media content, as shown in Figure 2, it may display the text list 20 of the first media content in a predetermined area of ​​the media content display page, adjust the display range of the media content playback area according to the display range of the predetermined area, and adjust the first media content to play outside the predetermined area.

[0048] In addition, the display position of the text display control 40 can be flexibly set. For example, when playing the first media content, the first media content may be displayed at a predetermined position on the media content display page, as shown in Figure 4, or the text display control 40 may be displayed on a predetermined panel of the first media content. Therefore, when a user attempts to perform a text display operation, they may directly trigger the text display control 40 in the first media content, or they may perform a panel display operation to instruct the current application to display a predetermined panel of the first media content, and then trigger the text display control 40 displayed on that predetermined panel.

[0049] Furthermore, a text display operation for the first media content may further include a mute playback operation for the first media content. For example, if it is detected that the user has adjusted the playback volume of the first media content to mute, it can be determined that a text display operation for the first media content has been received. A list of text statements for the first media content is then displayed in a predetermined area of ​​the media content display page, and the first media content is adjusted to play outside the predetermined area to facilitate the user's acquisition of information in the target audio data of the first media content.

[0050] In another embodiment, the text display operation includes a media content jump operation to a third media content which is media content generated based on a second text sentence in a first media content, and further includes playing the third media content before receiving the text display operation for the first media content, and displaying the text sentence list of the first media content in a predetermined area includes playing the first media content outside the predetermined area and displaying the text sentence list of the first media content in the predetermined area.

[0051] In the embodiments described above, the user may perform a text display operation on the first media content when viewing the third media content generated based on the first media content.

[0052] The third media content may be media content generated based on the second text in the first media content, for example, media content generated by a media content generation operation on the second text, and the third media content may be video content or graphic text content, for example, graphic text content. The second text may include one or more texts in the text list of the first media content, and may be the same as or different from the first text. The third media content may be media content delivered by the current user or another user viewing the first media content. A media content jump operation may be a trigger operation to instruct a jump to the original media content (e.g., the first media content) corresponding to the media content currently being played (e.g., the third media content), for example, if the media content currently being played is media content generated based on text in another media content, the media content jump operation may be a trigger operation to jump to that other media content, for example, an operation that triggers the media content jump control of the media content currently being played.

[0053] As shown in Figure 5, the current application may display a third media content generated based on a second text statement in the first media content. Therefore, when a user attempts to check the first media content and / or the text statement list of the first media content, they may perform a text display operation on the first media content, for example, by triggering a media content jump control 50 corresponding to the third media content. Accordingly, when the current application receives a text display operation on the first media content, as shown in Figure 2, it may play the first media content outside a predetermined area and display the text statement list 20 of the first media content in a predetermined area, for example, by switching the current page to the media content display page of the first media content, playing the first media content outside a predetermined area on the media content display page, and then displaying the text statement list 20 of the first media content in a predetermined area on the media content display page.

[0054] In the embodiments described above, the timing of displaying the media content jump control corresponding to the third media content may be flexibly set. For example, the media content jump control corresponding to the third media content may be displayed when the third media content is displayed. Alternatively, the media content jump control corresponding to the third media content may be displayed after receiving a control display operation for the third media content. In this case, optionally, the media content jump control corresponding to the third media content may be displayed in response to the control display operation for the third media content, before receiving a text display operation for the first media content, to trigger the execution of the media content jump operation. The control display operation for the third media content may be a trigger operation to instruct the display of the media content jump control corresponding to the third media content. For example, it may be a display pause operation for the third media content. In this case, exemplary, upon receiving a display pause operation for the third media content, the display of the third media content may be paused, and the media content jump control corresponding to the third media content may be displayed.

[0055] In the embodiments described above, playing the first media content outside the predetermined area may include playing the first media content outside the predetermined area with the start point of the first media content as the playback start point, or playing the first media content outside the predetermined area with the time node corresponding to the start point of the second text sentence in the first media content as the playback start point. That is, when playing the first media content outside the predetermined area in response to a media content jump operation for the third media content, the first media content may be played with the start point of the first media content as the playback start point, or the first media content may be played with the time node corresponding to the start point of the second text sentence in the first media content as the playback start point, and this may be set as needed.

[0056] The interaction method provided in this embodiment receives a text display operation on a first media content including video content, and in response to the text display operation, displays a list of text sentences of the first media content in a predetermined area. The list of text sentences includes at least two text sentences, and each text sentence has a corresponding audio sentence in the target audio data of the first media content. By adopting the above-described technical proposal, this embodiment allows users to obtain information from a video not only by viewing the video screen or listening to the video audio, but also by checking the list of text sentences of the video content. This enriches the ways in which information can be obtained from a video, improves the efficiency of obtaining information from a video, and enhances the user experience.

[0057] Figure 6 is a schematic flowchart of another interaction method provided by an embodiment of the present disclosure. The technical proposals in this embodiment may be combined with one or more optional technical proposals in the embodiments described above. Optionally, this further includes, after displaying a list of text sentences of the first media content in a predetermined area, identifying the first text sentence selected by a word selection operation in response to a word selection operation acting within the predetermined area, copying the first text sentence in response to a copy operation on the first text sentence, sharing the first text sentence in response to a share operation on the first text sentence, playing the first media content in response to a play operation on the first text sentence, with the time node corresponding to the start point of the first text sentence in the first media content as the playback start point, and generating second media content containing the first text sentence in response to a media content generation operation on the first text sentence.

[0058] Optionally, the system further includes identifying the first text sentence selected by the word selection operation and then displaying the first text sentence in a selected state.

[0059] Accordingly, as shown in Figure 6, the interaction method provided by this embodiment includes the following steps.

[0060] S201: Receives a text display operation on a first media content, including video content.

[0061] S202: In response to the text display operation, a list of text statements of the first media content is displayed in a predetermined area, the list of text statements containing at least two text statements, the text statements having corresponding audio statements in the target audio data of the first media content.

[0062] S203: In response to a word selection operation performed within the predetermined area, the first text sentence selected by the word selection operation is identified, the first text sentence is displayed as selected, and at least one of S204 to S207 is executed.

[0063] Among these, the word selection operation may be an operation to select one or more text sentences in a list of text sentences of the first media content, for example, a click operation performed within a predetermined area. The first text sentence may be the text sentence selected by the word selection operation, and may include one or more text sentences.

[0064] For example, when a word selection operation is performed within a predetermined area, the first text sentence selected by the word selection operation can be identified, and the first text sentence can be displayed as selected. As shown in Figure 7, a copy control 70, a sharing control (not shown in Figure 7), a playback control 71, and / or a media content generation control 72, etc., corresponding to the first text sentence may be displayed.

[0065] For example, when a click operation is received within a predetermined area, the text displayed at the trigger position of the click operation is designated as the first text operation corresponding to the click operation, and the first text operation is displayed in a selected state. Copy controls, sharing controls, playback controls, and / or media content generation controls corresponding to the first text operation may then be displayed. Furthermore, the user can adjust the selected range by adjusting the position of the start point indicator and / or end point indicator of the selected first text operation, thereby adjusting the content contained within the first text operation.

[0066] S204: In response to the copy operation on the first text statement, the first text statement is copied.

[0067] For example, when a copy operation is received on the first text, the first text may be copied, or prompt information may be displayed to indicate to the user that the copy of the first text was successful. The user can then input the first text into another page of the current application or another application by pasting. The copy operation on the first text may also be a trigger operation that instructs the copying of the first text, for example, an operation that triggers a copy control 70 (shown in Figure 7) corresponding to the first text.

[0068] S205: In response to the sharing operation for the first text statement, the first text statement is shared.

[0069] For example, upon receiving a share operation on a first text document, the first text document may be shared. For instance, a sharing panel may be displayed for the current user to select other users to whom the first text document will be shared, and after the current user has made their selection, the first text document may be shared with the other users selected by the current user. Alternatively, the first text document may be directly shared with a target user pre-associated with the current user. In this scenario, a copy operation on the first text document may be a trigger operation to instruct the display of a sharing panel for sharing the first text document, or a trigger operation to instruct the sharing of the first text document, such as an operation that triggers a sharing control corresponding to the first text document.

[0070] S206: In response to the playback operation on the first text sentence, the first media content is played back, with the time node corresponding to the start point of the first text sentence in the first media content as the playback start point.

[0071] For example, when a playback operation is received for the first text sentence, the playback progress of the first media content may be adjusted to the playback progress indicated by the time node in the first media content corresponding to the start point of the first text sentence. That is, in order to make it easier for the user to listen to the audio sentence corresponding to the first text sentence, the first media content may be played with the time node in the first media content corresponding to the start point of the first text sentence as the playback start point. Among these, the playback operation for the first text sentence may be a trigger operation that instructs the playback of the audio sentence corresponding to the first text sentence, for example, an operation that triggers the playback control 71 (shown in Figure 7) corresponding to the first text sentence.

[0072] S207: In response to a media content generation operation for the first text statement, a second media content containing the first text statement is generated.

[0073] Among these, the media content generation operation for the first text sentence may be a trigger operation to instruct the generation of new media content using the first text sentence, for example, an operation that triggers a media content generation control 72 (shown in Figure 7) corresponding to the first text sentence. The second media content may be media content generated using the first text sentence, and the second media content may be video content or graphic text content, for example, graphic text content. The graphic text content may be understood as media content that uses a picture as content and text as introductory information for the media content.

[0074] For example, when a media content generation operation is performed on a first text sentence, a second media content containing the first text sentence may be generated.

[0075] Taking the second media content as graphic text content, when generating the second media content containing the first text, a card containing the second text may be generated, the card may be used as the content of the media content, and the music corresponding to the second text and / or the card may be used as background music for the media content, and the second media content may be generated. In this case, optionally, generating the second media content containing the first text includes generating a target card containing the first text, using the target card as content, and generating the second media content using the music corresponding to the first text and / or the target card as background music.

[0076] The target card may be a picture containing a first text sentence. The music corresponding to the first text sentence / target card may be music associated with the first text sentence / target card, and the music associated with multiple text sentences / cards may be determined by a pre-trained model or may be pre-set. The target card may be determined randomly or selected from a plurality of pre-set cards based on the first text sentence, and different cards may have different display modes, but this embodiment is not limited thereto.

[0077] In this embodiment, regardless of the length of the first text, when a media content generation operation is received for the first text, a second media content including the first text may be generated. Alternatively, the length of the first text may be taken into consideration, and only when the length of the first text is within a predetermined length range (e.g., less than 150 characters), when a media content generation operation is received for the first text, a second media content including the first text may be generated. However, if the length of the first text is outside the predetermined length range, the second media content may not be generated based on the first text. For example, if the length of the first text is outside the predetermined length range, the media content generation control corresponding to the first text may not be displayed, and this may be flexibly configured as needed.

[0078] In this embodiment, derivative works may be created based on the text in the text list of the first media content to generate new media content. Thus, a new method for creating media content can be provided, reducing the difficulty of creating media content and improving the user's creative experience.

[0079] In one embodiment, in response to a media content generation operation for the first text sentence, a second media content containing the first text sentence may be generated and delivered. For example, the media content generation operation may be a trigger operation to instruct the generation and delivery of the second media content, and may include, for example, a delivery operation for the second media content containing the first text sentence. Thus, upon receiving a media content generation operation for the first text sentence, the second media content containing the first text sentence may be generated and the second media content delivered. In this case, optionally, generating the second media content containing the first text sentence includes generating the second media content containing the first text sentence and delivering the second media content.

[0080] In another embodiment, upon receiving a media content generation operation for the first text statement, only the generation of a second media content containing the first text statement may be performed, and the second media content may be distributed after receiving a distribution operation for the second media content. For example, upon receiving a media content generation operation for the first text statement, the second media content containing the first text statement may be generated, and as shown in Figure 8, an editing page for the second media content may be displayed for the user to edit the second media content, and the second media content may be distributed after receiving a distribution operation acting on the editing page of the second media content (for example, an operation that triggers the distribution control 80 on the editing page of the second media content) or a distribution operation acting within the distribution page of the second media content. In this case, optionally, the system further includes distributing the second media content in response to a distribution operation for the second media content after generating the second media content containing the first text statement.

[0081] The display of the text list of the first media content and the creation of derivative works based on the text in the text list of the first media content are both permitted only with the permission of the distributor of the first media content. For example, the distributor of the first media content may, on the settings page, turn on or off permission to display the text list or to create derivative works for all media content they distribute (including the first media content), or they may turn on or off permission to display the text list or to create derivative works for the first media content when distributing the first media content, or on the settings panel of the first media content. Furthermore, if the distributor of the first media content has not permitted the display of the text list of the first media content, the text list of the first media content will not be displayed in response to the user's text display operation. Similarly, if the distributor of the first media content has not permitted derivative works based on the text in the text list of the first media content, the system will not respond to the user's media content creation operation based on the text in the text list of the first media content.

[0082] The interaction method provided in this embodiment can support users in copying, sharing, and / or playing text displayed in a text list, or creating new media content based on text displayed in a text list, thereby meeting the diverse needs of users, reducing the difficulty of creating media content, and improving the user experience.

[0083] Figure 9 is a block diagram of the configuration of an interaction device provided by an embodiment of the present disclosure. The device can be implemented by software and / or hardware, may be configured in an electronic device, for example, a mobile phone or a tablet computer, and can display a list of text sentences for a video by performing an interaction method. As shown in Figure 9, the interaction device provided by this embodiment may include an operation reception module 901 and a list display module 902.

[0084] The operation reception module 901 is configured to receive text display operations for a first media content, including video content.

[0085] The list display module 902 is configured to display a list of text statements of the first media content in a predetermined area in response to the text display operation, wherein the list of text statements includes at least two text statements, and each text statement has a corresponding audio statement in the target audio data of the first media content.

[0086] The interaction device provided in this embodiment receives a text display operation for a first media content, including video content, via an operation reception module 901. In response to the text display operation, a list display module 902 displays a list of text statements of the first media content in a predetermined area. The list of text statements includes at least two text statements, and each text statement has a corresponding audio statement in the target audio data of the first media content. By adopting the above-described technology, this embodiment not only allows users to obtain information in the video by displaying the list of text statements of the video content and viewing the video screen or listening to the video audio, but also allows users to obtain information in the video (especially the video audio) by checking the list of text statements of the video content. This enriches the methods of obtaining information in the video, improves the efficiency of obtaining information in the video, and enhances the user experience.

[0087] In the proposed technology described above, the at least two text sentences may include the current text sentence and other text sentences other than the current text sentence that have a different display mode from the current text sentence, and the current text sentence may be the text sentence currently being played.

[0088] In the above-described technical proposal, the current text may be displayed at a designated location in the predetermined area, and the interaction device provided in this embodiment may further include a phrase movement module configured to move and display the modified current text at the designated location when the current text is changed.

[0089] Exemplary, the interaction device provided in this embodiment may further include a phrase switching module configured to display a position control for triggering a switch in the text displayed in the predetermined area, after displaying a text list of the first media content in a predetermined area, in response to a phrase switching operation performed within the predetermined area, to move the current text to a set position in the predetermined area and display it.

[0090] Exemplary, the interaction device provided in this embodiment may further include a word selection module configured to call at least one of the following: a word copy module configured to display a list of text sentences of the first media content in a predetermined area, then, in response to a word selection operation acting within the predetermined area, identify the first text sentence selected by the word selection operation, and, in response to a copy operation on the first text sentence, copy the first text sentence; a word sharing module configured to share the first text sentence in response to a share operation on the first text sentence; a word playback module configured to play the first media content in response to a playback operation on the first text sentence, using a time node in the first media content corresponding to the start point of the first text sentence as the playback start point; and a media content generation module configured to generate a second media content including the first text sentence in response to a media content generation operation on the first text sentence.

[0091] In the proposed technology described above, the media content generation module may include a card generation unit configured to generate a target card containing the first text statement, and a media content generation unit configured to generate a second media content using the target card as content and the first text statement and / or music corresponding to the target card as background music.

[0092] In the above-described technical proposal, the media content generation module may be configured to generate a second media content including the first text sentence and to distribute the second media content, or the interaction device provided in this embodiment may further include a media content distribution module configured to generate the second media content including the first text sentence and then distribute the second media content in response to a distribution operation for the second media content.

[0093] In the above-described technical proposal, the word selection module may be configured to display the first text sentence as selected after identifying the first text sentence selected by the word selection operation.

[0094] In the proposed technology described above, the search-waiting keywords included in the at least two text sentences may have a different display manner from other content in the at least two text sentences other than the search-waiting keywords, and the search-waiting keywords may be used to trigger the display of search results that match the triggered search-waiting keywords.

[0095] Exemplary, the interaction device provided in this embodiment may further include a first playback module configured to play the first media content before receiving a text display operation for the first media content, and the list display module 902 may be configured to display a text list of the first media content in a predetermined area and adjust the first media content outside the predetermined area for playback.

[0096] In the above-described technical proposal, the text display operation may include a media content jump operation for a third media content, and the interaction device provided in this embodiment may further include a second playback module configured to play a third media content, which is media content generated based on a second text sentence in the first media content, before receiving a text display operation for the first media content, and the list display module 902 may be configured to play the first media content outside a predetermined area and display a list of text sentences of the first media content in the predetermined area.

[0097] Exemplary, the interaction device provided in this embodiment may further include a control display module configured to display a jump control for the third media content to trigger the execution of the media content jump operation, in response to a control display operation for the third media content, before receiving a text display operation for the first media content.

[0098] In the above-described technical proposal, the list display module 902 may be configured to play the first media content outside the predetermined area, with the starting point of the first media content as the playback start point, or it may be configured to play the first media content outside the predetermined area, with the time node corresponding to the starting point of the second text sentence as the playback start point.

[0099] The interaction devices provided in the embodiments of this disclosure can perform the interaction methods provided in any embodiment of this disclosure and include functional modules and effects corresponding to the performance of the interaction methods. For technical details not described in detail in these embodiments, refer to the interaction methods provided in any embodiment of this disclosure.

[0100] Referring below to Figure 10, which shows a schematic diagram of the configuration of an electronic device (e.g., terminal device) 1000 suitable for realizing an embodiment of the present disclosure. The terminal device in the embodiment of the present disclosure may include, for example, mobile devices such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablets (Portable Android Devices, PADs), portable media players (PMPs), and in-vehicle terminals (e.g., car navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. The electronic device shown in Figure 10 is merely an example and should not limit the functions or scope of use of the embodiment of the present disclosure.

[0101] As shown in Figure 10, the electronic device 1000 may include a processing unit (e.g., a central processing unit, a graphics text processor, etc.) 1001 capable of performing various appropriate operations and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1008 into a random access memory (RAM) 1003. The RAM 1003 also stores various programs and data necessary for the operation of the electronic device 1000. The processing unit 1001, ROM 1002, and RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.

[0102] For example, input devices 1006, including touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc., output devices 1007, including liquid crystal displays (LCDs), speakers, vibrators, etc., storage devices 1008, including magnetic tape, hard disks, etc., and communication devices 1009 may be connected to the I / O interface 1005. The communication devices 1009 may allow the electronic device 1000 to communicate wirelessly or wired with other devices to exchange data. Figure 10 shows an electronic device 1000 with various devices, but it is not necessary to implement or include all the devices shown. More or fewer devices may be implemented or included instead.

[0103] According to embodiments of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product which includes a computer program mounted on a non-temporary computer-readable medium, which includes program code for performing the method shown in the flowchart. In such embodiments, the computer program may be downloaded and installed from a network by a communication device 1009, installed from a storage device 1008, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-described functions limited to the methods of embodiments of the present disclosure are performed.

[0104] The computer-readable media described above in this disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of both. The computer-readable storage medium may be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. The computer-readable storage medium may include an electrical connection having one or more wires, a portable computer disk, a hard disk, RAM, ROM, Erasable Programmable Read Only Memory (EPROM), flash memory, optical fiber, Compact Disc Read Only Memory (CD-ROM), optical memory device, magnetic memory device, or any suitable combination of the above. In this disclosure, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, apparatus, or device. On the other hand, in this disclosure, the computer-readable signal medium may include a data signal propagated in the baseband or as part of a carrier wave, on which computer-readable program code is carried. Such propagated data signals may take various forms, including electromagnetic signals, optical signals, or any suitable combination described above. Furthermore, the computer-readable signal medium may be any computer-readable medium other than a computer-readable storage medium on which a program can be transmitted, propagated, or transmitted by or in combination with an instruction execution system, apparatus, or device. The program code contained in the computer-readable medium may be transmitted by any suitable medium, including wires, optical cables, radio frequencies (RF), or any suitable combination described above.

[0105] In some embodiments, clients and servers can communicate using any network protocol currently known or to be developed in the future, such as the Hypertext Transfer Protocol (HTTP), and can connect with each other to digital data communications (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (LANs), wide area networks (WANs), extranets (e.g., the Internet), and end-to-end networks (e.g., ad-hoc end-to-end networks), and any network currently known or to be developed in the future.

[0106] The computer-readable medium described above may be included in the electronic device described above, or it may be a separate entity not incorporated into the electronic device.

[0107] The computer-readable medium described above contains one or more programs, and when the one or more programs described above are executed by the electronic device, the electronic device receives a text display operation for a first media content including video content, and in response to the text display operation, displays a list of text statements of the first media content in a predetermined area, the list of text statements containing at least two text statements, the text statements having corresponding audio statements in the target audio data of the first media content.

[0108] Computer program code for performing the operations disclosed herein may be written in one or more programming languages, or a combination thereof, and such programming languages ​​include, for example, object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as, for example, the "C" language or similar programming languages. The program code may be executed entirely on a user computer, partially on a user computer, executed as a standalone software package, partially on a user computer and partially on a remote computer, or fully executed on a remote computer or server. Where a remote computer is involved, the remote computer may be connected to the user computer via any type of network, including a LAN or WAN, or it may be connected to an external computer (for example, connected via the Internet using an Internet service provider).

[0109] The flowcharts and block diagrams in the drawings illustrate architectures, functions, and operations that can be implemented according to the systems, methods, and computer program products in various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, program segment, or part of code containing one or more executable instructions for implementing a given logic function. In some alternative implementations, the functions represented in a block may occur in a different order than shown in the drawings. For example, two blocks shown consecutively may actually be executed substantially in parallel, or they may be executed in reverse order depending on the functions involved. Each block in a block diagram and / or flowchart, and combinations of blocks in a block diagram and / or flowchart, may be implemented in a dedicated hardware-based system that performs a given function or operation, or in a combination of dedicated hardware and computer instructions.

[0110] The units relating to the embodiments of this disclosure described herein may be implemented in software or in hardware. In some cases, the names of the units do not constitute a limitation on the unit itself.

[0111] The functions described above in this specification may be performed by at least partially one or more hardware logic components. For example, non-limiting exemplary types of hardware logic components that may be used include field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard parts (ASSPs), systems on chips (SOCs), and complex programmable logic devices (CPLDs).

[0112] In the context of this disclosure, machine-readable media may be tangible media that contain or can store programs used by or in combination with instruction execution systems, devices, or equipment. Machine-readable media may be machine-readable signal media or machine-readable storage media. Machine-readable media may include electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or equipment, or any appropriate combination of the above. Machine-readable storage media may include electrical connections based on one or more wires, portable computer disks, hard disks, RAM, ROM, EPROM, flash memory, optical fibers, portable CD-ROMs, optical storage devices, magnetic storage devices, or any appropriate combination of the above. Storage media may be non-transitory storage media.

[0113] According to one or more embodiments of this disclosure, Example 1 receives a text display operation on first media content, which includes video content, The process includes, in response to the text display operation, displaying a list of text statements of the first media content in a predetermined area, The text list includes at least two text sentences, each of which has a corresponding audio sentence in the target audio data of the first media content, providing an interaction method.

[0114] According to one or more embodiments of the present disclosure, Example 2 is the method of Example 1, wherein the at least two text sentences include a current text sentence and other text sentences other than the current text sentence having a different display manner from the current text sentence, the current text sentence being the text sentence currently being played.

[0115] According to one or more embodiments of the present disclosure, Example 3 is the method of Example 2, wherein the current text is displayed at a designated location in the predetermined area, and the method is If the current text is changed, the modified current text is moved to the specified position and displayed.

[0116] According to one or more embodiments of the present disclosure, Example 4 is a method of Example 2, wherein after displaying the text list of the first media content in a predetermined area, The further includes displaying a position control to trigger a switch in the text displayed in the predetermined area in response to a word switching operation performed within the predetermined area, thereby moving the current text to a set position in the predetermined area for display.

[0117] According to one or more embodiments of the present disclosure, Example 5 is the method of Example 1, wherein after displaying the text list of the first media content in a predetermined area, In response to a word selection operation performed within the predetermined area, the first text sentence selected by the word selection operation is identified. The functionality further includes performing at least one of the following: copying the first text in response to a copy operation on the first text; sharing the first text in response to a share operation on the first text; playing the first media content in response to a play operation on the first text, using the time node in the first media content corresponding to the start point of the first text as the playback start point; and generating a second media content containing the first text in response to a media content generation operation on the first text.

[0118] According to one or more embodiments of this disclosure, Example 6 is a method of Example 5, wherein generating a second media content including the first text sentence is: To generate a target card containing the first text statement, This includes generating a second media content using the target card as content and the first text and / or music corresponding to the target card as background music.

[0119] According to one or more embodiments of the present disclosure, Example 7 is a method of Example 5, wherein generating a second media content containing the first text sentence includes the steps of generating a second media content containing the first text sentence and distributing the second media content, or The process further includes generating a second media content containing the first text statement, and then distributing the second media content in response to a distribution operation for the second media content.

[0120] According to one or more embodiments of the present disclosure, Example 8 is the method of Example 5, wherein after identifying the first text sentence selected by the word selection operation, This further includes displaying the first text statement as selected.

[0121] According to one or more embodiments of the present disclosure, Example 9 is the method of Example 1, wherein the search-waiting keywords contained in the at least two text sentences have a different display manner from other content in the at least two text sentences other than the search-waiting keywords, and the search-waiting keywords are used to trigger the display of search results that match the triggered search-waiting keywords.

[0122] According to one or more embodiments of the present disclosure, Example 10 is a method of any of Examples 1 to 9, wherein before receiving a text display operation on the first media content, Further including playing the first media content, Displaying the text list of the first media content in a predetermined area is: This includes displaying a text list of the first media content in a predetermined area, and adjusting the first media content to be outside the predetermined area for playback.

[0123] According to one or more embodiments of the present disclosure, Example 11 is a method of any of Examples 1 to 9, wherein the text display operation includes a media content jump operation to a third media content which is media content generated based on a second text sentence in a first media content, before receiving the text display operation to the first media content. This further includes playing third media content, Displaying the text list of the first media content in a predetermined area is: This includes playing the first media content outside a predetermined area and displaying a list of text statements of the first media content in the predetermined area.

[0124] According to one or more embodiments of the present disclosure, Example 12 is a method of any of Examples 1-11, wherein before the first media content receives a text display operation, The further includes, in response to a control display operation for the third media content, displaying a jump control corresponding to the third media content to trigger the execution of the media content jump operation.

[0125] According to one or more embodiments of this disclosure, Example 13 is the method of Example 11, wherein the first media content is played outside a predetermined area. A step of playing the first media content outside the predetermined area, with the starting point of the first media content as the playback start point, or The process includes the step of playing the first media content outside the predetermined area, with the time node corresponding to the start point of the second text sentence in the first media content as the playback start point.

[0126] According to one or more embodiments of this disclosure, Example 14 is, An operation receiving module configured to receive text display operations for a first media content including video content, The system includes a list display module configured to display a list of text statements of the first media content in a predetermined area in response to the text display operation, The text list includes at least two text sentences, each of which has a corresponding audio sentence in the target audio data of the first media content, providing an interaction device.

[0127] According to one or more embodiments of this disclosure, Exemplary 15 is, One or more processors, A memory configured to store one or more programs, The present invention provides an electronic device that, when one or more of the aforementioned programs are executed by one or more of the aforementioned processors, causes one or more of the aforementioned processors to implement the interaction method described in any of Examples 1 to 13.

[0128] According to one or more embodiments of the present disclosure, Example 16 provides a computer-readable storage medium that stores a computer program which, when executed by a processor, implements the interaction method described in any of Examples 1 to 13.

[0129] Furthermore, although multiple operations are described in a specific order, this should not be understood as requiring that these operations be performed in or in a specific order as indicated. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although the above discussion includes several implementation details, these should not be construed as limiting the scope of this disclosure. Some features described in the context of individual embodiments may be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may be implemented in multiple embodiments individually or in any suitable subcombination.

[0130] While this subject matter is described using language specific to structural features and / or method logic operations, it should be understood that the subject matter limited in the attached claims is not necessarily limited to the specific features or operations described above. Conversely, the specific features and operations described above are merely exemplary forms of realizing the claims.

Claims

1. Receiving text display operations on the first media content, including video content, The process includes, in response to the text display operation, displaying a list of text statements of the first media content in a predetermined area, The text list includes at least two text sentences, and each of the at least two text sentences has a corresponding audio sentence in the target audio data of the first media content. The text display operation includes a media content jump operation to a third media content, which is media content generated based on the second text sentence in the first media content. Before receiving a text display operation on the first media content, Further including playing the aforementioned third media content, Displaying the text list of the first media content in a predetermined area is: This includes playing the first media content outside the predetermined area and displaying a text list of the first media content in the predetermined area. Interaction method.

2. The at least two text sentences include the current text sentence and other text sentences other than the current text sentence that have a different display mode from the current text sentence, wherein the current text sentence is the text sentence currently being played. The interaction method according to claim 1.

3. The current text is displayed at the designated location in the predetermined area, and the interaction method is as follows: If the current text is changed, the modified current text is moved to the specified position and displayed. The interaction method according to claim 2.

4. After displaying the text list of the first media content in a predetermined area, The further includes displaying a position control to trigger a switch in the text displayed in the predetermined area in response to a word switching operation performed within the predetermined area, thereby moving and displaying the current text at a set position in the predetermined area. The interaction method according to claim 2.

5. After displaying the text list of the first media content in a predetermined area, In response to a word selection operation performed within the predetermined area, the first text sentence selected by the word selection operation is identified. The further includes performing at least one of the following in response to a copy operation on the first text statement: copying the first text statement; sharing the first text statement in response to a share operation on the first text statement; playing the first media content in response to a play operation on the first text statement, using the time node corresponding to the start point of the first text statement in the first media content as the playback start point; and generating a second media content including the first text statement in response to a media content generation operation on the first text statement. The interaction method according to claim 1.

6. Generating a second media content containing the first text is: To generate a target card containing the first text statement, The process includes generating a second media content using the target card as content and music corresponding to at least one of the first text and the target card as background music, The interaction method according to claim 5.

7. Generating a second media content containing the first text statement includes generating the second media content containing the first text statement and distributing the second media content, or The process further includes generating a second media content containing the first text statement, and then distributing the second media content in response to a distribution operation for the second media content. The interaction method according to claim 5.

8. After identifying the first text sentence selected by the aforementioned word selection operation, The first text statement is further displayed as selected. The interaction method according to claim 5.

9. The search-waiting keywords contained in the at least two text sentences have a different display pattern from other content in the at least two text sentences other than the search-waiting keywords, and the search-waiting keywords are used to trigger the display of search results that match the triggered search-waiting keywords. The interaction method according to claim 1.

10. Before receiving a text display operation on the first media content, Further including playing the aforementioned first media content, Displaying the text list of the first media content in a predetermined area is: This includes displaying a text list of the first media content in the predetermined area, and adjusting the first media content to play it outside the predetermined area. The interaction method according to claim 1.

11. Before receiving a text display operation on the first media content, The further includes, in response to a control display operation for the third media content, displaying a jump control corresponding to the third media content to trigger the execution of the media content jump operation, The interaction method according to claim 1.

12. Playing the first media content outside the predetermined area is, Playing the first media content outside the predetermined area, with the starting point of the first media content as the playback start point, or This includes playing the first media content outside the predetermined area, with the time node corresponding to the start point of the second text sentence in the first media content as the playback start point, The interaction method according to claim 1.

13. An operation receiving module configured to receive text display operations for a first media content including video content, A list display module configured to display a list of text statements of the first media content in a predetermined area in response to the text display operation, The system includes a second playback module configured to play a third media content, which is media content generated based on a second text sentence in the first media content, before receiving a text display operation on the first media content, The text list includes at least two text sentences, and each of the at least two text sentences has a corresponding audio sentence in the target audio data of the first media content. The text display operation includes a media content jump operation for the third media content, The list display module plays the first media content outside the predetermined area and displays a list of text statements of the first media content in the predetermined area. Interaction device.

14. At least one processor, The system comprises at least one processor and a memory that is communicated with it, The memory stores a computer program that can be executed by the at least one processor, and when the computer program is executed by the at least one processor, the memory causes the at least one processor to execute the interaction method described in any one of claims 1 to 12. electronic equipment.

15. When executed by a processor, it stores computer instructions for implementing the interaction method described in any one of claims 1 to 12. Computer-readable storage medium.

16. A computer program, when executed by a computing device, causes the computing device to implement the interaction method described in any one of claims 1 to 12.