Text display method, storage medium and electronic device

By obtaining video audio data and converting it into text information, combining time information, the problem of synchronizing text information and video in video notes is solved, and the text information is synchronized with the video screen is realized, improving the user experience.

CN117707394BActive Publication Date: 2025-08-08HONOR DEVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310861547.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-12
Publication Date
2025-08-08
Estimated Expiration
2043-07-12

AI Technical Summary

Technical Problem

When users play video notes, the text information cannot be displayed in association with the video, resulting in the inability to intuitively position or adjust the playback progress according to the video playback progress, and the user experience is low.

Method used

By obtaining the voice information and time information of the video audio data, converting it into text information, and keeping the text information synchronized with the video screen during the video playback, using preset modification strategies to adjust the editing operation to ensure the visual and text synchronization effect.

Benefits of technology

This enables text information to change with the change of video screen, ensuring that the user can still maintain the visual and text synchronization effect after editing operations, and improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117707394B_ABST
    Figure CN117707394B_ABST
Patent Text Reader

Abstract

The present application relates to the field of computer technology, and discloses a text display method, storage medium, and electronic device. The method includes: the electronic device first obtains the voice information and time information of the audio data in the video, and then after converting the voice information into text information, the converted text information is matched with the obtained time information. In this way, during the video playback process, the text information changes accordingly with the change of the video screen, achieving a visual-text synchronization effect. At the same time, when the user adds text to the text information, no new correspondence is added between the added text and the time information, and the correspondence between the original text and the time information is retained; when the user deletes text in the text information, the correspondence between the deleted text and the time information is deleted, and the correspondence between the original text and the event information is retained, thereby ensuring the visual-text synchronization effect between the text information and the video screen after the user edits the text information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a text display method, a storage medium, and an electronic device. Background Art

[0002] Users can use electronic devices with voice-to-text capabilities to convert the audio information of videos on the electronic device into text information, creating video notes based on the video and text information. Users can view the video content recorded in the text information and can also align, annotate, and edit the text information based on the video content.

[0003] However, when a user plays a video in a video note, the text information and the text information edited by the user are usually not displayed in association with the video. For example, the text in the text information cannot change accordingly with changes in the video screen. Summary of the Invention

[0004] Embodiments of the present application provide a text display method, a storage medium, and an electronic device.

[0005] In a first aspect, an embodiment of the present application provides a text display method, which is applied to an electronic device, the method comprising: in response to a first operation of a user, displaying a first interface and playing a first video on the first interface; in response to a second operation of the user, starting to record the first video played on the first interface; in response to a third operation of the user, ending the recording of the first video and displaying a second interface, wherein the second interface comprises a first display area and a second display area, the first display area being used to play the recorded first video, and text information corresponding to the voice information of the recorded first video being displayed in the second display area; in response to a user clicking on a first character in the text information, a video frame corresponding to the first character is displayed in the first display area; after adding a second character to the text information, receiving a user click on the second character in the text information, the video frame displayed in the first display area does not change; in response to a user clicking on a third character in the text information, a video frame corresponding to the third character is displayed in the first display area, wherein the position of the third character is after the second character.

[0006] In an embodiment of the present application, after the electronic device completes recording the first video, it obtains the voice information and time information of the audio data in the recorded first video, converts the voice information into text information, and the converted text information corresponds to the obtained time information. In this way, when the user clicks on the original character (such as the first character) in the text information, the video screen can jump to the video screen corresponding to the original character, achieving a video-text synchronization effect. When the user adds a second character in the text information, a new correspondence between the second character and the time information is not added, and the correspondence between the original character and the time information is retained. In this way, after the user adds the second character, when the user clicks on the original character in the text information (such as the third character after the second character), the video screen can still jump to the video screen corresponding to the original character, ensuring the video-text synchronization effect after the user adds the text.

[0007] In a possible implementation of the first aspect above, the method also includes: in response to the user's fourth operation, playing the recorded first video in the first display area of the second interface; in response to the user clicking the first character in the text information, the video frame corresponding to the first character is displayed in the first display area, including: in response to the user clicking the first character in the text information, the video playback screen in the first display area jumps to the video frame corresponding to the first character.

[0008] In a possible implementation of the first aspect above, the method further includes: registering a focus acquisition event monitor, a content change event monitor, and a focus loss event monitor; after adding a second character to the text message, it includes: based on the focus acquisition event, the content change event, and the focus loss event, modifying the character and the corresponding video frame so that when the user clicks the character, the video frame having a first corresponding relationship with the character can be displayed in the first display area, and the multiple characters in the text information corresponding to the voice information of the recorded first video respectively have a first corresponding relationship with the multiple video frames of the recorded first video.

[0009] In an embodiment of the present application, multiple characters in the text information corresponding to the voice information of the recorded first video are used to represent the original characters obtained by text conversion based on the voice information of the recorded first video; modifying the characters and the corresponding video frames can be understood as modifying the characters and the first corresponding relationship between the characters and the corresponding video frames.

[0010] In a possible implementation of the first aspect above, the method also includes: when the user clicks on the second display area, a focus acquisition event is monitored, the first cursor is displayed in the second display area, and the position of the first cursor when the first cursor is displayed in the second display area is recorded; when the user edits the text information, a content change event is monitored, and the editing character corresponding to the editing operation is recorded; when the user clicks on the third display area, a focus loss event is monitored, the first cursor in the second display area is hidden, and the position of the second cursor when the first cursor is hidden in the second display area is recorded, wherein the third display area is not the second display area; based on the focus acquisition event, the content change event and the focus loss event, the characters and the corresponding video frames are modified, including: based on the first cursor position, the editing character and the second cursor position, the characters and the corresponding video frames are modified.

[0011] In a possible implementation of the first aspect above, after adding a second character to the text information; the above-mentioned modification of the character and the corresponding video frame based on the first cursor position, the editing character and the second cursor position includes: the first cursor moves backward, and the second cursor position is located after the first cursor position; the character position of the character before the first cursor position, and the first corresponding relationship with the video frame remain unchanged; the character position of the character after the first cursor position moves backward by the number of character bits of the second character; the first corresponding relationship between the character after the first cursor position and the video frame remains unchanged; the second character is between the first cursor position and the second cursor position, wherein the second character has no corresponding relationship with the video frame.

[0012] In a possible implementation of the first aspect above, after deleting the fourth character in the text information by using the backspace key; the above-mentioned modification of the character and the corresponding video frame based on the first cursor position, the editing character and the second cursor position includes: the first cursor moves forward, the second cursor position is located before the first cursor position, and the fourth character is located between the first cursor position and the second cursor position; the character position of the character before the second cursor position, and the first corresponding relationship with the video frame remain unchanged; the fourth character and the first corresponding relationship between the fourth character and the video frame are deleted; the character position of the character after the first cursor position moves forward by the number of character bits of the fourth character; the first corresponding relationship between the character after the first cursor position and the video frame remains unchanged.

[0013] In a possible implementation of the first aspect above, after deleting the fourth character in the text information by using the delete key; the above-mentioned modification of the character and the corresponding video frame based on the first cursor position, the editing character and the second cursor position includes: setting the first cursor position to move backward by the number of character bits of the fourth character, so that the fourth character is located between the first cursor position and the second cursor position; the character position of the character before the second cursor position, and the first corresponding relationship with the video frame remain unchanged; deleting the fourth character, and the first corresponding relationship between the fourth character and the video frame; the character position of the character after the first cursor position moves forward by the number of character bits of the fourth character; the first corresponding relationship between the character after the first cursor position and the video frame remains unchanged.

[0014] In this embodiment of the present application, the fourth character is the original character in the text message. When the user deletes the fourth character in the text message, the correspondence between the fourth character and the time information is deleted, while the correspondence between the undeleted characters in the text message and the time information is retained. Thus, after deleting the fourth character, if the user clicks on an undeleted character in the text message, the video screen can still jump to the video screen corresponding to the undeleted character, ensuring the video and text synchronization effect after the user adds text.

[0015] In a possible implementation of the first aspect above, after selecting the fifth character in the text information and replacing the fifth character with the sixth character, wherein the number of character bits of the fifth character is less than the first digit of the character bit of the sixth character; the above-mentioned modification of the character and the corresponding video frame based on the first cursor position, the editing character and the second cursor position includes: the first cursor position includes a forward position and a backward position, the fifth character is located between the forward position and the backward position, and the backward position of the first cursor position is moved forward by the first digit to obtain the second cursor position; the character position of the character before the forward position, and the first corresponding relationship with the video frame remain unchanged; the fifth character, and the first corresponding relationship between the fifth character and the video frame remain unchanged; the fifth character is deleted, and the first corresponding relationship between the fifth character and the video frame is added between the forward position and the second cursor position, wherein the sixth character has no corresponding relationship with the video frame; the character position of the character after the backward position is moved forward by the first digit; the first corresponding relationship between the character after the backward position and the video frame remains unchanged.

[0016] In a possible implementation of the first aspect above, after selecting the fifth character in the text information and replacing the fifth character with the sixth character, wherein the number of character digits of the fifth character is one digit more than the number of character digits of the sixth character; the above-mentioned modification of the characters and the corresponding video frames based on the first cursor position, the editing character and the second cursor position includes: the first cursor position includes a forward position and a backward position, the fifth character is located between the forward position and the backward position, and the backward position of the first cursor position is moved forward by the first digit to obtain the second cursor position; the character position of the character before the forward position and the first corresponding relationship with the video frame remain unchanged; the fifth character and the first corresponding relationship between the fifth character and the video frame are deleted; the sixth character is added between the forward position and the second cursor position, wherein the sixth character has no corresponding relationship with the video frame; the character position of the character after the backward position is moved backward by the first digit; the first corresponding relationship between the character after the backward position and the video frame remains unchanged.

[0017] In a possible implementation of the first aspect above, after selecting the fifth character in the text information and replacing the fifth character with the sixth character, wherein the number of character bits of the fifth character is equal to the number of character bits of the sixth character; the above-mentioned modification of the character and the corresponding video frame based on the first cursor position, the editing character and the second cursor position includes: the first cursor position includes a forward position and a backward position, the fifth character is located between the forward position and the backward position, and the forward position of the first cursor position is the second cursor position; the character position of the character before the forward position, and the first corresponding relationship with the video frame remain unchanged; the fifth character, and the first corresponding relationship between the fifth character and the video frame are deleted; the sixth character is added between the forward position and the second cursor position, wherein the sixth character has no corresponding relationship with the video frame; the character position of the character after the backward position, and the first corresponding relationship with the video frame remain unchanged.

[0018] In this embodiment of the present application, the fifth character is the original character in the text message. When the user replaces the fifth character in the text message with the sixth character, the correspondence between the fifth character and the time information is deleted, no new correspondence between the sixth character and the time information is added, and the correspondence between the unreplaced character and the time information is retained. In this way, after the user replaces the text, when the user clicks on the unreplaced character in the text message, the video screen can still jump to the video screen corresponding to the unreplaced character, ensuring the video and text synchronization effect after the user adds text.

[0019] In a second aspect, an embodiment of the present application provides a readable storage medium having instructions stored thereon, which, when executed on an electronic device, enables the electronic device to implement any one of the text display methods provided by the first aspect and various possible implementations of the first aspect.

[0020] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a memory for storing instructions executed by one or more processors of the electronic device; and a processor, which is one of the processors of the electronic device, for executing the instructions stored in the memory to implement any one of the text display methods provided by the above-mentioned first aspect and various possible implementations of the above-mentioned first aspect.

[0021] In a fourth aspect, an embodiment of the present application provides a program product, which includes instructions. When the instructions are executed by an electronic device, the electronic device can implement any text display method provided by the first aspect and various possible implementations of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1A and 1B According to some embodiments of the present application, some interface schematic diagrams of the tablet 100 are shown;

[0023] Figure 2 According to some embodiments of the present application, a schematic diagram of an interface for adding text is shown;

[0024] Figure 3 According to some embodiments of the present application, a schematic diagram of an interface for deleting text is shown;

[0025] Figure 4A and 4B According to some embodiments of the present application, some schematic diagrams of interfaces for adjusting video progress are shown;

[0026] Figure 5A According to some embodiments of the present application, a schematic diagram of a display interface 510 is shown;

[0027] Figure 5B According to some embodiments of the present application, another interface schematic diagram of the display interface 510 is shown;

[0028] 5C (a) and (b) show some schematic diagrams of the display interface 530 according to some embodiments of the present application;

[0029] Figure 5D According to some embodiments of the present application, a schematic diagram of a display interface 540 is shown;

[0030] Figure 5E According to some embodiments of the present application, a schematic diagram of a display interface 550 is shown;

[0031] Figure 5F According to some embodiments of the present application, a schematic diagram of a display interface 560 is shown;

[0032] Figure 5G According to some embodiments of the present application, a schematic diagram of an interface displaying a cursor is shown;

[0033] Figure 5H According to some embodiments of the present application, a schematic diagram of an interface for adding text is shown;

[0034] Figure 5I According to some embodiments of the present application, a schematic diagram of an interface for hiding a cursor is shown;

[0035] Figure 5J According to some embodiments of the present application, a schematic diagram of an interface for selecting text is shown;

[0036] Figure 6A According to some embodiments of the present application, a software structure diagram of the tablet 100 is shown;

[0037] Figure 6B According to some embodiments of the present application, another software structure diagram of the tablet 100 is shown;

[0038] Figure 7A According to some embodiments of the present application, a flow chart of a text display method is shown;

[0039] Figure 7B According to some embodiments of the present application, another flow chart of a text display method is shown;

[0040] Figure 8 According to some embodiments of the present application, a hardware structure diagram of the tablet 100 is shown. DETAILED DESCRIPTION

[0041] The illustrative embodiments of the present application include, but are not limited to, a text display method, a storage medium, and an electronic device.

[0042] The technical solution of this application is introduced below with reference to Figures 1 to 8.

[0043] When a user plays a video using a video application in an electronic device, the played video can be stored by recording the screen or downloading the video, or the user can shoot and store the video using the camera and video recording function of the electronic device.

[0044] In some embodiments, users can use the speech-to-text function of an electronic device to convert voice information such as dialogue and narration in a video into text information, and form video notes based on the video and the converted text information. Users can then review the video content recorded in the video based on the text information in the video notes.

[0045] For example, Figure 1AAs shown, the user uses the voice-to-text function of the tablet computer (hereinafter referred to as the tablet) 100 to convert the voice information in the video 11 into text information, and display the text information in the text display area 12 below the video 11. Based on the video 11 and the text information in the text display area 12, a video note corresponding to the video 11 is formed.

[0046] As mentioned above, users can edit the text in the video notes by calibrating, annotating, and performing other operations based on the video content. For example, if a word is missing or extra in the electronic device's speech-to-text conversion, the user can click on the missing or extra word to add or remove it.

[0047] For example, Figure 1B As shown, the user can add the word "ice and snow" between the word "the earth is covered" and the word "covered" in the text information in the text display area 12; or the user can delete the word "all" after the word "these" in the text information in the text display area 12.

[0048] It can be understood that the user's modification of the text information in the text display area can also be any other type of editing operation, such as font modification, font size modification, color modification, etc.; the text in the text information can be any text type, such as Chinese characters, English letters, punctuation marks, etc. This application does not limit the editing operation type and text type of the text information.

[0049] In some embodiments, when a user plays a video in a video note, the text information in the video note text display area and the text information modified by the user are usually not displayed in association with the video.

[0050] For example, Figure 1B As shown, when the user plays the video 11 of the video note, the text information in the text display area 12 and the text information modified by the user cannot be displayed in association with the playback progress of the video 11. Figure 1B When the progress bar of the video 11 is at the progress position 11A, 11B or 11C, the display style of the text information is the same, and Figure 1B The display style of the text information when the video 11 is playing is the same as Figure 1A The text information displayed when the video 11 is not playing is displayed in the same style. That is, the text information displayed at the 0th second and the 10th second of the video 11 is displayed in the same style. The text information display style does not change with the progress of the video 11, such as the font, size, color, etc.

[0051] In this way, when the user plays the video in the video note, the user cannot intuitively locate the text information according to the playback progress of the video, nor can the user adjust the playback progress of the video according to the interesting part in the text information, resulting in a relatively low user experience.

[0052] For this reason, the present application proposes a text display method applied to an electronic device. The method includes: the electronic device first obtains the speech information and time information of the audio data in the video, and then after converting the speech information into text information, associates the obtained text information with the obtained time information. In this way, during the video playback process, the text information changes correspondingly as the video picture changes, achieving the effect of video-text synchronization. At the same time, for the user's editing operation on the text information, a preset modification strategy can be used to adjust the modified text to ensure the video-text synchronization effect between the text information and the video picture after the editing operation.

[0053] The following gives an example explanation of different editing operations.

[0054] (1) Adding text

[0055] In some embodiments, for the user's editing operation of adding text in the text information, the modification strategy includes: the electronic device retains the corresponding relationship between the original text and the time information in the text information, and does not add the corresponding relationship with the time information to the added text of the editing operation. At the same time, the electronic device associates the added text with the adjacent text of the added text, so that the added text and the adjacent text change correspondingly as the video picture changes.

[0056] For example, as Figure 1A shown, the corresponding relationship between the text "The earth is" and the time information in the text information of the text display area 12 is the 6th second, and the corresponding relationship between the text "covered" and the time information is the 7th second; when the user adds the text "by ice and snow" between the text "The earth is" and the text "covered" in the text information of the text display area 12, the text information of the text display area 12 as Figure 2 shown is obtained.

[0057] In some embodiments, the added text "by ice and snow" in the text information of the text display area 12 and the adjacent text "The earth is" change correspondingly as the video picture changes. For example, as Figure 2 shown, when the playback progress bar of the video 11 is at the progress position 21A, that is, when the video 11 is played to the 6th second. Figure 2In the text information in the text display area 12, the text at and before the 6th second of video 11 is bold, while the text after the 6th second of video 11 is non-bold. Furthermore, the text "Ice and Snow" and the text "Earth Cover" corresponding to the 6th second are bold. In other words, at the 6th second of video 11, the text "Ice and Snow" and the text before "Ice and Snow" are bold, while the text after "Ice and Snow" is non-bold.

[0058] In other embodiments, the text "Ice and Snow" added to the text information in the text display area 12 and the adjacent text "Overlay" change accordingly as the video screen changes. For example, when video 11 is played to the 7th second, the text in the text information in the text display area 12 corresponding to the 7th second and before the 7th second of video 11 is bold, and the text after the 7th second of video 11 is non-bold, and the text "Ice and Snow" is bolded along with the text "Overlay" corresponding to the 6th second. In other words, when video 11 is played to the 7th second, the text "Overlay" and the text before the text "Overlay" are bolded, and the text after the text "Overlay" is non-bold.

[0059] (2) Delete text

[0060] In some embodiments, when a user deletes text from a text message, the modification strategy includes: the electronic device deletes the correspondence between the deleted text and the time information in the text message, while retaining the correspondence between the undeleted text and the time information in the text message. Simultaneously, the electronic device associates the deleted text with its adjacent text, ensuring that the video corresponding to the deleted text is synchronized with the video of the undeleted text.

[0061] For example, Figure 1A As shown, the corresponding relationship between the word "these" in the text information of the text display area 12 and the time information is the 25th second, the corresponding relationship between the word "all" and the time information is the 26th second, and the corresponding relationship between the word "all" and the time information is the 27th second; when the user deletes the word "all" in the text information of the text display area 12, the result is Figure 3 The text display area 12 shows text information.

[0062] In some embodiments, the deleted word "all" is associated with the adjacent word "these", so that the video frame at the 26th second corresponding to the deleted word "all" is associated with the adjacent word "these" at the 25th second. Figure 3 As shown, when the progress bar of video 11 is at progress position 31, that is, video 11 is played to the 26th second. Figure 3In the text information in the text display area 12, the text corresponding to the 25th second and before of the video 11 is in bold, and the text corresponding to the 27th second and after of the video 11 is in non-bold. In other words, at the 26th second of the video 11, the text "these" and the text before "these" are in bold, and the text after "these" is in non-bold.

[0063] In other embodiments, the deleted text "all" is associated with the adjacent text "all in", so that the video frame at the 26th second corresponding to the deleted text "all in" is associated with the adjacent text "all in" at the 27th second. For example, when the video 11 is played to the 26th second, the text information in the text display area 12 corresponding to the 27th second and before the 27th second of the video 11 is in bold style, and the text after the 27th second of the video 11 is in non-bold style. In other words, when the video 11 is played to the 26th second, the text "all in" and the text before the text "all in" are in bold style, and the text after the text "all in" is in non-bold style.

[0064] (3) Replace text

[0065] In some embodiments, after a user deletes text from a text message, an editing operation of adding text is simultaneously performed at the deleted position; or after a user adds text to a text message, an editing operation of deleting the original text in the text message at the added position is simultaneously performed, the modification strategy includes: for deleting text, the above-mentioned modification strategy for deleting text can be adopted; for adding text, the above-mentioned modification strategy for adding text can be adopted; the specific implementation method can be referred to the description of the above-mentioned modification strategy for deleting / adding text, which will not be repeated here.

[0066] The following is an example of the video-text synchronization effect.

[0067] In some embodiments, when the user adjusts the video screen, the text information in the corresponding text display area can change accordingly with the video screen change. Figure 4A As shown, when the user clicks on the progress position 41A of the progress bar of video 11, video 11 jumps to the video screen at the 6th second, and the text information in the corresponding text display area 12 corresponding to the 6th second of video 11 and the part before the 6th second is bold, and the text after the 6th second of video 11 is non-bold. In other words, when the user clicks on the progress position 41A of the progress bar of video 11, video 11 jumps to the video screen at the 6th second, and the text information in the text display area 12 corresponding to the 6th second of video 11 and the text before the text "Ice and Snow" is bold, and the text after the text "Ice and Snow" is non-bold.

[0068] In other embodiments, when the user clicks on a certain text in the text information of the text display area, the video can jump to the video screen corresponding to the time of the clicked text. Figure 4B As shown, when the user clicks on the word "all are here" in the text information of the text display area 12, the progress bar of the video 11 jumps to the progress position 41B, the video 11 jumps to the video screen at the 27th second, and the text information of the corresponding text display area 12 at the 27th second of the video 11 and before the 27th second is in bold style, and the text after the 27th second of the video 11 is in non-bold style. That is to say, when the user clicks on the word "all are here" in the text display area 12, the video 11 jumps to the video screen at the 27th second, and the text information of the corresponding text display area 12 at the 27th second of the video 11 and before the 27th second is in bold style, and the text after the text "all are here" is in non-bold style.

[0069] In some embodiments, electronic devices include but are not limited to mobile phones, tablet computers, wearable devices, in-vehicle devices, augmented reality (AR) / virtual reality (VR) devices, laptop computers, ultra-mobile personal computers (UMPCs), netbooks, personal digital assistants (PDAs), or specialized cameras (e.g., SLR cameras, card cameras), etc. The embodiments of the present application do not impose any restrictions on the specific types of electronic devices.

[0070] The following uses the tablet 100 as an example of an electronic device, combines video excerpts to generate video notes, and adds text scenes to the text information of the video notes to explain some embodiments of the present application.

[0071] (1) Video recording and excerpting process

[0072] In some embodiments, while using tablet 100, a user can use the video application in tablet 100 to open a playing video and watch the playing video. The user can also use the note application in tablet 100 to excerpt the playing video to obtain an excerpt video and convert the voice information in the excerpt video into text information. The note application can generate a video note corresponding to the playing video based on the excerpt video and the converted text information. It is understood that note applications include but are not limited to memo applications, video note applications, etc., and are not specifically limited.

[0073] In some embodiments, as Figure 5A As shown, when a user uses the tablet 100 to open a video application and plays a video 500 through the video application, the tablet 100 displays a video playing interface 510 .

[0074] In some embodiments, as Figure 5A and Figure 5B As shown, when the user wants to Figure 5A When playing the video 500 shown in the figure, the user can touch the screen of the tablet 100 to extract the video based on the video playback interface 510 of the tablet 100. Figure 5B The trigger operation 511 of “slide down with three fingers” is shown to trigger the tablet 100 to start the video excerpt function and perform video excerpt on the played video 500 .

[0075] In some embodiments, as Figure 5B As shown in FIG5C(a), when the tablet 100 detects the user triggering operation 511, the tablet 100 starts the video excerpt function. Figure 5B The video playback interface 510 shown jumps to the display interface 530 shown in FIG5C(a).

[0076] In some embodiments, as shown in FIG5C(a), after the tablet 100 detects that the user has triggered operation 511, it starts the note application, activates the video excerpt function by starting the note application, and displays the interfaces of the video application and the note application in a left-right split screen manner. For example, the left side of the display interface 530 displays the interface of the video application, and the right side displays the interface of the note application. When the video excerpt function is activated, the note application displays an operation bar 531 for user operations. The operation bar 531 may include operation buttons such as pause / stop, mark, favorite, screenshot, voice-to-text, and close.

[0077] In some embodiments, as shown in FIG5C(a), when the video excerpting function is enabled in the note-taking application, the speech-to-text function is automatically enabled. The video note-taking application can simultaneously obtain the voice information in the excerpted video during the video excerpting process and convert the voice information into text information through the speech-to-text function.

[0078] In other embodiments, as shown in Figure 5C(b), when the video excerpt function is turned on in the note application, the user can turn on the speech-to-text function by clicking the speech-to-text button 534 in the operation bar 531, so that when the video excerpt function is turned on in the note application, the speech-to-text function is turned on at the same time.

[0079] In some embodiments, as shown in FIG. 5C(a) and Figure 5D As shown, when the note application detects that the user clicks the stop button 532 in the operation bar 531, the note application stops extracting and jumps to the stop extracting interface. The tablet 100 jumps from the display interface 530 shown in Figure 5C (a) to Figure 5D Display interface 540 is shown.

[0080] In some embodiments, such as Figure 5D and Figure 5E shown, when the note application detects that the user clicks the close button 533 in the operation bar 531, the note application closes the video excerpt function and hides the operation bar 531, and jumps to the closed excerpt interface. The tablet 100 jumps from the Figure 5D shown display interface 540 to the Figure 5E shown display interface 550.

[0081] In some embodiments, such as Figure 5E shown, the note application generates an excerpt video 500A of the playing video 500 according to the Figure 5B and the video excerpt operation shown in FIG. 5C(a), and displays the text information obtained by the voice-to-text function of the excerpt video 500A in the text display area 551 below the excerpt video 500A.

[0082] In some other embodiments, the note application can only excerpt the playing video 500 during the video excerpt process. After generating the excerpt video 500A of the playing video 500, the voice information in the excerpt video 500A is obtained, and the voice information is converted into text information through the voice-to-text function and displayed in the text display area 551 below the excerpt video 500A, and the specific is not limited.

[0083] In some embodiments, such as Figure 5E and 5F shown, when the user wants to display the note application in full screen mode, the user can click the Figure 5E shown full screen button 552. When the note application detects that the user clicks the full screen button 552, it jumps to the full screen display interface. For example, the tablet 100 jumps from the Figure 5E shown display interface 550 to the Figure 5F shown display interface 560 of the note application in full screen display.

[0084] (2) Video note generation process

[0085] In some embodiments, after the note application generates the excerpt video 500A, it first obtains the voice information, time information of the audio data in the excerpt video 500A, and the association relationship between the voice information and the time information.

[0086] For example, as Figure 5F shown, the voice information of the pronunciation of "the earth is" corresponds to the 6th second of the time information, the voice information of the pronunciation of "covered" corresponds to the 7th second of the time information, etc. It can be understood that the association relationship between the voice information and the time information can be the correspondence between the voice information of a single word pronunciation and the time information, or the correspondence between the voice information of multiple word pronunciations and the time information, and the specific is not limited.

[0087] In some embodiments, after the note application obtains the voice information, time information, and the association relationship between the voice information and the time information, it converts the voice information into text information, and associates the converted text information with the time information corresponding to the voice information, to obtain the correspondence between the text information and the time information.

[0088] For example, as Figure 5E shown, the voice information of the pronunciation "the earth is" can be converted into the text "the earth is" in the text information, and the voice information of the pronunciation "covered" can be converted into the text "covered" in the text information, etc. The note application obtains the correspondence between the text information and the time information based on the association relationship between the voice information and the time information. For example, the text "the earth is" in the text information corresponds to the 6th second of the time information, and the text "covered" in the text information corresponds to the 7th second of the time information, etc.

[0089] In some embodiments, Figure 5F the correspondence between the text information and the time information shown above can be represented by the following structure A in JSON format as follows:

[0091] {

[0092] "index":4,

[0093] "name":"1234",

[0094] "content":"the earth is covered",

[0095] "list":

[0096] {

[0097] "startOffset":4,

[0098] "endOffset":6,

[0099] "second":6

[0100] },

[0101] {

[0102] "startOffset":7,

[0103] "endOffset":8,

[0104] "second":7

[0105] }

[0107] } ​​​

[0109] In some embodiments, the structure A includes an index member 11, a name member 12, a content member 13, and a relationship list member 14. The list member 14 includes two arrays. The first array 14A includes a startOffset element 1, an endOffset element 2, and a time axis information (second) element 3. The second array 14B includes a startOffset element 4, an endOffset element 5, and a second element 6.

[0110] In some embodiments, the index member 11 represents the starting index position of the text, the name member 12 represents the name, the content member 13 represents the text information in the text display area 551, and the list member 14 represents the correspondence between the text information and the time information; the startOffset element 1 / 4 represents the starting position of the text, the endOffset element 2 / 5 represents the ending position of the text, and the second element 3 / 6 represents the time corresponding to the text between the starting position and the ending position of the text.

[0111] In some embodiments, in combination with the above Figure 5F description and the description of the structure A:

[0112] In the structure A: the starting index position of the text of the index member 11 is the fourth character "大" in the text display area 551; the name member 12 is the video name "1234" of the excerpted video 500A; the content member 13 of the text information is the text "大地被覆盖" in the text display area 551; the list member 14 represents the correspondence between the text "大地被覆盖" and the time information.

[0113] In the list member 14: the first array 14A: {"startOffset":4,"endOffset":6,"second":6}, which represents Figure 5F the 6th second of the time information corresponding to the fourth to sixth characters "大地被" in the shown text display area 551; the second array 14B: {"startOffset":7,"endOffset":8,"second":7}, which represents Figure 5F the 7th second of the time information corresponding to the seventh and eighth characters "覆盖" in the shown text display area 551.

[0114] (3) Video note editing process

[0115] In some embodiments, when the user discovers that there are missing characters in the text information in the text display area 551 based on the video content of the excerpted video 500A, the user can click on the text display area 551 to trigger the note application to open the modification function, enabling the user to perform an editing operation of adding text at the corresponding position where characters are missing.

[0116] In some embodiments, for example Figure 5G As shown, when the user discovers that the character "ice and snow" is missing between the characters "the earth is covered by" in the text information of the text display area 551, the user can click on the position 51A between the characters "the earth is covered by" and "covered" in the text display area 551 to trigger the note application to open the modification function.

[0117] In some embodiments, when the note application detects that the user clicks on the position 51A in the text display area 551, a cursor 52 is displayed between the characters "the earth is covered by" and "covered", enabling the user to determine the position for adding text based on the position of the cursor 52.

[0118] In some embodiments, for example Figure 5H As shown, after the user determines that the position for adding text is the position of the cursor 52, the missing characters "ice and snow" are edited and added between the characters "the earth is covered by" and "covered". At this time, the cursor 52 moves to the position 51B according to the user's editing operation. It can be understood that relative to Figure 5G the position 51A shown, Figure 5H the position 51B shown moves two positions backward.

[0119] In some embodiments, after the user completes the editing operation, the user can click on any position outside the text display area 551 to end the editing operation. For example Figure 5I As shown, when the note application detects that the user clicks on the position 53 outside the text display area 551, the note application hides the cursor 52 in the text display area 551, and the user ends the editing operation.

[0120] In some embodiments, since the user adds two characters "ice and snow" between the characters "the earth is covered by" and "covered", all the characters after the character "covered" in the text display area 551 move two positions backward as a whole. For example, before adding the characters, Figure 5G as shown, the character positions of the character "covered" are the seventh and eighth; after adding the characters, Figure 5H as shown, the character positions of the character "covered" are the ninth and tenth.

[0121] In some embodiments, when the user triggers the modification function of the note application, the note application first backs up the text information before modification. For example, the video note pairs Figure 5FThe text information in the text display area 551 is backed up. Secondly, the note application determines the position and number of the added text according to the change in the cursor position before and after the user's modification. For example, before modification Figure 5G The cursor 52 is shown between the sixth and seventh characters. Figure 5H The cursor 52 is shown to be located between the eighth and ninth characters, i.e., the position of the cursor 52 after modification is moved two positions backward relative to the position before modification, and the number of characters can be determined to be two characters added based on the change in the cursor position. Then, the note application modifies the backed-up text information based on the position and number of characters added, and adjusts the position of the characters after the characters added in the text information accordingly. For example, Figure 5G After adding two more words "ice and snow" to the words "the earth is covered" and "covered", Figure 5H The text "overlay" is shown relative to Figure 5G The text position of the text "overlay" shown is moved back two places. Finally, the note application adjusts the text based on the modified text information. Figure 5F The text information in the text display area 551 is updated to obtain Figure 5H The text information of the text display area 551 is shown.

[0122] In some embodiments, Figure 5H The correspondence between the text information and the time information can be expressed as follows using the following JSON format structure B: [

[0124] {

[0125] "index":4,

[0126] "name":"1234",

[0127] "content":"The earth is covered with ice and snow",

[0128] "list":[

[0129] {

[0130] "startOffset":4,

[0131] "endOffset":6,

[0132] "second":6

[0133] },

[0134] {

[0135] "startOffset":7,

[0136] "endOffset": 8,

[0137] "second": -1

[0138] },

[0139] {

[0140] "startOffset": 9,

[0141] "endOffset": 10,

[0142] "second": 7

[0143] }

[0145] }

[0147] In some embodiments, structure B includes member index 21, member name 22, member content 23, and member list 24. Member list 24 includes three arrays. The first array 24A includes element startOffset 01, element endOffset 02, and element second 03; the second array 24B includes element startOffset 04, element endOffset 05, and element second 06; the third array 24C includes element startOffset 07, element endOffset 08, and element second 09.

[0148] In some embodiments, in combination with the above Figure 5H description and the description of structure B:

[0149] In structure B: The starting index position of the text of member index 21 is the fourth character "big" in the text display area 551; member name 22 is the video name "1234" of the excerpted video 500A; the text information of member content 23 is the text "The earth is covered with ice and snow" in the text display area 551; member list 24 represents the correspondence between the text "The earth is covered" and the time information.

[0150] In member list 24: The first array 24A: {"startOffset": 4, "endOffset": 6, "second": 6}, indicating Figure 5H that the fourth to sixth characters "The earth is" in the shown text display area 551 correspond to the 6th second of the time information; the second array 24B: {"startOffset": 7, "endOffset": 8, "second": -1}, indicating Figure 5H ​​In the seventh to eighth characters "ice and snow" in the text display area 551 shown, there is no corresponding time; the third array 24C: {"startOffset": 9, "endOffset": 10, "second": 7}, indicating Figure 5H the ninth and tenth characters "cover" in the text display area 551 shown correspond to the 7th second of the time information.

[0151] In some embodiments, Figure 5H Although there is no corresponding time for the seventh to eighth characters "ice and snow" in the text display area 551 shown, during the playback of the excerpt video 500A, the characters "ice and snow" can change correspondingly with the adjacent characters as the video screen changes. For the specific process, refer to the above description of the modification strategy for adding text and the video-text synchronization effect, which will not be elaborated here.

[0152] The following gives an example of the cursor position change for different editing operations.

[0153] In some embodiments, the cursor position before the user's editing operation includes the forward position before the cursor before editing (before start) and the backward position after the cursor before editing (before end); the cursor position after the user's editing operation includes the forward position after the cursor after editing (after start) and the backward position after the cursor before editing (after end).

[0154] For example, as Figure 5G shown, for an editing operation where the user does not select text, before start is the same as before end, and both before start and before end of the cursor before the user's editing operation are at position 51A.

[0155] Again, for example, as Figure 5J shown, for an editing operation where the user does not select text, before start is not the same as before end. Before start of the cursor before the user's editing operation is the third position corresponding to the character "de", and before end is the seventh position corresponding to the character "fu".

[0156] (1) The cursor position moves forward (after end < before end)

[0157] In some embodiments, when the user performs an editing operation of deleting text by the backspace key without selecting text, before start is equal to before end, and after end is less than before end, that is, end moves forward, indicating that the cursor position moves forward. Among them, the changes of the cursor and the text are as follows:

[0158] a. All JSON objects before after end remain unchanged. In other words, the text position of all text information before the cursor's backward position after the user's editing operation, as well as the correspondence between the text information and the time information, remain unchanged;

[0159] b. All JSON objects between after end and before end are deleted. In other words, all text information between the cursor's backward position after the user's editing operation and the cursor's forward position before the editing operation, as well as the correspondence between the text information and time information, are deleted;

[0160] c. All JSON objects after the "before end" are shifted forward by the difference between the "after end" and "before end" values. In other words, the text positions of all text after the cursor's backward position before the user edits are shifted forward by the number of characters in the deleted text, but the correspondence between the text and the time information remains unchanged.

[0161] In other embodiments, when a user performs an editing operation to replace a selected text with a larger number of characters, if before start is not equal to before end and after start is greater than before start, it indicates that the larger number of characters replaces the smaller number of characters; and after end is less than before end, that is, end is forward, indicating that the cursor position moves forward. The changes of the cursor and the text are as follows:

[0162] a. All JSON objects before the before start do not change. In other words, the text position of all text information before the cursor's forward position before the user's editing operation, as well as the correspondence between the text information and the time information, do not change;

[0163] b. All JSON objects between before start and before end are deleted. In other words, all text information between the cursor's forward and backward positions before the user edits the data, as well as the correspondence between the text information and the time information, are deleted.

[0164] c. All JSON objects after "before end" are shifted forward by the difference between "after end" and "before end". In other words, the text positions of all text information after the cursor's backward position before the user's editing operation are shifted forward by the difference between the deleted text information and the added text information, but the correspondence between the shifted text information and the time information does not change;

[0165] d. Add a JSON object between before start and after start. That is, add text information between the cursor's forward position before the user's editing operation and the cursor's forward position after the user's editing operation to replace the deleted text information, but do not add a new correspondence between the added text information and the time information.

[0166] (2) The cursor position moves backward (after end>before end)

[0167] In some embodiments, when the user performs an editing operation to add text to unselected text, before start is equal to before end, and after end is greater than before end, that is, end is backward, indicating that the cursor position moves backward. The changes of the cursor and text are as follows:

[0168] a. All JSON objects before the before end remain unchanged. In other words, the text position of all text information before the cursor's backward position before the user's editing operation, as well as the correspondence between the text information and the time information, remain unchanged;

[0169] b. All JSON objects after the before end are shifted backward by the difference between the after end and before end. In other words, the text position of all text information after the cursor's backward position before the user's editing operation will be shifted backward by increasing the number of text digits, but the correspondence between the moved text information and the time information does not change;

[0170] c. Add a JSON object between before end and after end. That is, add text information between the cursor's backward position before and after the user's editing operation, but do not add a new correspondence between the added text information and the time information.

[0171] In other embodiments, when a user performs an editing operation to replace a larger text with a smaller text, before start is not equal to before end, after start is less than before start, indicating that the smaller text replaces the larger text; and after end is greater than before end, i.e., end is backward, indicating that the cursor position moves backward. The changes of the cursor and text are as follows:

[0172] a. All JSON objects before the before start do not change. In other words, the text position of all text information before the cursor's forward position before the user's editing operation, as well as the correspondence between the text information and the time information, do not change;

[0173] b. All JSON objects between before start and before end are deleted. In other words, all text information between the cursor's forward and backward positions before the user edits the data, as well as the correspondence between the text information and the time information, are deleted.

[0174] c. All JSON objects after the before end are shifted backward by the difference between the after end and before end. In other words, the text positions of all text information after the cursor's backward position before the user's editing operation are shifted backward by the difference between the deleted text information and the added text information, but the correspondence between the shifted text information and the time information does not change;

[0175] d. Add a JSON object between before start and after end. That is, add text information between the forward position of the cursor before the user's editing operation and the backward position of the cursor after the user's editing operation to replace the deleted text information, but do not add a new correspondence between the added text information and the time information.

[0176] (3) The cursor position remains unchanged

[0177] In some embodiments, when the user performs an editing operation to delete unselected text by pressing the delete key, before start is equal to before end, and after end is equal to before end, that is, end remains unchanged, indicating that the cursor position remains unchanged.

[0178] The cursor and text changes are as follows:

[0179] a. Set before end to before end plus the number of characters in the deleted text message;

[0180] b. All JSON objects before after end remain unchanged. In other words, the text position of all text information before the cursor's backward position after the user's editing operation, as well as the correspondence between the text information and the time information, remain unchanged;

[0181] c. All JSON objects between after end and before end are deleted. In other words, all text information between the cursor's backward position after the user's editing operation and the cursor's forward position before the editing operation, as well as the correspondence between the text information and time information, are deleted;

[0182] d. All JSON objects after the before end are shifted forward by the difference between the after end and before end. In other words, the text positions of all text after the cursor's backward position before the user edits are shifted forward by the number of characters in the deleted text, but the correspondence between this text and the time information does not change.

[0183] In other embodiments, when the user performs an editing operation to replace a character with the same number of digits in the selected text, before start is not equal to before end, and after end is equal to before end, that is, end remains unchanged. The cursor and text change as follows:

[0184] a. All JSON objects before the before start do not change. In other words, the text position of all text information before the cursor's forward position before the user's editing operation, as well as the correspondence between the text information and the time information, do not change;

[0185] b. All JSON objects between before start and before end are deleted. In other words, all text information between the cursor's forward and backward positions before the user edits the data, as well as the correspondence between the text information and the time information, are deleted.

[0186] c. Add a JSON object between before start and after start. That is, add text information between the cursor's forward position before the user's editing operation and the cursor's forward position after the user's editing operation to replace the deleted text information, but do not add a new correspondence between the added text information and the time information;

[0187] d. All JSON objects after before end remain unchanged. In other words, the text position of all text information before the cursor's backward position before the user's editing operation, as well as the correspondence between this text information and time information, remain unchanged.

[0188] Figure 6A According to some embodiments of the present application, a software structure diagram of the tablet 100 is shown. Figure 6AAs shown, the layered architecture divides the software into several layers, each with a clear role and division of labor. Layers communicate with each other through software interfaces. In some embodiments, the Android system applied to tablet 100 can be divided into four layers: from top to bottom, the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.

[0189] In some embodiments, the application layer may include a series of application packages. Figure 6A As shown, the application package can include applications such as camera, gallery, calendar, call, map, WLAN, music, short message, video, and video notes. The application framework layer provides the application programming interface (API) and programming framework for the applications in the application layer. The application framework layer includes some predefined functions.

[0190] In some embodiments, as Figure 6B As shown, the module of the note application 600 in the application layer includes at least a video display area 601, a text information display area 602, a text display area 603, and a timestamp area 604. Figure 5H As shown, the video display area 601 is used to present the interface of the excerpt video 500A, the text information display area 602 is used to present the interface of the text display area 551, and the text display area 603 is used to present the text information of the text display area 551; the timestamp area 604 is used to save the correspondence between the text information and the time information.

[0191] In some embodiments, the application framework layer may include a text display area editing module (VideoVoiceEdit), a text editing module (EditerUtils), an editing control module (EditorControl), and the like.

[0192] In some embodiments, the text display area editing module is used to back up the text information before modification, for example Figure 5F The text information in the text display area 551 is backed up. The text display area editing module is also used to monitor the cursor position changes and text modification changes, such as Figure 5G The cursor and text shown are relative Figure 5H The text editing module is used to modify the backed up text information according to the user's editing operation, for example, Figure 5G After adding two more words "ice and snow" to the words "the earth is covered" and "covered", Figure 5H The text "overlay" is shown relative to Figure 5GThe text position of the word "cover" shown is moved back two positions. The editing control module is used to update the text information of the text display area according to the modified text information, for example, Figure 5F The text information in the text display area 551 is updated to obtain Figure 5H The text information of the text display area 551 is shown.

[0193] In some embodiments, the Android Runtime includes core libraries and a virtual machine. The Android runtime is responsible for scheduling and management of the Android system. The core library consists of two parts: one for Java language functions and the other for the Android core library. The application layer and application framework layer run in the virtual machine. The virtual machine executes the Java files in the application layer and application framework layer as binary files. The virtual machine is responsible for performing functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.

[0194] In some embodiments, the system library may include multiple functional modules, such as a surface manager, a media library, a 3D graphics processing library (such as OpenGL ES), a 2D graphics engine (such as SGL), etc.

[0195] The surface manager is used to manage the display subsystem and provide fusion of 2D and 3D layers for multiple applications.

[0196] The media library supports playback and recording of a variety of common audio and video formats, as well as static image files. The media library can support a variety of audio and video encoding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.

[0197] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.

[0198] A 2D graphics engine is a drawing engine for 2D drawings.

[0199] The kernel layer is the layer between hardware and software. The kernel layer includes at least display driver, camera driver, audio driver, and sensor driver.

[0200] I understand. Figure 6A The software structure of the tablet 100 shown is only an example. In other embodiments, the software structure of the tablet 100 may include more or fewer layers, each layer may include more or fewer modules, or may be another operating system structure, such as Android. TM 、iOS TM, Windows TM etc., not limited here.

[0201] The following combination Figures 5F to 5H The scene shown and Figure 6A and 6B The structure shown is used to illustrate some embodiments of the present application.

[0202] Figure 7A According to some embodiments of the present application, a flowchart of a text display method is shown. Figure 7A As shown, the text display method includes the following steps:

[0203] S721: Extract the played video and generate video notes.

[0204] In some embodiments, as Figures 5A to 5E As shown, after the user opens the video 500 played through the video application in the tablet 100, he can use the note application in the tablet 100 to excerpt the video 500 to obtain an excerpt video 500A. At the same time, the note application converts the voice information in the excerpt video 500A into text information, and generates a video note corresponding to the played video 500 based on the excerpt video 500A and the converted text information.

[0205] S722: Back up the initial text information and the initial relationship list.

[0206] In some embodiments, after the note application generates the video note corresponding to the played video 500, the text information in the video note and the correspondence between the text information and the time information are backed up. Figure 5F As shown in structure A, after the note application generates a video note, the note application edits the text display area 701 to Figure 5F The text information (content) member 13 and the relationship list (list) member 14 in the structure A are backed up to serve as the initial text information and the initial relationship list.

[0207] S723: Detecting a user editing operation, obtaining updated text information and an updated relationship list based on the editing operation.

[0208] In some embodiments, the user can edit the text information in the video note, such as adding text, deleting text, and replacing text. The note application determines the type of user editing operation based on the changes in the cursor and text during the user editing operation, and obtains the new text information and relationship list after the user editing operation. For example, Figure 5HAs shown in structure B, based on the user's editing operation, the text information (content) member 23 and the relationship list (list) member 24 in structure B are obtained.

[0209] In some embodiments, when the note application detects that the cursor position moves forward, the user editing operation includes two cases:

[0210] 1. The user deletes unselected text by pressing the backspace key. In response to this edit, the note-taking app leaves all JSON objects before the "after end" unchanged, deletes all JSON objects between the "after end" and "before end" sections, and shifts all JSON objects after the "before end" section forward by the number of digits between the "after end" and "before end" sections.

[0211] That is to say, the text position of all text information before the backward position of the cursor after the user's editing operation, and the correspondence between the text information and the time information do not need to change; all text information between the backward position of the cursor after the user's editing operation and the forward position of the cursor before the editing operation, and the correspondence between the text information and the time information are deleted; the text position of all text information after the backward position of the cursor before the user's editing operation is moved forward by the number of character bits of the deleted text information, but the correspondence between the text information and the time information does not need to change.

[0212] 2. The user selects a text and performs an edit operation to replace a longer text with a shorter text. In response to this edit operation, the note-taking application controls all JSON objects before before start to remain unchanged, deletes all JSON objects between before start and before end, moves all JSON objects after before end forward by the number of digits between afterend and before end, and adds a JSON object between before start and after start.

[0213] That is to say, the text position of all text information before the forward position of the cursor before the user's editing operation, and the correspondence between the text information and the time information do not need to be changed; all text information between the forward position and backward position of the cursor before the user's editing operation, and the correspondence between the text information and the time information are deleted; the text position of all text information after the backward position of the cursor before the user's editing operation is moved forward by the number of text bits that differ between the deleted text information and the added text information, but the correspondence between the moved text information and the time information does not need to be changed; text information is added between the forward position of the cursor before the user's editing operation and the forward position of the cursor after the user's editing operation to replace the deleted text information, but no new correspondence with the time information is added to the added text information.

[0214] In some other embodiments, when the note application detects that the cursor position moves backward, the user editing operation includes two cases:

[0215] 1. The user edits unselected text by adding text. In response to this edit, the note-taking app leaves all JSON objects before the "before end" unchanged, shifts all JSON objects after the "before end" backward by the number of digits between the "after end" and "before end" sections, and adds the JSON object between the "before end" and "after end" sections.

[0216] That is to say, the text positions of all text information before the backward position of the cursor before the user's editing operation, and the correspondence between the text information and the time information do not need to be changed; the text positions of all text information after the backward position of the cursor before the user's editing operation, moving backward requires increasing the number of character bits of the text information, but the correspondence between the moved text information and the time information does not need to be changed; text information is added between the backward position of the cursor before the user's editing operation and the backward position of the cursor after the user's editing operation, but no new correspondence between the added text information and the time information is added.

[0217] 2. The user selects a text and performs an edit operation to replace a larger text with a smaller text. In response to this edit operation, the note-taking application controls all JSON objects before before start to remain unchanged, deletes all JSON objects between before start and before end, shifts all JSON objects after before end backward by the number of digits between afterend and before end, and adds a JSON object between before start and after end.

[0218] That is to say, the text position of all text information before the forward position of the cursor before the user's editing operation, and the correspondence between the text information and the time information do not need to be changed; all text information between the forward position and backward position of the cursor before the user's editing operation, and the correspondence between the text information and the time information are deleted; the text position of all text information after the backward position of the cursor before the user's editing operation is moved backward by the number of text bits that differ between the deleted text information and the added text information, but the correspondence between the moved text information and the time information does not need to be changed; text information is added between the forward position of the cursor before the user's editing operation and the backward position of the cursor after the user's editing operation to replace the deleted text information, but no new correspondence with the time information is added to the added text information.

[0219] In some other embodiments, when the note application detects that the cursor position remains unchanged but the text changes, the user editing operation includes two cases:

[0220] 1. The user deletes unselected text using the Delete key. In response to this edit, the Notes app sets the "before end" to the value of "before end" plus the number of characters in the deleted text. All JSON objects before "after end" remain unchanged, all JSON objects between "after end" and "before end" are deleted, and all JSON objects after "before end" are shifted forward by the difference in the number of characters between "after end" and "before end."

[0221] That is to say, the text position of all text information before the backward position of the cursor after the user's editing operation, and the correspondence between the text information and the time information do not need to change; all text information between the backward position of the cursor after the user's editing operation and the forward position of the cursor before the editing operation, and the correspondence between the text information and the time information are deleted; the text position of all text information after the backward position of the cursor before the user's editing operation is moved forward by the number of character bits of the deleted text information, but the correspondence between the text information and the time information does not need to change.

[0222] 2. The user performs an edit operation to replace the selected text with the same number of characters. In response to this edit operation, the note-taking application controls all JSON objects before "before start" to remain unchanged, deletes all JSON objects between "before start" and "before end", adds a JSON object between "before start" and "after start", and leaves all JSON objects after "before end" unchanged.

[0223] That is to say, the text position of all text information before the forward position of the cursor before the user's editing operation, and the correspondence between the text information and the time information do not need to be changed; all text information between the forward position and backward position of the cursor before the user's editing operation, and the correspondence between the text information and the time information are deleted; text information is added between the forward position of the cursor before the user's editing operation and the forward position of the cursor after the user's editing operation to replace the deleted text information, but no new correspondence with the time information is added to the added text information; the text position of all text information before the backward position of the cursor before the user's editing operation, and the correspondence between the text information and the time information do not need to be changed.

[0224] S724: Detect whether the user continues the editing operation.

[0225] In some embodiments, when the user continues to perform editing operations on the text display area 551, the cursor position of the cursor 52 in the text display area 551 and the text of the text information will change. The note application determines that the user continues to perform editing operations based on the cursor position and text changes, and jumps to step S723.

[0226] In other embodiments, when the user clicks anywhere outside the text display area 551, for example Figure 5I When the position shown is 53, jump to step S725.

[0227] S725: Update the initial text information to the updated text information, and update the initial relationship list to the updated relationship list.

[0228] In some embodiments, for example Figure 5I As shown, when the user wants to end the editing operation, the user can click on a position 53 outside the text display area 551. Based on the click operation at position 53, the note application hides the cursor 52 in the text display area 551 and determines that the user has ended the editing operation. After the note application determines that the user has ended the editing operation, the initial text information in the text display area editing module 701 is updated to the updated text information, and the initial relationship list is updated to the updated relationship list. For example, the text information (content) member 13 and the relationship list (list) member 14 in structure A are updated to: the text information (content) member 23 and the relationship list (list) member 24 in structure B, respectively, so that Figure 5H The edited text "ice and snow" can change accordingly with the adjacent text as the video picture changes.

[0229] The following example uses the editing operation scenario of adding text to unselected text as an example. Figure 7A S722 to S725 are further illustrated.

[0230] Figure 7B According to some embodiments of the present application, an interaction flowchart of a text display method is shown. As Figure 7B shown, the text display method includes the following steps:

[0231] S711: Back up the initial text information and the initial relationship list.

[0232] In some embodiments, as Figure 5F and the structure A shown, after the note application generates a video note, the note application uses the text display area editing module 701 to Figure 5F back up the text information (content) member 13 and the relationship list (list) member 14 in the shown structure A as the initial text information and the initial relationship list.

[0233] S712: Register for focus acquisition event listening, content change event listening, and focus loss event listening.

[0234] In some embodiments, the text display area editing module 701 registers for focus acquisition event listening, content change event listening, and focus loss event listening.

[0235] In some embodiments, the focus acquisition event can be used to represent an operation where the user clicks on any position in the text display area 551, causing the text display area 551 to display a cursor at the corresponding position. For example, as Figure 5G shown, when the user clicks on the position 51A between the text "The earth is covered by" and the text "covered" in the text display area 551, a cursor 52 is displayed between the text "The earth is covered by" and the text "covered".

[0236] The content change event can be used to represent an editing operation by the user on the text information in the text display area 551. For example, as Figure 5G and 5H shown, the user edits and adds the text "ice and snow" between the text "The earth is covered by" and the text "covered".

[0237] The focus loss event can be used to represent an operation where the user clicks on any position outside the text display area 551, causing the text display area 551 to hide the cursor. For example, as Figure 5I shown, when the user clicks on the position 53 outside the text display area 551, the cursor 52 in the text display area 551 disappears.

[0238] S713: When the focus acquisition event is monitored, record the cursor position before the content change.

[0239] In some embodiments, when the text display area editing module 701 detects that the user clicks on any position in the text display area 551, a focus acquisition event is triggered. The text display area editing module 701 displays a cursor at the position corresponding to the user's click in the text display area 551 and records the cursor position of the cursor. For example, as Figure 5G shown, when the user clicks on the position 51A between the text "The earth is covered by" and the text "covered" in the text display area 551, the text display area editing module 701 displays the cursor 52 between the text "The earth is covered by" and the text "covered", and records the current cursor position of the cursor 52 as the position 51A.

[0240] It can be understood that since the user has not selected the text at this time, that is, the cursor positions before the user's editing operation, including the forward and backward positions of the cursor before editing, are the same, both being the position 51A.

[0241] S714: When a content change event is monitored, record the cursor position after the content change, and obtain the updated text information corresponding to the content change event.

[0242] In some embodiments, when the text display area editing module 701 detects an editing operation of the user on the text information in the text display area 551, a content change event is triggered. The text display area editing module 701 records the changed cursor position according to the position change of the cursor, and obtains the text information in the text display area 551 after the user's editing operation as the updated text information. For example, as Figure 5G and 5H shown, after the user clicks on the position 51A between the text "The earth is covered by" and the text "covered", the user edits and adds the text "ice and snow" between the text "The earth is covered by" and the text "covered", so that the cursor 52 changes to the position 51B according to the user's editing operation. At this time, the changed cursor position recorded by the text display area editing module 701 is the position 51B, and the obtained updated text information is Figure 5H the text information shown in the text display area 551.

[0243] It can be understood that since the user edits and adds the text "ice and snow" between the text "The earth is covered by" and the text "covered", the cursor 52 moves backward from the position 51A to the position 51B, that is, the cursor moves backward, and the position difference between the position 51A and the position 51B is the number of digits of the two characters "ice and snow".

[0244] S715: Send the recorded cursor positions before and after the content change and the updated text information.

[0245] In some embodiments, the text display area editing module 701 sends the recorded cursor positions before and after the content change and the updated text information to the text editing module 702.

[0246] S716: Based on the cursor position before and after the content change and the updated text information, obtain and send an update relationship list.

[0247] In some embodiments, the text editing module 702 calculates an update relationship list based on the cursor position before and after the content changes and the updated text information, and sends the update relationship list to the text display area editing module 701. For example, Figure 5H As shown in structure B, the text editing module 702 calculates the relationship list member 24, ie, updates the relationship list, based on the cursor position before and after the content change and the updated text information.

[0248] In some embodiments, when the user continues to perform editing operations on the text display area 551, the cursor position of the cursor 52 in the text display area 551 and the text of the text information will change. The text display area editing module 701 determines that the user continues to perform editing operations based on the cursor position and text changes, and jumps to step S714.

[0249] In other embodiments, when the user clicks anywhere outside the text display area 551, for example Figure 5I When the position shown is 53, jump to step S717.

[0250] S717: A focus loss event is detected.

[0251] In some embodiments, for example Figure 5I As shown, when the user wants to end the editing operation, the user can click a position 53 outside the text display area 551. The text display area editing module 701 hides the cursor 52 in the text display area 551 according to the user's click operation on the position 53 and determines that the user ends the editing operation.

[0252] S718: Send update instruction.

[0253] In some embodiments, after the text display area editing module 701 triggers a focus loss event, the text display area editing module 701 sends an update instruction to the editing control module 703 .

[0254] S719: Update the initial text information to the updated text information, and update the initial relationship list to the updated relationship list.

[0255] In some embodiments, the editing control module 703 receives the update instruction and updates the initial text information in the text display area editing module 701 to the updated text information and the initial relationship list to the updated relationship list. For example, the text information (content) member 13 and the relationship list (list) member 14 in the structure A are updated to the text information (content) member 23 and the relationship list (list) member 24 in the structure B, respectively, so that Figure 5H The text "Ice and Snow" in the editing operation can change accordingly with the adjacent text as the video picture changes. The specific process can be found in the description of adding text and the video-text synchronization effect above, which will not be repeated here.

[0256] Figure 8 According to some embodiments of the present application, a schematic diagram of the hardware structure of the tablet 100 is shown. It should be understood that the structure shown in the embodiments of the present application does not constitute a specific limitation on the tablet 100. In other embodiments of the present application, the tablet 100 may include more or fewer components than shown, or some components may be combined, separated, or arranged differently. The components shown in the diagram may be implemented in hardware, software, or a combination of software and hardware.

[0257] like Figure 8 As shown, the tablet 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, and the like.

[0258] The processor 110 may include one or more processing units, for example: the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units can be independent devices or integrated into one or more processors. The controller can generate an operation control signal based on the instruction opcode and the timing signal to complete the control of instruction fetching and execution. In some embodiments, the processor 110 can execute the instructions corresponding to the text display method provided in the aforementioned embodiments.

[0259] Processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 110 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 110. If processor 110 needs to use the same instruction or data again, it can directly access the memory. This avoids duplicate accesses, reduces processor 110 latency, and thus improves system efficiency.

[0260] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface.

[0261] The I2C interface is a bidirectional synchronous serial bus that includes a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 110 may include multiple I2C bus lines. The processor 110 can be coupled to the touch sensor 180K, the charger, the flash, the camera 193, and the like via different I2C bus interfaces. For example, the processor 110 can be coupled to the touch sensor 180K via the I2C interface, enabling communication between the processor 110 and the touch sensor 180K via the I2C bus interface, thereby implementing the touch function of the tablet 100.

[0262] The I2S interface can be used for audio communication. In some embodiments, the processor 110 can include multiple I2S buses. The processor 110 can be coupled to the audio module 170 via the I2S bus to enable communication between the processor 110 and the audio module 170. In some embodiments, the audio module 170 can transmit audio signals to the wireless communication module 160 via the I2S interface, enabling the function of answering calls through a Bluetooth headset.

[0263] The PCM interface can also be used for audio communication, sampling, quantizing, and encoding analog signals. In some embodiments, the audio module 170 and the wireless communication module 160 can be coupled via a PCM bus interface. In some embodiments, the audio module 170 can also transmit audio signals to the wireless communication module 160 via the PCM interface, enabling the function of answering calls via a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication.

[0264] The UART interface is a universal serial data bus used for asynchronous communication. This bus can be a bidirectional communication bus. It converts the data to be transmitted between serial communication and parallel communication. In some embodiments, the UART interface is generally used to connect the processor 110 and the wireless communication module 160. For example, the processor 110 communicates with the Bluetooth module in the wireless communication module 160 via the UART interface to implement Bluetooth functionality. In some embodiments, the audio module 170 can transmit audio signals to the wireless communication module 160 via the UART interface, implementing the function of playing music through a Bluetooth headset.

[0265] The MIPI interface can be used to connect the processor 110 to peripheral devices such as the display 194 and the camera 193. MIPI interfaces include the camera serial interface (CSI) and the display serial interface (DSI). In some embodiments, the processor 110 and the camera 193 communicate via the CSI interface to enable the tablet 100's camera function. The processor 110 and the display 194 communicate via the DSI interface to enable the tablet 100's display function.

[0266] The GPIO interface can be configured through software. The GPIO interface can be configured as a control signal or a data signal. In some embodiments, the GPIO interface can be used to connect the processor 110 to the camera 193, the display 194, the wireless communication module 160, the audio module 170, the sensor module 180, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.

[0267] USB interface 130 is an interface that complies with USB standards and specifications, and may be a Mini USB interface, a Micro USB interface, a USB Type-C interface, or the like. USB interface 130 can be used to connect a charger to charge tablet 100, or to transfer data between tablet 100 and peripheral devices. It can also be used to connect headphones to play audio. This interface can also be used to connect other electronic devices, such as AR devices.

[0268] It should be understood that the interface connection relationship between the modules illustrated in the embodiments of the present application is merely illustrative and does not constitute a structural limitation on the tablet 100. In other embodiments of the present application, the tablet 100 may also adopt different interface connection methods from the above embodiments, or a combination of multiple interface connection methods.

[0269] The charging management module 140 is configured to receive charging input from a charger. The charger can be either a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 can receive charging input from the wired charger via the USB interface 130. In some wireless charging embodiments, the charging management module 140 can receive wireless charging input via the wireless charging coil of the tablet 100. While charging the battery 142, the charging management module 140 can also provide power to the tablet 100 via the power management module 141.

[0270] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, and provides power to the processor 110, the internal memory 121, the display 194, the camera 193, and the wireless communication module 160. The power management module 141 can also be used to monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage, impedance). In some other embodiments, the power management module 141 can also be set in the processor 110. In other embodiments, the power management module 141 and the charging management module 140 can also be set in the same device.

[0271] The wireless communication functionality of tablet 100 can be implemented using antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, a modem processor, and a baseband processor. Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in tablet 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In other embodiments, the antennas can be used in conjunction with a tuning switch.

[0272] The mobile communication module 150 can provide solutions for wireless communications including 2G / 3G / 4G / 5G / 6G applied on the tablet 100. The mobile communication module 150 may include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves from the antenna 1, and filter, amplify, and process the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves for radiation through the antenna 1. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the processor 110. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the same device as at least some of the modules of the processor 110.

[0273] The modem processor may include a modulator and a demodulator. The modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is passed to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 170A, the receiver 170B, etc.) or displays an image or video through the display screen 194. In some embodiments, the modem processor may be an independent device. In other embodiments, the modem processor may be independent of the processor 110 and be provided in the same device as the mobile communication module 150 or other functional modules.

[0274] The wireless communication module 160 can provide wireless communication solutions applied on the tablet 100, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, frequency modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 160 can also receive the signal to be sent from the processor 110, frequency modulate it, amplify it, and convert it into electromagnetic waves for radiation through the antenna 2.

[0275] In some embodiments, antenna 1 of tablet 100 is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, so that tablet 100 can communicate with a network and other devices via wireless communication technologies. The wireless communication technologies may include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), 5G and subsequent evolution standards, BT, GNSS, WLAN, NFC, FM, and / or IR technology. GNSS may include the global positioning system (GPS), the global navigation satellite system (GLONASS), the Beidou navigation satellite system (BDS), the quasi-zenith satellite system (QZSS) and / or the satellite-based augmentation system (SBAS).

[0276] Tablet 100 implements display functionality through a GPU, display screen 194, and an application processor. The GPU is a microprocessor for image processing that connects display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs that execute program instructions to generate or modify display information.

[0277] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. In some embodiments, the tablet 100 may include one or N display screens 194, where N is a positive integer greater than one.

[0278] Tablet 100 can implement shooting functions through the ISP, camera 193, video codec, GPU, display 194, and application processor. The ISP is used to process data fed back by camera 193. For example, when taking a photo, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, which is then transmitted to the ISP for processing and converted into an image visible to the naked eye. The ISP can also perform algorithmic optimization on image noise, brightness, and skin color. The ISP can also optimize parameters such as exposure and color temperature of the captured scene. In some embodiments, the ISP can be set in camera 193.

[0279] The camera 193 is used to capture still images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then passes the electrical signal to the ISP to be converted into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV or other format. In some embodiments, the tablet 100 may include 1 or N cameras 193, where N is a positive integer greater than 1.

[0280] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the tablet 100 is selecting a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy.

[0281] Video codecs are used to compress or decompress digital video. Tablet 100 may support one or more video codecs. This allows tablet 100 to play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, and MPEG4.

[0282] The NPU is a neural network (NN) computing processor. Drawing on the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it rapidly processes input information and can continuously self-learn. The NPU enables intelligent cognitive applications such as image recognition, face recognition, voice recognition, and text comprehension on the tablet 100.

[0283] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 via the external memory interface 120 to implement data storage functions. For example, files such as music and videos can be stored on the external memory card.

[0284] The internal memory 121 can be used to store computer executable program code, which includes instructions. The internal memory 121 may include a program storage area and a data storage area. The program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc. The data storage area may store data created during the use of the tablet 100 (such as audio data, a phone book, etc.), etc. In addition, the internal memory 121 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc. The processor 110 executes various functional applications and data processing of the tablet 100 by running instructions stored in the internal memory 121 and / or instructions stored in a memory provided in the processor. For example, in some embodiments, the internal memory 121 may be used to temporarily store instructions for the text display method provided in the aforementioned embodiments.

[0285] The tablet 100 can implement audio functions such as music playback and recording through the audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor.

[0286] The SIM card interface 195 is used to connect a SIM card. The SIM card can be connected to and disconnected from the tablet 100 by inserting or removing it from the SIM card interface 195. The tablet 100 can support one or N SIM card interfaces, where N is a positive integer greater than one. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, and the like. Multiple cards can be inserted into the same SIM card interface 195 at the same time. The multiple cards can be of the same or different types. The SIM card interface 195 can also be compatible with different types of SIM cards. The SIM card interface 195 can also be compatible with external memory cards. The tablet 100 interacts with the network through the SIM card to implement functions such as calls and data communications. In some embodiments, the tablet 100 uses an eSIM, or embedded SIM card. The eSIM card can be embedded in the tablet 100 and cannot be separated from the tablet 100.

[0287] In the accompanying drawings, some structural or method features may be shown in a particular arrangement and / or order. However, it should be understood that such a particular arrangement and / or order may not be required. Rather, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the illustrative drawings. In addition, the inclusion of a structural or method feature in a particular figure does not imply that such feature is required in all embodiments, and in some embodiments, such features may not be included or may be combined with other features.

[0288] It should be noted that the units / modules mentioned in the various device embodiments of the present application are all logical units / modules. Physically, a logical unit / module can be a physical unit / module, or a part of a physical unit / module, or can be implemented as a combination of multiple physical units / modules. The physical implementation of these logical units / modules themselves is not the most important. The combination of functions implemented by these logical units / modules is the key to solving the technical problems raised by this application. In addition, in order to highlight the innovative part of this application, the above-mentioned device embodiments of this application do not introduce units / modules that are not closely related to solving the technical problems raised by this application. This does not mean that other units / modules do not exist in the above-mentioned device embodiments.

[0289] It should be noted that in the examples and description of this patent, relational terms such as first and second, etc. are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "including a" does not exclude the presence of other identical elements in the process, method, article or device that includes the element.

[0290] While the present application has been shown and described with reference to certain preferred embodiments thereof, it will be understood by those skilled in the art that various changes in form and details may be made therein without departing from the scope of the present application.

Claims

1. A text display method, applied to an electronic device, characterized in that: include: In response to a first user operation, displaying a first interface, and playing a first video of the video application on the first interface; In response to a second user operation, starting a note application, After starting the note-taking application, a third interface is displayed, and the first video is recorded through the note-taking application, wherein the third interface includes a first window of the video application, a second window of the note-taking application, and an operation bar, the first video is played in the first window, the first window and the second window are respectively displayed in two different areas of the third interface, and the operation bar includes operation buttons for pause / stop, voice-to-text, and close; In response to a click operation on the speech-to-text operation button, enabling a speech-to-text function during recording of the first video by the note application, wherein the speech-to-text function is used to convert speech information in the recorded first video into text information; In response to a user clicking the stop operation button in the operation bar, the recording of the first video is stopped, and a fourth interface is displayed, the fourth interface including the first window, a third window of the note-taking application, and the operation bar, the third window including the recorded first video and text information corresponding to the voice information of the recorded first video; In response to a third operation on the close operation button in the operation bar, displaying a fifth interface, wherein the fifth interface includes the first window and the third window, and does not include the operation bar; In response to the user clicking the full-screen button, a second interface is displayed, in which the note application is displayed in full screen, wherein the second interface includes a first display area and a second display area, the first display area is used to play the first recorded video, and text information corresponding to the voice information of the first recorded video is displayed in the second display area; In response to a user clicking a first character in the text information, a video frame corresponding to the first character is displayed in the first display area; After the second character is added to the text message, when the user clicks the second character in the text message, the video frame displayed in the first display area does not change; In response to a user clicking a third character in the text information, a video frame corresponding to the third character is displayed in the first display area, wherein the third character is located after the second character; The first recorded video is played in the first display area. In response to a user clicking a first adjacent character adjacent to the second character, a first video frame corresponding to the first adjacent character is displayed in the first display area, and character styles of characters corresponding to the first video frame and a video frame preceding the first video frame change from a first character style to a second character style in the second display area, and the character style of the second character changes together with the character style of the first adjacent character. After the target character in the text information is deleted, the video frame corresponding to the target character is not deleted.

2. The method according to claim 1, characterized in that The method further comprises: In response to a fourth user operation, playing the recorded first video in the first display area of the second interface; In response to the user clicking the first character in the text information, the video frame corresponding to the first character is displayed in the first display area, including: In response to the user clicking on the first character in the text information, the video playback screen in the first display area jumps to the video frame corresponding to the first character.

3. The method according to claim 1, characterized in that The method further comprises: Register to listen for focus events, content change events, and focus loss events; After adding the second character to the text message, it includes: Based on the focus acquisition event, content change event and focus loss event, the characters and the corresponding video frames are modified so that when the user clicks the characters, the video frames having a first corresponding relationship with the characters are displayed in the first display area, wherein the multiple characters in the text information corresponding to the voice information of the recorded first video respectively have a first corresponding relationship with the multiple video frames of the recorded first video.

4. The method according to claim 3, characterized in that The method further comprises: When the user clicks the second display area, the focus acquisition event is monitored, a first cursor is displayed in the second display area, and a position of the first cursor when the first cursor is displayed in the second display area is recorded; When the user edits the text information, the content change event is monitored and the editing characters corresponding to the editing operation are recorded; When the user clicks the third display area, the focus loss event is monitored, the first cursor in the second display area is hidden, and the position of the second cursor in the second display area when the first cursor is hidden is recorded, wherein the third display area is not the second display area; The modifying of the character and the corresponding video frame based on the focus acquisition event, the content change event, and the focus loss event includes: modifying the character and the corresponding video frame based on the first cursor position, the edit character, and the second cursor position; The modifying of the characters and the corresponding video frames includes: The character position of the character in the corresponding structure is modified, and the first corresponding relationship between the character and the corresponding video frame is modified.

5. The method according to claim 4, characterized in that After adding a second character to the text message; The modifying the character and the corresponding video frame based on the first cursor position, the editing character, and the second cursor position includes: The first cursor moves backward, and the second cursor position is located after the first cursor position; The character position of the character before the first cursor position and the first corresponding relationship with the video frame remain unchanged; The character position of the character after the first cursor position is shifted backward by the number of character bits of the second character; The first correspondence between the character after the first cursor position and the video frame remains unchanged; The second character is between the first cursor position and the second cursor position, and the second character has no corresponding relationship with the video frame.

6. The method according to claim 4, characterized in that After deleting the fourth character in the text message by using the backspace key; The modifying the character and the corresponding video frame based on the first cursor position, the editing character, and the second cursor position includes: The first cursor moves forward, the second cursor position is located before the first cursor position, and the fourth character is located between the first cursor position and the second cursor position; The character position of the character before the second cursor position and the first corresponding relationship with the video frame remain unchanged; deleting the fourth character and the first corresponding relationship between the fourth character and the video frame; The character position of the character after the first cursor position is moved forward by the number of character bits of the fourth character; The first corresponding relationship between the characters after the first cursor position and the video frame remains unchanged.

7. The method according to claim 4, characterized in that After deleting the fourth character in the text message by using the delete key; The modifying the character and the corresponding video frame based on the first cursor position, the editing character, and the second cursor position includes: moving the first cursor position backward by the number of character digits of the fourth character so that the fourth character is located between the first cursor position and the second cursor position; The character position of the character before the second cursor position and the first corresponding relationship with the video frame remain unchanged; deleting the fourth character and the first corresponding relationship between the fourth character and the video frame; The character position of the character after the first cursor position is moved forward by the number of character bits of the fourth character; The first corresponding relationship between the characters after the first cursor position and the video frame remains unchanged.

8. The method according to claim 4, characterized in that After selecting the fifth character in the text information and replacing the fifth character with the sixth character, wherein the number of character digits of the fifth character is less than the number of character digits of the sixth character by a first digit; The modifying the character and the corresponding video frame based on the first cursor position, the editing character, and the second cursor position includes: The first cursor position includes a forward position and a backward position, the fifth character is located between the forward position and the backward position, and the second cursor position is obtained by moving the backward position of the first cursor position forward by the first digit; The character position of the character before the forward position and the first corresponding relationship with the video frame remain unchanged; deleting the fifth character and the first corresponding relationship between the fifth character and the video frame; adding the sixth character between the forward position and the second cursor position, wherein the sixth character has no corresponding relationship with the video frame; The character position of the character after the backward position is shifted forward by the first digit; The first corresponding relationship between the character after the backward position and the video frame remains unchanged.

9. The method according to claim 4, characterized in that After selecting the fifth character in the text information and replacing the fifth character with the sixth character, wherein the number of character digits of the fifth character is greater than the number of character digits of the sixth character by a first digit; The modifying the character and the corresponding video frame based on the first cursor position, the editing character, and the second cursor position includes: The first cursor position includes a forward position and a backward position, the fifth character is located between the forward position and the backward position, and the second cursor position is obtained by moving the backward position of the first cursor position forward by the first digit; The character position of the character before the forward position and the first corresponding relationship with the video frame remain unchanged; deleting the fifth character and the first corresponding relationship between the fifth character and the video frame; adding the sixth character between the forward position and the second cursor position, wherein the sixth character has no corresponding relationship with the video frame; The character position of the character after the backward position is shifted backward by the first digit; The first corresponding relationship between the character after the backward position and the video frame remains unchanged.

10. The method according to claim 4, characterized in that After selecting the fifth character in the text message and replacing the fifth character with the sixth character, wherein the number of character digits of the fifth character is equal to the number of character digits of the sixth character; The modifying the character and the corresponding video frame based on the first cursor position, the editing character, and the second cursor position includes: The first cursor position includes a forward position and a backward position, the fifth character is located between the forward position and the backward position, and the forward position of the first cursor position is the second cursor position; The character position of the character before the forward position and the first corresponding relationship with the video frame remain unchanged; deleting the fifth character and the first corresponding relationship between the fifth character and the video frame; adding the sixth character between the forward position and the second cursor position, wherein the sixth character has no corresponding relationship with the video frame; The character position of the character after the backward position and the first corresponding relationship with the video frame remain unchanged.

11. A computer-readable storage medium, characterized in that The readable storage medium stores instructions, which, when executed on an electronic device, enable the electronic device to implement the method according to any one of claims 1 to 10.

12. An electronic device, characterized in that: include: a memory for storing instructions to be executed by one or more processors of the electronic device; and a processor, which is one of the processors of the electronic device, configured to execute instructions stored in the memory to implement the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Subtitle generation method and device, computer storage medium and electronic equipment

    CN110798733A

  • Editing system

    CN209089103U