Audio text processing methods, devices, storage media and electronic devices
By acquiring structured data of audio-recognized text and storing and displaying the audio-recognized text in a shared storage area, the problem of not being able to listen back after copying and pasting audio text was solved, thus realizing the audio playback function and improving the user experience.
Patent Information
- Application Number
- CN202311118066.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-30
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2043-08-30
AI Technical Summary
In existing technologies, copied and pasted audio text cannot be replayed, resulting in a poor user experience.
By acquiring the structured data corresponding to the audio-recognized text, including audio description information, and storing and displaying the audio-recognized text in a shared storage area, the audio playback function is realized.
This feature enables audio playback for copied and pasted audio text, enriching the text display functionality and enhancing the user experience.
Smart Images

Figure CN117076705B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of computer technology, and in particular relates to an audio text processing method, apparatus, storage medium and electronic device. Background Technology
[0002] With the development of computer technology, the forms of meetings are becoming more and more diverse. Meetings are no longer limited to participants gathering in a unified conference room. They can be held across regions through remote audio and video network conferencing, which facilitates people's work and life.
[0003] Currently, there are some audio and video conferencing applications on the market that can conduct online meetings. These applications can record meeting audio data and obtain meeting text by performing text recognition on the audio data. Users can copy and paste the content of the meeting text to extract parts of the meeting text or merge the content of multiple meeting texts to obtain a new text different from the original meeting text. However, the copied and pasted text obtained by the current technology does not have a playback function, that is, it is impossible to listen to the corresponding meeting audio in each text content of the new text, resulting in a poor user experience. Summary of the Invention
[0004] This application aims to address at least one of the technical problems existing in the prior art. To this end, this application proposes an audio text processing method, apparatus, storage medium, and electronic device that enables copied and pasted audio text to have a playback function.
[0005] Firstly, this application provides an audio text processing method, including:
[0006] In response to a first operation on the target audio recognition text, target structured data corresponding to the target audio recognition text is obtained. The target structured data includes the target audio recognition text and audio description information. The audio description information is used to obtain target audio data corresponding to the target audio recognition text.
[0007] The target structured data is stored in a shared storage area;
[0008] In response to a second operation performed at the target location, the target structured data is retrieved from the shared storage area, and the target audio recognition text is displayed at the target location based on the target structured data, so as to retrieve the target audio data based on the displayed target audio recognition text.
[0009] In some embodiments, storing the target structured data in a shared storage area includes:
[0010] Generate the field content corresponding to the target field based on the target structured data;
[0011] The target field and its content are stored together in the shared storage area.
[0012] The step of obtaining the target structured data from the shared storage area includes: obtaining the field content corresponding to the target field from the shared storage area.
[0013] In some embodiments, generating the field content corresponding to the target field based on the target structured data includes:
[0014] The target structured data is encrypted to obtain encrypted data;
[0015] Use the encrypted data as the field content corresponding to the target field;
[0016] The step of obtaining the target structured data from the shared storage area further includes: decrypting the obtained field content to obtain the target structured data.
[0017] In some embodiments, the shared storage area includes a first storage area and a second storage area, and the step of storing the target field and the field content together in the shared storage area includes:
[0018] The target field and its content are stored together in the first storage area.
[0019] Storing the target structured data in the shared storage area further includes:
[0020] The target audio recognition text from the target structured data is stored in the second storage area.
[0021] In some embodiments, displaying the target audio-recognized text at the target location based on the target structured data includes:
[0022] Determine the text editing window where the target location is located;
[0023] Determine the first structured data storage area corresponding to the text editing window, and determine the shared storage area corresponding to the target location in the first structured data storage area;
[0024] The target structured data is stored in the shared storage area to update the first structured data storage area;
[0025] The content in the text editing window is updated based on the updated first structured data storage area to display the target audio recognition text at the target location.
[0026] In some embodiments, before determining the text editing window where the target location is located, the method further includes:
[0027] In response to a text editing instruction for the first audio-recognized text, all structured data in the second structured data storage area corresponding to the first audio-recognized text is obtained;
[0028] The first structured data storage area is created based on all the acquired structured data;
[0029] The text editing window is generated based on the structured data in the first structured data storage area.
[0030] In some embodiments, the audio text processing method further includes:
[0031] In response to an editing end operation applied to the text editing window, determine the structured data currently stored in the first structured data storage area;
[0032] Based on the structured data currently stored in the first structured data storage area, the second audio-recognized text is displayed through a browsing window.
[0033] In some embodiments, prior to the first operation on the target audio-recognized text, the method further includes:
[0034] Identify the audio recognition text corresponding to the target audio source data, and determine the audio description information corresponding to each text unit in the audio recognition text. The audio recognition text includes at least one text unit, the target audio source data includes the target audio data, and the target audio recognition text is one of the text units.
[0035] Structured data corresponding to each text unit is created based on each text unit and its corresponding audio description information;
[0036] The audio-recognized text is displayed based on the structured data created.
[0037] In some embodiments, the audio text processing method further includes:
[0038] In response to an audio playback command for the target audio-recognized text, the audio description information is extracted from the target structured data;
[0039] Obtain the target audio data corresponding to the audio description information;
[0040] Play the acquired target audio data.
[0041] In some embodiments, the audio description information includes an audio source identifier and an audio start and end time, and obtaining the target audio data corresponding to the audio description information includes:
[0042] Search for the target audio source data corresponding to the audio source identifier from the stored audio source dataset;
[0043] Extract the content corresponding to the start and end times of the audio from the target audio source data, and use it as the target audio data.
[0044] In some embodiments, the audio text processing method further includes:
[0045] In response to the determination of the target audio-recognized text, multiple functional controls are displayed, including a target functional control, which is used to instruct audio playback of the determined text;
[0046] In response to the confirmation operation of the target function control, an audio playback instruction for the target audio recognition text is generated.
[0047] Secondly, this application provides an audio text processing apparatus, comprising:
[0048] The acquisition module is configured to, in response to a first operation on the target audio recognition text, acquire the target structured data corresponding to the target audio recognition text, wherein the target structured data includes the target audio recognition text and audio description information, and the audio description information is used to acquire the target audio data corresponding to the target audio recognition text;
[0049] A storage module for storing the target structured data in a shared storage area;
[0050] The display module is configured to respond to a second operation performed at the target location, retrieve the target structured data from the shared storage area, and display the target audio recognition text at the target location based on the target structured data, so as to retrieve the target audio data based on the displayed target audio recognition text.
[0051] In some embodiments, the storage module is specifically used for:
[0052] Generate the field content corresponding to the target field based on the target structured data;
[0053] The target field and its content are stored together in the shared storage area.
[0054] The display module is specifically used to: obtain the field content corresponding to the target field from the shared storage area.
[0055] In some embodiments, the storage module is specifically used for:
[0056] The target structured data is encrypted to obtain encrypted data;
[0057] Use the encrypted data as the field content corresponding to the target field;
[0058] The display module is further configured to: decrypt the acquired field content to obtain the target structured data.
[0059] In some embodiments, the shared storage area includes a first storage area and a second storage area, and the storage module is specifically used for:
[0060] The target field and its content are stored together in the first storage area.
[0061] The storage module is further configured to: store the target audio recognition text in the target structured data within the second storage area.
[0062] In some embodiments, the display module is specifically used for:
[0063] Determine the text editing window where the target location is located;
[0064] Determine the first structured data storage area corresponding to the text editing window, and determine the shared storage area corresponding to the target location in the first structured data storage area;
[0065] The target structured data is stored in the shared storage area to update the first structured data storage area;
[0066] The content in the text editing window is updated based on the updated first structured data storage area to display the target audio recognition text at the target location.
[0067] In some embodiments, before determining the text editing window where the target location is located, the display module is further configured to:
[0068] In response to a text editing instruction for the first audio-recognized text, all structured data in the second structured data storage area corresponding to the first audio-recognized text is obtained;
[0069] The first structured data storage area is created based on all the acquired structured data;
[0070] The text editing window is generated based on the structured data in the first structured data storage area.
[0071] In some embodiments, the display module is further configured to:
[0072] In response to an editing end operation applied to the text editing window, determine the structured data currently stored in the first structured data storage area;
[0073] Based on the structured data currently stored in the first structured data storage area, the second audio-recognized text is displayed through a browsing window.
[0074] In some embodiments, prior to the first operation of recognizing text from the target audio, the acquisition module is further configured to:
[0075] Identify the audio recognition text corresponding to the target audio source data, and determine the audio description information corresponding to each text unit in the audio recognition text. The audio recognition text includes at least one text unit, the target audio source data includes the target audio data, and the target audio recognition text is one of the text units.
[0076] Structured data corresponding to each text unit is created based on each text unit and its corresponding audio description information;
[0077] The audio-recognized text is displayed based on the structured data created.
[0078] In some embodiments, the audio text processing device further includes a playback module for:
[0079] In response to an audio playback command for the target audio-recognized text, the audio description information is extracted from the target structured data;
[0080] Obtain the target audio data corresponding to the audio description information;
[0081] Play the acquired target audio data.
[0082] In some embodiments, the audio description information includes an audio source identifier and audio start and end times, and the playback module is specifically used for:
[0083] Search for the target audio source data corresponding to the audio source identifier from the stored audio source dataset;
[0084] Extract the content corresponding to the start and end times of the audio from the target audio source data, and use it as the target audio data.
[0085] In some embodiments, the playback module is further configured to:
[0086] In response to the determination of the target audio-recognized text, multiple functional controls are displayed, including a target functional control, which is used to instruct audio playback of the determined text;
[0087] In response to the confirmation operation of the target function control, an audio playback instruction for the target audio recognition text is generated.
[0088] Thirdly, this application provides a non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that the computer program, when executed by a processor, implements the audio text processing method described in any of the above claims.
[0089] Fourthly, this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the program to implement the audio text processing method described in any of the preceding claims.
[0090] Fifthly, this application provides a computer program product, including a computer program, characterized in that, when the computer program is executed by a processor, it implements the audio text processing method described in any of the above claims.
[0091] The audio-text processing method, apparatus, computer program, storage medium, and electronic device provided in this application, in response to a first operation on target audio-recognized text, acquire target structured data corresponding to the target audio-recognized text. The target structured data includes target audio-recognized text and audio description information, the audio description information being used to acquire target audio data corresponding to the target audio-recognized text. The target structured data is stored in a shared storage area. In response to a second operation applied to a target location, the target structured data is acquired from the shared storage area, and the target audio-recognized text is displayed at the target location based on the target structured data. This enables copied and pasted audio text to carry audio description information, thereby enabling copied and pasted audio text to have audio playback functionality, enriching text display functions, and providing a good user experience. Attached Figure Description
[0092] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which:
[0093] Figure 1 This is a schematic diagram illustrating an application scenario of the audio-text processing method provided in the embodiments of this application;
[0094] Figure 2 This is a flowchart illustrating the audio-text processing method provided in an embodiment of this application;
[0095] Figure 3 This is another schematic diagram of the audio text processing system provided in the embodiments of this application;
[0096] Figure 4This is a schematic diagram illustrating the generation process of the second meeting text provided in an embodiment of this application;
[0097] Figure 5 This is a schematic diagram of the structure of the audio text processing device provided in the embodiments of this application;
[0098] Figure 6 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application;
[0099] Figure 7 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0100] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting this application.
[0101] This application provides an audio text processing method, apparatus, computer program, storage medium, and electronic device.
[0102] Please see Figure 1 , Figure 1 This is a schematic diagram illustrating an application scenario of the audio-text processing method provided in this application embodiment. This audio-text processing method can be applied to a hardware environment comprised of electronic devices such as terminal 101 and server 102. Figure 1 In this context, server 102 connects to terminal 101 via a network and can provide services to clients pre-installed on the terminal, web-based clients accessed by the terminal, or clients embedded in third-party applications. The client can be a client of an audio / video conferencing application, and the services include online conferencing services and audio recognition services, etc.
[0103] A database can be set up on server 102 or on other devices independent of server 102 to provide data storage services for server 102, such as storing conference audio data. Terminal 101 can provide a graphical user interface to present conference text to the user and obtain user operation instructions. The aforementioned network includes, but is not limited to, wide area networks, metropolitan area networks, or local area networks, and terminal 101 is not limited to personal computers (PCs), mobile phones, tablets, etc. The audio text processing method provided in this application embodiment can be executed by terminal 101, server 102, or jointly by server 102 and terminal 101.
[0104] The terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, etc., but is not limited to these. The terminal and the server can be connected directly or indirectly through wired or wireless communication, which is not limited herein. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.
[0105] based on Figure 1 The following describes the audio-text processing method provided in the embodiments of this application, illustrating the application scenarios shown. Please refer to [link / reference]. Figure 2 , Figure 2 This is a flowchart illustrating the audio text processing method provided in this application embodiment. This audio text processing method can be applied to a client running on an electronic device. The method specifically includes the following steps 201-203, wherein:
[0106] 201. In response to a first operation on the target audio recognition text, obtain target structured data corresponding to the target audio recognition text, the target structured data including the target audio recognition text and audio description information, the audio description information being used to obtain target audio data corresponding to the target audio recognition text.
[0107] The target audio recognition text can be the text content within a complete audio recognition text. This complete audio recognition text can be divided into multiple text units, such as sentences, paragraphs, or even words. This complete audio recognition text can be conference text. The text content within the conference text can be the original text obtained through semantic recognition of the conference audio data, or it can be new text obtained after editing operations such as deletion, addition, merging, or modification of the original text; there are no restrictions here.
[0108] Each complete audio recognition text contains audio description information. The text content and audio description information are stored in the form of structured data. Different complete audio recognition texts typically correspond to different structured data. This structured data is displayed in the graphical user interface (GUI). Besides the text content and audio description information, the structured data may also include formatting information such as font color, font type, and font size. The GUI can display one or more complete audio recognition texts simultaneously. Users can specify the desired text content as the target audio recognition text and can simultaneously specify one or more target audio recognition texts for copying and pasting. For example, a user can specify one sentence in three different locations from a meeting text, or specify three sentences in the same location as the target audio recognition text for copying and pasting.
[0109] When a user performs the first operation on the target audio recognition text, the client system retrieves the corresponding structured data from the data storage area that stores structured data, using it as the target structured data. Structured data refers to data organized according to rules that can be queried, retrieved, and statistically analyzed. For example, structured data can be stored and managed in tabular form, with different tables corresponding to different types of structured data.
[0110] In some embodiments, the first operation may include a copy-triggered operation, a cut-triggered operation, or a drag-and-drop operation. For example, the graphical user interface may automatically display function keys such as copy or cut, or the user may display these function keys on the graphical user interface by right-clicking the mouse. Then, when the user selects the copy function key (for example, clicking the copy function key is considered selection), it can be considered that a copy-triggered operation has been performed. When the user selects the cut function key, it can be considered that a cut-triggered operation has been performed. When the user directly drags the target audio recognition text, it can be considered that a drag-and-drop operation has been performed. For example, the text content can be dragged by holding down the left mouse button and moving the mouse.
[0111] It is readily understood that, prior to performing the first operation described above, the target structured data should be created and displayed to the user in advance; that is, in some embodiments, please refer to [link to relevant documentation]. Figure 3 , Figure 3 This is another schematic flowchart of the audio text processing method provided in this application embodiment. Before step 201 above, the audio text processing method may further include:
[0112] 204. Identify the audio recognition text corresponding to the target audio source data, and determine the audio description information corresponding to each text unit in the audio recognition text. The audio recognition text includes at least one text unit, the target audio source data includes target audio data, and the target audio recognition text is one of the text units.
[0113] 205. Create structured data corresponding to each text unit based on each text unit and its corresponding audio description information;
[0114] 206. The structured data created shows the audio-recognized text.
[0115] The target audio data refers to the portion of the target audio source data corresponding to the target audio recognition text. This target audio source data can be conference audio data, which can be local to the terminal, uploaded to the server by the user via the terminal, or generated through an online conferencing service provided by the server. When the conference audio data is generated through an online conferencing service, semantic recognition can be performed on the conference audio data during or after the meeting to obtain the corresponding conference text and the audio description information for each text unit within the conference text. This semantic recognition can be achieved based on speech recognition technology in Artificial Intelligence (AI). Artificial intelligence is the theory, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0116] The audio description information can include an audio source identifier and the start and end times of the audio. The audio source identifier is used to uniquely locate the target audio source data, and the start and end times are the audio times corresponding to the target audio recognition text within that target audio source data. After all text units and their audio description information have been identified, the text units and audio description information can be stored in the form of structured data. For example, the text units and their corresponding audio description information can be stored as rows in a data storage table corresponding to the entire meeting text. In this case, different rows in the table represent different types of structured data. This structured data can be stored locally on the terminal for local user use, uploaded to a server for use by other devices, or directly stored on the server for use by various devices on the client side.
[0117] 202. Store the target structured data in the shared storage area.
[0118] Shared storage is a shared storage area provided by an electronic device, such as a clipboard, that can be used by various applications within the device. One or more target structured data items can be stored simultaneously in shared storage.
[0119] In some instances, please see [link to relevant documentation]. Figure 3 Step 202 above may specifically include:
[0120] 2021. Generate the field content corresponding to the target field based on the target structured data;
[0121] 2022. The target field and its content are associated and stored in the shared storage area.
[0122] The target field is a manually defined field, such as "copydata". Since text data is typically stored in the shared storage area using text tags, when the target structured data needs to be stored in the shared storage area, a target field can be added to the text tag, and the target structured data can be stored together with the target field as its content in the shared storage area.
[0123] In some embodiments, considering that clients of other applications can also use the shared storage area in addition to the client of the audio / video conferencing application, to prevent data leakage in the shared storage area from posing a security risk to the server, such as malicious actors directly obtaining the original conference audio data from the server based on the audio description information in the target structured data, the target structured data can be encrypted before being stored in the shared storage area. That is, step 2021 above can specifically include:
[0124] The target structured data is encrypted to obtain encrypted data;
[0125] Use the encrypted data as the field content corresponding to the target field.
[0126] One approach is to use symmetric encryption algorithms, such as the Advanced Encryption Standard (AES), where the same key is used for both encryption and decryption.
[0127] In some embodiments, a shared storage area typically provides multiple storage areas to store the same text content. While different storage areas store the same text content, they also store different formatting information of that text content, thus enabling the text content to be presented in different formats based on the data stored in different storage areas. For example, if the shared storage area includes a first storage area and a second storage area, then step 2022 above may specifically include:
[0128] The target field and its content are stored together in the first storage area.
[0129] Meanwhile, step 202 above may also include:
[0130] The target audio recognition text is stored in the target structured data in the second storage area.
[0131] The first storage area can be the storage area corresponding to the "html" field, mainly used to store formatted text content. Typically, the text content and its formatting information are stored in the first storage area in the form of text tags. In this embodiment, a target field is added to the text tags, and the target structured data is stored in the first storage area as the content of this target field. The second storage area can be the storage area corresponding to the "text" field, mainly used to store unformatted text content, i.e., plain text. The second storage area can be considered as being used to present text content with the default format.
[0132] In other words, the data structure of the shared storage area can be represented as: {text:'pasted plain text content', html:'data containing text tags'}, where {text:'pasted plain text content'} corresponds to the second storage area, and {html:'data containing text tags'} corresponds to the first storage area. For example, the data in the first storage area can be... <span fontcolor=”#f00”copydata=”data”> Paste data Here, `fontColor` is a text label field representing the text color, and "#f00" is the specific color (red). `copydata` is the target field mentioned above, which is also a text label field. "data" is the field content of the target field, that is, the encrypted data mentioned above, or the unencrypted target structured data mentioned above.
[0133] 203. In response to a second operation performed at a target location, the target structured data is retrieved from the shared storage area, and the target audio recognition text is displayed at the target location based on the target structured data, so as to retrieve the target audio data based on the displayed target audio recognition text.
[0134] The second operation corresponds to the first operation and can include a paste trigger operation or a drag-and-release operation. For example, a user can display a paste function button on the graphical user interface by right-clicking the target location. When the user selects the paste function button, it can be considered that a paste trigger operation has been performed. Alternatively, a user can directly press and hold the left mouse button to drag text content to the target location. When the user releases the left mouse button, it can be considered that a drag-and-release operation has been performed.
[0135] In some embodiments, when the target structured data is stored unencrypted in a shared storage area, the above step "obtain the target structured data from the shared storage area" may specifically include: obtaining the field content corresponding to the target field from the shared storage area.
[0136] Furthermore, when the target structured data is encrypted and stored in a shared storage area, the above step "obtain the target structured data from the shared storage area" may also include: decrypting the obtained field content to obtain the target structured data.
[0137] The decryption process is the reverse of the encryption process. It's important to emphasize that when a client system uses shared storage to copy and paste text within the system, it is essentially copying and pasting structured data. The copied and pasted text carries its corresponding audio description information, enabling subsequent audio playback. However, when copying text from outside the client system using shared storage—for example, copying text from another application's client via the clipboard—it is essentially copying and pasting the text and its formatting information. This allows for common copy-paste functions, such as copying and pasting plain text or text with common formatting (including but not limited to font color, font type, and font size). However, the copied and pasted content does not carry audio description information, making subsequent audio playback impossible.
[0138] Meanwhile, when the target structure data is encrypted and stored in the shared storage area, even if other applications outside the system obtain all the data in the html area from the shared storage area, the audio description information is encrypted and stored in the content of the "copyData" field. Therefore, other applications still cannot obtain the actual audio description information and will not pose a security risk to the server's data. However, since the text area also stores unencrypted plain text content, other applications can perform normal text pasting functions, but cannot perform text pasting functions with audio description information.
[0139] In some embodiments, the above step of "displaying the target audio recognition text at the target location based on the target structured data" may specifically include:
[0140] Determine the text editing window where the target location is located;
[0141] Determine the first structured data storage area corresponding to the text editing window, and determine the shared storage area corresponding to the target location within the first structured data storage area;
[0142] The target structured data is stored in the shared storage area to update the first structured data storage area;
[0143] The content of the text editing window is updated based on the updated first structured data storage area to display the target audio recognition text at the target location.
[0144] The text editing window is a window that is triggered when a user performs an editing operation. Users can edit the text displayed in this window, such as adding, merging, deleting, or modifying content. The initial content displayed in this text editing window can be existing audio-recognition text, such as meeting transcripts (i.e., editing based on existing meeting transcripts), or it can be a completely new blank text file; there are no restrictions here.
[0145] When the client system detects that a user has triggered an editing operation, it can generate a corresponding text editing window and specify the first structured data storage area for that window, as well as the initial stored content within that first structured data storage area. For example, if editing existing meeting text, to preserve the original text, a new structured data storage area can be specified as the first structured data storage area, instead of directly using the original meeting text's structured data storage area. Simultaneously, the stored content from the original meeting text's structured data storage area can be copied to the first structured data storage area as its initial stored content, thus creating the first structured data storage area. Subsequently, when the user performs corresponding editing operations in the text editing window, the stored content in the first structured data storage area can be updated in real time. For example, if new text content needs to be added to the text editing window, a new storage area can be added in the first structured data storage area to store the corresponding content, and the content displayed in the text editing window will be updated accordingly as the stored content in the first structured data storage area is updated.
[0146] It should be noted that the target audio recognition text that has undergone the first operation described above can be existing text content in the text editing window. That is, the first and second operations can be performed within the text editing window to move the target audio recognition text from one location to another by cutting or dragging, or to display the target audio recognition text in multiple locations within the text editing window by copying. In other embodiments, the target audio recognition text that has undergone the first operation described above can be text content in another window. That is, the first operation can be performed in another window, and the second operation can be performed in this text editing window to display the text content from the other window in this text editing window. This other window could be, for example, a browsing window for other meeting texts.
[0147] In some embodiments, when the aforementioned text editing window is generated by editing existing meeting text, the audio text processing method may further include the following before the step "determining the text editing window where the target location is located":
[0148] In response to a text editing instruction for the first audio-recognized text, all structured data in the second structured data storage area corresponding to the first audio-recognized text is obtained;
[0149] The first structured data storage area is created based on all the acquired structured data;
[0150] The text editing window is generated based on the structured data in the first structured data storage area.
[0151] The first audio-recognized text can be an existing meeting transcript. The second structured data storage area stores the structured data corresponding to the first audio-recognized text. The text content in the meeting transcript can be the original text obtained by semantic recognition of the meeting audio data, or it can be new text obtained after editing operations such as deletion, addition, merging, or modification of the original text. The text editing window is mainly used for editing text and can provide text editing functions such as adding, modifying, and deleting text.
[0152] In some embodiments, the user can end the editing process and output the transcript at any time while editing the content in the text editing window; that is, the audio text processing method may further include:
[0153] In response to the end-of-edit operation applied to the text editing window, determine the structured data currently stored in the first structured data storage area;
[0154] Based on the structured data currently stored in the first structured data storage area, the second audio-recognized text is displayed through a browsing window.
[0155] The second audio-recognized text is a new text created by editing the first audio-recognized text. The first structured data storage area is used to store the structured data corresponding to the second audio-recognized text. Typically, the browsing window is mainly used for viewing text. Unlike the text editing window, the browsing window generally cannot be used for text editing and can only provide non-editing functions such as text viewing and copying.
[0156] For example, see Figure 4 , Figure 4 This is a schematic diagram illustrating the generation process of the second meeting text provided in this application embodiment. Figure 4This example uses two audio-recognized texts as conference text 1 and conference text 2, respectively. The user can click the "Edit" button in the conference text 1 browsing window to trigger the text editing command. The system automatically retrieves all structured data from the second structured data storage area and uses this data as the initial structured data for conference text 2 to create the first structured data storage area for conference text 2. Once the first structured data storage area is created, it can be displayed in a text editing window on the graphical user interface. The text editing window provides a "Finish" button. Clicking this button completes the editing process, and the text editing window can be closed. The final generated conference text 2 will then be displayed in a browsing window on the graphical user interface.
[0157] In some embodiments, the target audio-recognized text displayed on the graphical user interface, such as the target audio-recognized text after copying and pasting, or the target audio-recognized text before copying and pasting, can be played back. That is, the audio-text processing method may further include:
[0158] In response to an audio playback command for the target audio-recognized text, the audio description information is extracted from the target structured data;
[0159] Obtain the target audio data corresponding to the audio description information;
[0160] Play the acquired target audio data.
[0161] The graphical user interface can provide an audio playback button, such as a small speaker-shaped button. When the user selects the target audio recognition text and clicks the audio playback button, the audio playback command can be triggered. Alternatively, the user can trigger the audio playback command by double-clicking the target audio recognition text or other agreed-upon methods.
[0162] In some embodiments, the audio text processing method may further include:
[0163] In response to the determination of the target audio-recognized text, multiple function controls are displayed, including a target function control, which is used to instruct audio playback of the determined text;
[0164] In response to the confirmation operation of the target function control, an audio playback instruction is generated for the target audio recognition text.
[0165] The "select" operation allows users to select desired text by holding down the left mouse button and dragging the mouse. The selected text is highlighted with a dark shading. Releasing the left mouse button reveals multiple function buttons on the graphical user interface, allowing for further actions on the selected text. These actions include playing back the audio of the selected text, the text preceding the selected text, or the text following the selected text. Clicking any function button is considered selection, and the system will execute the corresponding function.
[0166] In some embodiments, when the audio description information includes the aforementioned audio source identifier and audio start and end times, the aforementioned step "obtaining the target audio data corresponding to the audio description information" may specifically include:
[0167] Find the target audio source data corresponding to the audio source identifier from the stored audio source dataset;
[0168] Extract the content corresponding to the start and end times of the audio from the target audio source data, and use it as the target audio data.
[0169] Each audio source in the audio source dataset is assigned an audio source identifier during storage for later retrieval. The audio source dataset can reside locally on the terminal, allowing local users to listen to audio from the text, or it can reside on a server, such as conference audio data uploaded from the terminal or conference audio data generated by the server itself. When speech recognition technology is used to recognize the audio source data, both the recognized text and the start and end times of the audio can be obtained simultaneously.
[0170] As described above, the audio-text processing method provided in this application, in response to a first operation on the target audio recognition text, obtains the target structured data corresponding to the target audio recognition text. The target structured data includes the target audio recognition text and audio description information, and the audio description information is used to obtain the target audio data corresponding to the target audio recognition text. The target structured data is stored in a shared storage area. In response to a second operation applied to the target location, the target structured data is obtained from the shared storage area, and the target audio recognition text is displayed at the target location based on the target structured data. That is, by storing the audio text and its audio description information in the form of structured data, and copying and pasting the audio text using structured data, the copied and pasted audio text can carry audio description information, thereby enabling the copied and pasted audio text to have audio playback functionality, enriching the text display function, and providing a good user experience.
[0171] Based on the methods described in the above embodiments, this application also provides an audio text processing apparatus for performing the steps in the above audio text processing method. Please refer to... Figure 5 , Figure 5 This is a schematic diagram of the structure of the audio text processing device provided in an embodiment of this application. The audio text processing device 300 is applied to an electronic device and may include an acquisition module 301, a storage module 302, and a display module 303, wherein:
[0172] The acquisition module 301 is used to acquire target structured data corresponding to the target audio recognition text in response to the first operation on the target audio recognition text. The target structured data includes the target audio recognition text and audio description information. The audio description information is used to acquire the target audio data corresponding to the target audio recognition text.
[0173] Storage module 302 is used to store the target structured data in a shared storage area;
[0174] Display module 303 is configured to respond to a second operation performed at the target location, retrieve the target structured data from the shared storage area, and display the target audio recognition text at the target location based on the target structured data.
[0175] In some embodiments, the storage module 302 is specifically used for:
[0176] Generate the field content corresponding to the target field based on the target structured data;
[0177] The target field and its content are stored together in this shared storage area.
[0178] The display module 303 is specifically used to: retrieve the field content corresponding to the target field from the shared storage area.
[0179] In some embodiments, the storage module 302 is specifically used for:
[0180] The target structured data is encrypted to obtain encrypted data;
[0181] Use the encrypted data as the field content corresponding to the target field;
[0182] The display module 303 is also used to: decrypt the acquired field content to obtain the target structured data.
[0183] In some embodiments, the shared storage area includes a first storage area and a second storage area, and the storage module 302 is specifically used for:
[0184] The target field and its content are stored together in the first storage area.
[0185] The storage module 302 is also used to: store the target audio recognition text in the target structured data in the second storage area.
[0186] In some embodiments, the display module 303 is specifically used for:
[0187] Determine the text editing window where the target location is located;
[0188] Determine the first structured data storage area corresponding to the text editing window, and determine the shared storage area corresponding to the target location within the first structured data storage area;
[0189] The target structured data is stored in the shared storage area to update the first structured data storage area;
[0190] The content of the text editing window is updated based on the updated first structured data storage area to display the target audio recognition text at the target location.
[0191] In some embodiments, before determining the text editing window where the target location is located, the display module 303 is further configured to:
[0192] In response to a text editing instruction for the first audio-recognized text, all structured data in the second structured data storage area corresponding to the first audio-recognized text is obtained;
[0193] The first structured data storage area is created based on all the acquired structured data;
[0194] The text editing window is generated based on the structured data in the first structured data storage area.
[0195] In some embodiments, the display module 303 is further configured to:
[0196] In response to the end-of-edit operation applied to the text editing window, determine the structured data currently stored in the first structured data storage area;
[0197] Based on the structured data currently stored in the first structured data storage area, the second audio-recognized text is displayed through a browsing window.
[0198] In some embodiments, prior to the first operation of recognizing text from the target audio, the acquisition module 301 is further configured to:
[0199] Identify the audio recognition text corresponding to the target audio source data, and determine the audio description information corresponding to each text unit in the audio recognition text. The audio recognition text includes at least one text unit, the target audio source data includes target audio data, and the target audio recognition text is one of the text units.
[0200] Create structured data corresponding to each text unit based on each text unit and its corresponding audio description information;
[0201] The structured data generated is used to identify text from audio.
[0202] In some embodiments, the audio text processing device further includes a playback module for:
[0203] In response to an audio playback command for the target audio-recognized text, the audio description information is extracted from the target structured data;
[0204] Obtain the target audio data corresponding to the audio description information;
[0205] Play the acquired target audio data.
[0206] In some embodiments, the audio description information includes an audio source identifier and audio start and end times, and the playback module is specifically used for:
[0207] Find the target audio source data corresponding to the audio source identifier from the stored audio source dataset;
[0208] Extract the content corresponding to the start and end times of the audio from the target audio source data, and use it as the target audio data.
[0209] In some embodiments, the playback module is further configured to:
[0210] In response to the determination of the target audio-recognized text, multiple function controls are displayed, including a target function control, which is used to instruct audio playback of the determined text;
[0211] In response to the confirmation operation of the target function control, an audio playback instruction is generated for the target audio recognition text.
[0212] It should be noted that the specific details of each module unit in the audio text processing device 300 have been described in detail in the embodiments of the audio text processing method, and will not be repeated here.
[0213] In some embodiments, the audio-text processing device in this application can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or a server. Exemplary examples include mobile phones, tablets, laptops, handheld computers, in-vehicle electronic devices, mobile internet devices (MIDs), augmented reality (AR) / virtual reality (VR) devices, robots, wearable devices, ultra-mobile personal computers (UMPCs), netbooks, or personal digital assistants (PDAs), etc. It can also be a server, network attached storage (NAS), a personal computer (PC), a television set (TV), an ATM, or a self-service machine, etc. This application does not impose specific limitations.
[0214] In some embodiments, such as Figure 6 As shown, this application embodiment also provides an electronic device 400, including a processor 401, a memory 402, and a computer program stored in the memory 402 and executable on the processor 401. When the program is executed by the processor 401, it implements the various processes of the above-described audio-text processing method embodiment and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0215] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.
[0216] Figure 7 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.
[0217] The electronic device 500 includes, but is not limited to, components such as: radio frequency unit 501, network module 502, audio output unit 503, input unit 504, sensor 505, display unit 506, user input unit 507, interface unit 508, memory 509, and processor 510.
[0218] Those skilled in the art will understand that the electronic device 500 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 510 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 7The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0219] It should be understood that, in this embodiment, the input unit 504 may include a graphics processing unit (GPU) 5041 and a microphone 5042. The GPU 5041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 506 may include a display panel 5061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 507 includes at least one of a touch panel 5071 and other input devices 5072. The touch panel 5071 is also called a touch screen. The touch panel 5071 may include a touch detection device and a touch controller. Other input devices 5072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.
[0220] The memory 509 can be used to store software programs and various data. The memory 509 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 509 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 509 in this embodiment includes, but is not limited to, these and any other suitable types of memory.
[0221] Processor 510 may include one or more processing units; processor 510 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 510.
[0222] This application also provides a non-transitory computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the various processes of the above-described audio-text processing method embodiments and achieves the same technical effect. To avoid repetition, it will not be described again here.
[0223] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0224] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described audio-text processing method.
[0225] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0226] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0227] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0228] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
[0229] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0230] In the description of this application, "multiple" means two or more.
[0231] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0232] It should also be understood that the terms “comprising” and / or “including” as used herein specify the presence of the stated features, integers, steps, operations, units and / or components, without excluding the presence or addition of one or more other features, integers, steps, operations, units, components and / or combinations thereof.
[0233] Although embodiments of this application have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of this application, the scope of which is defined by the claims and their equivalents.
Claims
1. An audio text processing method, characterized in that, include: In response to a first operation on the target audio recognition text, target structured data corresponding to the target audio recognition text is obtained. The target structured data includes the target audio recognition text and audio description information. The audio description information is used to obtain target audio data corresponding to the target audio recognition text. The target structured data is stored in a shared storage area; In response to a second operation performed at the target location, the target structured data is retrieved from the shared storage area, and the target audio recognition text is displayed at the target location based on the target structured data, so as to retrieve the target audio data based on the displayed target audio recognition text; Storing the target structured data in the shared storage area includes: Generate the field content corresponding to the target field based on the target structured data; The target field and its content are stored together in the shared storage area. The step of obtaining the target structured data from the shared storage area includes: obtaining the field content corresponding to the target field from the shared storage area; The step of generating the field content corresponding to the target field based on the target structured data includes: The target structured data is encrypted to obtain encrypted data; Use the encrypted data as the field content corresponding to the target field; The step of obtaining the target structured data from the shared storage area further includes: decrypting the obtained field content to obtain the target structured data; The shared storage area includes a first storage area and a second storage area. The step of storing the target field and its content in association within the shared storage area includes: The target field and its content are stored together in the first storage area. Storing the target structured data in the shared storage area further includes: The target audio recognition text from the target structured data is stored in the second storage area.
2. The audio text processing method according to claim 1, characterized in that, The step of displaying the target audio-recognized text at the target location based on the target structured data includes: Determine the text editing window where the target location is located; Determine the first structured data storage area corresponding to the text editing window, and determine the shared storage area corresponding to the target location in the first structured data storage area; The target structured data is stored in the shared storage area to update the first structured data storage area; The content in the text editing window is updated based on the updated first structured data storage area to display the target audio recognition text at the target location.
3. The audio text processing method according to claim 2, characterized in that, Before determining the text editing window where the target location is located, the following steps are also included: In response to a text editing instruction for the first audio-recognized text, all structured data in the second structured data storage area corresponding to the first audio-recognized text is obtained; The first structured data storage area is created based on all the acquired structured data; The text editing window is generated based on the structured data in the first structured data storage area.
4. The audio text processing method according to claim 2, characterized in that, Also includes: In response to an editing end operation applied to the text editing window, determine the structured data currently stored in the first structured data storage area; Based on the structured data currently stored in the first structured data storage area, the second audio-recognized text is displayed through a browsing window.
5. The audio text processing method according to any one of claims 1-4, characterized in that, Prior to the first operation in response to the target audio-to-text recognition, it also includes: Identify the audio recognition text corresponding to the target audio source data, and determine the audio description information corresponding to each text unit in the audio recognition text. The audio recognition text includes at least one text unit, the target audio source data includes the target audio data, and the target audio recognition text is one of the text units. Structured data corresponding to each text unit is created based on each text unit and its corresponding audio description information; The audio-recognized text is displayed based on the structured data created.
6. The audio text processing method according to any one of claims 1-4, characterized in that, Also includes: In response to an audio playback command for the target audio-recognized text, the audio description information is extracted from the target structured data; Obtain the target audio data corresponding to the audio description information; Play the acquired target audio data.
7. An audio text processing apparatus, suitable for employing the audio text processing method according to any one of claims 1-6, characterized in that, include: The acquisition module is configured to, in response to a first operation on the target audio recognition text, acquire the target structured data corresponding to the target audio recognition text, wherein the target structured data includes the target audio recognition text and audio description information, and the audio description information is used to acquire the target audio data corresponding to the target audio recognition text; A storage module for storing the target structured data in a shared storage area; The display module is configured to respond to a second operation performed at the target location, retrieve the target structured data from the shared storage area, and display the target audio recognition text at the target location based on the target structured data, so as to retrieve the target audio data based on the displayed target audio recognition text.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the audio text processing method as described in any one of claims 1-6.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the audio text processing method as described in any one of claims 1-6.
Citation Information
Patent Citations
Data processing method, mobile terminal and storage medium
CN113555002A
Information leakage prevention method and apparatus and program for the same
US20060117178A1