Content arrangement method and device, electronic equipment and storage medium

By converting learning content into podcast audio files and subtitles, the time and convenience problems of learners when memorizing and reviewing knowledge content are solved, and the knowledge memory efficiency and user experience are improved.

CN120029509APending Publication Date: 2025-05-23WANGYIYOUDAO INFORMATION TECH BEIJING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411944139.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

When learning to memorize professional subjects, users need to spend a lot of time reciting, and reviewing the knowledge content is inconvenient, and the existing technology cannot effectively solve this problem.

Method used

Receive original content through the content input interface, switch to the podcast type selection interface, generate podcast audio files based on the podcast type selected by the user, and display podcast subtitles, providing sound settings and multilingual subtitles functions.

Benefits of technology

Conveniently organize the memorized content into knowledge notes, improve users' knowledge memory efficiency, solve the problem of inconvenience of users carrying textbooks, and improve users' user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029509A_ABST
    Figure CN120029509A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a content arrangement method and device, electronic equipment and a storage medium. The content arrangement method comprises the steps of receiving input of original content through a content input interface; in response to the obtained original content, switching the content input interface into a player type selection interface; and in response to selection of a target broadcaster type in the broadcaster type selection interface, generating a broadcaster audio file based on the original content and the target broadcaster type, and switching to display a broadcaster subtitle display interface. According to the technical scheme, the recitation content can be conveniently arranged into the knowledge notes, the user can conveniently review the knowledge content, the knowledge memory efficiency of the user is improved, and the use experience of the user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of artificial intelligence. More specifically, the embodiments of the present application relate to a content organization method, device, electronic device and storage medium. Background Art

[0002] This section is intended to provide background or context for the embodiments of the present application set forth in the claims. The descriptions herein may include concepts that may be explored, but not necessarily concepts that have been previously thought of or explored. Therefore, unless otherwise noted herein, the content described in this section is not prior art with respect to the specification and claims of the present application, and is not admitted to be prior art by inclusion in this section.

[0003] In the scenario where users are learning and reciting, for some professional subjects with high recitation requirements, such as history, psychology, law, etc., users need to spend a lot of time on recitation, and users need to frequently review and browse related knowledge content in their spare time to achieve the purpose of consolidating knowledge content. However, it is not convenient to carry textbooks and reference materials, and making your own learning materials will take up a lot of time and energy. Purchasing learning materials outside also has problems such as not being able to meet one's own personalized needs and requiring a lot of money.

[0004] Some individuals or organizations currently record video or audio recitation materials on certain professional subjects to provide training for learners and help them better understand the knowledge. However, the recording of audio and video takes a lot of time and money. It also cannot solve the pain point of inconvenience and low efficiency of users in reviewing knowledge content.

[0005] In view of this, there is an urgent need to propose a content organization method so that, for example, recitation content can be easily organized into knowledge notes, which is convenient for users to review knowledge content, improve users' knowledge memory efficiency, and enhance users' usage experience. Summary of the invention

[0006] In order to overcome the problems existing in the related art, the embodiments of the present application hope to provide a content organization method, device, electronic device and storage medium. The content organization method can conveniently organize the recitation content into knowledge notes, facilitate users to review the knowledge content, improve the user's knowledge memory efficiency, and improve the user's user experience.

[0007] In a first aspect of the implementation scheme of the present application, a content organization method is provided, including: receiving input of original content through a content input interface; in response to obtaining the original content, switching the content input interface to a podcast type selection interface; in response to selection of a target podcast type in the podcast type selection interface, generating a podcast audio file based on the original content and the target podcast type, and switching to display a podcast subtitle display interface.

[0008] In one embodiment of the present application, the method also includes: in response to selection of a target podcast type in a podcast type selection interface, switching the podcast type selection interface to a sound setting interface; receiving input of sound setting information through the sound setting interface; wherein the sound setting information includes at least one of voice gender information, pronunciation language information, voice emotion information, voice speaking speed information and voice timbre information; and, also generating a podcast audio file based on the sound setting information.

[0009] In one embodiment of the present application, the podcast subtitle display interface includes a bilingual subtitle control, and the method further includes: in response to a first trigger operation on the bilingual subtitle control, switching the podcast subtitle display interface to a language selection interface; wherein the language selection interface includes a bilingual subtitle function start / stop button and a language list; in response to a second trigger operation on the bilingual subtitle function start / stop button, enabling the bilingual subtitle function; in response to a third trigger operation, determining a target language from the language list, switching the language selection interface back to the podcast subtitle display interface, and displaying the podcast subtitles corresponding to the original language and the podcast subtitles corresponding to the target language in the podcast subtitle display interface.

[0010] In one embodiment of the present application, after switching to display the podcast subtitle display interface, the method also includes: the podcast subtitle display interface includes a label option bar; in response to selection of a content summary label in the label option bar, switching the podcast subtitle display interface to a content summary interface; in response to clicking a summary audio generation control in the content summary interface, generating a summary audio file based on the original content and switching the content summary interface to a summary playback interface.

[0011] In one embodiment of the present application, after switching to display the podcast subtitle display interface, the method also includes: in response to selection of the original text display tag in the tag option bar, switching the podcast subtitle display interface to the original text display interface; in response to clicking on the original text playback control in the original text display interface, generating an original audio file based on the original content and switching the original text display interface to the original text playback interface.

[0012] In one embodiment of the present application, an audio playback control area is provided in the podcast subtitle display interface, the summary playback interface and the original text playback interface; wherein the audio playback control area includes at least one of a progress bar control, a speed control control, a playback start and stop control, a jump playback control and a file export control.

[0013] In one embodiment of the present application, after switching to display the podcast subtitle display interface, the method further includes: in response to selection of a mind map tag in the tag option bar, switching the podcast subtitle display interface to a mind map display interface; wherein the mind map display interface includes a mind map generated based on the original content.

[0014] In one embodiment of the present application, after switching to display the podcast subtitle display interface, the method also includes: in response to selection of a presentation tag in the tag option bar, switching the podcast subtitle display interface to a presentation interface; wherein the presentation interface includes a presentation generated based on the original content.

[0015] In a second aspect of the implementation manner of the present application, a content organization device is provided for executing a content organization method as described in any one of the first aspects, including: a content receiving module for receiving input of original content through a content input interface; an interface switching module for switching the content input interface to a podcast type selection interface in response to obtaining the original content; a file generation module for generating a podcast audio file based on the original content and the target podcast type in response to selection of a target podcast type in the podcast type selection interface, and switching the display of the podcast subtitle display interface through the interface switching module.

[0016] A third aspect of the present application provides an electronic device, comprising: a processor; and a memory, on which is stored an executable code for content organization, and when the executable code is executed by the processor, the processor executes the method described above.

[0017] A fourth aspect of the present application provides a non-temporary machine-readable storage medium having stored thereon an executable code for content organization, and when the executable code is executed by a processor of an electronic device, the processor is caused to execute the method as described above.

[0018] The technical solution provided by the implementation method of this application has the following beneficial effects:

[0019] The content organization method, device, electronic device and storage medium provided in the embodiments of the present application receive original content input by a user through a content input interface, and then in response to obtaining the original content, switch the content input interface to a podcast type selection interface, thereby providing a variety of podcast styles to users in different scenarios.

[0020] Furthermore, the present application can generate a podcast audio file based on the original content and the target podcast type in response to the target podcast type determined through the podcast type selection interface, and switch to display the podcast subtitle display interface, so that the original content that needs to be recited can be conveniently organized into a podcast audio file according to the user's style preferences, making it convenient for users to review knowledge content. Combined with the display of podcast subtitles, it can also solve the problem that it is inconvenient for users to carry textbooks and reference materials, thereby improving user convenience.

[0021] In general, this application can conveniently organize, for example, recitation content into knowledge notes, facilitate users to review knowledge content, improve users' knowledge memory efficiency, and enhance users' usage experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] By reading the detailed description below with reference to the accompanying drawings, the above and other purposes, features and advantages of the exemplary embodiments of the present application will become readily understood. In the accompanying drawings, several embodiments of the present application are shown in an exemplary and non-limiting manner, wherein:

[0023] Figure 1 Schematically shows a block diagram of an exemplary computing system 100 suitable for implementing the embodiments of the present application;

[0024] Figure 2 A flowchart 200 of a content arrangement method according to an embodiment of the present application is schematically shown;

[0025] Figure 3 A flowchart 300 of a content arrangement method according to another embodiment of the present application is schematically shown;

[0026] Figure 4 Schematically shows a flow chart 400 of a content arrangement method according to another embodiment of the present application;

[0027] Figure 5 A schematic diagram schematically shows a content input interface when the original content is input text in the content arrangement method of an embodiment of the present application;

[0028] Figure 6 The following schematically shows a schematic diagram of a content input interface when the original content is a file in the content arrangement method of an embodiment of the present application;

[0029] Figure 7 A schematic diagram schematically shows a podcast type selection interface in the content organization method of an embodiment of the present application;

[0030] Figure 8 A first schematic diagram of a podcast subtitle display interface in a content organization method according to an embodiment of the present application is schematically shown;

[0031] Fig. 9 A schematic diagram schematically shows a sound setting interface in the content arrangement method of an embodiment of the present application;

[0032] Fig.10 A schematic diagram schematically shows a language selection interface in the content organization method of an embodiment of the present application;

[0033] Fig.11 A second schematic diagram schematically shows a case where the podcast subtitle display interface displays podcast subtitles corresponding to the original language and podcast subtitles corresponding to the target language in the content organization method of an embodiment of the present application;

[0034] Fig.12 A schematic diagram schematically shows a content summary interface in a content organization method according to an embodiment of the present application;

[0035] Fig.13 A schematic diagram schematically shows a summary playback interface in a content arrangement method according to an embodiment of the present application;

[0036] Fig.14 A schematic diagram schematically shows an original text display interface in a content arrangement method according to an embodiment of the present application;

[0037] Fig.15 A schematic diagram schematically shows an original text playback interface in a content arrangement method according to an embodiment of the present application;

[0038] Fig.16 A schematic diagram schematically shows a mind map display interface in the content organization method of an embodiment of the present application;

[0039] Fig.17 A schematic diagram schematically shows a presentation interface in a content arrangement method according to an embodiment of the present application;

[0040] Fig.18 The structure diagram of a content arrangement device according to an embodiment of the present application is schematically shown;

[0041] Fig.19 A schematic block diagram of an electronic device according to an embodiment of the present application is schematically shown.

[0042] In the drawings, the same or corresponding reference numerals represent the same or corresponding parts. DETAILED DESCRIPTION

[0043] The principles and spirit of the present application will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and implement the present application, and are not intended to limit the scope of the present application in any way. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.

[0044] Figure 1 1 is a block diagram of an exemplary computing system 100 suitable for implementing the embodiments of the present application. Figure 1 As shown, the computing system 100 may include: a central processing unit (CPU) 101, a random access memory (RAM) 102, a read-only memory (ROM) 103, a system bus 104, a hard disk controller 105, a keyboard controller 106, a serial interface controller 107, a parallel interface controller 108, a display controller 109, a hard disk 110, a keyboard 111, a serial external device 112, a parallel external device 113, and a display 114. Among these devices, the CPU 101, the RAM 102, the ROM 103, the hard disk controller 105, the keyboard controller 106, the serial controller 107, the parallel controller 108, and the display controller 109 are coupled to the system bus 104. The hard disk 110 is coupled to the hard disk controller 105, the keyboard 111 is coupled to the keyboard controller 106, the serial external device 112 is coupled to the serial interface controller 107, the parallel external device 113 is coupled to the parallel interface controller 108, and the display 114 is coupled to the display controller 109. It should be understood that Figure 1 The structural block diagram is only for the purpose of illustration, and is not intended to limit the scope of the present application. In some cases, some devices may be added or reduced according to specific circumstances.

[0045] Those skilled in the art know that the embodiments of the present application can be implemented as a system, method or computer program product. Therefore, the present disclosure can be specifically implemented in the following forms, namely: complete hardware, complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software, generally referred to herein as "circuit", "module" or "system". In addition, in some embodiments, the present application can also be implemented in the form of a computer program product in one or more computer-readable media, which contains computer-readable program code.

[0046] Any combination of one or more computer-readable media may be used. A computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples (non-exhaustive examples) of computer-readable storage media may include, for example: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device or device.

[0047] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, which carry computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0048] The program code embodied on the computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0049] Computer program code for performing the operation of the present application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., using an Internet service provider to connect through the Internet).

[0050] The following will describe the implementation of the present application with reference to the flowchart of the method of the present application embodiment and the block diagram of the device (or system). It should be understood that each square frame of the flowchart and / or block diagram and the combination of each square frame in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer or other programmable data processing device, so as to produce a kind of machine, and these computer program instructions are executed by a computer or other programmable data processing device, and a device for realizing the function / operation specified in the square frame in the flowchart and / or block diagram is generated.

[0051] These computer program instructions may also be stored in a computer-readable medium that enables a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable medium produce a product that includes an instruction device that implements the functions / operations specified in the blocks in the flowchart and / or block diagram.

[0052] Computer program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby enabling the instructions executed on the computer or other programmable device to provide a process for implementing the functions / operations specified in the blocks in the flowchart and / or block diagram.

[0053] According to an embodiment of the present application, a content organization method and device are proposed.

[0054] It should be understood herein that any number of elements in the drawings is for illustration rather than limitation, and any naming is only for distinction rather than having any limiting meaning.

[0055] The principles and spirit of the present application are explained in detail below with reference to several representative implementations of the present application. SUMMARY OF THE INVENTION

[0057] The applicant has found that, illustratively, for example, in the scenario where users are learning to recite, for some professional subjects with relatively high requirements for recitation, such as history, psychology, law, etc., users need to spend a lot of time on recitation, and users need to review and browse related knowledge content frequently in their spare time to achieve the purpose of consolidating knowledge content. However, it is not convenient to carry textbooks and reference materials, and self-made learning materials will consume a lot of time and energy. Purchasing learning materials outside also has the problems of not being able to meet their own personalized needs and needing to spend a lot of money. Some existing individuals or structural organizations will record video or audio recitation materials about certain professional subjects, and conduct training for learners to help them better understand knowledge, but the recording of audio and video requires a lot of time and money. It is also impossible to solve the pain point of inconvenience and low efficiency of users in reviewing knowledge content.

[0058] Based on this, the present application receives the original content input by the user through the content input interface, and then switches the content input interface to the podcast type selection interface in response to obtaining the original content, so as to provide a variety of podcast styles to users in different scenarios. Furthermore, the present application can generate a podcast audio file based on the original content and the target podcast type in response to the target podcast type determined through the podcast type selection interface, and switch to display the podcast subtitle display interface, so that the original content that needs to be recited can be conveniently organized into a podcast audio file according to the user's style preferences, which is convenient for users to review knowledge content. In combination with the display of podcast subtitles, it can also solve the problem that it is inconvenient for users to carry textbooks and reference materials, thereby improving user convenience. In general, the present application can conveniently organize the recited content into knowledge notes, which is convenient for users to review knowledge content, improve the user's knowledge memory efficiency, and improve the user's user experience.

[0059] After introducing the basic principles of the present application, various non-limiting implementation methods of the present application are described in detail below.

[0060] Application Scenario Overview

[0061] The content organization method of the present application is applicable to various types of electronic devices, such as mobile devices such as mobile phones and tablets, and learning devices such as dictionary pens. For example, the content organization method of the present application is applicable to various types of electronic devices, such as mobile devices such as mobile phones and tablets, and learning devices such as dictionary pens. Figure 1 The electronic device of the computing system 100 shown in FIG. Using the method proposed in the present application on these devices, the recitation content can be conveniently organized into knowledge notes, which is convenient for users to review the knowledge content, improve the user's knowledge memory efficiency, and enhance the user's use experience.

[0062] Exemplary Methods

[0063] Reference below Figure 2 , Figures 5 to 8 To describe the content arrangement method according to the exemplary embodiment of the present application. It should be noted that the above application scenarios are only shown to facilitate understanding of the spirit and principle of the present application, and the embodiments of the present application are not limited in this respect. On the contrary, the embodiments of the present application can be applied to any applicable scenario.

[0064] In step S201, the input of original content is received through the content input interface. In the embodiment of the present application, the aforementioned original content may include but is not limited to text, files, pictures, URL links, voice, etc. Accordingly, the user may input original content in the content input interface by inputting text, uploading text files, uploading or taking pictures, inputting URL links, uploading voice audio files, etc., thereby covering a variety of application scenarios, such as uploading classroom recording files, uploading learning courseware, uploading / taking / scanning pictures of materials, inputting URL links to import materials obtained from other note apps, etc. Figure 5 and Figure 6 As shown, in Figure 5 Users can directly enter text in the content input interface. Figure 6 The user can select and upload the corresponding file in the content input interface. It can be understood that text input and file upload in the content input interface can be called by the user by setting a text input control and a file upload control thereon.

[0065] It can also be understood that the original content may be in various forms, and the way in which users input the original content may also be various. In actual applications, the form of the original content and the way in which users input the original content need to be determined based on the actual application situation. This application does not impose any restrictions in this regard.

[0066] In step S202, in response to obtaining the original content, the content input interface is switched to a podcast type selection interface. Podcast is a form of digital media that allows users to download or stream audio files, usually a series of programs, over the Internet, and users can listen to the podcast content on a personal computer or mobile device. The content of the podcast can include various topics, such as news, interviews, storytelling, educational lectures, comedy sketches, etc. In the embodiment of the present application, Figure 7 As shown, the podcast type selection interface may have multiple podcast types for users to choose from, and may provide users in different scenarios with a variety of podcast styles. The narration format of each podcast style may be different, for example, a popular science podcast may be a narrator telling the story alone, and a speech podcast may be a debate or argument between guests on a certain topic. In some application scenarios, the podcast type selection interface may also support users to audition each podcast type to enhance the user experience.

[0067] In step S203, in response to the selection of the target podcast type in the podcast type selection interface, a podcast audio file is generated based on the original content and the target podcast type, and the podcast subtitle display interface is switched. In an embodiment of the present application, the original content can be matched in the form of the original content, and at least one of OCR image recognition, file parsing, and ASR speech recognition is used to convert the original content into text content, and then call an artificial intelligence large model, such as LLM (Large Language Model, large language model) or BLOOM (BigScience Large Open-science Open-accessMultilingual Language Model, large-scale multilingual pre-training model), and the converted text content is intelligently sorted, for example, it can be interpreted, commented or discussed based on the original content, and the original content can also be enhanced through elements such as sound, intonation, music and guest interviews. The complex concepts of the original text can also be simplified to make it easier for the public to understand, or the theme of the original text can be explored in depth. It can be understood that the collected podcast content materials can be used as training data to pre-train the above-mentioned artificial intelligence large model, so that the target text obtained by the intelligent sorting of the artificial intelligence large model is more suitable for conversion into a podcast audio file. Then, based on the target podcast type and the target text obtained by intelligent arrangement, a podcast audio file is generated, and the podcast audio file is switched to the following format: Figure 8 The podcast subtitle display interface shown meets personalized demands.

[0068] The embodiment of the present application receives the original content input by the user through the content input interface, and then switches the content input interface to the podcast type selection interface in response to obtaining the original content, so as to provide a variety of podcast styles to users in different scenarios. Furthermore, the present application can generate a podcast audio file based on the original content and the target podcast type in response to the target podcast type determined through the podcast type selection interface, and switch to display the podcast subtitle display interface, so that the original content that needs to be recited, for example, can be conveniently organized into a podcast audio file according to the user's style preferences, which is convenient for users to review knowledge content. In combination with the display of podcast subtitles, it can also solve the problem that it is inconvenient for users to carry textbooks and reference materials, thereby improving user convenience. In general, the present application can conveniently organize content into knowledge notes, which is convenient for users to review knowledge content, improve the user's knowledge memory efficiency, and improve the user's user experience.

[0069] In some embodiments, the user can set the sound of the podcast audio file, and the user can also select different tags in the tag option bar to view different presentation content, thereby assisting the user in understanding and memorizing knowledge in a variety of ways. Figure 3 , Fig. 9 and Figures 12 to 17 The above interaction process is described in detail. Figure 3 , Fig. 9 and Figures 12 to 17 The content arrangement method shown in the embodiment of the present application may include:

[0070] In step S301, in response to the selection of the target podcast type in the podcast type selection interface, the podcast type selection interface is switched to the sound setting interface, and the input of sound setting information is received through the sound setting interface. In the embodiment of the present application, the sound setting interface can be as follows: Fig. 9 As shown, for example, the user can select the sound setting information by clicking or voice control in the sound setting interface. The sound setting information may include but is not limited to at least one of voice gender information, pronunciation language information, voice emotion information, voice speed information, and voice timbre information. In some application scenarios, the user can audition the sound setting information after setting it so as to select a comfortable sound.

[0071] In step S302, a podcast audio file is generated based on the original content, the target podcast type and the sound setting information, and the podcast subtitle display interface is switched to display. In an embodiment of the present application, the original content can be converted into text content by using at least one of OCR image recognition, file parsing, and ASR speech recognition, matching the form of the original content, and then calling the artificial intelligence large model to intelligently organize the converted text content, and then perform speech synthesis according to the podcast style corresponding to the target podcast type and the sound style of the sound setting information. For example, TTS (Text To Speech) technology can be used for speech synthesis to synthesize and generate a podcast audio file. During the synthesis process, a sound wave pattern and a prompt indicating speech synthesis can be displayed in the podcast subtitle display interface. After the synthesis is completed, it can be displayed as Figure 8 As shown, the podcast subtitles corresponding to the podcast audio file are displayed in the podcast subtitle display interface.

[0072] In step S303, in response to the selection of a tag in the tag option bar, the podcast subtitle display interface is switched to the interface content corresponding to the tag. In the embodiment of the present application, a tag option bar is set in the application interface, and after the user selects a tag in the tag option bar, the interface will switch to the interface content corresponding to the tag, which is conducive to assisting the user in understanding and memorizing knowledge in a variety of ways.

[0073] Among them, Fig.12 and Fig.13As shown, when the user needs to refine and summarize the original content, the podcast subtitle display interface can be switched to the content summary interface in response to the selection of the content summary tag in the tag option bar, and then the summary audio generation control (such as Fig.12 A "synthesized audio" button in the image is clicked to generate a summary audio file based on the original content and switch the content summary interface to the following: Fig.13 The summary playback interface shown. Exemplarily, the artificial intelligence big model can be used to refine and summarize the text content converted from the original content, for example, to extract effective information from the text content converted from the original content, such as the time, place and event overview, which is more refined than the podcast audio file, and the summarized content is displayed in the content summary interface. In some implementation scenarios, an area can also be set up to record the questions that the user thinks about when reading the original content and the answers to the corresponding questions that the user thinks about, which is conducive to the user's deepening impression of the original content.

[0074] Among them, Fig.14 and Fig.15 As shown, when the user needs to re-read the original text of the original content, the podcast subtitle display interface can be switched to the original text display interface in response to the selection of the original text display tag in the tag option bar, such as Fig.14 As shown, the text content converted from the original content can be displayed in the original text display interface. Then, in response to clicking the original text playback control in the original text display interface, an original text audio file is generated based on the original content and the original text display interface is switched to Fig.15 The original text playback interface shown in the figure also displays text content converted from the original content. In some special scenarios, if the original content is originally an audio file, the original content can be resynthesized into an original audio file according to the user's voice setting information to meet the user's personalized needs.

[0075] In some application scenarios, such as Figure 8 , Fig.13 and Fig.15 As shown, an audio playback control area can be provided in the podcast subtitle display interface, summary playback interface, and original text playback interface to facilitate the user to control the playback of the audio file. The audio playback control area includes at least one of a progress bar control, a speed control control, a playback start and stop control, a jump playback control, and a file export control.

[0076] Among them, Fig.16As shown, in response to the selection of the mind map label in the label option bar, the podcast subtitle display interface is switched to the mind map display interface. The mind map display interface includes a mind map generated based on the original content, and the mind map can be obtained by intelligently summarizing and arranging the text content converted from the original content using an artificial intelligence big model.

[0077] Among them, Fig.17 As shown, in response to selection of the presentation tag in the tag option bar, the podcast subtitle display interface is switched to a presentation interface, wherein the presentation interface includes a presentation generated based on the original content, and the presentation can utilize an artificial intelligence large model to intelligently summarize and organize the text content converted from the original content.

[0078] In some embodiments, a bilingual subtitle control can be set in the podcast subtitle display interface, and a bilingual subtitle control can also be set in the summary playback interface and the original text playback interface to meet the needs of bilingual learning users. Figure 4 , Fig.10 and Fig.11 A detailed description of the process of implementing bilingual subtitles. Figure 4 , Fig.10 and Fig.11 The content arrangement method shown in the embodiment of the present application may include:

[0079] In step S401, in response to a first trigger operation on a bilingual subtitle control, the podcast subtitle display interface is switched to a language selection interface. In the embodiment of the present application, the aforementioned first trigger operation may be any one of clicking a bilingual subtitle button, sliding a bilingual subtitle function start bar, voice triggering to enable the bilingual subtitle function, and shaking to enable the bilingual subtitle function. In actual applications, the specific form of the first trigger operation needs to be determined according to actual application conditions, and the present application does not impose any restrictions in this regard.

[0080] like Fig.10 As shown, the language selection interface includes a bilingual subtitle function start / stop button and a language list, wherein the bilingual subtitle function start / stop button can avoid the situation where the podcast subtitle display interface is mistakenly switched to the language selection interface when the first trigger operation is an erroneous operation, and the language list displays a variety of language names for selection.

[0081] In step S402, in response to the second trigger operation on the start / stop button of the bilingual subtitle function, the bilingual subtitle function is enabled. Fig.10The switch button shown in the figure, the second trigger operation is to toggle the switch button, the bilingual subtitle function start and stop button can also be a check button, and the second trigger operation is to click to check. It can be understood that in actual application, the specific form of the bilingual subtitle function start and stop button needs to be determined according to the actual application situation, and this application does not make any restrictions in this regard.

[0082] In step S403, in response to the third trigger operation, the target language is determined from the language list, the language selection interface is switched back to the podcast subtitle display interface, and the podcast subtitles corresponding to the original language and the podcast subtitles corresponding to the target language are displayed in the podcast subtitle display interface. In an embodiment of the present application, the aforementioned third trigger operation can be to click on the desired language to trigger the determination of the target language, or to drag the desired language to the bottom of the interface to trigger the determination of the target language, or to double-click the desired language to trigger the determination of the target language. It can be understood that in actual applications, the specific form of the third trigger operation needs to be determined according to the actual application situation, and the present application does not impose any restrictions in this regard. In addition, the podcast subtitles corresponding to the original language can be translated by a large artificial intelligence model to obtain the podcast subtitles corresponding to the target language.

[0083] In some application scenarios, if the user clicks on the display area of ​​the podcast subtitles corresponding to the target language, it means that the user may be more interested in the podcast subtitles corresponding to the target language. The display area of ​​the podcast subtitles corresponding to the target language can be placed on top of the display area of ​​the podcast subtitles corresponding to the original language, so that the podcast subtitles corresponding to the target language can be more prominently displayed. At the same time, the podcast audio file can also be updated to an audio file with the target language as the expression language, which is conducive to meeting the user's listening and reading needs for the target language.

[0084] Exemplary Devices

[0085] After introducing the method of the exemplary embodiment of the present application, next, refer to Fig.18 and Fig.19 The related products of the content organization method according to the exemplary embodiment of the present application are described.

[0086] Fig.18 The structure diagram of the content arrangement device according to an embodiment of the present application is schematically shown. Fig.18 The content arrangement device 1800 shown in the embodiment of the present application may include:

[0087] Content receiving module 1801, used for receiving original content input by a user;

[0088] An interface switching module 1802 is used to switch the content input interface to a podcast type selection interface in response to the original content reception completion information;

[0089] The file generation module 1803 is used to generate a podcast audio file based on the original content and the target podcast type in response to the user determining the target podcast type in the podcast type selection interface; and to switch to display the podcast subtitle display interface through the interface switching module 1802.

[0090] The content organization device shown in the present application receives the original content input by the user through the content input interface, and then switches the content input interface to the podcast type selection interface in response to obtaining the original content, so as to provide a variety of podcast styles to users in different scenarios. Furthermore, the present application can generate a podcast audio file based on the original content and the target podcast type in response to the target podcast type determined through the podcast type selection interface, and switch to display the podcast subtitle display interface, so that the original content that needs to be recited can be conveniently organized into a podcast audio file according to the user's style preferences, which is convenient for users to review knowledge content. In combination with the display of podcast subtitles, it can also solve the problem that it is inconvenient for users to carry textbooks and reference materials, thereby improving user convenience. In general, the present application can conveniently organize the recited content into knowledge notes, which is convenient for users to review knowledge content, improve the user's knowledge memory efficiency, and improve the user's user experience.

[0091] Since the specific functions implemented by the content organization device are the same as those of the content organization method described above, the specific details or further implementation methods can refer to the above description of the content organization method, and will not be described in detail here.

[0092] Fig.19 Schematically shows a schematic block diagram of an electronic device according to an embodiment of the present application. Fig.19 , the electronic device 1900 may include a processor 1901. Further, the electronic device may also include a memory 1902 storing computer instructions, which, when executed by the processor 1901, enables the electronic device 1900 to execute the method according to the above multiple embodiments or implementations.

[0093] In some implementation scenarios, the electronic device 1900 may include a server or a terminal device, such as a physical server, a cloud server, a server cluster, a data processing device, an application testing robot, a computer terminal, a smart terminal, a PC device, an Internet of Things terminal, and the like.

[0094] According to different implementation scenarios, the processor 1901 mentioned above may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or the processor may be any conventional processor, etc.

[0095] Based on the above, the present application further discloses a computer-readable storage medium, which includes program instructions. When the program instructions are executed by a processor, the methods described in the above multiple embodiments or implementations are implemented.

[0096] In some implementation scenarios, the above-mentioned computer-readable storage medium may be any appropriate magnetic storage medium or magneto-optical storage medium, such as resistive random access memory RRAM (Resistive Random Access Memory), dynamic random access memory DRAM (Dynamic Random Access Memory), static random access memory SRAM (Static Random-Access Memory), enhanced dynamic random access memory EDRAM (Enhanced Dynamic Random Access Memory), high-bandwidth memory HBM (High-Bandwidth Memory), hybrid memory cube HMC (Hybrid Memory Cube), etc., or any other medium that can be used to store the required information and can be accessed by the application, module or both. Any such computer storage medium may be part of the device or accessible or connectable to the device. Any application or module described in the present invention may be implemented using computer-readable / executable instructions that may be stored or otherwise maintained by such a computer-readable medium.

[0097] It should be noted that although several devices or sub-devices of the content arrangement device are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiment of the present application, the features and functions of two or more devices described above can be embodied in one device. Conversely, the features and functions of one device described above can be further divided into multiple devices to be embodied.

[0098] In addition, although the operations of the present method are described in a specific order in the accompanying drawings, this does not require or imply that the operations must be performed in this specific order, or that all the operations shown must be performed to achieve the desired results. On the contrary, the steps depicted in the flow chart can be performed in a different order. Additionally or alternatively, some steps can be omitted, multiple steps can be combined into one step, and / or one step can be decomposed into multiple steps.

[0099] The use of the verbs "comprise", "include" and their conjugations mentioned in the application documents does not exclude the presence of elements or steps other than those recorded in the application documents. The article "a" or "an" before an element does not exclude the presence of a plurality of such elements.

[0100] Although the spirit and principle of the present application have been described with reference to several specific embodiments, it should be understood that the present application is not limited to the disclosed specific embodiments, and the division of various aspects does not mean that the features in these aspects cannot be combined to benefit, and this division is only for the convenience of expression. The present application is intended to cover various modifications and equivalent arrangements included in the spirit and scope of the attached claims. The scope of the attached claims conforms to the broadest interpretation, thereby including all such modifications and equivalent structures and functions.

Claims

1. A content arrangement method, characterized in that: include: Receiving input of original content through a content input interface; In response to acquiring the original content, switching the content input interface to a podcast type selection interface; In response to selection of a target podcast type in the podcast type selection interface, a podcast audio file is generated based on the original content and the target podcast type, and a podcast subtitle display interface is switched for display.

2. The content arrangement method according to claim 1, further comprising: In response to selection of the target podcast type in the podcast type selection interface, switching the podcast type selection interface to a sound setting interface; Receiving input of sound setting information through the sound setting interface; The voice setting information includes at least one of voice gender information, pronunciation language information, voice emotion information, voice speech speed information and voice timbre information; And, the podcast audio file is also generated based on the sound setting information.

3. The content arrangement method according to claim 1, characterized in that: The podcast subtitle display interface includes a bilingual subtitle control, and the method further includes: In response to a first trigger operation on the bilingual subtitle control, the podcast subtitle display interface is switched to a language selection interface; wherein the language selection interface includes a bilingual subtitle function start / stop button and a language list; In response to a second trigger operation on the start / stop button of the bilingual subtitle function, enabling the bilingual subtitle function; In response to the third trigger operation, the target language is determined from the language list, the language selection interface is switched back to the podcast subtitle display interface, and the podcast subtitles corresponding to the original language and the podcast subtitles corresponding to the target language are displayed in the podcast subtitle display interface.

4. The content arrangement method according to claim 1, characterized in that: After the podcast subtitle display interface is switched to be displayed, the method further includes: The podcast subtitle display interface includes a label option bar; In response to a selection of a content summary tag in the tag option bar, switching the podcast subtitle display interface to a content summary interface; In response to a click on a summary audio generation control in the content summary interface, a summary audio file is generated based on the original content and the content summary interface is switched to a summary playback interface.

5. The content arrangement method according to claim 4, characterized in that: After the podcast subtitle display interface is switched to be displayed, the method further includes: In response to a selection of an original text display tag in the tag option bar, switching the podcast subtitle display interface to an original text display interface; In response to a click on an original text playback control in the original text display interface, an original text audio file is generated based on the original content and the original text display interface is switched to an original text playback interface.

6. The content arrangement method according to claim 5, characterized in that: An audio playback control area is provided in the podcast subtitle display interface, the summary playback interface and the original text playback interface; wherein the audio playback control area includes at least one of a progress bar control, a speed control control, a playback start and stop control, a jump playback control and a file export control.

7. The content arrangement method according to claim 4, characterized in that: After the podcast subtitle display interface is switched to be displayed, the method further includes: In response to selection of a mind map tag in the tag option bar, the podcast subtitle display interface is switched to a mind map display interface; wherein the mind map display interface includes a mind map generated based on the original content.

8. A content arrangement device, characterized in that: The method for executing the content arrangement method according to any one of claims 1 to 7 comprises: A content receiving module, used for receiving input of original content through a content input interface; An interface switching module, configured to switch a content input interface to a podcast type selection interface in response to acquiring the original content; A file generation module is used to generate a podcast audio file based on the original content and the target podcast type in response to the selection of the target podcast type in the podcast type selection interface, and to switch the display of the podcast subtitle display interface through the interface switching module.

9. An electronic device, characterized in that: include: processor; as well as A memory having executable codes for content arrangement stored thereon, wherein when the executable codes are executed by the processor, the processor is caused to execute the method according to any one of claims 1 to 7.

10. A non-transitory machine-readable storage medium having stored thereon an executable code for content organization, wherein when the executable code is executed by a processor of an electronic device, the processor is caused to execute the method according to any one of claims 1 to 7.