Method and system for providing audio records generated based on post-recording information
By integrating post-recording information, the method and system enhance voice-to-text conversion accuracy, allowing users to accurately convert voice recordings into text.
Patent Information
- Application Number
- JP2025112705
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-04-07
- Filing Date
- 2025-07-03
- Publication Date
- 2025-10-01
AI Technical Summary
Existing speech-to-text conversion technologies face reduced accuracy when converting voice recordings into text, leading to inaccurately converted text being provided to users.
A method and system that receive information associated with voice data after recording, allowing for voice-to-text conversion based on both the voice data and additional information related to the recording, such as notes or keywords, to improve accuracy.
Users can perceive the content of voice recordings both audibly and visually, with improved text conversion accuracy by incorporating post-recording information, enabling accurate text generation.
Smart Images

Figure 2025143399000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to methods and systems for providing an audio recording generated based on information received after the audio is recorded, and more particularly, to methods and systems for receiving information associated with audio data after the audio is recorded, and providing an audio recording generated based on at least some audio contained in the audio data and information associated with the audio data received after the audio is recorded. [Background technology]
[0002] Recently, with the development and widespread use of mobile electronic devices such as smartphones and tablet PCs, users can easily create and save records of voice conversations, text, images, etc., through their mobile electronic devices in their daily lives. For example, users can record audio and / or video of conferences, meetings, classes, interviews, etc., using a note application or a voice recording application. In addition, while recording audio and / or video using a mobile electronic device, users can input notes about the audio and / or video content by creating text about the audio and / or video content.
[0003] Furthermore, with the development of speech-to-text conversion technology (i.e., speech recognition technology), it is possible to convert the content of a voice recording, generated by audio and / or video recording, into text and provide it to a user. In this case, a user can recognize the content of the voice recording through the converted text without directly listening to the voice recording. However, when speech-to-text conversion is performed using only the information of the voice recording, there is a risk that the accuracy of speech recognition may be reduced. In other words, there is a risk that text that is inaccurately converted from the voice contained in the voice recording may be provided to a user. Summary of the Invention [Problem to be solved by the invention]
[0004] The present disclosure provides a method for providing audio recordings, a computer program stored on a recording medium, and an apparatus (system) for solving the above-mentioned problems. [Means for solving the problem]
[0005] The present disclosure can be embodied in numerous ways, including as a method, an apparatus (system), or a computer program stored on a readable storage medium.
[0006] According to one embodiment of the present disclosure, a method for providing a voice recording generated based on information after voice recording, performed by at least one computing device, includes receiving information associated with voice data after voice recording, receiving a request for voice-to-text conversion for the voice data, and outputting a voice recording generated based on at least a portion of the voice included in the voice data and information associated with the voice data received after voice recording in response to the voice-to-text conversion request.
[0007] A non-transitory computer-readable recording medium is provided that stores instructions for executing a method for providing an audio recording according to one embodiment of the present disclosure on a computer.
[0008] An information processing system according to one embodiment of the present disclosure includes a communication module, a memory, and at least one processor coupled to the memory and configured to execute at least one computer-readable program contained in the memory, the at least one program including instructions for receiving information associated with voice data after voice recording, receiving a request for voice-to-text conversion for the voice data, and generating a voice recording in response to the voice-to-text conversion request based on at least a portion of the voice included in the voice data and the information associated with the voice data received after voice recording. [Effects of the Invention]
[0009] In some embodiments of the present disclosure, a user may be provided with a voice recording corresponding to the content of the voice data contained in the voice recording, thereby allowing the user to perceive the content of the voice recording both audibly and visually. Furthermore, after voice recording or voice recognition, the voice recording may be converted based on information related to the voice data entered by the user, thereby providing text in which the voice data has been converted more accurately.
[0010] In some embodiments of the present disclosure, if a user has difficulty creating notes via a mobile phone or PC while recording, the user can create a note related to the recording after the recording is completed, extract keywords contained in the created note, and re-recognize the voice of the recording file, thereby improving the voice recognition rate.
[0011] The effects of the present disclosure are not limited to these, and other effects not mentioned should be clearly understood by a person with ordinary skill in the art to which the present disclosure pertains (hereinafter referred to as "a person skilled in the art") from the description of the claims. [Brief explanation of the drawings]
[0012] BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Embodiments of the present disclosure will now be described, without limitation, with reference to the accompanying drawings, in which like reference numerals refer to like elements and in which:
[0014] FIG. [Figure 1] FIG. 10 is a diagram illustrating an example of providing a voice recording generated based on information related to voice data created after voice recording according to an embodiment of the present disclosure. [Figure 2] FIG. 1 is a schematic diagram showing a configuration in which an information processing system is connected to multiple user terminals so as to be able to communicate with them in order to provide an audio recording provision service according to one embodiment of the present disclosure. [Figure 3] 1 is a block diagram showing an internal configuration of a user terminal and an information processing system according to an embodiment of the present disclosure. [Figure 4] 1 is a flowchart illustrating a method for providing an audio recording according to one embodiment of the present disclosure. [Figure 5]FIG. 10 illustrates an example in which a memo is created regarding the audio data after a first audio recording for the audio data is output according to an embodiment of the present disclosure. [Figure 6] FIG. 10 is a diagram showing an example of outputting a second voice recording as a result of reconversion of at least a portion of the voice contained in the voice data according to an embodiment of the present disclosure. [Figure 7] FIG. 10 is a diagram illustrating an example of outputting a second speech recording generated by correction information related to at least a portion of the text included in the first speech recording according to an embodiment of the present disclosure. [Figure 8] FIG. 10 is a diagram illustrating an example of outputting a voice recording generated based on notes regarding the voice data created after voice recording according to an embodiment of the present disclosure. [Figure 9] FIG. 10 illustrates an example of outputting an audio recording generated based on one or more participant information and / or topics related to audio data received after audio recording according to one embodiment of the present disclosure. [Figure 10] 1 is a flowchart illustrating a process for reconverting audio data and / or editing an audio recording according to one embodiment of the present disclosure. [Figure 11] FIG. 1 illustrates an example of an artificial neural network model according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0013] Hereinafter, specific implementations of the present disclosure will be described in detail with reference to the accompanying drawings. However, in the following description, specific descriptions of well-known functions and configurations will be omitted if they may unnecessarily obscure the gist of the present disclosure.
[0014] In the accompanying drawings, the same or corresponding components are denoted by the same reference numerals. In addition, in the following description of the embodiments, duplicated descriptions of the same or corresponding components may be omitted. However, even if a description of a component is omitted, it should not be intended that such a component is not included in a certain embodiment.
[0015] The advantages and features of the disclosed embodiments, and methods for achieving them, will become clearer with reference to the following examples in conjunction with the accompanying drawings. However, the present disclosure is not limited to the embodiments disclosed below, and may be embodied in various different forms. However, the present embodiments are provided solely for the purpose of completeness of the disclosure and to enable those skilled in the art to accurately recognize the scope of the invention.
[0016] The terms used in this specification will be briefly explained, and the disclosed embodiments will be specifically described. The terms used in this specification are currently commonly used and generally used terms, taking into consideration the function of the present disclosure. However, these terms may change depending on the intentions of engineers in the relevant field, legal precedents, the emergence of new technologies, etc. In addition, in specific cases, the applicant may arbitrarily select terms, and their meanings will be described in detail in the description of the invention. Therefore, the terms used in this disclosure should be defined based on the meanings of the terms and the overall content of the present disclosure, rather than simply by the names of the terms.
[0017] In this specification, unless otherwise clearly specified in the context, singular expressions can include plural expressions, and plural expressions can include singular expressions. Throughout this specification, when a part "comprises" a certain element, this does not exclude other elements, and means that other elements may also be included, unless otherwise specified to the contrary.
[0018] Additionally, the terms "module" and "module" used in this specification refer to software or hardware components, each of which performs a certain function. However, the terms "module" and "module" are not limited to software or hardware. A "module" or "module" may be configured to reside on an addressable storage medium or to execute on one or more processors. Thus, by way of example, a "module" or "module" may include components such as software components, object-oriented software components, class components, and task components, as well as at least one of processes, functions, attributes, procedures, subroutines, program code segments, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables. The components and "modules" or "modules" may be combined into fewer components and "modules" or "modules," or the functionality provided therein may be further separated into additional components and "modules" or "modules."
[0019] According to one embodiment of the present disclosure, a "module" or "unit" may be embodied with a processor and memory. "Processor" should be broadly interpreted to include a general-purpose processor, a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a controller, a microcontroller, a state machine, etc. In some environments, "processor" may also refer to an application-specific semiconductor (ASIC), a programmable logic device (PLD), a field-programmable gate array (FPGA), etc. "Processor" may also refer to a combination of processing devices, such as, for example, a combination of a DSP and a microprocessor, a combination of multiple microprocessors, a combination of one or more microprocessors in conjunction with a DSP core, or any other such configuration. Also, "memory" should be broadly interpreted to include any electronic component capable of storing electronic information. "Memory" can refer to various types of processor-readable media, such as RAM (Random Access Memory), ROM (Read Only Memory), NVRAM (Non-Volatile Random Access Memory), PROM (Programmable Read-Only Memory), EPROM (Erasable Programmable Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), flash memory, magnetic or optical data storage devices, registers, etc. Memory is said to be in electronic communication with a processor if the processor can read information from / record information read into the memory. Memory that is integrated into a processor is in electronic communication with the processor.
[0020] In this disclosure, "audio data" may include data generated / stored by audio recording. Here, audio recording may refer to audio data, and audio data may refer to audio recordings. In one embodiment, audio data may include one or more voices. Here, one or more voices may refer to data corresponding to at least one section of multiple sections of audio data. Alternatively or additionally, one or more voices may refer to the voice, voice data, utterances, and / or utterance data of each speaker included in the audio data. In this disclosure, audio data and / or information related to a voice may include the voice itself and / or data indicative of the voice (e.g., vector data).
[0021] In this disclosure, "audio recording" may refer to a recording generated by converting spoken content contained in an audio recording into text. Here, a first audio recording may refer to an audio recording generated without reflecting information related to audio data received after the audio recording, and a second audio recording may refer to an audio recording generated by reflecting information related to audio data received after the audio recording, but is not limited thereto.
[0022] 1 is a diagram illustrating an example of providing an audio transcript 122 generated based on information associated with audio data created after audio recording according to one embodiment of the present disclosure. The screen illustrated in FIG. 1 illustrates an example in which a user executes a recording application, such as an audio recording application, a memo application, and / or a note application, via a user terminal (e.g., a smartphone, a tablet PC, a desktop, etc.), and is provided with an audio transcript 122 related to the audio data. In one embodiment, the user can be provided with text corresponding to the audio data included in the audio recording via such a recording application.
[0023] A user terminal (e.g., at least one processor of the user terminal) can receive information associated with audio data after audio recording to provide an audio recording. For example, the user terminal can receive information associated with audio data input by a user via an input device (e.g., a keyboard, a mouse, a microphone, etc.) after audio recording. Additionally or alternatively, the user terminal can receive information associated with audio data stored in a storage device after audio recording from the storage device. Here, the information associated with audio data can refer to any information that can be included in the audio data or that can describe or characterize the audio data, and can include, for example, notes 112, 114 related to the audio data, a topic 116 related to the audio data, information 118 about one or more participants associated with the audio included in the audio data, etc. Additionally or alternatively, the information associated with the audio data can include one or more keywords extracted from the information associated with the audio data. Additionally or alternatively, the information associated with the audio data can include one or more keywords extracted from text included in an existing audio recording.
[0024] In response to a request for speech-to-text conversion of the speech data, the user terminal may output a speech recording 122 generated based on at least a portion of the speech included in the speech data and information associated with the speech data received after the speech recording. For example, the user terminal may receive a user input selecting an icon 120 indicating a request for conversion (or reconversion) of the speech recording, and in response, display at least one text included in the speech recording 122 on a display. Here, the speech recording 122 may be generated by at least one processor of the information processing system and / or at least one processor of the user terminal.
[0025] In one embodiment, the audio recording 122 may include text information that is output by inputting information about at least a portion of the audio and information associated with the audio data into a speech-to-text transcription model. Here, the speech-to-text transcription model may include a model trained to output text corresponding to the reference audio by inputting a reference audio and reference information associated with the reference audio. For example, the reference information associated with the reference audio may include one or more reference keywords associated with the reference audio. That is, the speech-to-text transcription model may include a model trained to output text corresponding to the reference audio by inputting a reference audio and one or more reference keywords associated with the reference audio.
[0026] In one embodiment, applying a speech recognition weighting to one or more keywords input to a speech-to-text transcription model causes the one or more keywords to be recognized as having a higher priority than the one or more other keywords. For example, applying a speech recognition weighting to the keyword "demo" input to a speech-to-text transcription model causes the keyword "demo" to be recognized as having a higher priority than the keyword "nemo." Thus, the speech-to-text transcription model can recognize at least a portion of the speech included in the input speech data as "demo" instead of "nemo," convert the speech to text, and output the text, allowing the user terminal to output a speech recording that includes "demo" instead of "nemo."
[0027] As shown in the figure, the user terminal may display information associated with the audio data on a display. For example, the information associated with the audio data may include a title of the audio data ("Demo Site Meeting") 116, information about one or more participants associated with the audio included in the audio data ("user1, user2, user3") 118, notes 112 created during audio recording, and notes 114 created after audio recording. Additionally or alternatively, the information associated with the audio data may include an audio recording (e.g., a first audio recording) that does not reflect information received after audio recording. Such an audio recording that does not reflect information received after audio recording may also be displayed on the display. Subsequently, in response to a user's touch input on a "reconversion" icon 120 indicating a request for conversion (or a request for reconversion) of the audio recording, the user terminal may display an audio recording (e.g., a second audio recording) 122 created based on information associated with the audio data received after audio recording. The user terminal may also output a pop-up message 124 including the message, "Audio recording reconversion completed."
[0028] According to the above-described embodiments, a user can receive a voice recording corresponding to the content of the voice data included in the voice recording, thereby audibly and visually perceiving the content of the voice recording. Furthermore, after the voice recording or voice recognition, the voice recording can be converted based on information related to the voice data entered by the user, thereby providing text in which the voice data has been converted more accurately.
[0029] 2 is a schematic diagram illustrating a configuration in which an information processing system 230 is communicably connected to multiple user terminals 210_1, 210_2, and 210_3 to provide a voice recording service according to one embodiment of the present disclosure. The information processing system 230 may include a system capable of providing a voice recording service, a system capable of providing recording services such as voice recording, memos, and notes, and / or a system capable of providing a voice-to-text conversion service. In one embodiment, the information processing system 230 may include computer-executable programs (e.g., downloadable applications) related to the voice recording service, the recording service, and / or the voice-to-text conversion service, one or more server devices and / or databases capable of storing, providing, and executing data, or one or more distributed computing devices and / or distributed databases based on a cloud computing service. For example, the information processing system 230 may include separate systems (e.g., servers) for the voice recording service, the recording service, and / or the voice-to-text conversion service.
[0030] The voice recording service, recording service, voice-to-text conversion service, etc. provided by the information processing system 230 are provided to users through a voice recording application, a memo application, a note application, a voice-to-text conversion application, etc. installed in each of the user terminals 210_1, 210_2, and 210_3. For example, the information processing system 230 can provide information corresponding to a voice-to-text conversion request received from the user terminals 210_1, 210_2, and 210_3 or perform a corresponding process through the voice recording application, etc.
[0031] A plurality of user terminals 210_1, 210_2, and 210_3 can communicate with the information processing system 230 via a network 220. The network 220 can be configured to enable communication between the plurality of user terminals 210_1, 210_2, and 210_3 and the information processing system 230. Depending on the installation environment, the network 220 can be composed of a wired network such as Ethernet (registered trademark), PLC (Power Line Communication), telephone line communication device, and RS-serial communication, a mobile communication network, a wireless network such as WLAN (Wireless LAN), Wi-Fi (registered trademark), Bluetooth (registered trademark), and ZigBee (registered trademark), or a combination thereof. The communication method is not limited and can include not only a communication method utilizing a communication network (e.g., a mobile communication network, a wired Internet, a wireless Internet, a broadcast network, a satellite network, etc.) that can include the network 220, but also short-range wireless communication between the user terminals 210_1, 210_2, and 210_3.
[0032] 2 illustrates a mobile phone terminal 210_1, a tablet terminal 210_2, and a PC terminal 210_3 as examples of user terminals, but is not limited thereto, and the user terminals 210_1, 210_2, and 210_3 may be any computing devices capable of wired and / or wireless communication and capable of installing and executing a voice recording application, etc. For example, the user terminal may include a smartphone, a mobile phone, a navigation system, a desktop computer, a laptop computer, a digital broadcasting terminal, a PDA (Personal Digital Assistant), a PMP (Portable Multimedia Player), a tablet PC, a game console, a wearable device, an IoT (Internet of Things) device, a VR (Virtual Reality) device, an AR (Augmented Reality) device, etc. Also, while FIG. 2 shows three user terminals 210_1, 210_2, and 210_3 communicating with the information processing system 230 via the network 220, this is not limited thereto, and a different number of user terminals may be configured to communicate with the information processing system 230 via the network 220.
[0033] In one embodiment, the information processing system 230 may receive a request for speech-to-text conversion of voice data from the user terminals 210_1, 210_2, and 210_3. The information processing system 230 may also receive at least one of voice data or information related to the voice data from the user terminals 210_1, 210_2, and 210_3. In response to the speech-to-text conversion request, the information processing system 230 may generate a voice recording based on at least a portion of the voice included in the voice data and information related to the voice data received after the voice recording, and provide the generated voice recording to the user terminals 210_1, 210_2, and 210_3. Alternatively, the user terminals 210_1, 210_2, and 210_3 may generate a voice recording based on at least a portion of the voice included in the voice data and information related to the voice data received after the voice recording.
[0034] FIG. 3 is a block diagram showing the internal configuration of a user terminal 210 and an information processing system 230 according to an embodiment of the present disclosure. The user terminal 210 may refer to any computing device capable of executing a voice recording application, a memo application, a note application, a voice-to-text conversion application, etc., and capable of wired / wireless communication, and may include, for example, the mobile phone terminal 210_1, the tablet terminal 210_2, and the laptop computer terminal 210_3 of FIG. 2 . As shown in the figure, the user terminal 210 may include a memory 312, a processor 314, a communication module 316, and an input / output interface 318. Similarly, the information processing system 230 may include a memory 332, a processor 334, a communication module 336, and an input / output interface 338. As shown in FIG. 3 , the user terminal 210 and the information processing system 230 may be configured to communicate information and / or data over the network 220 using their respective communication modules 316 and 336. Additionally, the input / output device 320 may be configured to input information and / or data to the user terminal 210 and output information and / or data generated by the user terminal 210 via the input / output interface 318 .
[0035] The memories 312 and 332 may include any non-transitory computer-readable recording medium. According to one embodiment, the memories 312 and 332 may include a permanent mass storage device such as a random access memory (RAM), a read only memory (ROM), a disk drive, a solid state drive (SSD), or a flash memory. As another example, a permanent mass storage device such as a ROM, an SSD, a flash memory, or a disk drive may be included in the user terminal 210 or the information processing system 230 as a separate permanent storage device distinct from the memory. The memories 312 and 332 may also store an operating system and at least one program code (e.g., code for a voice recording application, a memo application, a note application, a speech-to-text conversion application, etc.).
[0036] Such software components may be loaded from a computer-readable recording medium separate from the memories 312, 332. Such separate computer-readable recording medium may include a recording medium directly connectable to the user terminal 210 and the information processing system 230, but may also include computer-readable recording media such as a floppy drive, a disk, a tape, a DVD / CD-ROM drive, and a memory card. As another example, the software components may be loaded into the memories 312, 332 via the communication modules 316, 336 rather than from a computer-readable recording medium. For example, at least one program may be loaded into the memories 312, 332 based on a computer program (e.g., a voice recording application, a memo application, a note application, a speech-to-text conversion application, etc.) installed by a file provided via the network 220 by a developer or a file distribution system that distributes application installation files.
[0037] The processors 314, 334 may be configured to process computer program instructions by performing basic arithmetic, logic, and input / output operations. The instructions may be provided to the processors 314, 334 by the memory 312, 332 or the communication modules 316, 336. For example, the processors 314, 334 may be configured to execute instructions received by program code stored in a storage device, such as the memory 312, 332.
[0038] The communication modules 316 and 336 may provide configurations and functions for the user terminal 210 and the information processing system 230 to communicate with each other via the network 220, and may provide configurations and functions for the user terminal 210 and / or the information processing system 230 to communicate with other user terminals or other systems (e.g., another cloud system, etc.). As an example, a request or data (e.g., a voice-to-text conversion request for voice data) generated by the processor 314 of the user terminal 210 via program code stored in a storage device such as the memory 312 may be transmitted to the information processing system 230 via the network 220 under the control of the communication module 316. Conversely, a control signal or command provided under the control of the processor 334 of the information processing system 230 may be received by the user terminal 210 via the communication module 316 of the user terminal 210 via the communication module 336 and the network 220. For example, the user terminal 210 can receive, from the information processing system 230 via the communication module 316, at least a portion of the voice contained in the voice data and a voice recording generated based on information related to the voice data received after the voice recording.
[0039] The input / output interface 318 may be a means for interfacing with the input / output device 320. For example, the input device may include a device such as a camera including an audio sensor and / or an image sensor, a keyboard, a microphone, a mouse, etc., and the output device may include a device such as a display, a speaker, a haptic feedback device, etc. As another example, the input / output interface 318 may be a means for interfacing with a device that integrates components or functions for performing input and output, such as a touch screen. Although FIG. 3 illustrates the input / output device 320 as not being included in the user terminal 210, the present invention is not limited thereto and may be configured integrally with the user terminal 210. In addition, the input / output interface 338 of the information processing system 230 may be a means for connecting to the information processing system 230 or for interfacing with an input or output device (not shown) that may be included in the information processing system 230. In FIG. 3, the input / output interfaces 318 and 338 are shown as elements configured separately from the processors 314 and 334 , but this is not limiting, and the input / output interfaces 318 and 338 may also be configured to be included in the processors 314 and 334 .
[0040] The user terminal 210 and the information processing system 230 may include more components than those shown in FIG. 3 . However, it is not necessary to explicitly show most of the conventional components. According to one embodiment, the user terminal 210 may be embodied to include at least a portion of the input / output device 320 described above. The user terminal 210 may also include other components such as a transceiver, a global positioning system (GPS) module, a camera, various sensors, and a database. For example, if the user terminal 210 is a smartphone, it may include components typically found in smartphones. For example, the user terminal 210 may be embodied to further include various components such as an acceleration sensor, a gyro sensor, a microphone module, a camera module, various physical buttons, buttons using a touch panel, input / output ports, and a vibrator for vibration.
[0041] According to one embodiment, the processor 314 of the user terminal 210 may be configured to run a voice recording application, a memo application, a note application, a voice-to-text conversion application, etc. In this case, program code associated with the application may be loaded into the memory 312 of the user terminal 210. While the application is running, the processor 314 of the user terminal 210 may receive information and / or data provided by the input / output device 320 via the input / output interface 318 or from the information processing system 230 via the communication module 316, process the received information and / or data, and store it in the memory 312. Such information and / or data may also be provided to the information processing system 230 via the communication module 316.
[0042] While an audio recording application is running, the processor 314 can receive audio data, text, images, videos, etc. input or selected through an input device connected to the input / output interface 318, such as a touchscreen, keyboard, camera including an audio sensor and / or image sensor, microphone, etc., and can store the received audio data, text, images, videos, etc. in the memory 312 or provide them to the information processing system 230 via the communication module 316 and the network 220. In one embodiment, the processor 314 can receive information related to audio data, a request for speech-to-text conversion of the audio data, etc. through the input device 320, such as a touchscreen or mouse, and can provide the information related to the audio data, a request for speech-to-text conversion of the audio data, etc. to the information processing system 230 via the communication module 316 and the network 220.
[0043] The processor 314 of the user terminal 210 can transfer information and / or data to the input / output device 320 for output via the input / output interface 318. For example, the processor 314 of the user terminal 210 can output the processed information and / or data via the output device 320, such as a display-capable device (e.g., a touchscreen or display) or an audio-capable device (e.g., a speaker). In one embodiment, the processor 314 can display a voice recording of the audio data on the display of the user terminal 210. Additionally, the processor 314 can output at least a portion of the audio included in the audio data through the speaker of the user terminal 210.
[0044] The processor 334 of the information processing system 230 may be configured to manage, process, and / or store information and / or data received from the plurality of user terminals 210 and / or the plurality of external systems. The information and / or data processed by the processor 334 may be provided to the user terminal 210 via the communication module 336 and the network 220. In one embodiment, the processor 334 of the information processing system 230 may receive information associated with the voice data after recording the voice, receive a request for speech-to-text conversion for the voice data, generate a voice recording based on at least a portion of the voice included in the voice data and the information associated with the voice data received after recording the voice in response to the speech-to-text conversion request, and provide the generated voice recording to the user terminal 210 via the communication module 336 and the network 220.
[0045] For example, processor 334 may use notes about the audio data created after recording the audio (e.g., notes related to audio included in a specific section of the audio data) to perform speech-to-text conversion on the audio included in the specific section, thereby generating an audio recording that includes text information corresponding to the audio included in the specific section. Alternatively or additionally, processor 334 may input information about at least a portion of the audio and information related to the audio data into a speech-to-text transcription model, thereby generating an audio recording that includes text information corresponding to at least a portion of the audio information. Alternatively or additionally, processor 334 may apply speech recognition weightings to one or more keywords input into the speech-to-text transcription model, and recognize one or more keywords as having a higher priority than keywords that are different from the one or more keywords. Alternatively or additionally, the processor 334 may generate a first audio recording by speech-to-text conversion of at least some of the audio contained in the audio data before receiving information associated with the audio data, receive a request for speech-to-text reconversion of the audio data, and, in response to the speech-to-text reconversion request, generate a second audio recording based on at least some of the audio contained in the audio data and information associated with the audio data received after recording the audio.
[0046] FIG. 4 is a flowchart illustrating a method 400 for providing an audio recording according to one embodiment of the present disclosure. In one embodiment, the method 400 for providing an audio recording may be performed by a processor (e.g., at least one processor of a user terminal and / or an information processing system). As shown, the method 400 for providing an audio recording may begin by the processor receiving information associated with audio data after audio recording (S410). Here, the information associated with the audio data received after audio recording may include notes about the audio data created after audio recording. Additionally or alternatively, the information associated with the audio data may include information about one or more participants associated with the audio included in the audio data. Additionally or alternatively, the information associated with the audio data may include a topic related to the audio data.
[0047] The processor may receive a request for speech-to-text conversion of the voice data (S420). In response to the speech-to-text conversion request, the processor may output a voice recording generated based on at least a portion of the voice included in the voice data and information associated with the voice data received after the voice recording (S430).
[0048] In one embodiment, the audio recording may include text information that is output by inputting information about at least a portion of the audio and information related to the audio data into a audio-to-text transcription model. Here, the audio-to-text transcription model may be trained to output text corresponding to the reference audio by inputting reference audio and reference information related to the reference audio. In this case, the information related to the audio data may include one or more keywords extracted from the information related to the audio data, and the reference information related to the reference audio may include one or more reference keywords related to the reference audio.
[0049] In addition, by applying a speech recognition weighting to one or more keywords input into the speech-to-text transcription model, the one or more keywords can be recognized as having a higher priority than other keywords. Here, the one or more keywords may correspond to meaningful keywords extracted from information related to the speech data using an artificial neural network model, a machine learning model, a keyword extraction algorithm, etc. Examples of keyword extraction algorithms include, but are not limited to, an algorithm that extracts keywords that are frequently used in the current recording, keywords that are used more frequently in the current recording compared to other documents (recordings), keywords that are not used in other documents (recordings) and are used for the first time in the current recording, etc.
[0050] In one embodiment, notes about the audio data created after recording the audio may be associated with audio included in a specific section of the audio data, and the audio recording may include text information generated by speech-to-text conversion of the audio included in the specific section using the notes created in association with the specific section.
[0051] In one embodiment, the processor may output a first voice recording generated by speech-to-text conversion of at least a portion of the voice included in the voice data before receiving information related to the voice data. The processor may then receive information related to the voice data after generating the first voice recording and receive a request for speech-to-text reconversion of the voice data. In response to the request for speech-to-text reconversion, the processor may output, as a voice recording, a second voice recording generated based on at least a portion of the voice included in the voice data and information related to the voice data received after the voice recording. In this case, the information related to the voice data may include one or more keywords extracted from text included in the first voice recording. Alternatively or additionally, the information related to the voice data may include information on corrections to at least a portion of the text included in the first voice recording. In this case, speech recognition weights in the voice-to-text transcription model may be applied to the corrected text.
[0052] 5 illustrates an example in which notes 516 related to audio data are created after a first audio recording 510 for the audio data is output according to one embodiment of the present disclosure. A user can input information related to the audio recording (i.e., information related to the audio data) 512, 514 via a user terminal before or during audio recording. For example, a user can input information related to the subject of the audio recording and / or information related to participants in the audio recording 514 before starting audio recording. As another example, a user can input notes related to the audio recording (i.e., notes related to the audio data) 512 during audio recording. A processor (e.g., at least one processor of a user terminal) can receive the information related to the audio data 512, 514 input in this manner before or during audio recording.
[0053] The processor may output a first audio recording 510 generated by speech-to-text conversion of at least some of the speech contained in the audio data. In one embodiment, the first audio recording 510 may be generated based on information 512, 514 associated with the audio data received before or during the audio recording. To output the first audio recording 510, one or more keywords may be extracted from the information 512, 514 associated with the audio data received before or during the audio recording. The first audio recording 510 may include the extracted one or more keywords and text information output by inputting the audio data into a speech-to-text transcription model. For example, the first audio recording 510 may be generated by applying speech recognition weightings to one or more keywords extracted through the speech-to-text transcription model, and recognizing weighted keywords as a higher priority.
[0054] The keyword "plan" may be extracted from the notes 512 made during the audio recording, and the extracted keyword "plan" and at least some of the information about the audio included in the audio data may be input to a speech-to-text transcription model. A speech recognition weighting value may be applied to the keyword "plan" through the speech-to-text transcription model, and the weighted keyword "plan" may be recognized as having a higher priority than other keywords. This allows the speech-to-text transcription model to output a "proposal" corresponding to at least some of the information about the audio, generating a first audio recording 510 containing the "proposal."
[0055] A user can input information associated with an audio recording (i.e., information associated with audio data) 516 via a user terminal even after audio recording. For example, a user can input notes about the audio recording (i.e., notes associated with audio data) 516 after audio recording. As another example, a user can input (or additionally input) participant information about the audio recording after audio recording. As shown in FIG. 5 , a user can create / input notes about audio data 516 after audio recording and after the first audio recording 510 is generated / output. Thus, the processor can receive notes about audio data 516 created / input by the user after audio recording and after the first audio recording 510 is generated / output.
[0056] 5 illustrates, without limitation, an example in which the processor receives notes 516 regarding the audio data after the first audio recording 510 is generated and / or output. For example, the processor can receive notes regarding the audio data after the audio is recorded and before the first audio recording is generated and / or output.
[0057] 6 is a diagram illustrating an example of outputting a second audio recording 622 as a result of reconversion of at least some of the audio included in the audio data according to an embodiment of the present disclosure. In one embodiment, a processor (e.g., at least one processor of a user terminal) may receive information associated with the audio data after the audio is recorded. Here, the information associated with the audio data received after the audio is recorded may include notes 616 associated with the audio data created after the audio is recorded, a title related to the audio data created after the audio is recorded, information about one or more participants associated with the audio included in the audio data, the first audio recording (or one or more keywords extracted from text included in the first audio recording) 618, etc. For example, the processor may receive information associated with the audio data after the first audio recording 618 is generated / output.
[0058] The processor may receive a request for speech-to-text reconversion of the audio data. In response to the speech-to-text reconversion request, the processor may output, as a speech recording, a second audio recording 622 generated based on at least a portion of the audio included in the audio data and information 616 associated with the audio data received after the audio recording. Here, the information associated with the audio data may include one or more keywords extracted from the text included in the first audio recording 618. To generate the second audio recording 622, the keyword "demo" may be extracted from the text included in the first audio recording 618, and the keywords "web," "add," and "demo" may be extracted from the information 616 associated with the audio data received after the audio recording. Thereafter, the extracted keywords and information related to at least a portion of the audio included in the audio data may be input into a speech-to-text transcription model to generate the second audio recording 622 including output text information.
[0059] As shown in first operation 610, in response to, for example, a user's touch input on a "reconvert" icon 612 indicating a request for speech-to-text reconversion, the processor can output a pop-up message ("Reconvert speech recording?") 614 regarding whether or not to reconvert the speech recording. Based on the user's response to the output pop-up message 614, the processor can output a second speech recording 622 generated based on at least a portion of the speech included in the speech data and information associated with the speech data received after the speech recording. Thus, in first operation 610, due to an inaccurate speech-to-text conversion, a first speech recording 618 including the text "Web features have been exceeded in this demo," is displayed on the display, whereas in second operation 620, due to an accurate speech-to-text conversion, a second speech recording 622 including the text "Web features have been added in this demo." is displayed on the display.
[0060] FIG. 7 illustrates an example of outputting a second speech recording generated using correction information for at least a portion of text included in a first speech recording, according to an embodiment of the present disclosure. In one embodiment, the information associated with the speech data may include correction information for at least a portion of text included in the first speech recording. In this case, a speech recognition weighting value in a speech-to-text transcription model may be applied to the corrected text from the text included in the first speech recording based on the correction information. That is, the corrected text from the text included in the first speech recording may be extracted as a keyword, and information about at least a portion of the speech and the corrected text (i.e., the extracted keyword) may be input into a speech-to-text transcription model, thereby applying a speech recognition weighting value to the corrected text (i.e., the extracted keyword). In this case, the corrected text (i.e., the extracted keyword) may be recognized as having a higher priority than other keywords during speech-to-text conversion.
[0061] As shown in a first operation 710, a processor (e.g., at least one processor of a user terminal) can output a first voice recording generated by speech-to-text conversion of a voice recording. The first voice recording can include text in which the voice data is mistranslated, such as "Please share the Nemo site discussed last month" 712 and "This time, Nemo has exceeded the capabilities of Webbe" 714. A user can modify at least a portion of the text included in the first voice recording. For example, a user can select (e.g., click) an "Edit" icon 716 to modify at least a portion of the text included in the first voice recording. In response, the processor can switch to an edit mode to provide the user with an interface that allows the user to modify at least a portion of the text included in the first voice recording. Thereafter, the user can modify "Nemo" to "Demo" in "Please share the Nemo site discussed last month" 712 included in the first voice recording.
[0062] The user can then select the "Reconvert" icon 718 to request speech-to-text reconversion. In response to the user's speech-to-text reconversion request, the processor can output a second speech recording generated by speech-to-text reconversion based on correction information for at least a portion of the text included in the first speech recording. For example, the "demo" in the text corrected by the user in the first operation 710 can be extracted as a keyword, and information about at least a portion of the speech and the corrected text "demo" can be input into a speech-to-text transcription model, whereby a speech recognition weighting can be applied to the corrected text "demo." As a result, the speech converted to "In this demo, we've exceeded the functionality for use on Web" 714 in the first speech recording can be converted to "In this demo, we've added functionality for use on Web" 722 in the second speech recording. Therefore, as shown in the second operation 720, a second speech recording including "In this demo, we've added functionality for use on Web" 722 can be generated, and the processor can output the generated second speech recording.
[0063] 8 is a diagram illustrating an example of outputting a voice recording generated based on a note 814 about the voice data created after voice recording, according to an embodiment of the present disclosure. In one embodiment, information related to the voice data received after voice recording may include a note 814 about the voice data created after voice recording. Here, the note 814 about the voice data created after voice recording is associated with voice included in a specific section within the voice data. In this case, the voice recording may include text information generated by speech-to-text conversion of the voice included in the specific section using the note created in association with the voice included in the specific section.
[0064] As shown in first operation 810, a processor (e.g., at least one processor of a user terminal) can output a first audio recording including "This time, Nemo has exceeded the capabilities of Webe." 812. After recording the audio, a user can create / input a memo regarding a specific section of the audio data and / or a specific section of the first audio recording for the audio data. For example, a user can select a section (e.g., a section between a start point and an end point, or a specific time point) from the audio recording (or audio data) and create / input a memo regarding the section. As shown in the figure, a user can create / input a memo 814 regarding the audio data for a time point "01:07" in the audio recording. Here, the "01:07" time point can correspond to the "01:07" text section 812 in the first audio recording. In this case, the processor can output the created / inputted memo 814 together with time information ("01:07") indicating the corresponding section. The user can then select (eg, click-in) the "Reconvert" icon 816 to complete a voice-to-text reconversion request for the voice recording.
[0065] In response to the user's request for speech-to-text reconversion, the processor can output a second speech recording generated by speech-to-text reconversion based on a memo 814 about the speech data created at "01:07" after the speech recording (i.e., a memo about the speech data associated with the "01:07" time point in the speech recording). For example, keywords "demo," "web," and "add" can be extracted as keywords from the memo 814 about the speech data created by the user at "01:07" in the first operation 810. The extracted keywords "demo," "web," and "add" and a specific section of speech associated with the memo 814 can be input into a speech-to-text transcription model to apply speech recognition weights to the keywords "demo," "web," and "add." As a result, speech converted to "In this version of Nemo, we've exceeded the functionality used on Web" 812 in the first speech recording can be converted to "In this version of Nemo, we've added functionality used on Web" 822 in the second speech recording.
[0066] In contrast, for other sections of the audio data that are not related to the note 814, the keywords "demo," "web," and "add" extracted from the note 814 are not weighted. For example, a speech converted into "Please share the Nemo site discussed last month" in a first audio recording is not reconverted to "Please share the demo site discussed last month" in a second audio recording, but is instead converted directly into "Please share the Nemo site discussed last month" 824. In other words, when a note 814 related to audio data is created in relation to a specific section of the audio data, the keywords extracted from the note 814 related to the audio data are recognized as having a higher priority than other keywords only for that specific section, and are recognized in the same way as the existing priority for other sections. Therefore, as shown in second operation 820, a second audio recording can be generated that includes "In this demo, we added a feature for use on the web" 822 and "Please share the Nemo site discussed last month" 824, and the processor can output the generated second audio recording.
[0067] 9 illustrates an example of outputting an audio recording generated based on one or more participant information 918 and / or a topic 920 related to the audio data received after audio recording, according to one embodiment of the present disclosure. In one embodiment, the information associated with the audio data may include information about one or more participants associated with the audio included in the audio data. Additionally or alternatively, the information associated with the audio data may include a topic related to the audio data. In this case, a speech recognition weighting in a speech-to-text transcription model may be applied to one or more keywords extracted from the one or more participant information 918 and / or the topic 920 related to the audio data.
[0068] As shown in a first operation 910, a processor (e.g., at least one processor in a user terminal) can output a first audio recording generated by speech-to-text conversion of the audio recording. The first audio recording can include text in which the audio data is mistranslated, such as "Please share the Nemo site discussed last month" 912 and "This time, Nemo has exceeded the Web's usage limit" 914. After recording the audio, a user can create / input (e.g., newly enter, add, modify, etc.) information about one or more participants 918 and / or a title 920 related to the audio data. For example, a user can input participant information for the audio recording by selecting the "Add participant" icon 916 and selecting (or entering) information (e.g., name, job, age, job rank, location, etc.) for the participants "user1," "user2," and "user3" to be added.
[0069] The user may then request a speech-to-text reconversion of the audio recording by selecting (e.g., clicking) a "Reconvert" icon 922. In response to the user's speech-to-text reconversion request, the processor may output a second audio recording generated by speech-to-text reconversion based on information 918 about one or more participants entered after the audio recording and / or a topic 920 about the audio data.
[0070] For speech-to-text reconversion, one or more keywords can be extracted from the title 920 of the audio data created / input after audio recording and / or information about one or more participants 918. For example, if "user2" input as participant information by the user in the first operation 910 corresponds to a person who will be working on a demo site, "demo" and "site" can be extracted as keywords based on the information about "user2." Additionally, the keywords "web," "function," and "add" can be extracted from "Web Feature Addition Conference" 920, which is the title of the audio recording input by the user in the first operation 910.
[0071] The extracted keywords "demo," "site," "web," "feature," and "add" and at least some of the speech information are input into a speech-to-text transcription model, whereby speech recognition weights can be applied to the keywords "demo," "site," "web," "feature," and "add." As a result, speech converted into "Please share the Nemo site discussed last month" 912 in the first speech recording can be converted into "Please share the demo site discussed last month" 932 in the second speech recording. Similarly, speech converted into "This time, Nemo has exceeded the functionality used in Webe" 914 in the first speech recording can be converted into "This time, we added functionality to be used in the web" 934 in the second speech recording. Thus, as shown in second operation 930, a second speech recording can be generated that includes "Please share the demo site discussed last month" 932 and "This time, we added functionality to be used in the web" 934, and the processor can output the generated second speech recording.
[0072] 10 is a flowchart illustrating a process for reconverting voice data and / or editing a voice recording according to one embodiment of the present disclosure. In one embodiment, when speech-to-text conversion of a voice recording is completed and a first voice recording is generated / output (S1010), a processor (e.g., at least one processor of a user terminal) may output a message guiding the user to make a note of the voice data and / or reconvert the voice data to text (S1020). For example, the processor may output a message guiding the user to make a note of the voice data and / or reconvert the voice data to text, such as "Please make a note of the voice data and reconvert it," or "Creating a note of the voice data and reconverting it will improve the recognition rate."
[0073] The processor may then receive a user request for reconversion of the voice recording (S1022). In one embodiment, if there are notes created for the voice recording, the processor may, in response to the received user request for reconversion, output a first pop-up message (e.g., a reconversion confirmation pop-up) regarding whether or not to reconvert the voice recording to text, confirming whether or not to reconvert the voice recording to text (S1024). For example, the processor may output the first pop-up message including "Do you want to reconvert the voice recording?". The processor may then output a second voice recording generated by the voice-to-text reconversion and / or a second pop-up message indicating completion of the reconversion based on the user's input to the first pop-up message (S1026). For example, the processor may output a second voice recording generated by the voice-to-text reconversion and / or a second pop-up message including "Reconversion of the voice recording is complete" based on the user's affirmative input to the first pop-up message (i.e., a user input indicating a request for reconversion).
[0074] On the other hand, if no memo has been created for the voice recording, the processor may output a third pop-up message (e.g., a memo creation prompting pop-up) to prompt the user to create a memo in response to the received user's reconversion request (S1028). For example, the processor may output a third pop-up message including "Please create a memo and reconvert."
[0075] In another embodiment, when speech-to-text conversion of a voice recording is completed and a voice recording is generated / output (S1010), the processor may receive a request to edit the voice recording (S1030). If reconversion of the voice recording is not performed (i.e., if the generated / output voice recording corresponds to the first voice recording), the processor may output a fourth pop-up message regarding whether or not editing is permitted before reconversion in response to the received editing request (S1032). For example, the processor may output a fourth pop-up message including, "Please make a note and edit after reconversion" in response to the received editing request. Thereafter, based on the user's response to the fourth pop-up message indicating the editing request, the processor may switch the voice recording to an editing mode to provide the user with an interface for editing the voice recording.
[0076] On the other hand, if reconversion of the voice recording has already been performed (i.e., if the generated / output voice recording corresponds to the second voice recording), the processor can immediately switch the voice recording to an edit mode in response to the received edit request, thereby providing the user with an interface for editing the voice recording (S1034).The user can correct / edit at least some of the text that has been incorrectly converted among the multiple texts included in the voice recording through the interface for editing the voice recording.
[0077] 11 is a diagram illustrating an example of an artificial neural network model 1100 according to an embodiment of the present disclosure. The artificial neural network model 1100 may be, as an example of a machine learning model, a statistical learning algorithm implemented based on the structure of a biological neural network in machine learning technology and cognitive science, or a structure for executing such an algorithm.
[0078] According to one embodiment, the artificial neural network model 1100 may represent a machine learning model with problem-solving capabilities, in which nodes, which are artificial neurons forming a network through synaptic connections like a biological neural network, repeatedly adjust synaptic weights to learn to reduce the error between a correct output corresponding to a specific input and an inferred output. For example, the artificial neural network model 1100 may include any probability model, neural network model, etc. used in artificial intelligence learning methods such as machine learning and deep learning.
[0079] According to one embodiment, the artificial neural network model 1100 may include an artificial neural network model configured to receive information about at least a portion of speech included in the voice data and information related to the voice data, and to output text corresponding to at least a portion of the speech. Here, the information related to the voice data may include notes about the voice data, information about one or more participants associated with the speech included in the voice data, a title for the voice data, one or more keywords extracted from the information related to the voice data, one or more keywords extracted from text included in the first voice recording, correction information for at least a portion of the text included in the first voice recording, etc. Additionally or alternatively, the artificial neural network model 1100 may include an artificial neural network model configured to receive information about at least a portion of speech included in the voice data and one or more keywords, and to output text corresponding to at least a portion of the speech by applying a speech recognition weighting such that one or more keywords are recognized as having a higher priority than other keywords.
[0080] The artificial neural network model 1100 is implemented as a multilayer perceptron (MLP) composed of multiple nodes and connections between them. The artificial neural network model 1100 according to this embodiment can be implemented using one of various artificial neural network model structures, including an MLP. As shown in FIG. 11 , the artificial neural network model 1100 includes an input layer 1120 that receives an input signal or data 1110 from the outside, an output layer 1140 that outputs an output signal or data 1150 corresponding to the input data, and n (where n is a positive integer) hidden layers 1130_1 through 1130_n positioned between the input layer 1120 and the output layer 1140. The hidden layers 1130_1 through 1130_n receive signals from the hidden layers 1130_1 through 1130_n and output the signals to the outside.
[0081] The learning method of the artificial neural network model 1100 includes a supervised learning method in which the model learns to optimize problem solving by inputting a teacher signal (correct answer), and an unsupervised learning method in which a teacher signal is not required. In one embodiment, the information processing system can perform supervised and / or unsupervised learning on the artificial neural network model 1100 so that the model outputs text (or text information) corresponding to at least a portion of speech (or information related to speech) included in speech data. For example, the information processing system can perform supervised and / or unsupervised learning on the artificial neural network model 1100 by inputting a reference speech and reference information related to the reference speech, so that the model outputs text corresponding to the reference speech.
[0082] The artificial neural network model 1100 thus trained can be stored in a memory (not shown) of the information processing system and can output text (or text information) corresponding to at least some of the speech (or information related to the speech) included in the speech data in response to at least some of the speech (or information related to the speech) included in the speech data received from the communication module and / or memory and / or information associated with the speech data. Additionally or alternatively, the artificial neural network model 1100 can output a speech recording including text (or text information) corresponding to at least some of the speech (or information related to the speech) included in the speech data.
[0083] According to one embodiment, the input variables of the machine learning model, i.e., artificial neural network model 1100, that performs speech-to-text transcription, may be at least a portion of speech (or information about speech) contained in the audio data. For example, the input variables input to the input layer 1120 of the artificial neural network model 1100 may be vector 1110, with at least a portion of speech contained in the audio data organized as a single vector data element. In response to at least a portion of the speech input contained in the audio data, the output variable output from the output layer 1140 of the artificial neural network model 1100 may be vector 1150 that indicates or characterizes text (or text information) corresponding to at least a portion of the speech (or information about the speech). Additionally or alternatively, the output layer 1140 of the artificial neural network model 1100 may be configured to output a vector that indicates or characterizes a speech recording that includes text (or text information) corresponding to at least a portion of the speech (or information about the speech). In the present disclosure, the output variables of the artificial neural network model 1100 are not limited to the types described above, but may include any information / data indicative of text (or text information) and / or audio recordings corresponding to at least a portion of the audio (or information related to the audio).
[0084] Additionally, the output layer 1140 of the artificial neural network model 1100 can be configured to output a vector indicating the confidence and / or accuracy of the output speech-to-text (or re-conversion) result.
[0085] In this way, the input layer 1120 and output layer 1140 of the artificial neural network model 1100 are matched with a plurality of input variables and a plurality of corresponding output variables, and the synaptic values between the nodes included in the input layer 1120, hidden layers 1130_1 through 1130_n, and output layer 1140 are adjusted, thereby learning to extract a correct output corresponding to a specific input. Through this learning process, the hidden characteristics of the input variables of the artificial neural network model 1100 can be understood, and the synaptic values (or weights) between the nodes of the artificial neural network model 1100 can be adjusted to reduce the error between the output variables calculated based on the input variables and the target output. An information processing system and / or user terminal can input at least some information related to speech and information related to speech data to the trained artificial neural network model 1100, and generate and / or output a speech recording for the speech data using the output text information.
[0086] The above-described method may be provided as a computer program stored on a computer-readable recording medium for execution by a computer. The medium may continuously store the computer-executable program or temporarily store it for execution or download. The medium may also be various recording or storage means in the form of a single piece of hardware or multiple pieces of hardware combined together. The medium is not limited to media directly connected to a computer system but may also be distributed over a network. Examples of media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; ROMs, RAMs, and flash memories, which are configured to store program instructions. Other examples of media include recording or storage media managed by app stores that distribute applications and other sites or servers that provide or distribute various software.
[0087] The methods, operations, or techniques of the present disclosure can be implemented by a variety of means. For example, such techniques can be embodied in hardware, firmware, software, or a combination thereof. Those skilled in the art will appreciate that the various illustrative logical blocks, modules, circuits, and algorithm steps described in this disclosure can be embodied in electronic hardware, computer software, or a combination of both. To clearly illustrate this interchange between hardware and software, the various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is embodied as hardware or software will vary depending on the particular application and design requirements imposed on the overall system. Those skilled in the art may implement the described functionality in various ways for each particular application, but such implementations should not be interpreted as departing from the scope of the present disclosure.
[0088] In a hardware implementation, the processing units utilized to perform the techniques may be embodied in one or more ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, electronic devices, other electronic units designed to perform the functions described in this disclosure, computers, or any combination thereof.
[0089] Accordingly, the various exemplary logic blocks, modules, and circuits described in this disclosure may be embodied or performed by a general-purpose processor, a DSP, an ASIC, an FPGA or other programmable logic device, discrete gate and transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general-purpose processor may be a microprocessor, but alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be embodied by a combination of computing devices, such as a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other configuration.
[0090] In a firmware and / or software implementation, the techniques may be embodied with instructions stored on a computer-readable medium such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable PROM (EEPROM), flash memory, compact disc (CD), magnetic or optical data storage device, etc. The instructions are executable by one or more processors, enabling the processors to perform certain aspects of the functions described in this disclosure.
[0091] Although the above-described embodiments are described as utilizing aspects of the presently disclosed subject matter on one or more stand-alone computer systems, the present disclosure is not limited thereto and may be implemented in any computing environment, such as a network or distributed computing environment. Furthermore, aspects of the subject matter in the present disclosure may be implemented on multiple processing chips or devices, and storage may be similarly affected across multiple devices. Such devices may include PCs, network servers, and handheld devices.
[0092] Although the present disclosure has been described herein with reference to some examples, various modifications and alterations that would be understood by those of ordinary skill in the art to which the present disclosure pertains can be made without departing from the present disclosure, and such modifications and alterations should be understood to fall within the scope of the claims appended hereto.
Claims
1. 1. A method for providing an audio recording, performed by at least one computing device, comprising: obtaining information associated with the audio data after recording the audio; receiving a request for speech-to-text conversion of the voice data; and outputting a voice recording generated based on at least a portion of the voice included in the voice data and information related to the voice data obtained after the voice recording in response to the voice-to-text conversion request; the information associated with the audio data includes notes about the audio data made at least one of before, during, or after the audio recording; The memo on the audio data is associated with audio included in a specific section of the audio data; The method of providing an audio recording, wherein the audio recording includes text information generated by speech-to-text conversion of the audio contained in the specified section using notes created in association with the specified section.
2. 10. The method of claim 1, wherein the information associated with the audio data includes information about one or more participants associated with audio contained in the audio data.
3. 10. The method of claim 1, wherein the information associated with the audio data includes a subject related to the audio data.
4. the audio recording includes text information output by inputting information about the at least some of the audio and information associated with the audio data into a speech-to-text transcription model; 2. The method of claim 1, wherein the speech-to-text transcription model is trained to input a reference speech and reference information associated with the reference speech to output text corresponding to the reference speech.
5. the information associated with the voice data includes one or more keywords extracted from the information associated with the voice data; The method of claim 4 , wherein the reference information associated with the reference audio comprises one or more reference keywords associated with the reference audio.
6. 6. The method of claim 5, wherein speech recognition weightings are applied to one or more keywords input to the speech-to-text transcription model such that the one or more keywords are recognized as having a higher priority than keywords different from the one or more keywords.
7. before receiving information associated with the audio data, outputting a first audio recording generated by speech-to-text conversion of at least a portion of the audio included in the audio data; receiving a request for speech-to-text conversion of the voice data includes receiving a request for speech-to-text reconversion of the voice data; 2. The method of claim 1, wherein the outputting step includes a step of outputting, in response to the voice-to-text reconversion request, a second voice recording generated based on at least a portion of the voice included in the voice data and information related to the voice data obtained after the voice recording as the voice recording.
8. 8. The method of claim 7, wherein the information associated with the audio data includes one or more keywords extracted from text included in the first audio recording.
9. the information associated with the audio data includes text correction information for at least a portion of the text included in the first audio recording; 8. The method of claim 7, wherein the correction information applies speech recognition weightings in a speech-to-text transcription model to the corrected text in the first speech recording.
10. A non-transitory computer-readable recording medium having recorded thereon instructions for executing the method of claim 1 on a computer.
Citation Information
Patent Citations
Caption generator, retrieval device, method for integrating document processing and speech processing together, and program
JP2006178087A
In-vehicle device
JP2009092975A
Method and system for providing audio recordings generated based on post audio recording information - Patents.com
JP2024514260A
Utterance presentation device, utterance presentation method, and program
WO2016163028A1