Method and System for Providing an Audio Recording Generated Based on Information after Audio Recording
The method and system enhance speech-to-text conversion accuracy by generating audio recordings based on post-recording information, allowing for accurate audio perception and text conversion.
Patent Information
- Application Number
- JP2023561860
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-04-07
- Filing Date
- 2022-03-22
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-03-22
AI Technical Summary
Existing speech-to-text conversion technologies face accuracy issues when converting audio recordings, leading to inaccurately converted text.
A method and system that receives information related to audio data after recording, applies a voice-to-text conversion request, and generates an audio recording based on the audio data and additional information to improve accuracy.
Enables users to perceive audio content both auditorily and visually, with improved text conversion accuracy by incorporating additional information post-recording.
Smart Images

Figure 0007708879000001 
Figure 0007708879000002 
Figure 0007708879000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to a method and system for providing an audio recording generated based on information after audio recording. Specifically, the present disclosure relates to a method and system for receiving information related to audio data after audio recording and providing an audio recording generated based on at least a part of the audio included in the audio data and the information related to the audio data received after audio recording.
Background Art
[0002] Recently, with the development and popularization of mobile electronic devices such as smartphones and tablet PCs, users can easily generate / store recordings of voice conversations, texts, images, etc. through mobile electronic devices in their daily lives. For example, users can use a note application, an audio recording application, etc. to record and / or record videos of meetings, conferences, classes, interviews, etc. In addition, while recording and / or recording videos through a mobile electronic device, users can input a memo regarding the content being recorded and / or recorded by creating a text regarding the content being recorded and / or recorded.
[0003] In addition, with the development of speech-to-text conversion technology (i.e., speech recognition technology), the content included in an audio recording generated by recording and / or recording videos can be converted into text and provided to users. At this time, users can recognize the content of the audio recording through the converted text without directly listening to the audio recording. However, when performing speech-to-text conversion using only the information of the audio recording, there is a risk of reducing the accuracy of speech recognition. That is, there is a risk of providing users with text in which the audio included in the audio recording is inaccurately converted.
Summary of the Invention
Problems to be Solved by the Invention
[0004] The present disclosure provides a method for providing an audio recording for solving the above problems, a computer program stored in a recording medium, and an apparatus (system).
Means for Solving the Problems
[0005] The present disclosure can be embodied in various ways including a method, an apparatus (system), or a computer program stored in a readable storage medium.
[0006] According to an embodiment of the present disclosure, a method for providing an audio recording performed by at least one computing device includes receiving information related to audio data after audio recording, receiving a voice-to-text conversion request for the audio data, and outputting at least a part of the audio included in the audio data and an audio recording generated based on the information related to the audio data received after audio recording in response to the voice-to-text conversion request.
[0007] A computer-readable non-transitory recording medium recording instructions for executing, by a computer, a method for providing an audio recording according to an embodiment of the present disclosure is provided.
[0008] An information processing system according to an embodiment of the present disclosure includes a communication module, a memory, and at least one processor coupled to the memory and configured to execute at least one computer-readable program included in the memory, and the at least one program receives information related to audio data after audio recording, receives a voice-to-text conversion request for the audio data, and includes instructions for generating an audio recording based on at least a part of the audio included in the audio data and the information related to the audio data received after audio recording in response to the voice-to-text conversion request.
Advantages of the Invention
[0009] In some embodiments of the present disclosure, a user can receive a voice recording corresponding to the content of the voice data included in the voice recording, so that the content of the voice recording can be perceived both auditorily and visually. Also, after the voice recording or voice recognition, by converting the voice recording based on the information related to the voice data input by the user, text in which the voice data is more accurately converted can be provided.
[0010] In some embodiments of the present disclosure, when it is difficult for a user to create a memo via a mobile device or a PC during recording, if a memo related to the recording is created after the recording is finished, by extracting the keywords included in the created memo and re-recognizing the voice of the recording file, the recognition rate of the voice can be improved.
[0011] The effects of the present disclosure are not limited to this, and other effects not mentioned should be clearly understood by those with ordinary knowledge in the technical field to which the present disclosure belongs from the description of the claims (hereinafter referred to as "persons skilled in the art").
Brief Description of the Drawings
[0012] Embodiments of the present disclosure are described based on the following attached drawings. Here, similar reference numerals indicate similar elements, but are not limited thereto.
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
DETAILED DESCRIPTION OF THE INVENTION
[0013] Hereinafter, specific contents for implementing the present disclosure will be described in detail based on the accompanying drawings. However, in the following description, specific descriptions regarding well-known functions and configurations may be omitted if there is a risk of unnecessarily obscuring the gist of the present disclosure.
[0014] In the accompanying drawings, the same reference numerals are assigned to the same or corresponding components. Also, in the following description of the embodiments, duplicate descriptions regarding the same or corresponding components may be omitted. However, even if the description regarding a component is omitted, such a component should not be construed as not being included in a certain embodiment.
[0015] The advantages and features of the disclosed embodiments, and the methods for achieving them, will become clear by referring to the embodiments described below with reference to the accompanying drawings. However, the present disclosure is not limited to the embodiments disclosed below and can be embodied in various different forms. However, these embodiments are only provided to make the present disclosure complete and to enable those skilled in the art to accurately recognize the category of the invention.
[0016] The terms used in this specification will be briefly explained, and the embodiments of the disclosure will be specifically described. The terms used in this specification are selected as general terms that are currently widely used as much as possible while considering their functions in the present disclosure. However, this may change due to the intentions or precedents of those skilled in the relevant field, the emergence of new technologies, etc. In addition, in certain cases, there may be terms arbitrarily selected by the applicant, and their meanings will be described in detail in the part of the description of the invention. Therefore, the terms used in the present disclosure should be defined based on the meanings of the terms and the overall content of the present disclosure, rather than simply by the names of the terms.
[0017] In this specification, unless specifically specified in the context, singular expressions include plural expressions, and plural expressions can include singular expressions. Throughout the specification, when a certain part states that a certain component "includes", this means that, unless otherwise stated to the contrary, it does not exclude other components and can further include other components.
[0018] Also, the terms "module" or "section" used in the specification mean software or hardware components, and the "module" or "section" performs a certain role. However, the "module" or "section" is not meant to be limited to software or hardware. The "module" or "section" may be configured to be on an addressable storage medium or may be configured to cause one or more processors to execute. Thus, by way of example, the "module" or "section" can include at least one of software components, object-oriented software components, class components, task components such as these, and processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays or variables. Components and "modules" or "sections" can be combined with a smaller number of further components and "modules" or "sections" that provide internal functions, or can be further separated into additional components and "modules" or "sections".
[0019] According to an embodiment of the present disclosure, a "module" or "unit" can be implemented by a processor and a memory. The "processor" should be broadly interpreted to include a general-purpose processor, a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a controller, a microcontroller, a state machine, and the like. In some environments, the "processor" can also refer to an application-specific integrated circuit (ASIC), a programmable logic device (PLD), a field-programmable gate array (FPGA), etc. The "processor" can also refer to a combination of processing devices such as, for example, a combination of a DSP and a microprocessor, a combination of multiple microprocessors, a combination of one or more microprocessors coupled with a DSP core, or any other such configuration. Also, the "memory" should be broadly interpreted to include any electronic component capable of storing electronic information. The "memory" can also refer to various types of processor-readable media such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage devices, registers, and the like. When a processor can read information from a memory or record the information read into the memory, the memory is said to be in electronic communication with the processor. Memory integrated with a processor is in electronic communication with the processor.
[0020] In the present disclosure, "voice data" can include data generated / saved by voice recording. Here, voice recording can refer to voice data, and voice data can refer to voice recording. In one embodiment, voice data can include one or more voices. Here, one or more voices can refer to data corresponding to at least one section of a plurality of sections of voice data. Alternatively or additionally, one or more voices can refer to each voice of the speakers included in the voice data, the voice data, the utterance, and / or the utterance data. In the present disclosure, information regarding voice data and / or voice can include the voice itself and / or data indicating the voice (for example, vector data).
[0021] In the present disclosure, "voice record" can refer to a record generated by converting the utterance content included in the voice recording into text. Here, the first voice record can refer to a voice record generated without reflecting information related to the voice data received after the voice recording, and the second voice record can refer to a voice record generated by reflecting information related to the voice data received after the voice recording, but is not limited thereto.
[0022] FIG. 1 is a diagram showing an example of providing a voice record 122 generated based on information related to voice data created after voice recording according to an embodiment of the present disclosure. The screen shown in FIG. 1 shows an example in which a user executes a recording application such as a voice recording application, a memo application, and / or a note application via a user terminal (for example, a smartphone, a tablet PC, a desktop, etc.) and receives the provision of the voice record 122 regarding the voice data. In one embodiment, the user can receive the provision of text corresponding to the voice data included in the voice recording by such a recording application.
[0023] A user terminal (e.g., at least one processor of the user terminal, etc.) can receive information related to voice data after voice recording in order to provide voice recording. For example, the user terminal can receive information related to voice data input by the user via an input device (e.g., keyboard, mouse, microphone, etc.) after voice recording. Additionally or alternatively, the user terminal can receive from the storage device information related to the voice data stored in the storage device after voice recording. Here, the information related to the voice data can refer to any information that can be included in, indicate, or characterize the voice data. For example, it can include memos 112, 114 related to the voice data, a title 116 related to the voice data, information 118 related to one or more participants related to the voice included in the voice data, etc. Additionally or alternatively, the information related to the voice data can include one or more keywords extracted from the information related to the voice data. Additionally or alternatively, the information related to the voice data can include one or more keywords extracted from the text included in the existing voice recording.
[0024] The user terminal can output a voice recording 122 generated based on at least some of the voices included in the voice data and the information related to the voice data received after voice recording in response to a voice-text conversion request for the voice data. For example, the user terminal receives a user input that selects an icon 120 indicating a conversion request (or re-conversion request) for the voice recording, and in response, can display at least one text included in the voice recording 122 on the display. Here, the voice recording 122 can be generated by at least one processor of the information processing system and / or at least one processor of the user terminal.
[0025] In one embodiment, the voice recording 122 can include text information output by inputting information regarding at least a portion of the voice and information associated with the voice data into a speech-to-text transcription model. Here, the speech-to-text transcription model can include a model that is learned to output text corresponding to a reference voice by inputting the reference voice and reference information associated with the reference voice. For example, the reference information associated with the reference voice can include one or more reference keywords associated with the reference voice. That is, the speech-to-text transcription model can include a model that is learned to output text corresponding to the reference voice by inputting the reference voice and one or more reference keywords associated with the reference voice.
[0026] In one embodiment, by applying a voice recognition weight value to one or more keywords input into the speech-to-text transcription model, the one or more keywords are recognized as having a higher priority than keywords different from the one or more keywords. For example, by applying a voice recognition weight value to the keyword "demo" input into the speech-to-text transcription model, the keyword "demo" is recognized as having a higher priority than the keyword "nemo". Thus, the speech-to-text transcription model can recognize at least a portion of the voice included in the input voice data as "demo" rather than "nemo", convert it into text, and output it, and the user terminal can output a voice recording including "demo" rather than "nemo".
[0027] As shown in the figure, the user terminal can display information related to the voice data on the display. For example, the information related to the voice data can include the title of the voice data ("Demo Site Meeting") 116, information about one or more participants related to the voice included in the voice data ("user1, user2, user3") 118, a memo 112 created during the voice recording, and a memo 114 created after the voice recording. Additionally or alternatively, the information related to the voice data can include a voice recording (e.g., the first voice recording) that does not reflect the information received after the voice recording. Such a voice recording that does not reflect the information received after the voice recording can also be displayed on the display. Then, in response to a user's touch input on the "Re-conversion" icon 120 indicating a request for conversion (or re-conversion) of the voice recording, the user terminal can display on the display a voice recording (e.g., the second voice recording) 122 generated based on the information related to the voice data received after the voice recording. Also, the user terminal can output a pop-up message 124 including "The re-conversion of the voice recording is completed."
[0028] According to the embodiments described above, the user can receive the provision of a voice recording corresponding to the content of the voice data included in the voice recording, whereby the content of the voice recording can be perceived both aurally and visually. Also, after the voice recording or voice recognition, by converting the voice recording based on the information related to the voice data input by the user, a text in which the voice data is more accurately converted can be provided.
[0029] FIG. 2 is a schematic diagram showing a configuration in which an information processing system 230 is connected to be communicable with a plurality of user terminals 210_1, 210_2, and 210_3 in order to provide a voice recording providing service according to an embodiment of the present disclosure. The information processing system 230 can include a system capable of providing a voice recording providing service, a system capable of providing recording services such as voice recording, memo, note, etc., and / or a system capable of providing a voice-to-text conversion service. In one embodiment, the information processing system 230 includes computer-executable programs (e.g., downloadable applications) related to the voice recording providing service, recording service, and / or voice-to-text conversion service, one or more server devices and / or databases capable of storing, providing, and executing data, and one or more distributed computing devices and / or distributed databases of a cloud computing service infrastructure. For example, the information processing system 230 can include another system (e.g., a server) for the voice recording providing service, recording service, and / or voice-to-text conversion service, etc.
[0030] The voice recording providing service, recording service, voice-to-text conversion service, etc. provided by the information processing system 230 are provided to the user via voice recording applications, memo applications, note applications, voice-to-text conversion applications, etc. installed on each of the plurality of user terminals 210_1, 210_2, and 210_3. For example, the information processing system 230 can provide information corresponding to a voice-to-text conversion request received from the user terminals 210_1, 210_2, and 210_3 or perform corresponding processing via a voice recording application or the like.
[0031] A plurality of user terminals 210_1, 210_2, 210_3 can communicate with an information processing system 230 via a network 220. The network 220 can be configured to enable communication between the plurality of user terminals 210_1, 210_2, 210_3 and the information processing system 230. The network 220 can be composed of a wired network such as, for example, Ethernet (registered trademark), PLC (Power Line Communication), a telephone line communication device, and RS-serial communication, a mobile communication network, a WLAN (Wireless LAN), Wi-Fi (registered trademark), Bluetooth (registered trademark), and ZigBee (registered trademark) and the like, or a combination thereof, depending on the installation environment. The communication method is not limited, and it can be not only a communication method that utilizes a communication network (for example, a mobile communication network, a wired Internet, a wireless Internet, a broadcast network, a satellite network, etc.) including the network 220, but also short-range wireless communication between the user terminals 210_1, 210_2, 210_3 may be included.
[0032] In FIG. 2, the mobile phone terminal 210_1, the tablet terminal 210_2, and the PC terminal 210_3 are shown as examples of user terminals. However, the present invention is not limited thereto, and the user terminals 210_1, 210_2, and 210_3 can be any computing device capable of wired and / or wireless communication and on which an audio recording application or the like is installed and can be executed. For example, the user terminal can include a smartphone, a mobile phone, a navigation device, a desktop computer, a laptop computer, a digital broadcast terminal, a PDA (Personal Digital Assistants), a PMP (Portable Multimedia Player), a tablet PC, a game console, a wearable device, an IoT (internet of things) device, a VR (virtual reality) device, an AR (augmented reality) device, and the like. Further, in FIG. 2, three user terminals 210_1, 210_2, and 210_3 are shown as communicating with the information processing system 230 via the network 220. However, the present invention is not limited thereto, and a different number of user terminals can be configured to communicate with the information processing system 230 via the network 220.
[0033] In one embodiment, the information processing system 230 can receive a voice-to-text conversion request for voice data from the user terminals 210_1, 210_2, and 210_3. Further, the information processing system 230 can receive at least one of voice data or information related to the voice data from the user terminals 210_1, 210_2, and 210_3. In response to the voice-to-text conversion request, the information processing system 230 can generate a voice record based on at least a part of the voice included in the voice data and information related to the voice data received after the voice recording, and provide the generated voice record to the user terminals 210_1, 210_2, and 210_3. Alternatively, the user terminals 210_1, 210_2, and 210_3 can generate a voice record based on at least a part of the voice included in the voice data and information related to the voice data received after the voice recording.
[0034] FIG. 3 is a block diagram showing the internal configurations of user terminal 210 and information processing system 230 according to an embodiment of the present disclosure. User terminal 210 can execute voice recording applications, memo applications, note applications, voice-text conversion applications, etc., and can refer to any computing device capable of wired / wireless communication. For example, it can include mobile phone terminal 210_1, tablet terminal 210_2, and laptop computer terminal 210_3 in FIG. 2. As shown in the figure, user terminal 210 can include memory 312, processor 314, communication module 316, and input / output interface 318. Similarly, information processing system 230 can include memory 332, processor 334, communication module 336, and input / output interface 338. As shown in FIG. 3, user terminal 210 and information processing system 230 can be configured such that information and / or data can be communicated via network 220 using respective communication modules 316, 336. Also, input / output device 320 can be configured to input information and / or data to user terminal 210 or output information and / or data generated from user terminal 210 via input / output interface 318.
[0035] The memories 312 and 332 can include any non-transitory computer-readable recording medium. According to one embodiment, the memories 312 and 332 can include permanent mass storage devices such as RAM (random access memory), ROM (read only memory), disk drives, SSDs (solid state drives), and flash memories. As another example, permanent mass storage devices such as ROMs, SSDs, flash memories, and disk drives can be included in the user terminal 210 or the information processing system 230 as another separate permanent storage device distinct from the memories. Also, the memories 312 and 332 can store an operation system and at least one program code (for example, codes for a voice recording application, a memo application, a note application, a voice-text conversion application, etc.).
[0036] Such software components can be loaded from a computer-readable recording medium different from the memories 312 and 332. Such another computer-readable recording medium can include a recording medium directly connectable to such user terminals 210 and information processing systems 230, and can include, for example, computer-readable recording media such as floppy drives, disks, tapes, DVD / CD-ROM drives, and memory cards. As another example, software components, etc. can also be loaded into the memories 312 and 332 via the communication modules 316 and 336 instead of a computer-readable recording medium. For example, at least one program can be a computer program (for example, a voice recording application, a memo application, a note application, a voice-text conversion application, etc.) installed by a file provided by a file distribution system that distributes developer or application installation files via the network 220 into the memories 312 and 332.
[0037] The processors 314 and 334 can be configured to process instructions of a computer program by performing basic arithmetic, logic, and input / output operations. The instructions can be provided to the processors 314 and 334 by the memories 312 and 332 or the communication modules 316 and 336. For example, the processors 314 and 334 can be configured to execute instructions received by program code stored in a recording device such as the memories 312 and 332.
[0038] The communication modules 316 and 336 can provide configurations and functions for the user terminal 210 and the information processing system 230 to communicate with each other via the network 220, and the user terminal 210 and / or the information processing system 230 can provide configurations and functions for communicating with other user terminals or other systems (e.g., another cloud system, etc.). As an example, a request or data (e.g., a speech-text conversion request for speech data, etc.) generated by the processor 314 of the user terminal 210 by program code stored in a recording device such as the memory 312 can be transmitted to the information processing system 230 via the network 220 under the control of the communication module 316. Conversely, a control signal or instruction provided under the control of the processor 334 of the information processing system 230 can be received by the user terminal 210 via the communication module 336 of the information processing system 230, the network 220, and the communication module 316 of the user terminal 210. For example, the user terminal 210 can receive, via the communication module 316 from the information processing system 230, a voice recording generated based on at least a part of the voice included in the voice data and information related to the voice data received after voice recording.
[0039] The input / output interface 318 can be means for interfacing with an input / output device 320. As an example, the input device can include devices such as a camera including an audio sensor and / or an image sensor, a keyboard, a microphone, a mouse, etc., and the output device can include devices such as a display, a speaker, a haptic feedback device, etc. As another example, the input / output interface 318 can be means for interfacing with a device in which a configuration or function for performing input and output is integrated into one, such as a touch screen. In FIG. 3, although the input / output device 320 is shown as not being included in the user terminal 210, it is not limited thereto and can also be configured integrally with the user terminal 210. Also, the input / output interface 338 of the information processing system 230 can be means for interfacing with the information processing system 230 or with a device (not shown) for input and output that the information processing system 230 can include. In FIG. 3, the input / output interfaces 318, 338 are shown as elements configured separately from the processors 314, 334, but it is not limited thereto and the input / output interfaces 318, 338 can also be configured to be included in the processors 314, 334.
[0040] The user terminal 210 and the information processing system 230 can include more components than those shown in FIG. 3. However, it is not necessary to clearly show most of the conventional components. According to one embodiment, the user terminal 210 can be embodied to include at least a part of the input / output device 320 described above. Further, the user terminal 210 can further include other components such as a transceiver, a GPS (Global Positioning System) module, a camera, various sensors, and a database. For example, when the user terminal 210 is a smartphone, it can include components generally possessed by a smartphone. For example, various components such as an acceleration sensor, a gyro sensor, a microphone module, a camera module, various physical buttons, buttons using a touch panel, an input / output port, and a vibrator for vibration can be embodied to be further included in the user terminal 210.
[0041] According to one embodiment, the processor 314 of the user terminal 210 can be configured to operate voice recording applications, memo applications, note applications, voice-text conversion applications, and the like. At this time, the program code related to the application can be loaded into the memory 312 of the user terminal 210. While the application is operating, the processor 314 of the user terminal 210 can receive information and / or data provided from the input / output device 320 and / or receive information and / or data from the information processing system 230 via the communication module 316, process the received information and / or data, and save it in the memory 312. Further, such information and / or data can be provided to the information processing system 230 via the communication module 316.
[0042] While an audio recording application or the like is operating, the processor 314 can receive audio data, text, images, video, etc. input or selected by an input device such as a touch screen, keyboard, audio sensor, and / or a camera including an image sensor, microphone, etc. connected to the input / output interface 318, and can store the received audio data, text, images, and / or video, etc. in the memory 312, or provide them to the information processing system 230 via the communication module 316 and the network 220. In one embodiment, the processor 314 can receive information related to the audio data, a voice-text conversion request for the audio data, etc. from an input device 320 such as a touch screen or a mouse, and can provide the information related to the audio data, the voice-text conversion request for the audio data, etc. to the information processing system 230 via the communication module 316 and the network 220.
[0043] The processor 314 of the user terminal 210 can transfer and output information and / or data to the input / output device 320 via the input / output interface 318. For example, the processor 314 of the user terminal 210 can output the processed information and / or data via an output device 320 such as a display output-capable device (e.g., a touch screen or a display), a voice output-capable device (e.g., a speaker). In one embodiment, the processor 314 can display a voice recording of the audio data on the display of the user terminal 210. Additionally, the processor 314 can output at least a part of the audio included in the audio data via the speaker of the user terminal 210.
[0044] The processor 334 of the information processing system 230 can be configured to manage, process, and / or store information and / or data received from a plurality of user terminals 210 and / or a plurality of external systems. The information and / or data processed by the processor 334 can be provided to the user terminal 210 via the communication module 336 and the network 220. In one embodiment, the processor 334 of the information processing system 230 receives information related to the audio data after audio recording, receives a request for audio-text conversion for the audio data, and in response to the request for audio-text conversion, generates an audio record based on at least a portion of the audio included in the audio data and the information related to the audio data received after audio recording, and the generated audio record can be provided to the user terminal 210 via the communication module 336 and the network 220.
[0045] For example, the processor 334 can utilize a memo related to the audio data created after audio recording (e.g., a memo related to the audio included in a specific section within the audio data) to perform audio-text conversion on the audio included in the specific section, thereby generating an audio record including text information corresponding to the audio included in the specific section. Alternatively or additionally, the processor 334 can input information related to at least a portion of the audio and information related to the audio data into an audio-text transcription model, thereby generating an audio record including text information corresponding to the information related to at least a portion of the audio. Alternatively or additionally, the processor 334 can apply an audio recognition weight value to one or more keywords input into the audio-text transcription model, and recognize the one or more keywords as having a higher priority than keywords different from the one or more keywords. Alternatively or additionally, the processor 334 generates a first audio record by performing audio-text conversion on at least a portion of the audio included in the audio data before receiving the information related to the audio data, receives a request for audio-text re-conversion for the audio data, and in response to the request for audio-text re-conversion, generates a second audio record based on at least a portion of the audio included in the audio data and the information related to the audio data received after audio recording.
[0046] Figure 4 is a flowchart showing a method 400 for providing an audio recording according to an embodiment of the present disclosure. In one embodiment, the method 400 for providing an audio recording can be performed by a processor (e.g., at least one processor of a user terminal and / or an information processing system). As shown in the figure, the method 400 for providing an audio recording can start when the processor receives information related to the audio data after the audio recording (S410). Here, the information related to the audio data received after the audio recording can include a memo regarding the audio data created after the audio recording. Additionally or alternatively, the information related to the audio data can include information regarding one or more participants related to the audio included in the audio data. Additionally or alternatively, the information related to the audio data can include a topic regarding the audio data.
[0047] The processor can receive a voice-to-text conversion request for the audio data (S420). In response to the voice-to-text conversion request, the processor can output an audio recording generated based on at least a part of the audio included in the audio data and the information related to the audio data received after the audio recording (S430).
[0048] In one embodiment, the audio recording can include text information output by inputting information regarding at least a part of the audio and information related to the audio data into a voice-to-text transcription model. Here, the voice-to-text transcription model can be learned to output text corresponding to the reference audio by inputting the reference audio and the reference information related to the reference audio. At this time, the information related to the audio data can include one or more keywords extracted from the information related to the audio data, and the reference information related to the reference audio can include one or more reference keywords related to the reference audio.
[0049] Also, by applying a voice recognition weight value to one or more keywords input to the voice-text transcription model, the one or more keywords can be recognized with a higher priority than keywords different from the one or more keywords. Here, the one or more keywords can correspond to meaningful keywords extracted from information related to the voice data by using an artificial neural network model, a machine learning model, a keyword extraction algorithm, or the like. As an example of the keyword extraction algorithm, keywords frequently used in the current recording, keywords more frequently used in the current recording compared to other documents (recordings), keywords not used in other documents (recordings) and first used in the current recording, etc. are used as meaningful keywords, but the algorithm is not limited to this.
[0050] In one embodiment, a memo regarding voice data created after voice recording is associated with the voice included in a specific section within the voice data. At this time, the voice recording can include text information generated by voice-text conversion for the voice included in the specific section by using the memo created in association with the specific section.
[0051] In one embodiment, before receiving information related to voice data, the processor can output a first voice recording generated by voice-text conversion for at least a part of the voices included in the voice data. Thereafter, after generating the first voice recording, the processor can receive information related to the voice data and receive a voice-text reconversion request for the voice data. In response to the voice-text reconversion request, the processor can output, as a voice recording, a second voice recording generated based on at least a part of the voices included in the voice data and the information related to the voice data received after the voice recording. At this time, the information related to the voice data can include one or more keywords extracted from the text included in the first voice recording. Alternatively or additionally, the information related to the voice data can include information on corrections related to at least a part of the text included in the first voice recording. At this time, the voice recognition weight value in the voice-text transcription model can be applied to the corrected text.
[0052] FIG. 5 is a diagram showing an example in which a memo 516 related to voice data is created after a first voice recording 510 for the voice data according to an embodiment of the present disclosure is output. The user can input information 512, 514 related to the voice recording (i.e., information related to the voice data) before or during the voice recording via the user terminal. For example, before starting the voice recording, the user can input information regarding the title of the voice recording and / or information 514 about the participants in the voice recording. As another example, during the voice recording, the user can input a memo related to the voice recording (i.e., a memo related to the voice data) 512. The processor (e.g., at least one processor of the user terminal) can receive the information 512, 514 related to the voice data input in this way before or during the voice recording.
[0053] The processor can output a first voice recording 510 generated by voice-text conversion for at least some of the voices included in the voice data. In one embodiment, the first voice recording 510 can be generated based on information 512, 514 related to the voice data received before or during the voice recording. To output the first voice recording 510, one or more keywords can be extracted from the information 512, 514 related to the voice data received before or during the voice recording. The first voice recording 510 can include the extracted one or more keywords and the text information output by inputting the voice data into a voice-text transcription model. For example, the first voice recording 510 can be generated by applying a voice recognition weight value to one or more keywords extracted via the voice-text transcription model and recognizing the keyword with the applied weight value as having a higher priority than other keywords.
[0054] The keyword "plan" is extracted from the memo 512 created during the voice recording, and the extracted keyword "plan" and information regarding at least some of the voices included in the voice data can be input into the voice-text transcription model. A voice recognition weight value is applied to the keyword "plan" via the voice-text transcription model, and the keyword "plan" with the applied weight value can be recognized as having a higher priority than other keywords. As a result, a "plan proposal" corresponding to the information regarding at least some of the voices can be output from the voice-text transcription model, and the first voice recording 510 including the "plan proposal" can be generated.
[0055] The user can input information related to the voice recording (i.e., information related to the voice data) 516 via the user terminal even after the voice recording. For example, the user can input a memo related to the voice recording (i.e., a memo related to the voice data) 516 after the voice recording. As another example, the user can input (or add input) information about the participants related to the voice recording after the voice recording. As shown in FIG. 5, after the first voice recording 510 is generated / output, the user can create / input a memo 516 related to the voice data. Therefore, the processor can receive the memo 516 related to the voice data created / input by the user after the first voice recording 510 is generated / output after the voice recording.
[0056] FIG. 5 shows an example in which the processor receives a memo 516 related to the voice data after the first voice recording 510 is generated and / or output, but it is not limited thereto. For example, the processor can receive a memo related to the voice data after the voice recording and before the first voice recording is generated and / or output.
[0057] FIG. 6 is a diagram showing an example of outputting a second voice recording 622 as a re-conversion result for at least a part of the voices included in the voice data according to an embodiment of the present disclosure. In one embodiment, the processor (e.g., at least one processor of the user terminal) can receive information related to the voice data after the voice recording. Here, the information related to the voice data received after the voice recording can include a memo 616 related to the voice data created after the voice recording, a title related to the voice data created after the voice recording, information about one or more participants related to the voices included in the voice data, the first voice recording (or one or more keywords extracted from the text included in the first voice recording) 618, and the like. For example, the processor can receive information related to the voice data after the first voice recording 618 is generated / output.
[0058] The processor can receive a voice-text reconversion request for voice data. In response to the voice-text reconversion request, the processor can output, as a voice recording, a second voice recording 622 generated based on at least a part of the voice included in the voice data and information 616 related to the voice data received after the voice recording. Here, the information related to the voice data can include one or more keywords extracted from the text included in the first voice recording 618. To generate the second voice recording 622, the keyword "demo" is extracted from the text included in the first voice recording 618, and the keywords "web", "add", and "demo" can be extracted from the information 616 related to the voice data received after the voice recording. Then, a second voice recording 622 including text information output by inputting the extracted keywords and information regarding at least a part of the voice included in the voice data into a voice-text transcription model can be generated.
[0059] As shown in the first operation 610, in response to a user's touch input to a "reconversion" icon 612 indicating a voice-text reconversion request, etc., the processor can output a pop-up message ("Do you want to reconvert the voice recording?") 614 regarding the possibility of reconverting the voice recording. Based on the user's response to the output pop-up message 614, the processor can output a second voice recording 622 generated based on at least a part of the voice included in the voice data and information related to the voice data received after the voice recording. Therefore, in the first operation 610, due to inaccurate voice-text conversion, the first voice recording 618 including the text "In this Nemo, the functions used on the web have been exceeded." is displayed on the display, while in the second operation 620, due to accurate voice-text conversion, the second voice recording 622 including the text "In this demo, functions used on the web have been added." is displayed on the display.
[0060] FIG. 7 is a diagram showing an example of outputting a second audio recording generated by correction information regarding at least a part of the text included in the first audio recording according to an embodiment of the present disclosure. In one embodiment, the information associated with the audio data can include correction information regarding at least a part of the text included in the first audio recording. At this time, based on the correction information, an audio recognition weight value in an audio-text transcription model can be applied to the corrected text among the texts included in the first audio recording. That is, the corrected text among the texts included in the first audio recording is extracted as a keyword, and by inputting at least part of the information regarding the audio and the corrected text (i.e., the extracted keyword) into the audio-text transcription model, the audio recognition weight value can be applied to the corrected text (i.e., the extracted keyword). At this time, in audio-text conversion, the corrected text (i.e., the extracted keyword) can be recognized with a higher priority than other keywords.
[0061] As shown in the first operation 710, a processor (e.g., at least one processor of a user terminal) can output a first audio recording generated by audio-text conversion for an audio recording. The first audio recording can include text in which the audio data is mis-converted, such as "Please share the Nemo site discussed last month." 712 and "In this Nemo, the functions used on the web exceeded." 714. The user can correct at least a part of the text included in the first audio recording. For example, the user can select (e.g., click input) an "Edit" icon 716 to correct at least a part of the text included in the first audio recording. In response to this, the processor can provide the user with an interface that can correct at least a part of the text included in the first audio recording by switching to an edit mode. Thereafter, the user can correct "Nemo" to "Demo" in "Please share the Nemo site discussed last month." 712 included in the first audio recording.
[0062] Thereafter, the user can initiate a voice-text reconversion request by selecting the "Reconversion" icon 718. The processor can output a second voice recording generated by voice-text reconversion based on correction information regarding at least a part of the text included in the first voice recording in response to the user's voice-text reconversion request. For example, the text "demo" corrected by the user in the first operation 710 is extracted as a keyword, and the voice recognition weighting value can be applied to the "demo" of the corrected text by inputting information regarding at least a part of the voice and the "demo" of the corrected text into the voice-text transcription model. As a result, the voice converted to "In this Nemo, the functions used on the web have been exceeded." 714 in the first voice recording can be converted to "In this demo, functions used on the web have been added." 722 in the second voice recording. Therefore, as shown in the second operation 720, a second voice recording including "In this demo, functions used on the web have been added." 722 can be generated, and the processor can output the generated second voice recording.
[0063] FIG. 8 is a diagram showing an example of outputting a voice recording generated based on a memo 814 regarding voice data created after voice recording according to an embodiment of the present disclosure. In one embodiment, information related to the voice data received after voice recording can include a memo 814 regarding the voice data created after voice recording. Here, the memo 814 regarding the voice data created after voice recording is associated with the voice included in a specific section within the voice data. At this time, the voice recording can include text information generated by voice-text conversion for the voice included in the specific section using the memo created in relation to the voice included in the specific section.
[0064] As shown in the first operation 810, a processor (e.g., at least one processor of a user terminal) can output a first voice recording including "The function used on the web in this Nemo has exceeded the limit." 812. After the voice recording, the user can create / input a memo regarding a specific section of the voice data and / or a specific section of the first voice recording with respect to the voice data. For example, the user can select a section (e.g., a section between the start time and the end time, a specific time point) from the voice recording (or the voice data) and create / input a memo regarding the section. As shown in the figure, the user can create / input a memo 814 regarding the voice data at the time point of "01:07" in the voice recording. Here, the time point of "01:07" can correspond to the text section 812 of "01:07" in the first voice recording. At this time, the processor can output the created / input memo 814 together with the time information ("01:07") indicating the corresponding section. Then, the user can perform a voice-text reconversion request for the voice recording by selecting (e.g., click input, etc.) the "Reconversion" icon 816.
[0065] In response to the user's voice-text reconversion request, the processor can output a second voice recording generated by voice-text reconversion based on the memo 814 regarding the voice data created at the time point of "01:07" after the voice recording (i.e., the memo regarding the voice data related to the time point of "01:07" in the voice recording). For example, keywords "demo", "web", and "add" are extracted as keywords from the memo 814 regarding the voice data created by the user at the time point of "01:07" in the first operation 810, and the extracted keywords "demo", "web", and "add" and the voice of a specific section related to the memo 814 are input into the voice-text transcription model, so that the voice recognition weight value can be applied to the keywords "demo", "web", and "add". As a result, the voice converted to "The function used on the web in this Nemo has exceeded the limit." 812 in the first voice recording can be converted to "In this demo, the function used on the web has been added." 822 in the second voice recording.
[0066] On the other hand, for other sections in the voice data not related to the memo 814, the voice recognition weighting values are not applied to the keywords "demo", "web", and "add" extracted from the memo 814. For example, the voice converted to "Please share the Nemo site discussed last month." in the first voice recording is not reconverted to "Please share the demo site discussed last month." in the second voice recording, but remains converted to "Please share the Nemo site discussed last month." 824 as it is. That is, when the memo 814 regarding the voice data is created in relation to a specific section within the voice data, the keywords extracted from the memo 814 regarding the voice data are recognized as having a higher priority than other keywords only for the specific section, and for other sections, they are recognized in the same way as the existing priority. Therefore, as shown in the second operation 820, a second voice recording including "In this demo, we added functions to be used on the web." 822 and "Please share the Nemo site discussed last month." 824 can be generated, and the processor can output the generated second voice recording.
[0067] FIG. 9 is a diagram showing an example of outputting a voice recording generated based on information 918 of one or more participants received after voice recording and / or a topic 920 regarding the voice data according to an embodiment of the present disclosure. In one embodiment, the information related to the voice data can include information regarding one or more participants related to the voice included in the voice data. Additionally or alternatively, the information related to the voice data can include a topic regarding the voice data. At this time, voice recognition weighting values in the voice-text transcription model can be applied to one or more keywords extracted from the information 918 regarding one or more participants and / or the topic 920 regarding the voice data.
[0068] As shown in the first operation 910, a processor (e.g., at least one processor of a user terminal) can output a first voice recording generated by voice-text conversion for a voice recording. The first voice recording can include text in which voice data is mis-converted, such as "Please share the Nemo site discussed last month." 912 and "The function used in the web exceeded the limit in this Nemo." 914. After the voice recording, the user can create / input information 918 about one or more participants and / or a topic 920 about the voice data (e.g., new input, additional input, modified input, etc.). For example, the user can select the "Add Participant" icon 916 and select (or input) information (e.g., name, business, age, rank, location, etc.) of the participants "user1", "user2", "user3" to be added, thereby inputting information about the participants related to the voice recording.
[0069] Thereafter, the user can perform a voice-text re-conversion request for the voice recording by selecting (e.g., click input, etc.) the "Re-convert" icon 922. The processor can output a second voice recording generated by voice-text re-conversion based on information 918 about one or more participants and / or a topic 920 about the voice data input after the voice recording in response to the user's voice-text re-conversion request.
[0070] For voice-text re-conversion, one or more keywords can be extracted from the topic 920 about the voice data created / input after the voice recording and / or information 918 about one or more participants. For example, if "user2" input by the user as participant information in the first operation 910 corresponds to a person who performs work on the demo site, based on such information of "user2", "demo" and "site" can be extracted as keywords. Additionally, keywords "web", "function", and "addition" can be extracted from the topic "Web Function Addition Meeting" 920 of the voice recording input by the user in the first operation 910.
[0071] By inputting information regarding the extracted keywords "demo", "site", "web", "function", "addition", and at least a portion of the voice into the speech-to-text transcription model, a voice recognition weight value can be applied to the keywords "demo", "site", "web", "function", "addition". As a result, the voice converted to "Please share the Nemo site discussed last month." 912 in the first voice recording can be converted to "Please share the demo site discussed last month." 932 in the second voice recording. Also, the voice converted to "In this Nemo, the functions used on the web were exceeded." 914 in the first voice recording can be converted to "In this demo, the functions used on the web were added." 934 in the second voice recording. Therefore, as shown in the second operation 930, a second voice recording including "Please share the demo site discussed last month." 932 and "In this demo, the functions used on the web were added." 934 can be generated, and the processor can output the generated second voice recording.
[0072] FIG. 10 is a flowchart showing a process of re-converting voice data and / or editing a voice recording according to an embodiment of the present disclosure. In one embodiment, when the speech-to-text conversion for a voice recording is completed and a first voice recording is generated / output (S1010), a processor (for example, at least one processor of a user terminal) can output a message guiding memo re-conversion (S1020). For example, the processor can output a guiding message inducing memo creation and / or speech-to-text re-conversion regarding voice data, such as "Please create and re-convert a memo.", "Creating and re-converting a memo related to the recording will increase the recognition rate."
[0073] Thereafter, the processor can receive a user's request for re-conversion of the voice recording (S1022). In one embodiment, if there is a memo created for the voice recording, the processor can output a first pop-up message regarding the feasibility of re-conversion (e.g., re-conversion confirmation pop-up) to confirm the feasibility of voice-text re-conversion in response to the received user's re-conversion request (S1024). For example, the processor can output a first pop-up message including "Do you want to re-convert the voice recording?". Thereafter, the processor can output a second voice recording generated by voice-text re-conversion and / or a second pop-up message indicating the completion of re-conversion based on the user's input to the first pop-up message (S1026). For example, the processor can output a second voice recording generated by voice-text re-conversion and / or a second pop-up message including "The re-conversion of the voice recording is completed." based on an affirmative user input to the first pop-up message (i.e., a user input indicating a re-conversion request).
[0074] On the contrary, if there is no memo created for the voice recording, the processor can output a third pop-up message (e.g., memo creation guidance pop-up) to induce memo creation in response to the received user's re-conversion request (S1028). For example, the processor can output a third pop-up message including "Please create a memo and then re-convert."
[0075] In another embodiment, when the voice-text conversion for the voice recording is completed and the voice recording is generated / output (S1010), the processor can receive an editing request for the voice recording (S1030). If the re-conversion for the voice recording is not performed (i.e., if the generated / output voice recording corresponds to the first voice recording), the processor can output a fourth pop-up message regarding the permission of editing before re-conversion according to the received editing request (S1032). For example, the processor can output a fourth pop-up message including "Please create a memo and edit it after re-conversion" according to the received editing request. Then, based on the response indicating the user's editing request for the fourth pop-up message, the processor can provide the user with an interface for editing the voice recording by switching to the editing mode for the voice recording.
[0076] On the contrary, if the re-conversion for the voice recording has already been performed (i.e., if the generated / output voice recording corresponds to the second voice recording), the processor can provide the user with an interface for editing the voice recording by immediately switching to the editing mode for the voice recording according to the received editing request (S1034). The user can correct / edit at least a part of the text in which the voice is mis-converted among the multiple texts included in the voice recording through the interface for editing the voice recording.
[0077] FIG. 11 is a diagram showing an example of an artificial neural network model 1100 according to an embodiment of the present disclosure. The artificial neural network model 1100 can be a statistical learning algorithm embodied based on the structure of a biological neural network in machine learning technology and cognitive science as an example of a machine learning model, or a structure for executing the algorithm.
[0078] According to one embodiment, the artificial neural network model 1100 can be a machine learning model with problem-solving capabilities by having nodes, which are artificial neurons forming a network through synaptic connections like a biological neural network, repeatedly adjust synaptic weights so that the error between the correct output corresponding to a specific input and the inferred output decreases. For example, the artificial neural network model 1100 can include any probability model, neural network model, etc. used in artificial intelligence learning methods such as machine learning and deep learning.
[0079] According to one embodiment, the artificial neural network model 1100 can include an artificial neural network model configured to output text corresponding to at least a part of the voice when information regarding at least a part of the voice included in the voice data and information related to the voice data are input. Here, the information related to the voice data can include a memo regarding the voice data, information regarding one or more participants related to the voice included in the voice data, a topic regarding the voice data, one or more keywords extracted from the information related to the voice data, one or more keywords extracted from the text included in the first voice recording, correction information regarding at least a part of the text included in the text of the first voice recording, and the like. Additionally or alternatively, the artificial neural network model 1100 can include an artificial neural network model configured to output text corresponding to at least a part of the voice by applying voice recognition weights so that one or more keywords are recognized as having a higher priority than other keywords when information regarding at least a part of the voice included in the voice data and one or more keywords are input.
[0080] The artificial neural network model 1100 is implemented as a multilayer perceptron (MLP) composed of multiple layers of nodes and the like and the connections between them. The artificial neural network model 1100 according to this embodiment can be implemented using one of various artificial neural network model structures including an MLP. As shown in FIG. 11, the artificial neural network model 1100 includes an input layer 1120 that receives an input signal or data 1110 from the outside, an output layer 1140 that outputs an output signal or data 1150 corresponding to the input data, and is located between the input layer 1120 and the output layer 1140, receives a signal from the input layer 1120, extracts features, and transmits them to the output layer 1140. It consists of n (where n is a positive integer) hidden layers 1130_1 to 1130_n. Here, the output layer 1140 receives a signal from the hidden layers 1130_1 to 1130_n and outputs it to the outside.
[0081] The learning method of the artificial neural network model 1100 includes a supervised learning method that learns to optimize the solution of the problem by inputting a teacher signal (correct answer), and an unsupervised learning method that does not require a teacher signal. In one embodiment, the information processing system can cause the artificial neural network model 1100 to perform supervised learning and / or unsupervised learning so as to output text (or text information) corresponding to at least a part of the voice (or information related to the voice) included in the voice data. For example, the information processing system can cause the artificial neural network model 1100 to perform supervised learning and / or unsupervised learning so as to output text corresponding to the reference voice by inputting the reference voice and reference information related to the reference voice.
[0082] The artificial neural network model 1100 thus learned can be stored in the memory (not shown) of the information processing system, and outputs text (or text information) corresponding to at least a part of the voice (or information related to the voice) and / or information related to the voice data included in the voice data received from the communication module and / or the memory according to at least a part of the voice (or information related to the voice) included in the voice data. Additionally or alternatively, the artificial neural network model 1100 can output an audio recording including text (or text information) corresponding to at least a part of the voice (or information related to the voice) included in the voice data.
[0083] According to one embodiment, the input variable of the machine learning model that performs voice-text transcription, i.e., the artificial neural network model 1100, can be at least a part of the voice (or information related to the voice) included in the voice data. For example, the input variable input to the input layer 1120 of the artificial neural network model 1100 can be a vector 1110 that constructs at least a part of the voice included in the voice data as one vector data element. According to at least a part of the voice input included in the voice data, the output variable output from the output layer 1140 of the artificial neural network model 1100 can be a vector 1150 that indicates or characterizes text (or text information) corresponding to at least a part of the voice (or information related to the voice). Additionally or alternatively, the output layer 1140 of the artificial neural network model 1100 can be configured to output a vector that indicates or characterizes an audio recording including text (or text information) corresponding to at least a part of the voice (or information related to the voice). In the present disclosure, the output variable of the artificial neural network model 1100 is not limited to the types described above, and can include any information / data indicating text (or text information) and / or an audio recording corresponding to at least a part of the voice (or information related to the voice).
[0084] Furthermore, the output layer 1140 of the artificial neural network model 1100 can be configured to output a vector indicating the reliability and / or accuracy of the output voice-text conversion (or re-conversion) result.
[0085] In this way, a plurality of input variables and a plurality of output variables corresponding thereto are respectively matched to the input layer 1120 and the output layer 1140 of the artificial neural network model 1100, and the synaptic values among nodes and the like included in the input layer 1120, the hidden layers 1130_1 to 1130_n, and the output layer 1140 are adjusted, so that learning can be performed to extract the correct output corresponding to a specific input. Through such a learning process, the hidden characteristics of the input variables of the artificial neural network model 1100 can be grasped, and the synaptic values (or weight values) among nodes and the like of the artificial neural network model 1100 can be adjusted so that the error between the output variable calculated based on the input variable and the target output is reduced. The information processing system and / or the user terminal inputs at least part of the information related to the voice and the information related to the voice data into the learned artificial neural network model 1100, and can generate and / or output a voice recording for the voice data by using the output text information.
[0086] For execution by a computer, the above-described method can be provided as a computer program stored in a computer-readable recording medium. The medium can continuously store a computer-executable program, or can temporarily store it for execution or download. Further, the medium can be various recording means or storage means in a form in which a single or multiple hardware components are combined, and is not limited to a medium directly connected to a certain computer system, and can be distributed and present on a network. Examples of the medium include magnetic media such as hard disks, floppy (registered trademark) disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, ROMs, RAMs, flash memories, etc., and those configured to store program instruction words. Further, examples of other media include app stores through which applications are distributed, and recording media or storage media managed by sites, servers, etc. that supply or distribute other various software.
[0087] The methods, operations, or techniques of the present disclosure can be implemented by various means. For example, such techniques can be implemented in hardware, firmware, software, or combinations thereof. Those of ordinary skill in the art should understand that the various exemplary logical blocks, modules, circuits, and algorithm steps described by the disclosure of this application can be implemented in electronic hardware, computer software, or combinations of both. To clearly illustrate such mutual substitution between hardware and software, various exemplary components, blocks, modules, circuits, and steps have been generally described from their functional perspectives. Whether such functions are implemented as hardware or as software varies depending on the design requirements imposed on the specific application and the overall system. Those of ordinary skill in the art can also implement the functions described in various ways for each specific application, but such implementations should not be construed as departing from the scope of the present disclosure.
[0088] In a hardware implementation, the processing unit utilized for the execution of the technique can also be implemented in one or more ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, electronic devices, other electronic units designed to perform the functions described in the present disclosure, computers, or combinations thereof.
[0089] Accordingly, the various illustrative logical blocks, modules, and circuits described by this disclosure can be implemented or performed by any combination such as a general-purpose processor, DSP, ASIC, FPGA, or other programmable logic device, discrete gates or transistor logic, discrete hardware components, or those designed to perform the functions described in this application. The general-purpose processor can be a microprocessor, but alternatively, the processor can be any conventional processor, controller, microcontroller, or state machine. The processor can also be implemented by a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors associated with a DSP core, or any other configuration combination.
[0090] In the implementation of firmware and / or software, the techniques can be implemented by instructions stored on a computer-readable medium such as RAM (random access memory), ROM (read-only memory), NVRAM (non-volatile random access memory), PROM (programmable read-only memory), EPROM (erasable programmable read-only memory), EEPROM (electrically erasable PROM), flash memory, CD (compact disc), magnetic or optical data storage devices, and the like. The instructions are executable by one or more processors, and the processors can perform specific aspects of the functions described by this disclosure.
[0091] The foregoing embodiments have been described as utilizing aspects of the presently disclosed subject matter in one or more stand-alone computer systems, but the present disclosure is not so limited and can be embodied by any computing environment, such as a network or a distributed computing environment. Further, aspects of the subject matter of the present disclosure can also be embodied on a plurality of processing chips or devices, and storage can be similarly affected across a plurality of devices. Such devices can also include PCs, network servers, and portable devices.
[0092] Although the present disclosure has been described by way of some embodiments herein, various modifications and changes can be made without departing from the present disclosure as would be understood by those of ordinary skill in the art to which the present disclosure pertains. Also, such modifications and changes should be understood to fall within the scope of the claims appended hereto.
Claims
1. In a method of providing an audio recording generated based on information after audio recording, which is performed by at least one computing device, receiving information related to audio data after audio recording; receiving a voice-to-text conversion request for the audio data; outputting an audio recording generated based on at least a part of the audio included in the audio data and information related to the audio data received after the audio recording in response to the voice-to-text conversion request, the method comprising: wherein the information related to the audio data received after the audio recording includes a memo regarding the audio data created after the audio recording; wherein the memo regarding the audio data created after the audio recording is associated with the audio included in a specific section within the audio data; the method of providing an audio recording, wherein the audio recording includes text information generated by voice-to-text conversion for the audio included in the specific section, using the memo created in relation to the specific section.
2. The method of providing an audio recording according to claim 1, wherein the information related to the audio data includes information regarding one or more participants related to the audio included in the audio data.
3. The method of providing an audio recording according to claim 1, wherein the information related to the audio data includes a title regarding the audio data.
4. The audio recording includes text information output by inputting information regarding the at least a part of the audio and information related to the audio data into a voice-to-text transcription model, the method of providing an audio recording according to claim 1, wherein the voice-to-text transcription model is trained to output text corresponding to a reference audio by inputting the reference audio and reference information related to the reference audio.
5. The information related to the audio data includes one or more keywords extracted from the information related to the audio data, the method of providing an audio recording according to claim 4, wherein the reference information related to the reference audio includes one or more reference keywords related to the reference audio.
6. The method of providing an audio recording according to claim 5, wherein a voice recognition weight value is applied to the one or more keywords input into the voice-to-text transcription model, so that the one or more keywords are recognized as having a higher priority than keywords different from the one or more keywords.
7. Further including the step of outputting a first voice record generated by voice-text conversion for at least a part of the voices included in the voice data before receiving information related to the voice data, The step of receiving a voice-text conversion request for the voice data includes the step of receiving a voice-text reconversion request for the voice data, The step of outputting includes the step of outputting, as the voice record, a second voice record generated based on at least a part of the voices included in the voice data and information related to the voice data received after the voice recording in response to the voice-text reconversion request. A method for providing a voice record according to claim 1.
8. The information related to the voice data includes one or more keywords extracted from the text included in the first voice record. A method for providing a voice record according to claim 7.
9. The information related to the voice data includes correction information regarding at least a part of the text included in the first voice record, With the correction information, a voice recognition weighting value in a voice-text transcription model is applied to the corrected text among the text included in the first voice record. A method for providing a voice record according to claim 7.
10. A computer-readable non-transitory recording medium recording instruction words for executing the method according to claim 1 by a computer.
11. An information processing system, A communication module, A memory, At least one processor coupled to the memory and configured to execute at least one computer-readable program included in the memory, The at least one program includes Receiving information related to voice data after voice recording, Receiving a voice-text conversion request for the voice data, In response to the voice-text conversion request, including instruction words for generating a voice record based on at least a part of the voices included in the voice data and information related to the voice data received after the voice recording, The information related to the voice data received after the voice recording includes a memo regarding the voice data created after the voice recording, The memo regarding the voice data created after the voice recording is associated with the voices included in a specific section within the voice data, The at least one program is An information processing system further including instruction words for generating an audio recording including text information corresponding to the audio included in the specific section by performing audio-text conversion on the audio included in the specific section by using a memo regarding the audio data created after the audio recording.
12. The information related to the audio data includes at least one of information regarding one or more participants related to the audio included in the audio data or a topic regarding the audio data, according to the information processing system described in Claim 11.
13. The at least one program is Further including instruction words for generating an audio recording including text information corresponding to the information regarding at least a part of the audio by inputting the information regarding at least a part of the audio and the information related to the audio data into an audio-text transcription model. The audio-text transcription model is learned to output text corresponding to the reference audio by inputting the reference audio and reference information related to the reference audio, according to the information processing system described in Claim 11.
14. The information related to the audio data includes one or more keywords extracted from the information related to the audio data. The reference information related to the reference audio includes one or more reference keywords related to the reference audio, according to the information processing system described in Claim 13.
15. The at least one program is Applying an audio recognition weight value to one or more keywords input to the audio-text transcription model. Further including instruction words for recognizing the one or more keywords as having a higher priority than keywords different from the one or more keywords, according to the information processing system described in Claim 14.
16. The at least one program is Generating a first audio recording by performing audio-text conversion on at least a part of the audio included in the audio data before receiving the information related to the audio data. Receiving an audio-text reconversion request for the audio data. Further including instruction words for generating a second audio recording based on at least a part of the audio included in the audio data and the information related to the audio data received after the audio recording in response to the audio-text reconversion request, according to the information processing system described in Claim 11.
Citation Information
Patent Citations
Caption generator, retrieval device, method for integrating document processing and speech processing together, and program
JP2006178087A
Voice recognition apparatus and voice recognition program
JP2008089825A
In-vehicle device
JP2009092975A
Transcription device and program
JP2014134640A
Speech-to-text converter, speech-to-text conversion method and speech-to-text conversion program
JP2020190671A