Speech-to-Text Conversion System and Speech-to-Text Conversion Program

The voice-to-character conversion system addresses the challenges of security and operability by using two terminals for short-distance communication, enabling secure and convenient real-time character information display in medical settings.

JP7689682B1Active Publication Date: 2025-06-09BOATRIP INC

Patent Information

Application Number
JP2024209087
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-11-29
Publication Date
2025-06-09
Estimated Expiration
2044-11-29

AI Technical Summary

Technical Problem

Existing voice-to-character conversion systems face challenges in ensuring information security and providing high convenience and operability for multiple users, especially in medical settings where confidentiality and ease of use are critical.

Method used

A voice-to-character conversion system that utilizes two terminals for short-distance communication, where the first terminal acquires voice information and converts it into character information, and the second terminal displays the character information in real-time, ensuring secure and convenient operation.

Benefits of technology

The system effectively separates voice input and information storage terminals, ensuring secure communication and real-time character information display, thereby enhancing user convenience and operability while maintaining information security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007689682000001_ABST
    Figure 0007689682000001_ABST
Patent Text Reader

Abstract

Provide a highly operable speech-to-text conversion system. 【Solution means】The speech-to-text conversion system 1 includes a first terminal that acquires voice information of a user, and a second terminal that displays in real time the character information obtained by converting the voice information. By performing short-distance communication between these terminals, a speech-to-text conversion system with high convenience and operability for many users is provided while ensuring the security of communication. Also, the user can select whether to acquire either the voice information (first voice information) acquired while the input operation unit is receiving the user's operation or the voice information (second voice information) for which acquisition is started when the input operation unit receives the user's operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a voice-to-character conversion system and a voice-to-character conversion program.

Background Art

[0002] In the daily operations of various industries, software for converting voice information into character information, so-called speech recognition, has come to be widely used. Due to such technological improvements, the convenience of using voice information, such as creating meeting minutes, has been significantly improved. In addition, since this information is stored as character data, it is also easy to search for and obtain necessary information later. Such technologies are useful for improving business productivity, but are essential technologies, especially in industries where business efficiency is required.

[0003] For example, there is a strong demand for digital transformation in medical settings such as hospitals. This is because as the birthrate declines and the population ages, the increasing burden on medical institutions has become a problem. That is, while an increase in the number of people in need of medical care, such as the elderly, is expected, the burden on doctors and the like increases, and there is a problem that appropriate medical services may not be provided.

[0004] Therefore, there is a need for software that can reduce the workload of doctors and the like in medical settings and is easy to introduce.

[0005] Examples of such software that reduces the workload of doctors and the like include electronic medical records that enable voice input.

[0006] Patent Document 1 discloses an electronic medical record creation device including voice analysis means for analyzing voice and converting it into character information. In addition, Patent Document 2 discloses an invention related to character input for mobile phones and the like.

Prior Art Documents

Patent Documents

[0007] [Patent Document 1] Japanese Unexamined Patent Application Publication No. 2015-035099 [Patent Document 2] Japanese Unexamined Patent Application Publication No. 2004-032275

[0008] However, when introducing such software and speech recognition devices, convenience and operability at the site become problems. For example, examinations are not always conducted in the examination room, and examinations may be conducted in the patient's room for patients who have difficulty moving, or there may be cases where records need to be left in special situations such as outdoors.

[0009] On the other hand, general speech recognition devices connect a microphone or the like to a desktop computer or a laptop computer and input voice from the microphone. However, there may be cases where a desktop computer or the like cannot be installed at the site, and a laptop computer cannot be brought in either. In particular, since terminals storing confidential information such as personal information should be isolated from accessible places by patients, there is also a problem that they are not suitable for carrying.

[0010] Thus, there is a need to separate the terminal for speech recognition input and the terminal for storing information. Also, it is more preferable that the result of speech recognition can be confirmed on the spot. In particular, there is a need to use a terminal such as a smartphone that one is familiar with as an input device.

[0011] On the other hand, when communicating, if an Internet network that anyone can connect to is used, problems such as wiretapping will occur. In particular, special attention is required when transmitting personal information.

[0012] Patent Document 1 does not describe the separation of terminals. In addition, Patent Document 2 does not display characters on the character information creation device 2, so the uses and purposes are different. Also, it does not touch on security other than encryption in the communication network.

[0013] In addition to the above, for those who are busy working on-site, there is resistance to adopting applications with poor operability. However, whether one feels that the operability of a certain application is good or bad depends on the individual, and individual differences are inevitable. That is, even if the operability is good for one user, if it is bad for another user, ultimately the adoption of that application will be postponed, which may cause a delay in the implementation of DX on-site.

Summary of the Invention

Problems to be Solved by the Invention

[0014] The problem to be solved is that it is impossible to provide a voice-to-character conversion system that ensures information security and has high convenience and operability for many users.

Means for Solving the Problems

[0015] The present invention includes a first terminal and a second terminal that perform short-distance communication. The first terminal includes a voice information acquisition unit that acquires the user's voice information. At least the first terminal or the second terminal includes an information conversion unit that converts the voice information into character information. The second terminal is mainly characterized by including a receiving-side character information display unit that displays the character information on the display unit of the second terminal in real time.

[0016] The present invention has been made in view of the above problems and, for example, adopts the following means. That is, a voice-to-character conversion system including a first terminal that acquires voice information and a second terminal that displays in real time the character information obtained by converting the voice information, The first terminal is A voice standby display unit that displays on the display unit of the first terminal that voice acquisition is possible. A voice information acquisition unit that acquires voice information of the user, and A short-distance communication transmission unit that transmits the voice information and / or character information obtained by converting the voice information by short-distance communication. The voice character conversion system is provided with the following components: The second terminal includes A short-distance communication reception unit that receives the voice information and / or character information obtained by converting the voice information by short-distance communication, and A reception-side character information display unit that displays the character information on the display unit of the second terminal in real time. Provided is a voice character conversion system, characterized in that at least one of the first terminal or the second terminal includes an information conversion unit that converts the voice information into character information.

Effect of the Invention

[0017] The voice character conversion system of the present invention separates a first terminal that acquires voice information of a user from a second terminal that displays in real time character information obtained by converting the voice information, and further, by performing short-distance communication between them, provides a voice character conversion system that is highly convenient and operable for many users while ensuring communication security.

Brief Description of the Drawings

[0018]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Mode for Carrying Out the Invention

[0019] Embodiments of the present invention will be described with reference to the drawings. In the following embodiments, the same or corresponding parts may be denoted by the same reference numerals, and the description may be omitted as appropriate. Also, the drawings used below are for explaining this embodiment, and may be different from the actual device configuration, user interface (UI), data configuration, etc.

[0020] (Outline of the Embodiment) The outline of this embodiment will be described with reference to FIG. 1. FIG. 1 is a diagram showing the outline of the voice-to-character conversion system 1 according to this embodiment.

[0021] The voice-to-character conversion system 1 includes a transmitting-side terminal 10 (10a to 10c in FIG. 1) and a receiving-side terminal 20 (20a to 20c in FIG. 1). The transmitting-side terminal 10 and the receiving-side terminal 20 perform short-distance communication in accordance with a standard such as Bluetooth (registered trademark). Note that the dotted line in FIG. 1 indicates that the transmitting-side terminal 10 is receiving a signal emitted by the receiving-side terminal 20.

[0022] Here, the transmitting-side terminal 10 and the receiving-side terminal 20 each include a transmitting-terminal-side application program P10 (hereinafter referred to as "app P10") and a receiving-terminal-side application program P20 (hereinafter referred to as "app P20").

[0023] The program used by the voice-to-text conversion system 1 to achieve the objective of this embodiment is referred to as the "voice-to-text conversion program P1". The voice-to-text conversion program P1 includes programs (such as app P10 and app P20) used by the transmitting terminal 10 and the receiving terminal 20.

[0024] Hereinafter, an example in which the voice-to-text conversion system 1 is used in a medical institution will be described. In this embodiment, the transmitting terminal 10 is a smartphone (hereinafter referred to as "smartphone"), and the receiving terminal 20 is a desktop or laptop personal computer (hereinafter referred to as "PC").

[0025] First, a user such as a doctor selects their own PC (receiving terminal 20a) equipped with an electronic medical record, which is the information transmission destination, from their own smartphone (for example, transmitting terminal 10a). During the examination, the content to be input into the electronic medical record is input by voice into the voice input unit (microphone) of the smartphone. The voice is, for example, the voice of a doctor or a patient.

[0026] At this time, the user can select the voice input method (transceiver mode, recording mode). For example, in the "transceiver mode" of this embodiment, the voice input while the input operation unit (button UI) is being operated by the user is acquired (first voice information acquisition). That is, the voice is acquired while the user is pressing (holding down) the button UI, and the acquisition of the voice stops when the user releases the button UI. Also, in the "recording mode" of this embodiment, the acquisition of the voice input into the microphone is started (second voice information acquisition) when the input operation unit (button UI) is operated (once) by the user. In the recording mode, when the input operation unit (button UI) is operated again by the user, the acquisition of the voice ends. That is, when the user presses the button UI once, the acquisition of the voice starts, and when the user presses the button UI again, the acquisition of the voice ends.

[0027] The app P10 on the smartphone converts voice information into character information and displays it on the display unit of the smartphone. In addition, the character information is transmitted to a PC (receiving terminal 20) by short-range communication. The PC of the receiving terminal 20 displays the acquired character information on the electronic medical record on the monitor.

[0028] (Details of the Embodiment) Hereinafter, the voice character conversion system 1 according to the present embodiment will be described in detail. The voice character conversion system 1 includes computers (transmitting terminal 10 and receiving terminal 20) equipped with a voice character conversion program P1, and provides a highly operable voice character conversion service (caption service) to the user. That is, in the voice character conversion system 1, the information processing by the voice character conversion program P1 is specifically realized using hardware resources. Hereinafter, the 1. user interface, 2. program processing, 3. data, and 4. hardware configuration, which constitute the voice character conversion system 1, will be described in order.

[0029] (Definition of Terms) Here, some terms will be defined. "Voice information" is information related to voice. In principle, sounds and voices acquired by a microphone or the like are regarded as voices, and those obtained by converting voices into electronic data are regarded as voice information, but they will not be strictly distinguished hereinafter. That is, sounds and voices acquired by a microphone or the like may also be referred to as voice information. "Short-range communication" refers to direct communication between terminals without going through a communication network such as an intranet or the Internet. It is a term used to distinguish from communication using an existing communication network to which an unspecified number of terminals are connected. It does not matter whether it is wireless or wired. If it is wireless, for example, Bluetooth (registered trademark) or BLE (Bluetooth (registered trademark) low energy) can be mentioned. Here, the term "short range" is a term used to distinguish from communication such as the Internet that can communicate over any long distance as long as a communication network is laid, and does not strictly limit the communication distance, but is generally within 100m. "Real-time" means being processed almost simultaneously. For example, it means a case where the input voice information is displayed as character information almost simultaneously. The "almost simultaneously" here means that due to the processing from the input of the voice to the display of the character information and the processing including communication, there is a time lag within several seconds, so it is not completely simultaneous but includes a time lag. Although the term "real-time" appears multiple times below, since the time related to each of these processes is different, the time represented by each "real-time" term is different. For example, in the embodiments described later, regarding the display of characters on the first terminal and the display of characters on the second terminal, although there should be a time lag until the display for the latter because communication is involved, it is expressed that both of these displays are performed in real-time. In addition to the above, in the text, the definitions of words are given using square brackets.

[0030] In the following, when it is described as "○○" processing, it means that the computer's processor executes the processing based on the "○○" program stored in the program storage unit. In this paragraph, the same word is entered in the "○○" part. That is, the "○○" program is a program that causes the computer to function as "○○" means by executing the "○○" processing. Also, at this time, the control unit including the processor also means that it functions as the "○○" unit (or "○○" device). In this case, the "○○" unit means executing the "○○" processing based on the "○○" program.

[0031] For example, the voice character conversion program P1 is a program that causes the computer to function as voice character conversion means by executing the voice character conversion processing. Also, at this time, the control unit of the computer (transmission-side terminal 10 and / or reception-side terminal 20) including the processor functions as the voice character conversion unit (or voice character conversion device).

[0032] Similarly, each of the terminals that execute the selection screen display process, user selection acquisition process, voice standby display process, voice acquisition process, voice information acquisition process, volume-driven voice acquisition start process, volume-driven voice acquisition stop process, information conversion process, technical term conversion process, short-distance communication process, short-distance communication transmission process, short-distance communication reception process, transmission log storage process, character information acquisition process, character information display process, transmission-side character information display process, reception-side character information display process, and character code conversion process, described below, functions as a selection screen display unit, user selection acquisition unit, voice standby display unit, voice acquisition unit, voice information acquisition unit, volume-driven voice acquisition start unit, volume-driven voice acquisition stop unit, information conversion unit, technical term conversion unit, short-distance communication unit, short-distance communication transmission unit, short-distance communication reception unit, transmission log storage unit, character information acquisition unit, character information display unit, transmission-side character information display unit, reception-side character information display unit, and character code conversion unit, respectively.

[0033] In the voice character conversion system 1, since each terminal (computer) such as the transmission-side terminal 10 and the reception-side terminal 20 is equipped with a processor, when simply referring to the "processor", it refers to the processor that performs the processing related to the voice character conversion program P1, that is, the processor 111 of the transmission-side terminal 10 and / or the processor 211 of the reception-side terminal 20.

[0034] In the following, for simplicity, the expression "the processor causes the storage unit (data storage unit) to store data" may be described as "the processor stores (causes to store) (data)". Similarly, the expression "the processor causes the input unit such as a voice device to acquire voice" may be described as "the processor acquires (causes to acquire) voice". Furthermore, the expression "the processor causes the output unit (browser, speaker, etc.) to output display, voice, etc." may be described as "the processor displays (causes to display)". As described above, even when the actually executed functional units are different, the execution entity may be described as the processor.

[0035] 1. User Interface (UI) First, the user interface that the speech-to-text conversion system 1 of the present embodiment causes to be displayed on the transmission-side terminal 10, for example, a smartphone, will be described with reference to the drawings. The interfaces described hereinafter are simplified versions of those that the application P10 causes to be displayed on the browser of the transmission-side terminal 10.

[0036] Also, only icons and the like related to functions necessary for the description will be displayed, and other well-known icons and the like will be omitted. For example, the back button for returning to the previously displayed page, the close button for closing the screen, or the end button for ending the system are omitted.

[0037] FIG. 2 is a diagram showing a connection destination selection screen. The receiving-side terminal 20 enables the transmitting-side terminal 10 to find a connection destination through its communication function. That is, the transmitting-side terminal 10 recognizes the presence of the receiving-side terminal 20 and displays the name of the receiving-side terminal 20 as a connection destination display icon UI-10a on the display unit 151.

[0038] The user operating the transmitting-side terminal 10 selects an appropriate connection destination (receiving-side terminal 20) from the connection destination display icon UI-10a. When the user makes a selection and communication is successfully established with the receiving-side terminal 20, the processor 111 of the transmitting-side terminal 10 displays the voice acquisition means selection screen in the next section.

[0039] FIG. 3 is a diagram showing a voice acquisition means selection screen. As shown in FIG. 3, the processor 111 displays a voice acquisition means selection icon UI-10b, and the user selects a voice acquisition means (here, "transceiver mode" or "recording mode"). That is, the processor 111 of the smartphone obtains the user's selection as to whether to acquire the first voice information in transceiver mode or the second voice information in recording mode. Note that FIG. 3 is also a drawing showing the transceiver mode screen (before voice acquisition starts). While the user presses the voice acquisition button UI-111 (input operation unit 141) shown in FIG. 3, the processor 111 acquires voice. Next, the screen during voice acquisition will be described.

[0040] FIG. 4 is a diagram showing a transceiver mode screen (during voice acquisition). The processor 111 of the transmitting terminal 10 acquires the voice information input to the voice input unit 142.

[0041] Also, as shown in FIG. 4, the processor 111 of the present embodiment converts the voice information input to the voice input unit 142 into character information in real time and displays it on the character information display unit UI-112.

[0042] That is, in the present embodiment, the conversion process from voice information to character information (hereinafter referred to as "information conversion process") is performed by the transmitting terminal 10. Note that an aspect in which this information conversion process is performed by the receiving terminal 20 will be shown in a modification example described later.

[0043] Returning to FIG. 4, here, voice information of a doctor or a patient such as "From last night to early morning, there has been pain in the abdomen" is displayed on the screen as character information. Note that the "|" after the sentence "~pain" in the figure is a cursor (the same applies hereinafter). In this way, by displaying the character information on the screen, the doctor can confirm whether the conversion from voice information to character information is correctly performed.

[0044] Also, in the present embodiment, so that the user can visually recognize that the voice acquisition button UI-111 is pressed, the displayed state of the pressed voice acquisition button UI-111 changes (the voice acquisition button (during operation) UI-111a in FIG. 4). The display change is not particularly limited as long as the display content changes, and includes changes in the color or shape of the button, or disappearance of the button. Also, the processor 111 displays an audio spectrum (UI-113) on the display unit 151 during voice acquisition.

[0045] FIG. 5 is a diagram showing a recording mode screen (before voice acquisition starts). This is a screen displayed by the processor 111 when the user selects "recording mode" on the voice acquisition means selection screen. When the user presses (taps once, presses and releases once) the voice acquisition start / end button UI-121 (input operation unit 141) shown in FIG. 5, the processor 111 starts voice acquisition. Also, when the user presses the voice acquisition start / end button UI-121 (input operation unit 141) again, the processor 111 stops voice acquisition. Excluding the button UI part, the screen UI during voice acquisition is the same as the UI in the transceiver mode, so the description is omitted.

[0046] FIG. 6 is a diagram showing a received-side character display screen. The processor 211 of the receiving-side terminal 20 describes the character information obtained by converting the voice information in the electronic medical record opened as an application. As shown in FIG. 6, there is a cursor in the symptom column, and the processor 211 describes the character information converted from the voice (the same character information as that displayed in FIG. 4) in real time.

[0047] As shown in FIG. 6, the electronic medical record of this embodiment includes a patient column, a symptom column, and a finding column. The patient column is a column containing the personal information of the patient. The personal information is, for example, name and age. It may also include other medical histories and the like. By displaying the personal information of the patient, it is possible to prevent misidentification of the patient and enable the doctor to make a more appropriate diagnosis based on information such as age and medical history.

[0048] The symptom column is a column for entering the symptoms of the patient. Character information can be input regardless of whether it is the doctor's voice information or the patient's voice information.

[0049] The finding column is a column for entering the doctor's findings. It can be input with the doctor's voice information. The voices of doctors and those of third parties (such as nurses and patients) may be identified so that only the voice of the doctor is input.

[0050] In this embodiment, as a rule, a cursor is arranged in the symptom column. However, it is not limited to this. The doctor can select the input target by clicking on each column. For example, by clicking on the findings column, the cursor in the findings column is activated, and input by voice or the like may be output as text in the findings column.

[0051] Also, the doctor as the user can select the input field by voice. For example, when there is a voice input of "patient input", the patient column is activated and the cursor is automatically placed at the end of the patient column. Similarly, if it is a voice input of "symptom input", the symptom column is activated, and if it is "findings input", the findings column is activated. That is, the user can select the input destination of the character information by inputting a predetermined term by voice (through the voice input unit 142).

[0052] FIG. 7 is a diagram showing a reception-side screen for displaying a plurality of electronic medical records. As shown in FIG. 7, the display unit 251 of the reception-side terminal 20 can display a plurality of electronic medical records. By selecting (activating) a window (input target), the user can input character information related to the voice-to-character conversion of this embodiment for the selected electronic medical record. Also, in this embodiment, the user can activate an arbitrary window by voice. For example, in FIG. 7, by voice-inputting "Electronic Medical Record A", the processor 211 activates the window of the Electronic Medical Record A system.

[0053] With the above configuration, a user such as a doctor can use the mobile terminal as a voice input device (microphone) and utilize the functions related to voice-to-character conversion while confirming whether the voice information is correctly converted into character information on-site. In particular, the user can select either the "transceiver mode" or the "recording mode" according to their preferences. The selection can be made for operability, or for example, in a noisy environment, the transceiver mode can be selected to acquire voice only when speaking. In this case, there is an advantage that noise can be prevented from being input as character information.

[0054] 2. Program Processing <Voice Character Conversion Processing> The program processing on the voice character conversion system 1 in this embodiment will be described.

[0055] In this embodiment, the processor performs voice character conversion processing based on the voice character conversion program P1. The voice character conversion program P1 includes at least a voice selection acquisition program P11, a conversion display program P12, a character conversion improvement program P13, and a speaker identification display program P14. Each of these will be described below.

[0056] <2-1. Voice Selection Acquisition Processing> The processor performs voice selection acquisition processing based on the voice selection acquisition program P11. That is, the voice selection acquisition program P11 causes the computer to function as a voice selection acquisition means (voice selection acquisition unit) by executing the voice selection acquisition processing by the processor.

[0057] In the voice selection acquisition processing of this embodiment, the processor 111 of the transmission-side terminal 10 displays a screen for selecting a voice acquisition means (that is, a screen for selecting either the transceiver mode or the recording mode), and acquires the selection by the user. That is, in the voice selection acquisition processing, the user's selection regarding whether to acquire voice information as first voice information or as second voice information is acquired, and the voice information is acquired in accordance with the selection.

[0058] Here, the first voice information is voice information obtained by the voice input unit acquiring the voice input while the input operation unit is receiving the user's operation. When the input operation unit is not receiving the user's operation, the voice input to the voice input unit is not acquired. That is, the first voice information in the present embodiment is voice information acquired in the transceiver mode. The second voice information is voice information obtained by the input operation unit starting to acquire the voice input to the voice input unit when receiving the user's operation. When the input operation unit receives the user's operation again, (the processor 111) ends the acquisition of the voice. That is, the second voice information in the present embodiment is voice information acquired in the recording mode.

[0059] Here, the process of the processor 111 displaying a screen for selecting either the first voice information or the second voice information (that is, a screen for selecting the transceiver mode or the recording mode) is referred to as the "selection screen display process", and the process of the processor 111 acquiring the selection by the user at this time is referred to as the "user selection acquisition process".

[0060] Also, the voice selection acquisition process of the present embodiment includes the conversion display process described later.

[0061] FIG. 8 is a flowchart showing the voice selection acquisition process. By starting the application P10 on the receiving terminal 10, or by returning to the voice acquisition means selection screen, the processor 111 starts the voice selection acquisition process. At this time, it is assumed that the application P20 on the transmitting terminal 20 and the application program related to the electronic medical record are also started.

[0062] The processor 111 displays a voice acquisition means selection screen for the user to select the transceiver mode or the recording mode (step 1).

[0063] At this time, the processor 111 displays the input operation unit 141 on the display unit 151 (see FIG. 3). In the present embodiment, the input operation unit 141 becomes the voice acquisition button UI-111 in the transceiver mode, and becomes the voice acquisition start / end button UI-121 in the recording mode. However, the display timing of the input operation unit 141 is not limited to this. After the branch in step 3 of the next item, that is, after selecting the transceiver mode or the recording mode, the processor 111 may display the input operation unit 141 (the voice acquisition button UI-111 in the transceiver mode, or the voice acquisition start / end button UI-121 in the recording mode).

[0064] To put it in parentheses, the processor 111 displays on the display unit (of the transmission-side terminal 10) that voice acquisition is possible (referred to as "voice standby display process"). In the present embodiment, the display that voice acquisition is possible is the display of the button UI (voice acquisition button UI-111 or voice acquisition start / end button UI-121) that becomes the input operation unit 141.

[0065] However, it is not limited to this, and a character indicating that it is waiting for voice input may be displayed. Also, the input operation is not limited to the button UI, and may be a physical button (for example, a volume adjustment button in the case of a smartphone).

[0066] Returning to FIG. 8, first, the case where the user selects the transceiver mode (step 2 Yes) will be described.

[0067] While the user presses the voice acquisition button UI-111 and continues in that state (step 3 Yes), the processor performs conversion display processing (step 4). The conversion display processing will be described later.

[0068] On the other hand, when the user stops pressing the voice acquisition button UI-111 after pressing it (No in step 3), and performs an end operation (for example, pressing a close button or a back button (not shown)) (Yes in step 5), various data such as the temporarily saved voice information, character information, and information including the transmission date and time information of short-range communication are saved (for example, in auxiliary storage) (step 6), and the voice selection acquisition process is terminated. Note that information including the transmission date and time information of short-range communication and the like is referred to as a "transmission log". Also, the data saved here will be exemplified in the data section.

[0069] Although omitted in FIG. 8, when the user performs an end operation, the processor 111 saves the information (step 6) and terminates the voice selection acquisition process regardless of which step it is in.

[0070] When the user stops pressing the voice acquisition button UI-111 (No in step 3) and does not perform an end operation (No in step 5), the processor 111 waits for the button to be pressed.

[0071] In this embodiment, the data saved in step 6 includes voice information, character information obtained by converting the voice, information about the patient (such as surname), and processing date and time information (for example, transmission log). Also, the processor stores the voice information and the character information in association with each other. Details will be described in the data section.

[0072] Note that the process of saving the transmission log including the transmission date and time information of short-range communication is referred to as the "transmission log saving process".

[0073] Next, the case where the user selects the recording mode (No in step 2) will be described.

[0074] In the recording mode, when the user presses the voice acquisition start / end button UI-121 once (Yes in step 7), the processor performs conversion display processing (step 8). In this embodiment, although there is no difference in the processing content between the conversion display processes of step 4 and step 8, the processor 111 distinguishes and records which mode the data is in.

[0075] Next, even if the voice acquisition start / end button UI-121 is not pressed (step 7 No), if the voice input unit 142 of the transmitting terminal 10 detects a volume equal to or greater than a certain level (step 9 Yes) and this continues for a certain period of time (x seconds) or more (step 10 Yes), the processor performs the conversion display process (step 8).

[0076] Thereby, even if the user forgets to perform the voice acquisition operation, the processor can automatically acquire the voice and perform conversion to character information and the like. Note that the above-mentioned volume equal to or greater than a certain level, that is, the volume at which voice acquisition starts (hereinafter referred to as the "first volume"), can be set by the user or the like. Similarly, the above-mentioned certain period of time for starting voice acquisition (hereinafter referred to as the "first time", corresponding to x seconds in FIG. 8) can also be set by the user or the like.

[0077] If the voice input unit 142 of the transmitting terminal 10 does not detect a volume equal to or greater than a certain level (step 9 No), or even if the voice input unit 142 detects a volume equal to or greater than a certain level but it does not continue for a certain period of time or more (step 10 No), the processor 111 returns to the determination of pressing the voice acquisition start / end button UI-121 (step 7).

[0078] To summarize, even when there is no operation by the user on the input operation unit (here, the voice acquisition start / end button UI-121), the processor 111 starts voice acquisition when the voice input unit detects a voice equal to or greater than a predetermined first volume and this state continues for a predetermined first time or more. This process is referred to as the "volume-driven voice acquisition start process".

[0079] Returning to Fig. 8, after the voice acquisition start / end button UI-121 is pressed and then pressed again (voice acquisition stop operation, Step 11 Yes), the processor 111 stops voice acquisition (Step 12), saves various data such as the acquired data (Step 6), and ends the voice selection acquisition process.

[0080] Here, even if the voice acquisition start / end button UI-121 is not pressed (Step 11 No), when the volume of the voice detected by the voice input unit 142 of the transmitting terminal 10 becomes equal to or less than a certain level (Step 13 Yes) and this continues for a certain period of time (y seconds) or more (Step 14 Yes), the processor 111 also stops voice acquisition (Step 12).

[0081] Thus, even if the user forgets to perform an operation to stop voice acquisition, the processor can automatically stop voice acquisition and store voice information, etc. Note that the above-mentioned volume equal to or less than a certain level, that is, the volume at which voice acquisition ends (hereinafter referred to as the "second volume"), can be set by the user or the like. Similarly, the above-mentioned certain period of time for stopping voice acquisition (hereinafter referred to as the "second time", corresponding to y seconds in Fig. 8) can also be set by the user or the like.

[0082] If the volume detected by the voice input unit 142 of the transmitting terminal 10 is equal to or greater than a certain level (Step 13 No), or even if the volume detected by the voice input unit 142 is equal to or less than a certain level but does not continue for a certain period of time or more (Step 14 No), the processor 111 continues the conversion display process (Step 8).

[0083] To summarize, after voice acquisition starts, the processor 111 stops voice acquisition when the state where the voice input unit does not detect a voice equal to or greater than a predetermined second volume continues for a predetermined second time or more. This process is referred to as the "volume-driven voice acquisition stop process".

[0084] In the present embodiment, the user can disable the functions related to the volume-driven voice acquisition start process and / or the volume-driven voice acquisition stop process. This is to prevent the voice acquisition from starting or stopping automatically. When the volume-driven voice acquisition start process is disabled, if the button is not pressed in step 7 (step 7 No), wait for the button to be pressed (return to step 7). When the volume-driven voice acquisition stop process is disabled, if there is no stop operation (step 11 No), continue the conversion display process (return to step 8).

[0085] With the above configuration, the user can appropriately use the transceiver mode and the recording mode according to the situation. In addition, the volume-driven voice acquisition start process can prevent forgetting to record, and the volume-driven voice acquisition stop process can prevent forgetting to stop while the voice is being acquired in the recording mode.

[0086] As a conventional problem, when using software for voice-character conversion or a speech recognition device, there is a problem that necessary information cannot be obtained due to forgetting to start the device or the like. For example, when starting recording by operating an operation unit (buttons, switches, etc.) provided in a recording device, if the operation is forgotten, the voice cannot be recorded. On the other hand, although starting recording when a volume above a certain level is detected can also be mentioned, in this case, there may be a problem that the device starts up without being noticed and the recording state continues for a long time, which may compress the storage capacity of the storage device.

[0087] Particularly, for a device related to the latter case, when the capacity of the storage device runs out and it is of the type that deletes old data and records new data (i.e., overwrites), even past data may be erased.

[0088] Also, when the capacity of the storage device runs out, if it is of a type that does not delete old data or record new data (if it is of a type that stops recording), recording will still not be possible. Particularly in situations such as medical examinations, information recording omissions can lead to a situation where appropriate medical treatment cannot be provided, so a foolproof system is required to prevent such occurrences.

[0089] That is, in software with a function to convert voice information into character information, there was a problem that it was impossible to prevent the omission of obtaining voice information.

[0090] The voice character conversion system 1 according to the present embodiment has the advantages of preventing the forgetting of storing voice because it has both the function of obtaining the voice input to the voice input unit by operating the input operation unit and the function of detecting the voice input from the voice input unit and automatically starting voice acquisition.

[0091] <2-2. Conversion display process> The processor performs conversion display processing based on the conversion display program P12. That is, the conversion display program P12 causes the computer to function as conversion display means (conversion display unit) by executing the conversion display processing by the processor.

[0092] In the conversion display processing of the present embodiment, the processor 111 of the transmission-side terminal 10 (1) Voice information acquisition processing for acquiring voice information as first voice information or second voice information based on the user's voice information, particularly the user's selection, (2) Information conversion processing for converting the voice information into character information, (3) Transmission-side character information display processing for displaying the character information on the display unit of the transmission-side terminal 10 in real time, and (4) Short-distance communication transmission processing for transmitting the character information by short-distance communication, is executed.

[0093] Also, in the voice selection acquisition processing, the processor 211 of the reception-side terminal 20 (5) Short-distance communication reception processing for receiving character information (obtained by converting the above voice information) by short-distance communication, (6) Character information acquisition processing for acquiring character information obtained by converting voice information acquired as first voice information or second voice information based on a user's selection, (7) Reception-side character information display processing for displaying the character information in real time on the display unit of the reception-side terminal 20 is executed.

[0094] FIG. 9 is a flowchart showing the conversion display processing. The conversion display processing is a subroutine of the above-described voice selection acquisition processing. The processor starts the conversion display processing by receiving an operation related to the start of voice acquisition by the user.

[0095] Note that in FIG. 9, for simplicity, the processing of the processor 111 and the processing of the processor 211 in the present embodiment are described together. That is, from step 21 to step 24 are the processing of the processor 111, and step 25 and step 26 are the processing of the processor 211. In other words, the smartphone is a segment of processing from voice acquisition (step 21) to transmission of character information (step 24), and the PC is a segment of processing from reception of character information in step 24 to display of the character information on the display unit 251 (step 28). The following describes each step.

[0096] The processor 111 acquires voice through the voice input unit 142 (microphone) (step 21).

[0097] The processor 111 converts the acquired voice information into character information (information conversion processing, step 22).

[0098] The processor 111 displays the character information obtained by converting the voice information on the display unit 151 (smartphone screen) in real time (referred to as "transmission-side character information display process". Step 23), and also transmits at least the above-mentioned character information to the receiving-side terminal 20 by short-distance communication (Step 24).

[0099] The processor 211 that acquires character information by short-distance communication acquires the position on the display unit 251 where the acquired character information is to be displayed (character display position acquisition process. Step 25). Here, the position on the display unit 251 is, for example, the position where the cursor is activated in an input field (such as the above-mentioned symptom field) in the electronic medical record.

[0100] Subsequently, the processor 211 displays the acquired character information on the display unit 251 (electronic medical record) of the receiving-side terminal 20 (PC) in real time (referred to as "receiving-side character information display process". Step 26).

[0101] After the character information is displayed on the display unit 151 (smartphone screen) of the transmitting-side terminal 10 and the display unit 251 (electronic medical record) of the receiving-side terminal 20, the processor returns.

[0102] As described above, in FIG. 9, for the sake of convenience, it is explained in one flowchart. However, since the terminals in Steps 23·24 and Steps 25·26 are different, the respective processes are executed separately. Also, since communication processing and the like are involved, the term "real time" in Step 23 and the term "real time" in Step 26 are different in terms of processing time (there is a time lag). However, as defined above, the term "real time" in this specification is a concept that includes cases where the processing times are different.

[0103] In Step 21, regardless of whether it is the first voice information or the second voice information, the process of acquiring voice information through an input unit such as the voice input unit 142 (microphone) is referred to as "voice acquisition process".

[0104] That is, the "voice information acquisition process" of acquiring voice information as the first voice information or the second voice information based on the user's selection (the above (1)) refers to the voice acquisition process performed according to the user's selection. The selection by the user mentioned here refers to the selection by the user as to whether to acquire the voice as the first voice information (voice information in the transceiver mode) or as the second voice information (voice information in the recording mode) (see the above "user selection acquisition process").

[0105] In step 22, a known method is used as the technique for converting voice information into character information. For example, a technique using a large language model (LLM) (such as generative AI) can be mentioned.

[0106] Also, when converting this voice information into character information, the processor 111 performs the conversion into character information while referring to the technical term dictionary database D30 provided in the storage unit 12. Thereby, incorrect conversion of uncommon technical terms can be prevented.

[0107] That is, the storage unit 12 of the transmitting-side terminal 10 is provided with a technical term dictionary database, and the processor 111 executes a technical term conversion process of converting the technical terms included in the voice information into character information based on the technical term dictionary database.

[0108] In step 24, the processor 211 of the receiving-side terminal 20 receives the character information by short-range communication. The process of the transmitting-side terminal 10 (processor 111) at this time is referred to as "short-range communication transmission process", and the process of the receiving-side terminal 20 (processor 211) is referred to as "short-range communication reception process".

[0109] Also, the process of performing short-range communication is referred to as "short-range communication process". That is, the short-range communication process includes the short-range communication transmission process and / or the short-range communication reception process.

[0110] Furthermore, in the present embodiment, the processor 211 acquires character information through short-range communication reception processing. That is, the processor 211 acquires character information obtained by converting the voice information acquired as the first voice information or the second voice information based on the user's selection. The processing on the side of this processor 211 is referred to as "character information acquisition processing".

[0111] Here, the "character information acquisition processing" which is the processing by the processor 211, that is, "acquiring the character information obtained by converting the voice information acquired as the first voice information or the second voice information based on the user's selection (the above (6))" means acquiring the character information obtained by the processing by the processor 111, that is, "voice information acquisition processing of acquiring voice information as the first voice information or the second voice information based on the user's selection (the above (1)), and information conversion processing of converting the voice information into character information (the above (2))". In the present embodiment, the processor 211 does not convert voice information into character information.

[0112] Note that in the present embodiment, the processor 111 transmits not only character information but also voice information. Data will be described in the data section.

[0113] At the position of step 25, the initial position of the cursor in the present embodiment is, in principle, the end of the symptom column and is predetermined. Since the most frequently used mode is entry into the symptom column, there is an advantage that input is thereby accelerated.

[0114] In step 26, the processing of displaying the character information obtained by converting the voice information is referred to as "character information display processing". The character information conversion processing includes transmission-side character information display processing and / or reception-side character information display processing.

[0115] Here, the voice character conversion system 1 performs conversion of the character code according to the application program of the electronic medical record for character information. For example, in the present embodiment, the processor 211 converts the acquired character information into "Unicode (UTF-8)" and attaches it to the electronic medical record. More precisely, the code point assigned in Unicode according to the character is encoded in UTF-8 format and used as the character information for attachment.

[0116] For example, by once pasting the converted voice information into the clipboard and using the character information, it is possible to support many electronic medical records.

[0117] Since the character codes used by applications related to electronic medical records (hereinafter referred to as "electronic medical record applications") differ depending on the manufacturer, if the character information created on the transmitting terminal 10 is directly pasted, character garbling will occur. Therefore, by performing the above-described processing, the processor 211 can perform correct display without character garbling regardless of the electronic medical record application where the character output destination is located.

[0118] The processor 211 automatically performs these processes, but it may also be made selectable by the user. For example, it may be possible to select a character code in the application P20. As selectable character codes, for example, ASCII, JIS, Shift-JIS, EUC, etc. can be selected. Note that what is described here is an example, and known ones such as EUC-JP, UTF-8, UTF-16, UTF-16LE, UTF-16BE, etc. can be used.

[0119] To summarize, the processor 211 executes a character code conversion process for converting the character code in order to conform to the display format of the output destination application program.

[0120] <2-3. Character Conversion Improvement Process> The processor performs a character conversion improvement process based on the character conversion improvement program P13. That is, the character conversion improvement program P13 causes the computer to function as a character conversion improvement means (character conversion improvement unit) (not shown) by executing character conversion improvement processing by the processor.

[0121] In the character conversion improvement processing of the present embodiment, the processor 111 of the transmission-side terminal 10 (1) Regarding the voice information in the voice character information correspondence database (described later), the character information corresponding to the voice information, and further the information related to the correctness thereof as learning data, for the machine learning model that learns the correspondence between the voice information and the character information, (2) Uses the voice information as an input and obtains the character information corresponding to the voice information as an output.

[0122] Further, the processor 111 displays the character information obtained as the output of the machine learning model on the display unit 151 and transmits it to the reception-side terminal 20.

[0123] Here, the information related to the correctness is, for example, teacher data that takes the incorrect part as it is as an error, teacher data that corrects the incorrect part to a correct answer, and the like.

[0124] These processes are processes related to learning and inference of machine learning. In the present embodiment, the machine learning itself is performed on a terminal other than the transmission-side terminal 10 and the reception-side terminal 20 (for example, a machine learning server or the like), and the machine learning model for character conversion improvement processing obtained by learning is stored in and used by the transmission-side terminal 10.

[0125] <2-4. Speaker identification display processing> The processor performs speaker identification display processing based on the speaker identification display program P14. That is, the speaker identification display program P14 causes the computer to function as a speaker identification display means (speaker identification display unit) (not shown) by executing speaker identification display processing by the processor.

[0126] The speaker identification display process is a process of identifying a speaker and changing the display mode on the display unit for each identified speaker. In the present embodiment, at the stage of the information conversion process, a speaker is identified (such a process is referred to as "speaker identification process"), and for example, on the display unit 251 of the receiving terminal 20, the display mode is changed for each identified speaker. Thereby, even after the voice information is converted into character information, it is possible to make it easier to read who made the speech.

[0127] Also, in the present embodiment, for such distinguishable display, when identifying a speaker, at least a machine learning model using the speaker's voice as learning data is used. For such a machine learning model, for example, a machine learning model having a cepstrum coefficient vector sequence as a feature amount can be used.

[0128] In the speaker identification display process of the present embodiment, the processor 111 of the transmitting terminal 10 (1) With respect to a machine learning model that is learning the relationship between a speaker and voice information using voice information labeled for each speaker as learning data, (2) Using the speaker's voice as an input, the speaker name of the voice is obtained as an output.

[0129] Also, the processor 111 displays the speaker name obtained as an output of the machine learning model on the display unit 151 and also transmits it to the receiving terminal 20.

[0130] The label referred to here is, for example, the name of the speaker. That is, the voice information of a certain person and the name of that person become learning data.

[0131] These processes are processes related to learning and inference of machine learning. In the present embodiment, the machine learning itself is performed on a terminal other than the transmitting terminal 10 and the receiving terminal 20 (for example, a machine learning server, etc.), and the machine learning model for the speaker identification process obtained by learning is stored in the transmitting terminal 10 and used.

[0132] The speaker identification display process enables the identification of a doctor's voice and a patient's voice in the input field (symptom field) of the electronic medical record on the receiving terminal 20 and displays them in a distinguishable manner (for example, different colors). For this purpose, a machine learning model for speaker identification processing may use, for example, at least the doctor's voice and the voice of a third party as learning data. For example, by inputting the voice of a certain patient, it is inferred which patient the voice belongs to and is explicitly indicated on the display unit 251.

[0133] 3. Data Hereinafter, the data handled by the voice character conversion system 1 of the present embodiment will be described with reference to the drawings. The voice character conversion system 1 includes a voice character conversion database D1. The data is stored in the storage unit (data storage unit 12b or data storage unit 22b) of the transmitting terminal 10 or the receiving terminal 20. The voice character conversion database D1 of the present embodiment includes a character information database D10, a patient database D20, and a technical term dictionary database D30.

[0134] In the present embodiment, the transmitting terminal 10 includes the character information database D10 and the technical term dictionary database D30 in the storage unit 12. Also, the storage unit 22 of the receiving terminal 20 includes the patient database D20. Both the transmitting terminal 10 and the receiving terminal 20 may include them.

[0135] The character information database D10 is a database including data (character data) related to character information.

[0136]

Table 1

[0137] Table 1 exemplifies the data included in the character information database D10. As shown in Table 1, the character information database D10 of the present embodiment includes, as items, a voice character ID, a transmission date and time, a voice information file name, a character information file name, and a type.

[0138] The voice character ID is a unique character string assigned to each piece of character information. The transmission date and time is the transmission date and time of the voice information and / or the character information. It is data related to the above-mentioned transmission log. The voice information file indicates the name of the voice information file storing the voice information. That is, the data storage unit 12b includes the voice information file indicated by the voice information file name. The voice information file may be of a known file format as appropriate. For example, AIFF, MP3, FLAC, WAVE, AAC (registered trademark), etc. The character information file indicates the name of the character information file storing the character information. That is, the data storage unit 12b includes the character information file indicated by the character information file name. The character information file may be of a known file format as appropriate. For example, a txt file, etc.

[0139] In the present embodiment, since the date and character information are stored in association with each other, the user can search for document information (medical records) by the date and characters (words, etc.).

[0140] Briefly, in the present embodiment, the transmission-side terminal 10 converts voice information into character information, and the storage unit 12 of the transmission-side terminal 10 includes a database (referred to as a "voice-character information correspondence database") that stores the voice information and the character information corresponding to each piece of voice information in association with each other.

[0141] The patient database D20 is a database including data related to patients (patient data). In the present embodiment, due to security reasons, the patient database D20 is stored only in the storage unit 22 (data storage unit 22b) of the reception-side terminal 20.

[0142]

Table 2

[0143] Table 2 exemplifies the data included in the patient database D20. As shown in Table 2, the patient database D20 of the present embodiment includes, as items, a patient ID, personal information of the patient (name, contact information, etc.), information related to diseases such as the admission date and disease name, and a voice character ID. The patient ID is a unique character string assigned to each patient. The voice character ID is data for associating a certain patient with the above character data. For example, two voice character IDs are associated with Mr. Suzuki ○ro with patient ID 000001, indicating that there are two pieces of character information in this system.

[0144] By including patient information in the patient database D20, for example, when a doctor speaks to the voice input unit 142 of the transmitting terminal 10 such as "Suzuki ○ro" or "Mr. ○ro", the medical record of that person is called up on the receiving terminal 20 and enters the input waiting state. That is, it is possible to call (activate) the input destination (electronic medical record application) of the character information on the receiving terminal 20 by voice.

[0145] Furthermore, the voice character conversion system 1 includes the above-mentioned specialized term dictionary database D30 in the storage unit 12 (data storage unit 12b) of the transmitting terminal 10. The specialized term dictionary database D30 of the present embodiment is a database in which voice information and character information related to medical specialized terms are stored in an associated form (not shown). The specialized term dictionary database D30 has the advantage that the voice input by a user such as a doctor is correctly converted into character information.

[0146] If there is an error in the character information obtained by converting the voice information, the user can correct it on the receiving terminal 20. In the case of such a correction, the processor 211 stores the voice information in association with the corrected character information. Since the specialized terminology dictionary database D30 accumulates more information as the voice character conversion system 1 is used, it is possible to provide a voice character conversion system 1 with less misrecording.

[0147] For example, the processor 111 receives a voice input related to a certain word, refers to the character string in the specialized terminology dictionary database D30, and selects an appropriate term.

[0148] In addition to the above, data such as the voice data of a doctor and data related to a machine learning model using the doctor's voice data as learning data may be included.

[0149] With the configuration as described above, the database of the present embodiment not only stores character information, but also stores voice information associated with the character information. Therefore, the character information that has been incorrectly converted can be corrected later. As a result, a machine learning model can be created using the parts where the voice information and the character information match and do not match as learning data, and by using the machine learning model, high-precision voice character conversion can be performed. Medical terms have high uniqueness and are insufficient, for example, in learning data for daily conversations. Therefore, such a data structure has the advantage that it can be particularly useful for improving the machine learning model.

[0150] 4. Hardware Configuration As shown in FIG. 1, the voice character conversion system 1 in the present embodiment includes a transmission-side terminal 10 and a reception-side terminal 20. Software (app P10 and app P20) for operating the voice character conversion system 1 according to the present embodiment is installed in the transmission-side terminal 10 and the reception-side terminal 20, respectively, and various processes are executed by the functions of the software. The app P10 and the app P20 function as a part of the voice character conversion program P1. Hereinafter, each hardware will be described.

[0151] <Transmitting-side terminal 10> The transmitting-side terminal 10 is an information processing device for a user to utilize the voice-to-character conversion system 1, and is equipped with the application P10. In particular, as an essential function, the transmitting-side terminal 10 is equipped with a function of acquiring voice information through the voice input unit 142.

[0152] In this embodiment, the transmitting-side terminal 10 is a smartphone. However, the transmitting-side terminal 10 is not limited to this, and it may be a portable terminal other than a smartphone, such as a tablet, or it may not be a portable terminal as long as it is equipped with necessary devices such as the voice input unit 142 (for example, a desktop personal computer, etc.). However, from the perspective of convenience, the transmitting-side terminal 10 is preferably a portable terminal.

[0153] In FIG. 1, one transmitting-side terminal 10 is illustrated per user, but the number is not limited to this.

[0154] FIG. 10 is a hardware configuration diagram of the transmitting-side terminal 10. As shown in FIG. 10, the transmitting-side terminal 10 includes a control unit 11, a storage unit 12, and a communication control unit 13. The control unit 11 further includes a processor 111, a ROM 112, a RAM 113, and a timing unit 114. In this embodiment, the processor 111 is a CPU (Central Processing Unit). The basic functions of each will be summarized and explained later.

[0155] The control unit 11 including the processor 111 and the control unit 22 including the processor 211 described later constitute and function as a voice-to-character conversion unit (not shown). The voice-to-character conversion unit executes the voice-to-character conversion program P1 to perform voice-to-character conversion processing.

[0156] Also, one program may include another program. For example, in this embodiment, the voice-to-character conversion program P1 includes a voice selection acquisition program P11, a conversion display program P12, and the like.

[0157] As shown in FIG. 10, the storage unit 12 includes a program storage unit 12a and a data storage unit 12b, and stores programs and data necessary for various processes. For example, in the program storage unit 12a, in addition to the application P10 according to the present embodiment, a control program for controlling devices connected to the transmission-side terminal 10, such as a communication control program for controlling the communication control unit 13, is stored. For example, when the user activates the application P10, various processes such as a process for allowing the user to select a communication destination are executed.

[0158] In the present embodiment, the application P10 is installed in the reception-side terminal 20 through Internet communication or a storage medium.

[0159] As shown in FIG. 10, the communication control unit 13 is a device that communicates with an external terminal.

[0160] In particular, in the present embodiment, the communication control unit 13 includes a transmission-side short-range communication unit 131 and performs short-range communication with the reception-side terminal 20. Also, the transmission-side terminal 10 and the reception-side terminal 20 correspond one-to-one, such that one reception-side terminal 20 for displaying character information corresponds to one of the transmission-side terminals 10 that acquires voice.

[0161] The transmission-side short-range communication unit 131 is a device that uses a communication method defined by IEEE802.15 (for example, Bluetooth (registered trademark), BLE). That is, the transmission-side short-range communication unit 131 is a device for performing peer-to-peer communication or broadcast communication with the reception-side terminal 20. Generally, the reception-side terminal 20 of the present embodiment is referred to as the parent, and the transmission-side terminal 10 is referred to as the child.

[0162] This is distinguished from communication via a network such as the Internet. Communication via a network is, for example, a communication method defined by IEEE802.3 or IEEE802.5 in the case of wired, or a wireless communication method (so-called Wi-Fi) defined by IEEE802.11 in the case of wireless.

[0163] That is, communication by the transmission-side short-distance communication unit 131 has the advantage that the transmission-side terminal 10 can communicate with the reception-side terminal 20 not connected to the network. That is, terminals in physically distant locations cannot be connected, and since they do not connect to a communication network to which unspecified terminals are connected like the Internet, high security can be maintained. Also, communication by the transmission-side short-distance communication unit 131 has the advantage that it is advantageous for carrying the terminal (transmission-side terminal 10) because connection is possible without a physical cable.

[0164] The input unit 14 is a device that receives input from the user. The input unit 14 of the present embodiment includes an input operation unit 141 and a voice input unit 142.

[0165] The input operation unit 141 is a functional unit that receives the input operation of the user. In particular, in the present embodiment, the processor 111 displays at least the input operation unit 141 for receiving the operation of the user and starting voice input on the display unit 151.

[0166] The input operation unit 141 in the above-described embodiment becomes the voice acquisition button UI-111 in the transceiver mode, and becomes the voice acquisition start / end button UI-121 in the recording mode.

[0167] The transmission-side terminal 10 of the present embodiment is a smartphone and includes a touch panel as the input unit 14. The input operation unit 141 is a UI (such as UI-10a, UI-10b, UI-111, or UI-121) such as icons and buttons arranged on the surface of the display unit 151 (smartphone screen).

[0168] In the present embodiment, the voice input unit 142 is a microphone built into the smartphone. The voice input unit 142 converts voice into an electrical signal (voice signal, voice data, voice information). The processor 111 acquires voice information that has been converted into an electrical signal through the voice input unit 142. For simplicity, there may be cases where the process of converting voice into an electrical signal is omitted from the description.

[0169] In addition to the above, the transmitting terminal 10 may be provided with devices that are additionally necessary for the purposes of this embodiment, or devices for improving convenience with respect to the purposes of this embodiment.

[0170] Briefly, the transmitting terminal 10 includes · a voice information acquisition unit that acquires voice information as first voice information or second voice information based on a user's selection, · an information conversion unit that converts the voice information into character information, · a transmitting-side character information display unit that displays the character information in real time on a display unit (of the transmitting terminal 10), and · a short-distance communication transmitting unit that transmits the acquired character information to a terminal (receiving terminal 20) that displays the character information in real time by short-distance communication.

[0171] <Receiving terminal 20> The receiving terminal 20 is an information processing device for a user to use the voice character conversion system 1, and includes the application P20. In addition, in this embodiment, the receiving terminal 20 includes, separately from the application P20, a commercially available application program related to an electronic medical record (electronic medical record application). In particular, as an essential function, the receiving terminal 20 has a function of displaying in real time the character information obtained by converting voice information.

[0172] In this embodiment, the receiving terminal 20 is a desktop-type or laptop-type personal computer (PC). However, the receiving terminal 20 is not limited to this, and may be a portable terminal such as a tablet. However, due to security reasons, the receiving terminal 20 is preferably a desktop PC that is difficult to carry out.

[0173] In FIG. 1, only one receiving terminal 20 is shown per user, but the number is not limited to one, and it may be realized by a plurality of computers.

[0174] FIG. 11 is a hardware configuration diagram of the receiving terminal 20. As shown in FIG. 11, the receiving terminal 20 includes a control unit 21, a storage unit 22, a communication control unit 23, an input unit 24, and an output unit 25. The control unit 21 further includes a processor 211, a ROM, a RAM, and a timing unit. In this embodiment, the processor 211 is a CPU (Central Processing Unit). Descriptions of content that overlaps with the described content and those related to basic functions to be described later are omitted.

[0175] The program storage unit 22a of the receiving terminal 20 stores (installs) the application P20 according to this embodiment, and the processor 211 executes various processes according to the functions of the software.

[0176] The various processes include outputs (such as screen displays) based on information obtained from the transmitting terminal 10, reception of user inputs, or various communications. For example, when the user starts the application P20, it resides in the task bar and starts receiving characters.

[0177] In this embodiment, the application P20 is installed in the receiving terminal 20 through Internet communication or a storage medium.

[0178] As shown in FIG. 11, the communication control unit 23 is a device that communicates with an external terminal.

[0179] In particular, in this embodiment, the communication control unit 23 includes a transmission-side short-range communication unit 231 and performs short-range communication with the transmission-side terminal 10. Since the standard of the transmission-side short-range communication unit 231 is as described above, it is omitted.

[0180] In this embodiment, the display unit 251 is a display attached to a PC. In particular, the acquired character information is displayed on the electronic medical record being displayed on the display unit 251.

[0181] Briefly, the receiving-side terminal 20 includes a short-range communication receiving unit that performs short-range communication with a terminal (transmission-side terminal 10) that acquires voice information, a character information acquisition unit that acquires character information obtained by converting voice information acquired as first voice information or second voice information based on a user's selection, a receiving-side character information display unit that displays the character information on the display unit of the receiving-side terminal 20 in real time. It is provided with.

[0182] The voice character conversion system 1 may include devices other than those described above. For example, a terminal for machine learning (machine learning server, data storage, etc.) may be separately provided.

[0183] (Explanation regarding the basic functions of a computer) Hereinafter, the control unit (processor, ROM, RAM, timer unit), storage unit, communication control unit, input unit, and output unit will be described. In addition, except for the short-range communication described above, the connection mode (network topology) between functional units is not particularly limited in any of the terminals of this embodiment. For example, it may be a bus type, a star type, a mesh type, or the like.

[0184] The processor performs information processing and control of various devices according to programs stored in the ROM, storage unit, etc. In this embodiment, the processor is a CPU (Central Processing Unit).

[0185] However, the processor is not limited to the CPU. Various processors such as a CPU, a DSP (Digital Signal Unit), a GPU (Graphics Processing Unit), a GPGPU (General Purpose computing on GPU), an ASIC (Application Specific Integrated Circuit), or an FPGA (Field Programmable Gate Array) may be used alone or in combination. For example, a processor integrating a CPU and a GPU is called an APU (Accelerated Processing Unit) or the like, and such a processor may be used.

[0186] The ROM is a read-only memory in which various programs and data for the processor to perform various controls and operations are stored in advance.

[0187] The RAM is a random access memory used as a working memory for the processor. Various areas for performing various processes of this embodiment can be secured in this RAM.

[0188] The timing unit performs timing processes related to acquisition of time information and the like. When the computer includes a communication control unit, time information may be acquired from the outside by NTP (Network Time Protocol).

[0189] The storage unit is a device for storing information such as programs and data. The storage unit is also referred to as a storage. The storage unit may be an internal type or an external type.

[0190] The storage unit includes a storage medium capable of reading and writing data and a drive for reading and writing to the storage medium. Examples of the storage medium include an internal type and an external type, such as an HD (hard disk), a CD-ROM, and a flash memory. Examples of the drive include an HDD (hard disk drive), an SSD (solid state drive), and the like.

[0191] The memory unit includes a program storage unit and a data storage unit as functional units. The program storage unit stores control programs for controlling various devices, such as a communication control program for controlling communication.

[0192] The communication control unit is a device for performing communication between terminals and the like. The communication control unit connects the terminal including the communication control unit to a network for Internet communication.

[0193] The communication method of the communication control unit is a known method, and a wired method or a wireless method is applied according to the device. For example, if the terminal is a desktop PC, both wired and wireless cases are conceivable. If the terminal is a smartphone, a wireless communication method is conceivable.

[0194] If it is wired, for example, a communication method defined by IEEE802.3 (for example, bus-type or star-type wired LAN) can be preferably used. However, in addition, a communication method defined by IEEE802.5 (for example, ring-type wired LAN) can also be used.

[0195] If it is wireless, for example, a communication method defined by IEEE802.11 (for example, Wi-Fi) can be preferably used. However, in addition, IEEE802.15 (for example, Bluetooth (registered trademark), BLE (Bluetooth (registered trademark) low energy), etc.), IEEE802.16 (for example, WiMAX), ZigBee (registered trademark), 920MHz band wireless (Wi-SUN, etc.), or a communication method defined by optical communication such as infrared communication can also be used.

[0196] The input unit and the output unit are devices responsible for input to and output from the terminal, respectively. The input unit and the output unit may be collectively referred to as the input / output unit. The input unit is a device that receives input from a user. Examples of such input units include a keyboard, a mouse as a pointing device, a trackpad, a tablet, or a touch panel.

[0197] When the terminal is a tablet, a smartphone, etc., and the input unit is a touch panel, the input unit is arranged on the surface of a display unit such as a touch screen that displays images, etc. In this case, the input unit identifies the touch position of the user corresponding to various operation icons displayed on the display unit and receives the input from the user.

[0198] The output unit is, for example, a device for outputting images, sounds, documents, etc. Examples of the output unit include display devices such as touch screens and displays (liquid crystal displays and organic EL displays), voice output devices such as speakers, and document output devices such as printers.

[0199] With the above configuration, since the transmitting terminal 10 and the receiving terminal 20 are separated, a user such as a doctor can bring the transmitting terminal 10 into the medical site. For example, when a doctor makes rounds in a patient's hospital room, etc., the doctor can carry the transmitting terminal 10 and input voice while listening to the patient's condition. After the rounds, when the doctor returns to the examination room or living room where the receiving terminal 20 is located and opens the screen, the input voice becomes character information and is input into the medical record for each patient, which can reduce the input burden on the doctor. Since the transmitting terminal 10 and the receiving terminal 20 communicate via short - distance communication and do not go through the Internet, there is an advantage that highly confidential personal information such as the patient's medical condition can be prevented from being leaked due to wiretapping, etc.

[0200] (Other embodiments) So far, the example of the voice - to - text conversion system 1 being used in a medical institution has been described, but other usage examples will be described.

[0201] The voice - to - text conversion system 1 may be used, for example, in an educational setting. In this case, the output destination of the character information will be a log created by the teacher instead of the electronic medical record. Since the voice information in schools and the like may include personal information of children and students, the voice-to-character conversion system 1 is preferably used.

[0202] In addition, the voice-to-character conversion system 1 may be used at a construction site. For example, when transmitting the structural evaluation of a building at the site from the transmitting terminal 10 such as a smartphone to the receiving terminal 20 installed in the office. When the evaluation at the site particularly includes technical secrets and the like, the voice-to-character conversion system 1 is preferably used.

[0203] (Modification example) The present invention is not limited to the above-described embodiments, and includes those obtained by making various changes to the above-described embodiments without departing from the gist of the present invention.

[0204] In the above-described embodiment, it is assumed that both the application P10 of the transmitting terminal 10 and the application P20 of the receiving terminal 20 are activated, but it is not limited thereto. That is, when the application P20 of the receiving terminal 20 (or the application program related to the electronic medical record) is not activated, the application P10 of the transmitting terminal 10 may store the acquired various information in the storage unit and perform necessary processing when the application P20 of the receiving terminal 20 is activated. Here, the storage unit mentioned here may be the storage unit of an external device in addition to the control unit 12, and for example, if communication is connected, it may be the storage unit of a server or the like.

[0205] In the above-described embodiment, the transmitting terminal 10 transmits voice information and character information to the receiving terminal 20 at the same time, but the transmitting terminal 10 may transmit only the character information. In this case, the voice information may not be sent, or may be sent at another timing. At this time, the data stored in the voice selection acquisition process also becomes the transmission log of character information instead of voice information. In this case, there is an advantage that the amount of transmission data to be transmitted in real time is reduced.

[0206] In the above-described embodiment, the conversion process from voice information to character information was performed by the transmitting terminal 10, but it is not limited to this, and the conversion process from voice information to character information may be performed by the receiving terminal 20. In this case, the flow of the above-described conversion display process changes. That is, the display of character information on the transmitting terminal 10 (step 23) is eliminated, and the order of the processes after voice acquisition also changes. For example, after the processor 111 of the transmitting terminal 10 acquires voice, it transmits the voice information to the receiving terminal 20 (by short-distance communication). Then, the processor 211 that acquires information (by short-distance communication) converts the voice information into character information and displays it after acquiring the display position of the character information. Also in this case, the transmission log is a transmission log of voice information, and the process related to speaker identification is also executed by the receiving terminal 20. However, when the conversion process from voice information to character information is performed by the transmitting terminal 10, there is an advantage that the character information can be displayed on the display unit 151 (smartphone screen), and the user inputting the voice information can confirm the character information more quickly. It is also conceivable to convert the voice information into character information at the receiving terminal 20 and then transmit it to the transmitting terminal 10, but in this case, the number of communications between the terminals increases. Therefore, it is considered faster to perform voice-character conversion at the transmitting terminal 10 and display the character information on the transmitting terminal 10.

[0207] In addition to the above-described embodiment, the voice-character conversion system 1 may include a relay device that relays communication between the transmitting terminal 10 and the receiving terminal 20. In this case, in order to maintain communication security, it is preferable that the relay device is independent from outside the voice-character conversion system 1 (does not communicate with terminals other than the terminals within the voice-character conversion system 1). In this case, there is an advantage that the communication distance is extended.

[0208] Aspects of the present invention including this embodiment have the following features, in other words. The following corresponds to the scope of the claims at the time of filing of this application. However, it may differ from the description of the scope of the claims after the amendment of the scope of the claims after filing. (1) In a first aspect, there is provided a voice-to-character conversion system comprising a first terminal that acquires voice information and a second terminal that displays in real time character information obtained by converting the voice information. The first terminal includes a voice standby display unit that displays on a display unit of the first terminal that voice acquisition is possible, a voice information acquisition unit that acquires voice information of a user, and a short-distance communication transmission unit that transmits the voice information and / or the character information obtained by converting the voice information by short-distance communication. The second terminal includes a short-distance communication reception unit that receives the voice information and / or the character information obtained by converting the voice information by short-distance communication, and a reception-side character information display unit that displays the character information in real time on a display unit of the second terminal. At least one of the first terminal or the second terminal includes an information conversion unit that converts the voice information into character information. (2) In a second aspect, there is provided a voice-to-character conversion system comprising a first terminal that acquires voice information and a second terminal that displays in real time character information obtained by converting the voice information. The first terminal includes a voice standby display unit that displays on a display unit of the first terminal that voice acquisition is possible, a voice information acquisition unit that acquires voice information of a user, an information conversion unit that converts the voice information into character information, and a short-distance communication transmission unit that transmits the character information by short-distance communication. The second terminal includes a short-distance communication reception unit that receives the character information by short-distance communication, and a reception-side character information display unit that displays the character information in real time on a display unit of the second terminal. In this case, since the voice information is converted into character information by the first terminal on the transmission side, the character information can be displayed on the first terminal side. (3) In the third aspect, the voice information acquisition unit that acquires the user's voice information is a voice information acquisition unit that acquires voice information as first voice information or second voice information based on the user's selection. The first voice information is voice information obtained by acquiring the voice input to the voice input unit while the input operation unit is receiving the user's operation. The second voice information is voice information obtained by starting the acquisition of the voice input to the voice input unit while the input operation unit is receiving the user's operation. Provided is the voice character conversion system according to the first or second aspect, characterized in that. In this case, since the recording method can be changed arbitrarily by the user, the convenience for the user is significantly improved. In the above-described embodiment, since each person may prefer either the transceiver mode or the recording mode, the hurdle for introduction can be lowered by covering such individual preferences. (4) In the fourth aspect, the first terminal further includes a volume-driven voice acquisition start unit that starts voice acquisition when a state in which the voice input unit detects a voice equal to or greater than a predetermined first volume continues for a predetermined first time or longer even when there is no operation of the input operation unit by the user. Provided is the voice character conversion system according to the first or second aspect, characterized in that. In this case, there is an advantage of preventing forgetting to record. In particular, since examinations and the like cannot be repeated, preventing forgetting to record improves the convenience for the user. (5) In the fifth aspect, the first terminal further includes a volume-driven voice acquisition stop unit that stops voice acquisition when a state in which the voice input unit does not detect a voice equal to or greater than a predetermined second volume continues for a predetermined second time or longer after voice acquisition starts. Provided is the voice character conversion system according to the first or second aspect, characterized in that. In this case, there is an advantage of preventing forgetting to stop recording. Preventing forgetting to stop recording reduces wasteful use of files. (6) In the sixth aspect, further provided is the voice-to-character conversion system according to the first or second aspect, characterized in that a storage unit of the first terminal or the second terminal further includes a technical term dictionary database, and a technical term conversion unit for converting technical terms included in the voice information into character information based on the technical term dictionary database. In this case, for example, even when a highly specialized conversation is made in the medical field or the like, the convenience for the user is improved in order to correctly display the character information from the voice information. (7) In the seventh aspect, further provided is the voice-to-character conversion system according to the first or second aspect, characterized in that the second terminal further includes a character code conversion unit for converting a character code to conform to the display format of the output destination application program. In this case, for example, when outputting to a plurality of electronic medical records with different manufacturers, even when the output destination application programs are different, there is an advantage that garbled characters are less likely to occur. (8) In the eighth aspect, further provided is the voice-to-character conversion system according to the first or second aspect, characterized in that the first terminal further includes a transmission log storage unit for storing a transmission log including the transmission date and time information of the short-distance communication. In this case, since the date and time of the voice can be accurately stored, there is an advantage that the voice information and the character information can be confirmed later. Also, at the time of information search, date and time search and the like can be performed. (9) In the ninth aspect, further provided is the voice-to-character conversion system according to the first or second aspect, characterized in that at least a storage unit of either the first terminal or the second terminal includes a voice-character information correspondence database for associating and storing the voice information and the character information corresponding to each voice information, and (1) using the voice information and the character information corresponding to the voice information in the voice-character information correspondence database, and further information related to its correctness as learning data, for a machine learning model that learns the correspondence relationship between the voice information and the character information, and (2) a character conversion improvement unit for obtaining, with the voice information as an input and the character information corresponding to the voice information as an output. In this case, even if the input voice includes technical terms, there is an advantage that the accuracy of voice recognition is improved. Also, for example, even when the voice is difficult to hear due to the speaker's active tongue or the like, the recognition accuracy is improved. (10) In a tenth aspect, further, at least one of the first terminal or the second terminal includes a speaker identification display unit that identifies a speaker and changes a display mode on the display unit for each identified speaker. The speaker identification display unit: (1) uses voice information labeled for each speaker as learning data to learn the relationship between the speaker and the voice information for a machine learning model; (2) takes the voice of the speaker as input and obtains the name of the speaker of the voice as output. Provided is a voice character conversion system according to the first or second aspect. In this case, since the speaker is clearly displayed, there is an advantage that usability is improved. (11) In an eleventh aspect, a voice standby display unit that displays on the display unit that voice acquisition is possible, a voice information acquisition unit that acquires voice information as first voice information or second voice information based on a user's selection, an information conversion unit that converts the voice information into character information, a transmission-side character information display unit that displays the character information on the display unit in real time, and a short-distance communication transmission unit that transmits the acquired character information to a terminal that displays the character information in real time by short-distance communication. The first voice information is voice information obtained by acquiring the voice input to the voice input unit while the input operation unit is receiving a user's operation. The second voice information is voice information obtained by starting the acquisition of the voice input to the voice input unit while the input operation unit is receiving a user's operation. Provided is a voice character conversion system. This relates to a first terminal (for example, a smartphone). (12) In the 12th aspect, a short - distance communication receiving unit that performs short - distance communication with a terminal for acquiring voice information and receives information, a character information acquisition unit that acquires character information obtained by converting voice information acquired as first voice information or second voice information based on a user's selection, and a receiving - side character information display unit that displays the character information on a display unit in real time are provided. The first voice information is voice information obtained by acquiring the voice input to the voice input unit while the input operation unit is receiving a user's operation. The second voice information is voice information obtained by starting the acquisition of the voice input to the voice input unit while the input operation unit is receiving a user's operation. A voice - character conversion system is provided, which is characterized by the above. This relates to a second terminal (for example, a PC). (13) In the 13th aspect, a first terminal for acquiring voice information functions as a voice standby display means for displaying on the display unit of the first terminal that voice acquisition is possible, a voice information acquisition means for acquiring a user's voice information, and a short - distance communication transmission means for transmitting the voice information and / or character information obtained by converting the voice information by short - distance communication. A second terminal for displaying in real time the character information obtained by converting the voice information functions as a short - distance communication reception means for receiving the voice information and / or character information obtained by converting the voice information by short - distance communication, and a reception - side character information display means for displaying the character information on the display unit of the second terminal in real time. At least one of the first terminal or the second terminal is provided with information conversion means for converting the voice information into character information. A voice - character conversion program is provided, which is characterized by the above.

Industrial Applicability

[0209] Since the transmitting - side terminal 10 is portable, it is particularly effective when it is desired to perform simple speech - to - text conversion by carrying the transmitting - side terminal 10 around in the same indoor area, etc. In particular, since it does not go through the Internet, it can be applied to applications where it is desired to prevent the leakage of confidential information such as personal information due to wiretapping, man - in - the - middle attacks, etc.

Explanation of Signs

[0210] 1 Voice-to-Text Conversion System 10 Transmitting Terminal (Smartphone (Smartphone)) 11 Control Unit 111 Processor 112 ROM 113 RAM 114 Timing Unit 12 Memory Unit 12a Program Storage Unit 12b Data Storage Unit 13 Communication Control Unit 131 Transmitting Short-Range Communication Unit 14 Input Unit 141 Input Operation Unit (Button) 142 Voice Input Unit (Microphone) 15 Output Unit 151 Display Unit (Smartphone Screen) 20 Receiving Terminal 21 Control Unit 211 Processor 22 Memory Unit 22a Program Storage Unit 22b Data Storage Unit 23 Communication Control Unit 231 Receiving Short-Range Communication Unit 24 Input Unit 25 Output Unit 251 Display Unit (Display, Electronic Medical Record Display Unit) UI-10 Transmitting Terminal Screen UI-10a Connection Destination Display Icon UI-10b Voice Acquisition Method Selection Icon UI-110 Transceiver Mode Screen UI-111 Voice Acquisition Button UI-111a Voice Acquisition Button (During Operation) UI-112 Character Information Display Section UI-113 Audio Spectrum UI-120 Recording Mode Screen UI-121 Voice Acquisition Start / End Button UI-20 Receiving Terminal Screen P1 Voice-to-Text Conversion Program P10 Transmitting Terminal Side Application Program (App P10) P11 Voice Selection Acquisition Program P12 Conversion Display Program P13 Character Conversion Improvement Program P14 Speaker Identification Program P20 Receiving Terminal Side Application Program (App P20) D1 Voice-to-Text Conversion Database D10 Character Information Database D20 Patient Database D30 Technical Terminology Dictionary Database

Claims

1. A speech-to-text conversion system comprising: a first terminal that acquires speech information; and a second terminal that displays text information obtained by converting the speech information in real time, The first terminal is a voice standby display unit that displays on a display unit of the first terminal that voice acquisition is possible; a selection screen display unit that displays a screen for allowing a user to select whether to acquire the first voice information or the second voice information; a user selection acquisition unit that acquires a selection by the user and determines whether to acquire the voice information as the first voice information or the second voice information; a voice information acquisition unit that acquires voice information as the first voice information or the second voice information through a voice input unit of a first terminal; an information conversion unit that converts the voice information into text information; and a short-distance communication transmitting unit that transmits the character information by short-distance communication; Equipped with The second terminal, a short-distance communication receiving unit that receives the character information by short-distance communication; and a receiver-side character information display unit for displaying the character information on a display unit of a second terminal in real time; the first voice information is voice information obtained by acquiring a voice input to the voice input unit only while an input operation unit of the first terminal continues to accept an operation by a user, A speech-to-text conversion system characterized in that the second voice information is voice information obtained when an input operation unit of a first terminal accepts a user's operation and starts acquiring the voice input into the voice input unit.

2. The speech-to-text conversion system as described in claim 1, further characterized in that it comprises a sender's character information display unit which displays the character information in real time on the display unit of the first terminal.

3. The speech-to-text conversion system described in Claim 1, characterized in that the selection screen display unit changes the display of the input operation unit on the first terminal depending on the selected speech acquisition means.

4. The speech-to-text conversion system of claim 1, wherein the first terminal further comprises a volume-driven speech acquisition start unit that starts speech acquisition when the speech input unit detects a voice having a volume equal to or greater than a predetermined first volume for a predetermined first period of time or more, even if the user does not operate the input operation unit.

5. The speech-to-text conversion system of claim 1, wherein the first terminal further comprises a volume-driven speech acquisition stop unit that stops speech acquisition when a state in which the speech input unit does not detect speech at or above a predetermined second volume continues for more than a predetermined second time after speech acquisition begins.

6. Furthermore, the storage unit of the first terminal or the second terminal includes a technical term dictionary database; 2. The speech-to-text conversion system according to claim 1, further comprising a technical term conversion unit that converts technical terms included in the speech information into text information based on the technical term dictionary database.

7. 2. The speech-to-text conversion system according to claim 1, wherein the second terminal further comprises a character code conversion unit that converts a character code to conform to a display format of an application program at an output destination.

8. 2. The speech-to-text conversion system according to claim 1, wherein the first terminal further comprises a transmission log storage unit for storing a transmission log including transmission date and time information of the short-distance communication.

9. Furthermore, at least one of the storage units of the first terminal and the second terminal includes a voice / text information correspondence database that stores the voice information and text information corresponding to each of the voice information in association with each other; (1) A machine learning model that learns about the correspondence between the voice information and the character information using, as learning data, the voice information of the voice-character information correspondence database, the character information corresponding to the voice information, and teacher data indicating the accuracy of the correspondence between the voice information and the character information, (2) A speech-to-text conversion system according to claim 1, further comprising a character conversion improvement unit that receives speech information as an input and obtains character information corresponding to the speech information as an output.

10. Furthermore, at least one of the first terminal and the second terminal is provided with a speaker identification display unit that identifies a speaker and changes a display mode on a display unit for each identified speaker, The speaker identification display unit is (1) For a machine learning model that learns the relationship between speakers and speech information using speech information labeled for each speaker as training data, (2) A speech-to-text conversion system according to claim 1, characterized in that a speaker's voice is input and the speaker's name of the voice is obtained as output.

11. A speech-to-text conversion system comprising a terminal for acquiring speech information, the terminal comprising: a voice standby display unit that displays on a display unit that voice acquisition is possible; a selection screen display unit that displays a screen for allowing a user to select whether to acquire the first voice information or the second voice information; a user selection acquisition unit that acquires a selection by the user and determines whether to acquire the voice information as the first voice information or the second voice information; a voice information acquisition unit that acquires voice information as the first voice information or the second voice information through a voice input unit; an information conversion unit for converting the voice information into text information; a transmitter character information display unit that displays the character information on a display unit in real time; and a short-distance communication transmitting unit that transmits the acquired character information to a terminal that displays the character information in real time by short-distance communication; Equipped with the first voice information is voice information obtained by acquiring a voice input to the voice input unit only while the input operation unit continues to accept a user operation, A speech-to-text conversion system characterized in that the second speech information is speech information obtained when an input operation unit accepts a user's operation and starts acquiring the speech input to the speech input unit.

12. A first terminal for acquiring voice information, a voice standby display means for displaying on a display unit of the first terminal a message indicating that voice acquisition is possible; a selection screen display means for displaying a screen for allowing a user to select whether to acquire the first voice information or the second voice information; a user selection acquisition means for acquiring a selection by the user and determining whether to acquire the voice information as the first voice information or the second voice information; a voice information acquiring means for acquiring voice information as the first voice information or the second voice information through a voice input unit of a first terminal; an information conversion means for converting the voice information into text information; and a short-distance communication transmitting means for transmitting the character information by short-distance communication; Function as a a second terminal which displays in real time character information obtained by converting the voice information; a short-distance communication receiving means for receiving the character information by short-distance communication; and a receiver-side character information display means for displaying the character information on a display unit of a second terminal in real time; Function as a the first voice information is voice information obtained by acquiring a voice input to a voice input unit only while an input operation unit of a first terminal continues to accept an operation by a user, A speech-to-text conversion program, characterized in that the second voice information is voice information obtained when an input operation unit of a first terminal accepts a user's operation and starts acquiring the voice input into the voice input unit.

Citation Information

Patent Citations

  • Medical data recorder

    JP1998011520A

  • Character information forming system and character information forming device

    JP2004032275A

  • Home visiting care support system and home visiting care support method

    JP2012073739A

  • Electronic clinical chart preparation apparatus

    JP2015035099A

  • Voice processor and voice processing method

    JP2015184487A

Cited By

  • Audio recording program and audio recording method

    JP7813500B1