Speaker identifying device, speaker identifying method, and program

The speaker identification device addresses the challenge of identifying speakers in simultaneous voice communications by using a voice section information collection, recognition, and display means to associate speaker IDs with text results, improving communication clarity in firefighting and emergency services.

JP2026022159APending Publication Date: 2026-02-12NEC PLATFROMS LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024123588
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-30
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

In simultaneous voice communication systems used by firefighting and emergency services, identifying the speaker for both uplink and downlink voices on a single radio channel is challenging, even after converting voices into text, as the same channel is used for switching between uplink and downlink voices, leading to insufficient speaker identification.

Method used

A speaker identification device that includes a voice section information collection means to detect the start and end of a press, a voice recognition means for recognizing the speaker, and a display means to associate and display the speaker ID with the recognition result text, enabling accurate identification of speakers in simultaneous voice communications.

Benefits of technology

The device effectively identifies the speaker for each voice text in radio channel communications, enhancing the utility of text conversion by associating speaker information with recognition results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026022159000001_ABST
    Figure 2026022159000001_ABST
Patent Text Reader

Abstract

To specify a speaker for each voice text of communication voice in a radio channel.SOLUTION: A speaker identification device that performs speech recognition and identifies a speaker who has uttered a speech, the speaker identification device including a speech segment information collection unit that detects a start of pressing and an end of pressing and outputs a speaker ID when pressing is started, a speech recognition unit that performs speech recognition on speech data collected from the start of pressing to the end of pressing and outputs a recognition result text, and a display unit that displays the recognition result text and a speaker corresponding to the speaker ID associated with the recognition result text.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a speaker identification device, a speaker identification method, and a program. [Background technology]

[0002] In recent years, digital radios that transmit digitally have come into use in firefighting and emergency services. In the simultaneous voice communication of digital radios used in firefighting and emergency services, two-way voice communication (upstream voice / downstream voice) is generated on a single radio channel. Figure 1 shows an image of simultaneous voice communication.

[0003] Here, the upstream voice is the voice from the mobile station to the command console, and the downstream voice is the voice from the command console to the mobile station. More specifically, the speaker of the upstream voice (mobile station) is mainly a fire engine or an ambulance, and the speaker of the downstream voice (command console) is mainly the command console of the fire command center.

[0004] Patent document 1 describes that when controlling calls between communication devices, in order to separate the spoken content that should be the processing report content, a software switching means, a hardware switching means, or a control unit of the terminal device is used to perform a distinguishing process, converts voice into text, and creates history information by adding the time when the processing report content was inputted into voice to the processing report content that has been converted from voice to text. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Japanese Patent Application Laid-Open No. 2013-152613 Summary of the Invention [Problem to be solved by the invention]

[0006] Simultaneous voice communication is communication in which multiple speakers take turns speaking on a single radio channel (one speaker at a time), and the speaker for both the uplink and downlink voices is not a single person. In other words, to make effective use of the radio channel, the same channel (frequency) is used to switch between uplink and downlink voices. For this reason, even if the voices (uplink and downlink voices) of a radio channel are converted into text using a voice recognition system, it is difficult to identify the speaker for each text.

[0007] For example, when fire engine A arrives at the scene of a fire, if it utters "arrived at the scene" to the fire command center via simultaneous voice communication, the speaker cannot be identified, and it is unclear from the text which vehicle has arrived. Therefore, even if the speech of the simultaneous voice communication is converted into text, there is insufficient information after the conversion, and the benefits of converting to text cannot be fully utilized. [Means for solving the problem]

[0008] The speaker identification device according to the present disclosure is a speaker identification device that performs voice recognition and identifies the speaker who spoke the voice, and includes: a voice section information collection means that detects the start and end of a press and outputs a speaker ID when a press starts; a voice recognition means that performs voice recognition on collected voice data from the start to the end of a press and outputs recognition result text; and a display means that displays the recognition result text and speaker information corresponding to the speaker ID associated with the recognition result text.

[0009] Furthermore, the speaker identification method disclosed herein is a speaker identification method that performs voice recognition and identifies the speaker who spoke the voice, detects the start and end of press, outputs a speaker ID when press starts, performs voice recognition on collected voice data from the start to the end of press, outputs recognition result text, and displays the recognition result text and speaker information corresponding to the speaker ID associated with the recognition result text.

[0010] In addition, the program disclosed herein is a program that performs voice recognition and identifies the speaker who spoke the voice, detects the start and end of press, outputs a speaker ID when press starts, performs voice recognition on collected voice data from the start to the end of press, outputs recognition result text, and displays the recognition result text and speaker information corresponding to the speaker ID associated with the recognition result text. [Effects of the Invention]

[0011] In simultaneous voice communications by fire and emergency digital radio, the speaker can be identified for each voice text of the communication voice on the radio channel. [Brief explanation of the drawings]

[0012] [Figure 1] FIG. 1 illustrates an example of voice communication between a command console and a mobile station according to the present disclosure. [Figure 2] FIG. 1 is a diagram illustrating an example of the configuration of a speaker identification device according to the present disclosure. [Figure 3] FIG. 1 is a diagram illustrating an example of the configuration of a speaker identification device according to the present disclosure. [Figure 4] FIG. 1 is a diagram illustrating an example of the configuration of a speaker identification device according to the present disclosure. [Figure 5] FIG. 10 is a diagram showing an example of correspondence between speaker IDs and speakers according to the present disclosure. [Figure 6] FIG. 10 is a diagram illustrating the contents of a recognition result information set queue according to the present disclosure. [Figure 7] 10A and 10B are diagrams illustrating a display of a speaker name and a recognition result according to the present disclosure. [Figure 8] FIG. 1 is a diagram illustrating an example of the configuration of a speaker identification device according to the present disclosure. [Figure 9] FIG. 2 is a diagram illustrating the contents of a speech segment information storage memory according to the present disclosure. [Figure 10] FIG. 1 is a diagram illustrating an example of the configuration of a speaker identification device according to the present disclosure. [Figure 11] FIG. 1 is a diagram illustrating an example of the configuration of a speaker identification device according to the present disclosure. [Figure 12] 1 is a diagram showing a time series of operations of a speaker identification device according to the present disclosure. [Figure 13] FIG. 2 is a diagram illustrating the contents of a speech segment information storage memory according to the present disclosure. [Figure 14] FIG. 10 is a diagram illustrating the contents of a recognition result information set queue according to the present disclosure. [Figure 15] 10A and 10B are diagrams illustrating a display of a speaker name and a recognition result according to the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0013] Embodiment 1 As a premise, the speaker identification device 1 will be described as being applied to an environment in which communication is carried out between a fire command center and a mobile station such as an ambulance via a base station, as shown in Figure 1, and in which simultaneous voice communication is carried out.

[0014] In other words, communication is performed between a firefighting command center and a mobile station such as an ambulance via a base station. Also, within the firefighting command center, voice communication and data communication are performed between a command control device 101, which transmits and receives information from a speaker at the firefighting command center, and a radio network control device 102, which inputs and outputs information from the mobile station.

[0015] Next, a configuration example of a speaker identification device 1 in wireless communication will be described with reference to Fig. 2. The speaker identification device 1 includes a voice section information collection means 11, a voice recognition means 12, and a recognition result display means 13. Note that in the following description, wireless communication refers to digital wireless communication, and both the mobile station and the command console will be described as having a wireless communication device that is a terminal for performing wireless communication. Furthermore, the recognition result display means 13 may be simply referred to as a display means.

[0016] The voice section information collecting means 11 detects the start and end of a press, and outputs a speaker ID when a press starts.

[0017] The voice recognition means 12 performs voice recognition on the voice data collected by the voice section information collection means 11 from the start of pressing to the end of pressing, and outputs voice text.

[0018] The display means 13 displays the speaker ID output by the voice section information collection means 11 and the voice text output by the voice recognition means 12 in association with each other.

[0019] The wireless communication device according to this embodiment is a press-to-talk type that enables transmission when the transmission button is pressed.

[0020] Whether it is a mobile station or a command console, the speaker presses the transmit button of the wireless communication device when conducting a simultaneous voice communication. That is, in the case of a mobile station or a command console, after the user starts pressing the button, the speaker can speak while pressing the button, and stops pressing the button when the speaker finishes speaking.

[0021] A control signal is sent each time a press starts and ends. This control signal stores signals indicating the start and end of a press and the speaker ID. As shown in Figure 1, voice and data communications are performed between the command control device 101 and the radio network control device 102, and in a configuration in which the radio network control device and the command control device collect these communications, the control signal is sent from the radio network control device to the command control device.

[0022] Therefore, the speaker identification device 1 can associate the voice being communicated in simultaneous voice communication with the speaker by using the speaker ID information included in the notification of the control signal. That is, the speaker identification device 1 can identify the speaker for each voice text output by the voice recognition means 12.

[0023] In addition, when the display means 13 detects a press end signal output from the voice section information collection means 11, it can terminate the association between the voice text output by the voice recognition means 12 and the speaker ID.

[0024] As an example, the display means 13 can terminate the association between the speech text generated by the speech recognition means 12 and the speaker ID in response to receiving a press end signal included in information sent to the display means 13 from the recognition result speech section information combination means 16, which will be described later. The determination of the end of this association can also be made by the recognition result speech section information combination means 16, rather than the display means 13.

[0025] Here, we will explain the detailed configuration and operation of the speaker identification device 1. Fig. 3 is a diagram showing the detailed configuration of the speaker identification device 1 shown in Fig. 2 and an example of information transmitted and received.

[0026] The speaker identification device 1 comprises a voice section information collection means 11, a voice recognition means 12, a display means 13, a voice signal collection and division means 14, a voice information aggregation means 15, a recognition result voice section information combination means 16, a voice information transmission means 17, a recognition result receiving means 18, and a recognition result information set queue 19. In the following description, the exchange of information between components is assumed to be the sending and receiving of signals containing the information, and a description of the signals themselves may be omitted.

[0027] The voice section information collecting means 11 collects voice section information when a speaker starts or ends a press on the wireless communication device. Here, the voice section information is information indicating the start or end of a press, and speaker ID information. The voice section information collecting means 11 sends the voice section information to the voice information summarizing means 15. In the following description, the speaker may be either a mobile station or a command console, and is not limited thereto.

[0028] When a speaker starts to press the wireless communication device and speak, digital voice data is input to the voice signal collecting and dividing means 14. In response to this, the voice signal collecting and dividing means 14 collects the digital voice data, divides this digital voice data into fixed size pieces, and sends them to the voice information collecting means 15.

[0029] The voice information aggregating means 15 aggregates information transmitted and received from the voice section information collecting means 11 and the voice signal collecting and dividing means 14. Specifically, the voice information aggregating means 15 receives the press state and speaker ID transmitted from the voice section information collecting means 11, and the voice data transmitted from the voice signal collecting and dividing means 14.

[0030] Furthermore, when the voice information collecting means 15 receives information about the start of press from the voice section information collecting means 11, it sends information about the speaker ID to the recognition result voice section information combining means 16. On the other hand, the voice information collecting means 15 sends information about the voice data and information about the speaker ID to the voice information transmitting means 17 during the period from the start of press to the end of press.

[0031] The voice information transmitting means 17 sends the voice data to the voice recognition means 12 .

[0032] The speech recognition means 12 performs speech recognition of the speech signal and generates speech text (hereinafter, "recognition result text"). The speech recognition means 12 also sends the generated recognition result text to the recognition result receiving means 18.

[0033] The recognition result receiving means 18 sends the recognition result text to the recognition result speech segment information combining means 16 .

[0034] The recognition result speech section information combining means 16 sends a recognition result information set, which is information associating the recognition result text input from the recognition result receiving means 18 with the speaker ID input from the speech information collecting means 15, to the recognition result display and storage means.

[0035] Here, the recognition result display and storage means can be one component having the recognition result information set queue 19 and the recognition result display means 13.

[0036] Therefore, when the recognition result information set is stored in the recognition result display storage means, the recognition result speech segment information combining means 16 stores the recognition result set in the recognition result information set queue 19 of the recognition result display storage means.

[0037] On the other hand, when it is desired to display the information of the recognition result information set, the recognition result display means 13 extracts the recognition result information set from the recognition result information set queue 19 and displays the recognition result together with the speaker ID.

[0038] As a result, the speaker identification device 1 can use the speaker ID information to output a recognition result information set that associates the recognition result text of the voice being communicated in the simultaneous voice communication with information about who the speaker is.

[0039] The speaker identification device 1 may be configured without the voice information transmitting means 17 and the recognition result receiving means 18. That is, the voice information collecting means 15 can transmit to the voice recognition means 12 without going through the voice information transmitting means 17, and the voice recognition means 12 can transmit the recognition result text to the recognition result voice section information combining means 16 without going through the recognition result receiving means 18.

[0040] Embodiment 2 Using the configuration described in embodiment 1 as a basic configuration, we will now describe a speaker identification device 2 that displays recognition result text in a more understandable manner by increasing the amount of information collected from data communication between the command control device and the radio line control device by the speech section information collection means 11. Note that, with regard to the components of the speaker identification device 2, those that perform the same functions as the components of the speaker identification device 1 described in embodiment 1 are assigned the same reference numerals, and their description will be omitted.

[0041] FIG. 4 is a diagram showing a configuration for real-time speech recognition of radio communication voices in simultaneous voice communication on a single radio channel (channel 01) in a fire and emergency digital radio system, and displaying the recognition result in text along with the speaker.

[0042] Specifically, Fig. 4 is a configuration based on Fig. 3, with the addition of a speaker DB 13a that is referenced by the recognition result display means 13 to convert a speaker ID into a speaker name. Note that the information collected by the voice section information collection means 11 from data communication between the command control device 101 and the radio line control device 102 includes information on the press start, press end, speaker ID, as well as downstream and upstream radio channel information. The data flow between each component is the same as in Fig. 3, so some of the explanation will be omitted.

[0043] Next, as an example of the operation of the speaker identification device 2, we will explain the operation of performing voice recognition of an uplink voice in which a fire engine is pressed and a downlink voice in which a command console is pressed. As a premise, it is assumed that the command console of the fire command center and the mobile stations (fire engines and ambulances) are capturing wireless channel 01. Also, the correspondence between speaker IDs and speakers is as shown in the table in Figure 5.

[0044] Here, the operation of identifying a speaker for speech recognition of uplink voice and downlink voice is common to part of the operation according to the third embodiment described later, and therefore the operation process will be described as steps A to F.

[0045] Voice recognition of fire engine press (upstream voice) (Step A) The voice section information collection means 11 collects press start information and speaker information [speaker ID: CAR-F-001, channel 01, uplink] as voice section information from data communication between the command control device 101 and the radio line control device 102.

[0046] Furthermore, the voice signal collection and division means 14 collects the uplink voice from the fire engine from the voice communication between the command control device 101 and the radio line control device 102, and transfers the voice data to the voice information collection means 15. This voice data is digital voice data, and the voice signal collection and division means 14 divides it at regular intervals and transfers it to the voice information collection means 15.

[0047] The voice section information collecting means 11 notifies the voice information aggregating means 15 of [press start, speaker ID: CAR-F-001, channel 01, uplink].

[0048] (Step B) In response to receiving the information of "press start" included in the voice section information from the voice section information collecting means 11, the voice information collecting means 15 notifies the recognition result voice section information combining means 16 of the speaker information [speaker ID: CAR-F-001, channel 01, uplink]. The recognition result voice section information combining means 16 holds the notified speaker information until the next time speaker information is notified. In addition, the voice information collecting means 15 passes the voice data received from the voice signal collecting and dividing means 14 to the voice information transmitting means 17.

[0049] The voice information transmitting means 17 passes the voice data received from the voice information collecting means 15 to the voice recognition means 12 .

[0050] The speech recognition means 12 passes the initial recognition result text ["From Fire Engine A"] to the recognition result receiving means 18.

[0051] The recognition result speech section information combining means 16 stores the recognition result information set [speaker ID: CAR-F-001, channel 01, uplink, recognition result: "From fire engine A"], which associates the speaker information [speaker ID: CAR-F-001, channel 01, uplink] that it holds with the recognition result received from the recognition result receiving means 18, in the recognition result information set queue 19. The contents of the recognition result information set queue 19 at this time are shown in Fig. 6(a).

[0052] Next, the speech recognition means 12 passes the next recognition result text ["Fire Department A"] to the recognition result receiving means 18.

[0053] The recognition result speech section information combining means 16 stores in the recognition result information set queue a recognition result information set [speaker ID: CAR-F-001, channel 01, uplink, recognition result: "Fire Department A"] that associates the speaker information [speaker ID: CAR-F-001, channel 01, uplink] that it holds with the recognition result received from the recognition result receiving means 18. The contents of the recognition result information set queue at this time are shown in Figure 6(b).

[0054] (Step C) After Fire Engine A finishes speaking on radio channel 01, it stops pressing the button.

[0055] At this time, the voice section information collecting means 11 collects voice section information [press end, speaker ID: CAR-F-001, channel 01, uplink] from the data communication between the command control device 101 and the radio line control device 102.

[0056] The voice section information collecting means 11 notifies the voice information aggregating means 15 of [press end, speaker ID: CAR-F-001, channel 01, uplink].

[0057] Upon receiving [Press End] from the voice section information collection means 11, the voice information collection means 15 terminates the transfer of the voice data received from the voice signal collection and division means 14 to the voice information transmission means 17 at that time.

[0058] Meanwhile, the recognition result display means 13 periodically monitors the recognition result information set queue 19. Here, a case will be described where the monitoring process is activated at the time of the queue in FIG.

[0059] First, the recognition result display means 13 extracts No. 1 [CAR-F-001, channel 01, uplink, "From Fire Engine A"] from the recognition result information set queue 19, and obtains the speaker name [Fire Engine A] from the speaker DB using the speaker ID [CAR-F-001] as a key. Here, when extracting a queue, the recognition result display means 13 stores the data number of the last extracted recognition result information set queue, and displays the speaker name and recognition result on a screen or the like. The screen image at this time is shown in Figure 7(a). Note that the recognition result display means 13 displays the recognition result for the uplink direction on the left side of the display area.

[0060] Next, the recognition result display means 13 extracts No. 2 [CAR-F-001, channel 01, uplink, "Fire Engine A"], which is the data next to the data number (No. 1) of the last extracted recognition result information set queue, and obtains the speaker name [Fire Engine A] from the speaker DB using the speaker ID [CAR-F-001] as a key, and displays the recognition result on a screen, etc. The screen image at this time is shown in Figure 7(b).

[0061] Since there is no change in the speaker, the speaker name is not displayed in the current recognition result display in FIG. 7(b), but the speaker name may be displayed.

[0062] Voice recognition of the press (downstream voice) from the control desk Next, we will explain the case where, after Fire Engine A has finished pressing, Command Console A starts voice communication with the press on radio channel 01. At this time, Command Console A utters "This is Fire Engine A, please come in" during the press.

[0063] (Step D) The voice section information collecting means 11 collects voice section information [press start, speaker ID: DAI-001, channel 01, downstream] from data communication between the command control device 101 and the radio line control device 102.

[0064] Furthermore, the voice signal collection and division means 14 collects the uplink voice from the fire engines from the voice communication between the command control device 101 and the radio line control device 102, and transfers the voice data to the voice information collection means 15. This voice data is digital voice data, and the voice signal collection and division means 14 divides it at regular intervals and transfers it to the voice information collection means 15.

[0065] The voice section information collecting means 11 notifies the voice information aggregating means 15 of [press start, speaker ID: DAI-001, channel 01, downstream].

[0066] (Step E) In response to receiving [press start] from the voice section information collection means 11, the voice information collection means 15 notifies the recognition result voice section information combination means 16 of the speaker information [speaker ID: DAI-001, channel 01, downstream]. Here, the recognition result voice section information combination means 16 discards the speaker information it received and held previously, and holds the newly received speaker information until the next speaker information is notified. In addition, the voice information collection means 15 passes the voice data received from the voice signal collection and division means 14 to the voice information transmission means 17.

[0067] The voice information transmitting means 17 passes the voice data received from the voice information collecting means 15 to the voice recognition means 12 .

[0068] The voice recognition means 12 passes the initial recognition result text ("This is Fire Department A") to the recognition result receiving means 18.

[0069] The recognition result speech section information combining means 16 associates the retained speaker information [speaker ID: DAI-001, channel 01, downstream] with the recognition result received from the recognition result receiving means 18 to create a recognition result information set [speaker ID: DAI-001, channel 01, downstream, recognition result: "This is Fire Department A"], and stores it in the recognition result information set queue 19. The contents of the recognition result information set queue 19 at this time are shown in Fig. 6(c).

[0070] Next, the speech recognition means 12 passes the next recognition result text ("please") to the recognition result receiving means 18.

[0071] The recognition result speech section information combining means 16 stores a recognition result information set [speaker ID: DAI-001, channel 01, downstream] that associates the retained speaker information [speaker ID: DAI-001, channel 01, downstream, recognition result: "Douzo"] with the recognition result received from the recognition result receiving means 18 in the recognition result information set queue 19. The contents of the recognition result information set queue 19 at this time are shown in Fig. 6(d).

[0072] (Step F) After finishing speaking on radio channel 01, command console A stops pressing the button.

[0073] The voice section information collecting means 11 collects voice section information [press end, speaker ID: DAI-001, channel 01, downstream] from data communication between the command control device 101 and the radio line control device 102.

[0074] The voice section information collecting means 11 notifies the voice information aggregating means 15 of [press end, speaker ID: DAI-001, channel 01, downstream].

[0075] Upon receiving [Press End] from the voice section information collection means 11, the voice information collection means 15 terminates the transfer of the voice data received from the voice signal collection and division means 14 to the voice information transmission means 17 at that time.

[0076] Meanwhile, the recognition result display means 13 periodically monitors the recognition result information set queue 19. A case will be described where the monitoring process is activated at the time of the queue shown in FIG.

[0077] Since the data number of the recognition result information set queue last extracted in step C is No. 2, the recognition result display means 13 extracts No. 3 of the recognition result information set queue [DAI-001, channel 01, outbound, "This is Fire Department A"]. Then, the recognition result display means 13 obtains the speaker name [Command Console A] from the speaker DB using the speaker ID [DAI-001] as a key, and displays the recognition result on a screen or the like. The screen image at this time is shown in Figure 7(c). The recognition result for the outbound direction is displayed on the right side of the display area.

[0078] Next, the recognition result display means 13 extracts the next data, No. 4 [DAI-001, channel 01, outbound, "Douzo"]. Then, the recognition result display means 13 uses the speaker ID [DAI-001] as a key to obtain the speaker name [command console] from the speaker DB, and displays the recognition result on a screen, etc. The screen image at this time is shown in Figure 7(d).

[0079] Since there is no change in the speaker, the speaker name is not displayed in this recognition result display, but the speaker name may be displayed.

[0080] As a result, the speaker identification device 2 can identify the speaker for each recognition result text generated from voice data by using information (speaker ID, wireless channel, downlink / uplink) exchanged in data communication between the command control device 101 and the wireless network control device 102. Therefore, the speaker identification device 2 can display speaker information associated with each recognition result text generated from voice data.

[0081] In other words, the voice section information collection means 11 collects at least the direction of communication and the communication channel number, and the recognition result voice section information combination means can use the direction of communication and the communication channel number to associate the recognition result text with the speaker ID.

[0082] Embodiment 3 Next, a speaker identification device 3 according to a third embodiment will be described with reference to Fig. 8. In the first and second embodiments, the speech recognition means 12 immediately returns a recognition result. However, depending on the specifications and connection configuration of the machine in which the speech recognition means 12 is implemented, there may be a time lag between when speech is input from the speech information transmitting means 17 to the speech recognition means 12 and when the speech recognition means 12 returns the recognition result to the recognition result receiving means 18.

[0083] If this time difference becomes large, for example, if ambulance A starts pressing immediately after fire engine A finishes pressing, there is a possibility that the recognition result uttered just before fire engine A finishes pressing may be notified from the voice recognition means 12 to the recognition result receiving means 18 after ambulance A starts pressing. In this case, it will be impossible to correctly associate the recognition result text with the speaker ID. The speaker identification device 3 used in the third embodiment below solves this problem.

[0084] Here, the speaker identification device 3 includes a voice section information collecting means 21, a voice recognition means 22, a display means 23, a voice signal collecting and dividing means 24, a voice information collecting means 25, a recognition result voice section information combining means 26, a voice information transmitting means 27, a recognition result receiving means 28, a recognition result information set queue 29, a speaker DB 23a, and a voice section information storage memory 26a. These components of the speaker identification device 3 each perform the same function as the voice section information collecting means 11, the voice recognition means 12, the display means 13, the voice signal collecting and dividing means 14, the voice information collecting means 15, the recognition result voice section information combining means 16, the voice information transmitting means 17, the recognition result receiving means 18, the recognition result information set queue 19, the speaker DB 13a, and the voice section information storage memory 16a shown in the first or second embodiment, and therefore description thereof will be omitted. In addition to this, the components of the speaker identification device 3 have the following functions.

[0085] As shown in FIG. 8, when the voice information aggregation means 25 is notified of the start of pressing from the voice section information collection means 21 and passes voice data to the voice information transmission means 27 for the first time after the start of pressing is notified, the voice information aggregation means 25 assigns a first-time flag indicating that it is the first voice after the start of pressing to the voice information transmission means 27. Along with this, the voice information aggregation means 25 passes the voice section information [speaker ID, radio channel, downlink / uplink] to the voice information transmission means 27.

[0086] When the voice data with the first-time flag attached is passed from the voice information aggregation means 25 to the voice information transmission means 27, the voice information transmission means 27 passes the voice section information [speaker ID, radio channel, downlink / uplink] together with the time [T0 (first transmission time)] when the voice data with the first-time flag attached was sent to the voice recognition means 22 to the recognition result voice section information combining means 26.

[0087] The recognition result voice section information combining means 26 stores [T0, speaker ID, radio channel, downlink / uplink] in the voice section information storage memory 26a.

[0088] When the voice recognition means 22 passes the recognition result to the recognition result receiving means 28, it also passes the time Tn (n: 1~) when the target voice for each recognition result was input from the voice information transmission means 27. Here, for the time, it is set as (T0≈T_{1}, T0<T_{2} and later).

[0089] The recognition result voice section information combining means​​​​​​​The difference between the speaker identification device 3 and the speaker identification device 2 shown in the second embodiment is the method of identifying a speaker ID by the recognition result speech section information combining means 26. That is, the method of identifying a speaker ID is different, and the speaker ID after identification is the same as in the second embodiment. Therefore, there is no difference in the contents of the recognition result information set queue 29 and the recognition result display by the recognition result display means 23, and therefore the contents of Figures 6(a) to 6(d) and Figures 7(a) to 7(d) are the same as those shown in the second embodiment.

[0092] Voice recognition of fire engine press (upstream voice) First, fire engine A starts a press and voice communication on radio channel 01. At this time, fire engine A utters "Fire engine A from Fire Department A" during the press.

[0093] (Step A) The speaker identification device 3 performs the same operation as in step A shown in the second embodiment.

[0094] (Step B) In response to receiving [press start] from the voice section information collecting means 21, the voice information collecting means 25 passes the voice data received from the voice signal collecting and dividing means 24 from that point on to the voice information transmitting means 27. At this time, the voice information collecting means 25 also passes to the voice information transmitting means 27 [speaker ID: CAR-F-001, channel 01, uplink] and information that the first flag is ON if it is the first voice data after the press start, and the first flag is OFF thereafter.

[0095] The voice information transmitting means 27 passes the voice data received from the voice information collecting means 25 to the voice recognition means 22. Here, if the data to be sent from the voice information transmitting means 27 to the voice recognition means 22 is voice data with the first time flag ON, the voice information transmitting means 27 passes the information [speaker ID: CAR-F-001, channel 01, uplink] to the recognition result voice section information combining means 26 together with the time (assumed to be time 1) at which it started sending the voice data to the voice recognition means 22.

[0096] The recognition result voice segment information combining means 26 stores [time 1, speaker ID: CAR-F-001, channel 01, uplink] in the voice segment information storage memory 26a. The contents of the voice segment information storage memory 26a at this time are shown in Figure 9(a).

[0097] The speech recognition means 22 passes to the recognition result receiving means 28 the initial recognition result text ("From fire engine A") and the time (time A) when the target speech of the recognition result text was input.

[0098] The recognition result voice segment information combining means 26 compares the voice segment start time stored in the voice segment information storage memory 26a with the time information (time A) included in the recognition result received from the recognition result receiving means 28. In this case, time 1<time A, so it can be determined that the recognition result text corresponds to [speaker ID: CAR-F-001, channel 01, uplink].

[0099] In this case, the recognition result speech section information combining means 26 stores [speaker ID: CAR-F-001, channel 01, uplink] and the recognition result: "From fire engine A" in the recognition result information set queue 29. The contents of the recognition result information set queue 29 at this time are shown in Fig. 6(a).

[0100] Next, the voice recognition means 22 passes to the recognition result receiving means 28 the next recognition result text ("Fire Department A") and the time (time B) when the target voice of the recognition result text was input.

[0101] The recognition result speech segment information combining means 26 compares the speech segment start time stored in the speech segment information storage memory 26a with the time information (time B) included in the recognition result received from the recognition result receiving means 28. In this case, time 1<time B, so the recognition result speech segment information combining means 26 can determine that the recognition result text corresponds to [speaker ID: CAR-F-001, channel 01, uplink].

[0102] In this case, the recognition result speech section information combining means 26 stores [speaker ID: CAR-F-001, channel 01, uplink] and the recognition result: "A Fire Department" in the recognition result information set queue 29. The contents of the recognition result information set queue 29 at this time are shown in FIG. 6(b).

[0103] (Step C) The speaker identification device 3 performs the same operation as in step C shown in the second embodiment.

[0104] Voice recognition of the press (downstream voice) from the control desk Next, we will explain the case where, after Fire Engine A has finished pressing, Command Console A starts pressing on radio channel 01 to start voice communication. At this time, Command Console A utters "This is Fire Engine A, please come in" during the press.

[0105] (Step D) The speaker identification device 3 performs the same operation as in step D shown in the second embodiment.

[0106] (Step E) In response to receiving [press start] from the voice section information collecting means 21, the voice information collecting means 25 passes the voice data received from the voice signal collecting and dividing means 24 from that point on to the voice information transmitting means 27. At this time, the voice information collecting means 25 also passes [speaker ID: DAI-001, channel 01, downlink] and information that the first flag is ON if it is the first voice data after the start of pressing, and that the first flag is OFF thereafter.

[0107] The voice information transmitting means 27 passes the voice data received from the voice information collecting means 25 to the voice recognition means 22. If the data received by the voice information transmitting means 27 is voice data with the first time flag ON, the voice information transmitting means 27 passes [speaker ID: DAI-001, channel 01, downstream] to the recognition result voice section information combining means 26 together with the time (assumed to be time 2) at which the voice data started to be transmitted to the voice recognition means 22.

[0108] The recognition result speech segment information combining means 26 stores [time 2, speaker ID: DAI-001, channel 01, downstream] in the speech segment information storage memory 26a. The contents of the speech segment information storage memory 26a at this time are shown in FIG. 9(b).

[0109] The voice recognition means 22 passes the initial recognition result text ("This is Fire Department A") and the time (time C) when the target voice of the recognition result text was input to the recognition result receiving means .

[0110] The recognition result speech segment information combining means 26 compares the speech segment start time stored in the speech segment information storage memory 26a with the time information (time C) included in the recognition result received from the recognition result receiving means 28. In this case, time 2<time C, so the recognition result speech segment information combining means 26 can determine that the recognition result text corresponds to [speaker ID: DAI-001, channel 01, downlink].

[0111] In this case, the recognition result speech section information combining means 26 stores [speaker ID: DAI-001, channel 01, downstream] and the recognition result: "This is Fire Department A" in the recognition result information set queue 29. The contents of the recognition result information set queue 29 at this time are shown in FIG. 7(c).

[0112] Next, the voice recognition means 22 passes to the recognition result receiving means 28 the next recognition result text ("Douzo") and the time (time D) when the target voice of the recognition result text was input.

[0113] The recognition result speech segment information combining means 26 compares the speech segment start time stored in the speech segment information storage memory 26a with the time information (time D) included in the recognition result received from the recognition result receiving means 28. In this case, time 2<time D, so the recognition result speech segment information combining means 26 can determine that the recognition result text corresponds to [speaker ID: DAI-001, channel 01, downlink].

[0114] In this case, the recognition result speech section information combining means 26 stores [speaker ID: DAI-001, channel 01, downstream] and the recognition result: "Douzo" in the recognition result information set queue 29. The contents of the recognition result information set queue 29 at this time are shown in Fig. 7(d).

[0115] (Step F) The speaker identification device 3 performs the same operation as in step F shown in the second embodiment.

[0116] As mentioned above, the timing at which the voice recognition means 22 returns the recognition results depends on the specifications of the voice recognition engine and the installed machine, so it is not necessarily possible to receive all the recognition results before the next speaker starts pressing the button.

[0117] In contrast, according to the configuration of the speaker identification device 3 shown in embodiment 3, the speaker can be correctly identified even if the recognition result receiving means 28 receives the recognition result information from the voice recognition means 22 with a delay.

[0118] For example, assume that the time when the voice recognition result information [time B, "Fire Engine A"] is received before the end of the press by Fire Engine A is after the start of the next press by Command Console A.

[0119] At this time, command console A has started pressing, the initial voice transmission to the voice recognition means 22 has been completed, and the state of the voice section information storage memory is as shown in Figure 9(b). Here, time B in the recognition result information is the time when the target voice of the recognition result text "Fire Department A" was input, and therefore it is a time earlier than time 2, which is the time of the initial voice transmission from command console A.

[0120] Therefore, time 1<time B<time 2, and therefore the speaker ID of the recognition result "Fire Department A" received by the recognition result receiving means can be identified as [CAR-F-001].

[0121] That is, the recognition result speech section information combining means 26 records the time when the press started or the time when the voice was collected, and if the time when the press started is earlier than the time when the voice was collected, associates the speaker ID with the recognition result text. On the other hand, if the time when the press started is later than the time when the voice was collected, the recognition result speech section information combining means 26 can operate so as not to associate the speaker ID with the recognition result text, assuming that the speech data is from a different speaker.

[0122] From the above, the speaker identification device 3 can more accurately identify the speaker of the target voice by having the function of referencing the input start time of the target voice, even if there is a delay in the input and output of information within the speaker identification device 3, such as when the specifications of each component, including the voice recognition means 22, are low.

[0123] Embodiment 4 In the third embodiment, real-time speech recognition of radio communication voices for simultaneous voice communication on a single radio channel (channel 01) in a fire and emergency digital radio system has been described. In the fourth embodiment, real-time speech recognition when simultaneous voice communication radio communication occurs simultaneously on multiple radio channels (channel 01, channel 02) will be described. The configuration of a speaker identification device 4 according to the fourth embodiment is shown in FIG. 10.

[0124] For simplicity of explanation, the number of wireless channels is two, but there is no limit to the number of wireless channels, and N channels can be realized with the configuration of FIG.

[0125] As shown in FIG. 10 , the speaker identification device 4 includes a voice section information collecting means 31, a voice recognition means 32, a display means 33, a first voice signal collecting and dividing means 341 (hereinafter referred to as the first voice signal collecting and dividing means # for channel 01 341), a second voice signal collecting and dividing means 342 (hereinafter referred to as the second voice signal collecting and dividing means # for channel 02 342), a voice information aggregating means 35, a recognition result voice section information combining means 36, and a first voice information transmitting means 371 ( The system includes a first voice information transmitting means 371 (hereinafter referred to as voice information transmitting means # for channel 01 371), a second voice information transmitting means 372 (hereinafter referred to as voice information transmitting means # for channel 02 372), a first recognition result receiving means 381 (hereinafter referred to as recognition result receiving means # for channel 01 381), a second recognition result receiving means 382 (hereinafter referred to as recognition result receiving means # for channel 02 382), a recognition result information set queue 39, a speaker DB 33a, and a voice section information storage memory 36a.

[0126] Each of these components of the speaker identification device 4 corresponds to the components of the speaker identification device 3 shown in embodiment 3, and detailed explanation of their functions will be omitted. Note that the voice signal collecting and dividing means, voice information transmitting means, and recognition result receiving means will be explained assuming that they have the same functions for both channel 01 and channel 02.

[0127] Specifically, voice communication between the command control device 101 and the radio network control device 102 is performed separately for each channel. Therefore, the speaker identification device 4 is configured to have a voice signal collection and division means for each of channels 01 and 02. Similarly, the speaker identification device 4 is configured to have a voice information transmission means and a recognition result reception means for each of channels 01 and 02.

[0128] The speech recognition results transmitted by the speech information transmitting means # for channel 01 371 are received by the recognition result receiving means # for channel 01 381, and the speech recognition results transmitted by the speech information transmitting means # for channel 02 372 are received by the recognition result receiving means # for channel 02 382.

[0129] Next, an example of the operation of the speaker identification device 4 will be described. Here, it is assumed that the command console A and fire engine A of the fire command center are capturing radio channel 01, and that the command console B and ambulance A are capturing radio channel 02. The correspondence between speaker IDs and speakers is as shown in the table in Fig. 5. It is also assumed that command console A, command console B, fire engine A, and ambulance A press and speak in the time sequence shown in Fig. 12.

[0130] 12 indicate the times when the speech information transmitting means transmits the first speech data to the speech recognition means 32 after each speaker starts pressing. Times A to L indicate the first times when the target speech of the recognition result text received by the recognition result receiving means from the speech recognition means 32 is input to the speech recognition means 32. The state of the speech section information storage memory 36a after time 4 in FIG. 12 is shown in FIG. 13.

[0131] As shown in the third embodiment, there is no problem even if the timing at which the recognition result receiving means receives the recognition result information from the speech recognition means 32 is delayed. Therefore, in order to simplify the description of the processing, the recognition result receiving means # for channel 01 381 and the recognition result receiving means # for channel 02 382 are assumed to receive all recognition results from the speech recognition means 32 after time 4 in FIG.

[0132] Next, we will explain an example of the actual operation of the speaker identification device 4. First, we will explain the operation of the speaker identification device 4 with respect to channel 01.

[0133] The recognition result receiving means # for channel 01 381 receives the recognition result information [time A, "From Fire Department A"] from the voice recognition means 32 from time 4 onwards, and passes the recognition result information to the recognition result voice section information combining means 36.

[0134] The recognition result speech section information combining means 36 compares the speech section start time stored in the speech section information storage memory 36a with the time information included in the recognition result received from the recognition result receiving means # for channel 01 381, using channel 01 as a key.

[0135] In this example, the times corresponding to channel 01 are time 1 and time 4, and the comparison result is time 1 < time A < time 4. Therefore, the recognition result speech section information combining means 36 can determine that the recognition result text corresponds to [speaker ID: DAI-001, channel 01, downlink].

[0136] In this case, the recognition result speech section information combining means 36 stores [speaker ID: DAI-001, channel 01, outbound] and the recognition result "This is Fire Department A" in the recognition result information set queue 39. The contents of this recognition result information set queue 39 are shown in Fig. 14(a).

[0137] Next, when the recognition result receiving means # for channel 01 381 receives the recognition result information [time B, "Fire Engine A"], the time comparison in the recognition result voice section information combining means 36 results in time 1 < time B < time 4. The recognition result voice section information combining means 36 stores [speaker ID: DAI-001, channel 01, downstream] and the recognition result [time B, "Fire Engine A"] in the recognition result information set queue 39. The contents of this recognition result information set queue 39 are shown in Fig. 14(b).

[0138] Next, the operation of the speaker identification device 4 related to channel 02 will be described.

[0139] The recognition result receiving means # for channel 02 382 receives the recognition result information [time C, "Ambulance A"] from the speech recognition means 32, and passes the recognition result information to the recognition result speech segment information combining means 36. The recognition result speech segment information combining means 36 compares the speech segment start time present in the speech segment information storage memory 36a with the time information included in the recognition result received from the recognition result receiving means # for channel 02 382, ​​using channel 02 as a key.

[0140] In this case, the times corresponding to channel 02 are time 2 and time 3, and the comparison result is time 2 < time C < time 3. Therefore, the recognition result speech section information combination means 36 can determine that the recognition result text corresponds to [speaker ID: CAR-A-001, channel 02, uplink].

[0141] In this case, the recognition result speech section information combining means 36 stores [speaker ID: CAR-A-001, channel 02, uplink] and the recognition result "ambulance A" in the recognition result information set queue 39. The contents of this recognition result information set queue 39 are shown in Fig. 14(c).

[0142] As described above, the recognition result speech segment information combining means 36 compares the recognition result information passed from the recognition result receiving means 381, 382 of each channel with the time information of each channel stored in the speech segment information storage memory 36a, using the channel number as a key. Here, the contents of the recognition result information set queue 39 after receiving the recognition results of all the utterance contents shown in Fig. 13 are shown in Fig. 14(d).

[0143] The recognition result display means 33 periodically monitors the recognition result information set queue 39. A case will be described where the monitoring process is activated at the time of the queue shown in FIG.

[0144] First, the recognition result display means 33 extracts No. 1 [DAI-001, channel 01, outbound, "From Fire Department A"] from the recognition result information set queue 39, and acquires the speaker name [Command Console A] from the speaker DB 33a using the speaker ID [DAI-001] as a key. Here, the recognition result display means 33 stores the data number of the recognition result information set queue 39 extracted last when extracting a queue.

[0145] The recognition result display means 33 displays the speaker name and the recognition result on the screen. The screen image at this time is shown in Figure 15(a). This screen is divided into an area for displaying the recognition results for channel 01 and an area for displaying the recognition results for channel 02, and the recognition result display means 33 determines the display area based on the wireless channel information in the data extracted from the recognition result information set queue 39. Furthermore, the recognition result display means 33 displays the recognition results for the downlink direction on the right side of the display area and the recognition results for the uplink direction on the left side of the display area.

[0146] Next, the recognition result display means 33 extracts the data No. 2 [DAI-001, channel 01, downbound, "Fire Engine A"], which is the data next to the data number (No. 1) of the last extracted recognition result information set queue 39, and obtains the speaker name [Command Desk A] from the speaker DB 33a using the speaker ID [DAI-001] as a key.

[0147] The recognition result display means 33 displays the recognition result in the display area of ​​wireless channel 01. The screen image at this time is shown in Fig. 15(b). Note that in this example, since there is no change in the speaker, the speaker name is not displayed in the recognition result display in Fig. 15(b), but the speaker name may be displayed.

[0148] Next, the recognition result display means 33 retrieves the next data, No. 3 [CAR-A-001, channel 02, upbound, "Ambulance A"], in the recognition result information set queue 39. After that, the recognition result display means 33 acquires the speaker name [Ambulance A] from the speaker DB 33a using the speaker ID [CAR-A-001] as a key.

[0149] The recognition result display means 33 displays this recognition result in the display area of ​​the wireless channel 02. The screen image at this time is shown in Fig. 15(c).

[0150] FIG. 15(d) shows the state when the recognition result display means 33 has finished displaying all the contents of the recognition result information set queue 39 shown in FIG. 14(d) on the screen in accordance with the above display rules.

[0151] As a result, even when simultaneous voice communication wireless communications occur simultaneously over a plurality of wireless channels, the speaker identification device 4 can identify and display the speaker for each recognition result text for each channel.

[0152] In other words, the speaker identification device 4 has multiple voice signal collecting and dividing means, multiple voice information transmitting means, and multiple recognition result receiving means, and when voice communication is carried out using multiple channels, each of the multiple voice signal collecting and dividing means, multiple voice information transmitting means, and multiple recognition result receiving means can be used corresponding to each channel.

[0153] As shown in Figures 10 and 11, in speaker identification device 4, rather than simply multiplexing the entire configuration of speaker identification device 2 shown in embodiment 2, by multiplexing only the necessary components, it is possible to achieve a multiplexed configuration while reducing machine resources, etc.

[0154] Although the present disclosure has been described above with reference to the embodiments, the present disclosure is not limited to the above-described embodiments. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present disclosure within the scope of the present disclosure. Furthermore, each embodiment can be combined with other embodiments as appropriate.

[0155] Embodiments of the present disclosure may be implemented in hardware or special purpose circuits, software, logic, or any combination thereof. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device.

[0156] The present disclosure also provides at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. The computer program product includes computer-executable instructions, such as instructions included in program modules, that execute on a target real or virtual processor or device to perform the processes or methods of the present disclosure. Program modules include routines, programs, libraries, objects, classes, components, data structures, etc. that perform particular tasks or implement particular abstract data types. The functionality of the program modules may be combined or divided among program modules as desired in various embodiments. The machine-executable instructions of the program modules may be executed in local or distributed devices. In a distributed device, the program modules may be located in both local and remote storage media.

[0157] The program code for executing the methods of the present disclosure may be written in any combination of one or more programming languages. The program code may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus. When the program code is executed by the processor or controller, the functions / acts in the flowcharts and / or implementing block diagrams are performed. The program code may be executed entirely on the machine, partly on the machine, as a standalone software package, partly on the machine and partly on a remote machine, or entirely on a remote machine or server.

[0158] The program can be stored and provided to a computer using various types of non-transitory computer-readable media. Non-transitory computer-readable media include various types of tangible recording media. Examples of non-transitory computer-readable media include magnetic recording media, magneto-optical recording media, optical disk media, and semiconductor memory. Magnetic recording media include, for example, flexible disks, magnetic tapes, and hard disk drives. Magneto-optical recording media include, for example, magneto-optical disks. Optical disk media include, for example, Blu-ray discs, CD (Compact Disc)-ROMs (Read Only Memory), CD-Rs (Recordable), and CD-RWs (Rewritable). Semiconductor memory includes, for example, solid-state drives, mask ROMs, PROMs (Programmable ROMs), EPROMs (Erasable PROMs), flash ROMs, and RAMs (Random Access Memory). The program may also be provided to a computer by various types of temporary computer-readable media. Examples of temporary computer-readable media include electrical signals, optical signals, and electromagnetic waves. The temporary computer-readable medium can supply the program to the computer via a wired communication path such as an electric wire or an optical fiber, or via a wireless communication path.

[0159] Each drawing is merely an example for describing one or more embodiments. Each drawing may relate not only to one particular embodiment, but also to one or more other embodiments. As will be understood by those skilled in the art, various features or steps described with reference to any one drawing can be combined with features or steps shown in one or more other drawings to create, for example, an embodiment not explicitly shown or described. Not all features or steps shown in any one drawing are necessary to describe an exemplary embodiment, and some features or steps may be omitted. The order of steps described in any drawing may be changed as appropriate.

[0160] A part or all of the above-described embodiments can be described as, but not limited to, the following supplementary notes. (Appendix 1) A speaker identification device that performs speech recognition and identifies a speaker who has spoken a speech, a voice section information collecting means for detecting a press start and a press end and outputting a speaker ID when a press start occurs; a voice recognition means for performing voice recognition on the collected voice data from the start of pressing to the end of pressing and outputting a recognition result text; a display means for displaying the recognition result text and speaker information corresponding to the speaker ID associated with the recognition result text, Speaker identification device. (Appendix 2) The display means When an end of the press is detected, the association between the speaker ID and the recognition result text is terminated. 2. A speaker identification device as described in claim 1. (Appendix 3) an audio signal collection and division means for collecting audio information; a voice information collecting means for receiving the voice data transmitted from the voice signal collecting and dividing means and the press state and speaker ID transmitted from the voice section information collecting means; a recognition result speech section information combining means for generating a recognition result information set by associating a speaker ID input from the speech information collecting means with a recognition result text generated by the speech recognition means based on the speech data sent from the speech information collecting means; a recognition result information set queue for storing the recognition result information set, the display means displays the recognition result information set stored in the recognition result information set queue. 3. A speaker identification device as described in appendix 2. (Appendix 4) a voice information transmitting means for receiving voice data from the voice signal collecting and dividing means and transmitting the voice data to the voice recognition means; a recognition result receiving means for receiving the recognition result text generated by the speech recognition means and sending it to the recognition result speech segment information combining means, 4. A speaker identification device as described in appendix 3. (Appendix 5) The speech segment information collecting means collecting at least the direction of the communication and the channel number of the communication; The recognition result speech segment information combining means The direction of the communication and the channel number of the communication are used to associate the recognition result text with the speaker ID. 5. A speaker identification device as described in appendix 4. (Appendix 6) The recognition result speech segment information combining means a time when the press started or a time when the voice was collected, transmitted from the voice information collecting means, is recorded, and if the time when the press started is earlier than the time when the voice was collected, the speaker ID is associated with the recognition result text, and if the time when the press started is later than the time when the voice was collected, the speaker ID is not associated with the recognition result text; 6. A speaker identification device as described in appendix 5. (Appendix 7) The recognition result speech segment information combining means acquiring information on the time when the press was started and the time when the voice was collected, which information is transmitted from the voice information collecting means via the voice information transmitting means; 7. A speaker identification device as described in appendix 6. (Appendix 8) A plurality of the audio signal collecting and dividing means; A plurality of the voice information transmitting means; a plurality of the recognition result receiving means; When voice communication is performed using a plurality of channels, each of the plurality of voice signal collecting and dividing means, the plurality of voice information transmitting means, and the plurality of recognition result receiving means is used correspondingly for each channel. 8. A speaker identification device as described in Appendix 7. (Appendix 9) A method for identifying a speaker who has spoken a voice, while performing speech recognition, comprising: Detects the start and end of a press, and outputs the speaker ID when a press starts. The system recognizes the collected voice data from the start of the press to the end of the press, and outputs the recognition result text. displaying the recognition result text and speaker information corresponding to the speaker ID associated with the recognition result text; Speaker identification methods. (Appendix 10) A program for performing speech recognition and identifying a speaker who has spoken a speech, Detects the start and end of a press, and outputs the speaker ID when a press starts. The system recognizes the collected voice data from the start of the press to the end of the press, and outputs the recognition result text. displaying the recognition result text and speaker information corresponding to the speaker ID associated with the recognition result text; program.

[0161] Some or all of the elements described in Supplementary Notes 2 to 8 that are dependent on Supplementary Note 1 may also be dependent on the method of Supplementary Note 9 and the program of Supplementary Note 10 in the same dependent relationship as Supplementary Notes 2 to 8. Some or all of the elements described in any Supplementary Note may be applied to various hardware, software, recording means for recording software, systems, and methods. [Explanation of symbols]

[0162] 1. Speaker identification device 2. Speaker Identification Device 3. Speaker Identification Device 4. Speaker Identification Device 101 Command and control device 102 Radio line control device 11. Voice section information collection means 12 Voice Recognition Methods 13 Recognition result display means 13a Speaker DB 14 Audio signal collection and division means 15. Audio information aggregation means 16. Recognition result speech interval information combining means 16a Voice section information storage memory 17 Voice information transmission means 18 Recognition result receiving means 19 Recognition result information set queue 21 Voice section information collection means 22 Voice Recognition Means 23 Recognition result display means 23a Speaker DB 24 Audio signal collection and division means 25 Voice information aggregation means 26 Recognition result speech interval information combination means 26a Voice section information storage memory 27 Audio information transmission means 28 Recognition result receiving means 29 Recognition result information set queue 31 Voice section information collection means 32 Voice recognition means 33 Recognition result display means 33a Speaker DB 341 first audio signal collection and division means 342 second audio signal collection and division means 35 Voice information aggregation means 36 Recognition result speech interval information combination means 36a Voice section information storage memory 371 first audio information transmission means 372 Second audio information transmission means 381 first recognition result receiving means 382 second recognition result receiving means 39 Recognition result information set queue

Claims

1. A speaker identification device that performs speech recognition and identifies a speaker who has spoken a speech, a voice section information collecting means for detecting a press start and a press end and outputting a speaker ID when a press start occurs; a voice recognition means for performing voice recognition on the voice data collected from the start of pressing to the end of pressing and outputting a recognition result text; a display means for displaying the recognition result text and information on a speaker according to the speaker ID associated with the recognition result text, Speaker identification device.

2. The display means When an end of the press is detected, the association between the speaker ID and the recognition result text is terminated. The speaker identification device according to claim 1 .

3. an audio signal collection and division means for collecting audio information; a voice information collecting means for receiving the voice data transmitted from the voice signal collecting and dividing means and the press state and speaker ID transmitted from the voice section information collecting means; a recognition result speech section information combining means for generating a recognition result information set by associating a speaker ID input from the speech information collecting means with a recognition result text generated by the speech recognition means based on the speech data sent from the speech information collecting means; a recognition result information set queue for storing the recognition result information set, the display means displays the recognition result information set stored in the recognition result information set queue. The speaker identification device according to claim 2 .

4. a voice information transmitting means for receiving voice data from the voice signal collecting and dividing means and transmitting the voice data to the voice recognition means; a recognition result receiving means for receiving the recognition result text generated by the speech recognition means and sending it to the recognition result speech segment information combining means, The speaker identification device according to claim 3 .

5. The speech segment information collecting means collecting at least the direction of the communication and the channel number of the communication; The recognition result speech segment information combining means The direction of the communication and the channel number of the communication are used to associate the recognition result text with the speaker ID. The speaker identification device according to claim 4.

6. The recognition result speech segment information combining means a time when the press started or a time when the voice was collected, transmitted from the voice information collecting means, is recorded, and if the time when the press started is earlier than the time when the voice was collected, the speaker ID is associated with the recognition result text, and if the time when the press started is later than the time when the voice was collected, the speaker ID is not associated with the recognition result text; The speaker identification device according to claim 5 .

7. The recognition result speech segment information combining means acquiring information on the time when the press was started and the time when the voice was collected, which information is transmitted from the voice information collecting means via the voice information transmitting means; The speaker identification device according to claim 6.

8. A plurality of the audio signal collecting and dividing means; A plurality of the voice information transmitting means; a plurality of the recognition result receiving means; When voice communication is performed using a plurality of channels, each of the plurality of voice signal collecting and dividing means, the plurality of voice information transmitting means, and the plurality of recognition result receiving means is used correspondingly for each channel. The speaker identification device according to claim 7.

9. A method for identifying a speaker who has spoken a voice, while performing speech recognition, comprising: Detects the start and end of a press, and outputs the speaker ID when a press is started; performing speech recognition on the collected voice data from the start of pressing to the end of pressing, and outputting a recognition result text; displaying the recognition result text and information on the speaker corresponding to the speaker ID associated with the recognition result text; Speaker identification methods.

10. A program for performing speech recognition and identifying a speaker who has spoken a speech, Detects the start and end of a press, and outputs the speaker ID when a press is started; performing speech recognition on the collected voice data from the start of pressing to the end of pressing, and outputting a recognition result text; displaying the recognition result text and information on the speaker corresponding to the speaker ID associated with the recognition result text; program.

Citation Information

Patent Citations

  • Information communication network system in emergency medical care

    JP2013152613A