Information processing system and information processing method

The information processing system enhances communication systems by identifying and marking important conversations with identifiers, addressing the challenge of distinguishing critical information in a sea of recorded interactions.

WO2025150416A1PCT designated stage expired Publication Date: 2025-07-17SONY GROUP CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/045663
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-10
Filing Date
2024-12-24
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

Existing communication systems, such as InCam, fail to efficiently distinguish important conversations from a large number of recorded interactions, making it difficult to locate and manage critical information.

Method used

An information processing system that includes a first terminal for converting speaker voice into voice packets, an information processing apparatus for attaching identifiers to voice packets based on predetermined identification information, and a second terminal for outputting these packets with identifiers, allowing easy discrimination of specific conversations.

Benefits of technology

The system effectively identifies and highlights important conversations, facilitating easier retrieval and management of critical information within a multitude of recorded interactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024045663_17072025_PF_FP_ABST
    Figure JP2024045663_17072025_PF_FP_ABST
Patent Text Reader

Abstract

According to the present disclosure, provided is an information processing system including: a first terminal comprising a control unit that executes processing of converting a speech of a speaker into a speech packet, and transmitting the speech packet to an information processing device; the information processing device comprising a control unit that executes processing of receiving the speech packet from the first terminal, imparting a predetermined identifier to the speech packet to store the speech packet when predetermined identification information is included in the speech packet, and transmitting, to a second terminal, the speech packet having the identifier imparted thereto; and the second terminal comprising a control unit that executes processing of receiving, from the information processing device, the speech packet having the identifier imparted thereto, and outputting the speech packet.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing system and information processing method

[0001] The present disclosure relates to an information processing system and an information processing method.

[0002] When having a conversation using intercommunication (intercom), the speaker presses the PTT (Push To Talk) button to turn on the microphone before speaking. The spoken voice is distributed to the members and played back, allowing members to talk to each other and have a conversation.

[0003] For example, in the case of smartphone intercom apps, functions are becoming increasingly common that allow users to record these utterances as audio or convert them into text, thereby keeping a history of past conversations that can be listened to or viewed later.

[0004] JP 2006-107044 A JP 2023-025464 A

[0005] However, because all utterances are treated equally and recorded, even if there is an important utterance among the many recorded utterances, it may be buried under the many other utterances, and it may take time to check it later, or it may be overlooked.

[0006] Therefore, the present disclosure proposes an information processing system and an information processing method that can more easily distinguish a specific utterance from multiple utterances via an intercom or the like.

[0007] According to the present disclosure, there is provided an information processing system including: a first terminal having a control unit that converts a speaker's voice into a voice packet and executes a process of transmitting the voice packet to an information processing device; an information processing device having a control unit that receives a voice packet from the first terminal, determines whether the voice packet contains predetermined identification information, and if the voice packet contains the identification information, assigns a predetermined identifier to the voice packet, stores it, and executes a process of transmitting the voice packet with the identifier assigned to a second terminal; and a second terminal having a control unit that receives the voice packet with the identifier assigned from the information processing device and executes a process of outputting the voice packet with the identifier assigned.

[0008] FIG. 1 is a diagram illustrating an example of the configuration of an information processing system 1 according to the present embodiment. FIG. 2 is a block diagram illustrating an example of the functional configuration of an information processing device 100 according to the present embodiment. FIG. 3 is a diagram illustrating an example of an important utterance according to the present embodiment. FIG. 4 is a diagram illustrating an example of a voice packet according to the present embodiment. FIG. 5 is a block diagram illustrating an example of the functional configuration of a user terminal 10 according to the present embodiment. FIG. 6 is a diagram illustrating an example of the processing configuration of an utterance sending side according to the present embodiment. FIG. 7 is a diagram illustrating an example of the processing configuration of an utterance receiving side according to the present embodiment. FIG. 8 is a diagram illustrating an example of the output of an important utterance according to the present embodiment. FIG. 9 is a flowchart illustrating the flow of a process for determining a specific utterance according to the present embodiment. FIG. 10 is a flowchart illustrating another example of the flow of a process for determining a specific utterance according to the present embodiment. FIG. 11 is a block diagram illustrating an example of the hardware configuration of an information processing device 100 according to the present embodiment.

[0009] The present embodiment will be described in detail below with reference to the drawings. In this specification and the drawings, substantially the same components are designated by the same reference numerals, and redundant description will be omitted.

[0010] The description will be given in the following order: 1. Embodiment 1.1. Example of functional configuration 1.2. Functional flow 2. Example of hardware configuration 3. Summary

[0011] <1. Embodiment> <<1.1. Functional Configuration Example>> Information processing implemented by an information processing device 100 and the like according to this embodiment will be described using FIG. 1 . FIG. 1 is a diagram illustrating an example configuration of an information processing system 1 according to this embodiment. In the example illustrated in FIG. 1 , the information processing system 1 includes user terminals 10-1 to 10-n (n is an arbitrary natural number; hereinafter, collectively referred to as "user terminals 10") and an information processing device 100. As illustrated in FIG. 1 , the user terminals 10 and the information processing device 100 are connected to each other via a network 50, for example, via a wired or wireless connection, allowing mutual communication. The network 50 is a communication network such as a local area network (LAN), a wide area network (WAN), a telephone network (such as a mobile phone network or a landline network), a regional Internet Protocol (IP) network, or the Internet. The network 50 may include a wired network or a wireless network.

[0012] The user terminal 10 shown in Fig. 1 is, for example, an information processing terminal used by a speaker and each user who converses with the speaker. The user terminal 10 may be, for example, an intercommunication device (intercom), or may be a desktop personal computer (PC), a notebook PC, a tablet terminal, a mobile phone, or the like, on which an intercom app is pre-installed. In the example shown in Fig. 1, the user terminal 10 is a smart device such as a smartphone or tablet terminal used by a user.

[0013] 1 receives a speech from a speaker in response to, for example, the speaker pressing a PTT button, converts the speech of the speaker into a voice packet, and transmits the voice packet to, for example, the information processing device 100.

[0014] Furthermore, the user terminal 10 receives, for example, a voice packet from a speaker and outputs the voice packet. The voice packet may be output by converting the voice packet into voice, or by reading or displaying the voice converted into text. Note that, as will be described in detail later, if the voice packet is a specific utterance such as an important utterance, the user terminal 10 can notify the user of this in a manner that the user can understand.

[0015] 1 is an information processing device such as a desktop PC, a notebook PC, or a server computer managed by a service provider that provides an intercom app, for example. Alternatively, the information processing device 100 may be a cloud computer managed by a cloud service provider that provides a cloud computing service, for example. While the information processing device 100 is shown as a single computer in FIG. 1, the information processing device 100 may be configured as a distributed computing system using multiple computers, for example.

[0016] The information processing device 100 also receives, for example, from the user terminal 10, voice packets into which the speaker's voice has been converted. The information processing device 100 also determines, for example, whether the received voice packets contain predetermined identification information. If the voice packets contain identification information, the information processing device 100 assigns a predetermined identifier to the voice packets and stores them. The information processing device 100 also transmits, for example, the voice packets with the assigned identifier to the user terminal 10.

[0017] Next, an information processing device 100 according to this embodiment will be described. Fig. 2 is a block diagram showing an example of the functional configuration of the information processing device 100 according to this embodiment. As shown in Fig. 2, the information processing device 100 according to this embodiment includes, for example, a communication unit 110, a storage unit 120, a transmission / reception unit 130, a determination unit 140, an assignment unit 150, and a control unit 190.

[0018] (Communication Unit 110) The communication unit 110 according to this embodiment is connected to various communication networks such as the Internet wirelessly or via a wire, and transmits and receives information to and from other devices such as the user terminal 10 on the network 50.

[0019] (Storage unit 120) The storage unit 120 according to this embodiment is a storage area for temporarily or permanently storing various programs and data. For example, the storage unit 120 can store programs and data for the information processing device 100 to execute various functions. As a specific example, the storage unit 120 may store voice packets into which the speaker's voice has been converted, which have been received from the speaker's user terminal 10. Note that these are merely examples, and the types of data stored in the storage unit 120 are not particularly limited.

[0020] (Transmitter / Receiver 130) The transmitter / receiver 130 according to this embodiment receives, for example, a voice packet obtained by converting the voice of a speaker from a first terminal which is the user terminal 10 used by the speaker.

[0021] Furthermore, the transmitting / receiving unit 130 transmits voice packets to which an identifier for identifying a specific utterance, such as an important utterance, has been assigned to a second terminal, which is the user terminal 10 used by the user having a conversation with the speaker. It should be noted that the transmitting / receiving unit 130 may also transmit voice packets other than the voice packets to which an identifier has been assigned to the second terminal. In particular, when determining whether or not predetermined identification information is included in a voice packet on the second terminal side, which is the user terminal 10 used by the user having a conversation with the speaker, the transmitting / receiving unit 130 transmits the voice packet received from the first terminal as is (without assigning an identifier) ​​to the second terminal.

[0022] (Determination Unit 140) The determination unit 140 according to this embodiment determines whether predetermined identification information is included in a voice packet received from a first terminal, which is, for example, the user terminal 10 used by the speaker. The identification information is, for example, a predetermined keyword (e.g., "important") during the utterance. FIG. 3 is a diagram illustrating an example of an important utterance according to this embodiment. As illustrated in FIG. 3, for example, in an important utterance, the speaker utters a predetermined specific keyword (e.g., "important") after turning on the PTT button of the intercom (user terminal 10) and before making a specific utterance. This allows the speaker to indicate that the utterance is important. Note that, for example, in FIG. 3, the timing for uttering the predetermined specific keyword is after turning on the PTT button and before making a specific utterance. However, this is merely an example, and the timing may be during the specific utterance, after the specific utterance, before turning off the PTT button, or the like. Therefore, the process executed by the determination unit 140 to determine whether or not an audio packet contains identification information may include, for example, a process to determine whether or not a keyword is contained in the audio during a predetermined period from the start of or before the end of the audio data in the audio packet.

[0023] Alternatively, the predetermined identification information in the voice packet received from the first terminal, which is the user terminal 10 used by the speaker, is, for example, a predetermined signal. The predetermined signal is, for example, a signal (data) transmitted in response to a predetermined user operation of a PTT button on the intercom (user terminal 10).

[0024] (Assignment Unit 150) When a voice packet received from a first terminal includes predetermined identification information, for example, the assignment unit 150 according to this embodiment assigns a predetermined identifier to the voice packet and stores the voice packet in the storage unit 120. Fig. 4 is a diagram showing an example of a voice packet according to this embodiment. As shown in Fig. 4, the assignment unit 150 assigns an identifier to voice data in the voice packet, for example.

[0025] Regarding the assignment of the identifier, for example, instead of using only one type of identifier, an identifier corresponding to the content of the identification information may be assigned to the voice packet. More specifically, if there are multiple types of identification information, such as keywords "important," "needs consideration," and "homework," individual identifiers corresponding to each of the keywords may be assigned to the voice packet. Alternatively, an identifier indicating the level of importance may be assigned to the voice packet.

[0026] Furthermore, for example, the assigned identifier may be a hierarchical identifier according to the content of the identification information. More specifically, for example, if the identification information has multiple types of keywords such as "important," "needs consideration," and "homework," and "needs consideration" and "homework" are included under the "important" hierarchical level, a hierarchical identifier such as "important > need to consider" may be assigned to the voice packet.

[0027] For example, when the first terminal or the second terminal determines whether or not predetermined identification information is included in a voice packet and assigns the predetermined identifier, similar processing by the determination unit 140 and the assignment unit 150 may not be performed. In this case, for example, the information processing device 100 may not include the determination unit 140 and the assignment unit 150.

[0028] (Control Unit 190) The control unit 190 according to this embodiment controls each component included in the information processing device 100. Note that the components shown in Fig. 2 are merely examples, and the components controlled by the control unit 190 are not limited to those shown in Fig. 2 .

[0029] Next, a user terminal 10 according to this embodiment will be described. Fig. 5 is a block diagram showing an example of the functional configuration of the user terminal 10 according to this embodiment. As shown in Fig. 5, the user terminal 10 according to this embodiment includes, for example, a communication unit 11, a storage unit 12, an input unit 13, a voice recognition unit 14, a PTT processing unit 15, a transmission / reception unit 16, an output unit 17, and a control unit 19.

[0030] (Communication Unit 11) The communication unit 110 according to this embodiment is connected to various communication networks such as the Internet wirelessly or via a wire, and transmits and receives information to and from other devices such as the information processing device 100 on the network 50.

[0031] (Storage Unit 12) The storage unit 12 according to this embodiment is a storage area for temporarily or permanently storing various programs and data. For example, the storage unit 12 can store programs and data for the user terminal 10 to execute various functions. As a specific example, the storage unit 12 may store voice packets converted from the speaker's voice. Note that the above is merely an example, and the type of data stored in the storage unit 12 is not particularly limited.

[0032] (Input Unit 13) The input unit 13 according to this embodiment accepts, for example, input of speech by a speaker. Fig. 6 is a diagram showing an example of a processing configuration on the speech transmission side according to this embodiment. As shown in Fig. 6, the input unit 13 accepts, for example, input of speech data of a speaker input via a microphone.

[0033] (Speech Recognition Unit 14) The speech recognition unit 14 according to this embodiment converts a speech of a speaker input via a microphone into a speech packet by speech recognition processing, as shown in Fig. 6. The speech packet conversion processing may include, for example, a process of transcribing the speech of the speaker by speech recognition and converting it into text as speech data in a speech packet.

[0034] (PTT Processing Unit 15) The PTT processing unit 15 according to the present embodiment receives a PTT button operation by a user and executes PTT processing in response to the operation (e.g., starting or ending a speech by turning the PTT button ON / OFF), as shown in Fig. 6. Furthermore, for example, when predetermined identification information is to be included in a voice packet by a predetermined user operation of the PTT button, the PTT processing unit 15 includes a predetermined signal (data) in response to the predetermined user operation as identification information in the voice packet.

[0035] 6, the transmitter / receiver 16 according to the present embodiment executes a voice transmission process and transmits voice packets into which the speaker's voice has been converted to the information processing device 100. In this way, for example, the user terminal 10 is an information processing terminal used by the speaker and can be a speech transmitting side that transmits the speaker's speech, but on the other hand, it can also be a user terminal 10 used by a user who converses with the speaker (naturally, the speaker may also be a listener), and can also be a speech receiving side that receives the speaker's speech.

[0036] Therefore, the transmitting / receiving unit 16 executes a voice reception process, for example, as shown in Fig. 7, and receives voice packets assigned with a predetermined identifier from the information processing device 100. Fig. 7 is a diagram showing an example of a processing configuration on the speech receiving side according to this embodiment. Naturally, the transmitting / receiving unit 16 may also receive voice packets other than voice packets assigned with an identifier from the information processing device 100.

[0037] (Output Unit 17) The output unit 17 according to the present embodiment outputs voice packets to which an identifier for identifying a specific utterance, such as an important utterance, is assigned. The output process of the voice packets may include, for example, a process of outputting information relating to the voice packets to which the identifier is assigned, distinguishing it from information relating to other voice packets.

[0038] More specifically, the output unit 17 executes an audio data display process, as shown in FIG. 7, and displays, on the display of the user terminal 10, a user interface (UI) selected for playing the audio data, a UI indicating the importance of the audio data, and the like.

[0039] Furthermore, in displaying the UI, the output unit 17, for example, displays information about voice packets to which an identifier has been assigned in a different color from information about other voice packets by highlighting or assigning a different color to the information. FIG. 8 is a diagram showing an example of output of important utterances according to this embodiment. For example, FIG. 8 shows a UI selected to play back the utterances of each user, and assumes that only the utterance of person B is important. In this case, the output unit 17 displays the UI for person B's utterance in a different color or by highlighting it, as shown in FIG. 8, to distinguish it from the UI for other normal utterances. Alternatively, the output unit 17 may, for example, hide information about the other voice packets to distinguish and display information about voice packets to which an identifier has been assigned.

[0040] For example, as shown in Figure 7, an identifier determination process may be performed in the user terminal 10 to determine whether or not a voice packet contains predetermined identification information, and if the identification information is included, a predetermined identifier for identifying a specific utterance, such as an important utterance, may be assigned to the voice packet, and information regarding the voice packet may be output (displayed).

[0041] Furthermore, the output of the voice packets into which the speaker's voice has been converted may be played back in real time, rather than being played back later. Therefore, the output process of the voice packets by the output unit 17 may include, for example, a process of outputting the voice of the speaker in the voice packets to which the identifier has been assigned and other voice packets in real time.

[0042] Furthermore, the output unit 17 outputs a predetermined sound or voice during a predetermined period before or after outputting the voice of the speaker of a voice packet to which an identifier for identifying a specific utterance such as an important utterance has been assigned. The predetermined sound or voice may be, for example, a predetermined beep such as "beep beep beep" for identifying a specific utterance such as an important utterance, or a predetermined voice such as "This is an important conversation."

[0043] (Control Unit 19) The control unit 19 according to this embodiment controls each component included in the user terminal 10. Note that the components shown in Fig. 5 are merely examples, and the components controlled by the control unit 19 are not limited to those shown in Fig. 5 .

[0044] <<1.2. Functional Flow>> Next, the procedure for the specific utterance determination process will be described with reference to Fig. 9. Fig. 9 is a flowchart showing the flow of the specific utterance determination process according to this embodiment. The determination process shown in Fig. 9 may be executed in response to, for example, the PTT button of the intercom (user terminal 10) being turned on.

[0045] First, as shown in FIG. 9, the user terminal 10 receives, for example, a voice input from a speaker and converts the voice into voice packets through a voice recognition process (step S101).

[0046] Next, the user terminal 10 transmits, for example, the voice packets converted in step S101 to the information processing device 100 (step S102).

[0047] Next, the information processing device 100 receives, for example, the voice packets transmitted in step S102 from the user terminal 10 used by the speaker (i.e., the speech originator) (step S103).

[0048] Next, the information processing device 100 determines whether or not predetermined identification information is included in the voice packet received in step S103 (step S104).

[0049] If the voice packet contains predetermined identification information (step S105: Yes), the information processing device 100 assigns, for example, a predetermined identifier to the voice packet (step S106). On the other hand, if the voice packet does not contain predetermined identification information (step S105: No), the information processing device 100 skips step S106 and proceeds to step S107.

[0050] Next, the information processing device 100 stores in the memory unit 120, for example, if an identifier was assigned in step S106, the voice packet to which the identifier was assigned, or if an identifier was not assigned, the voice packet received in step S103 (step S107).

[0051] Next, the information processing device 100 transmits the voice packet with the identifier assigned, for example, if an identifier was assigned in step S106, or the voice packet received in step S103, if no identifier was assigned, to the user terminal 10 (i.e., the voice receiving side) used by the user who is conversing with the speaker (step S108).

[0052] Next, the user terminal 10 receives, for example, the voice packets transmitted in step S108 from the information processing device 100 (step S109).

[0053] Next, the user terminal 10 outputs, for example, the voice packets received in step S109 (step S110). Note that the output of the voice packets may be, for example, outputting information about the voice packets to which an identifier has been assigned in step S106, distinguishing it from information about other voice packets. After step S110 is executed, the determination process shown in Fig. 9 ends. Note that, in the determination process shown in Fig. 9, an example is shown in which the determination of the identification information (steps S104 and S105) and the assignment of the identifier (step S106) are performed by the information processing device 100, but these processes may be performed by the user terminal 10 on the utterance originating side or the user terminal 10 on the utterance receiving side.

[0054] Next, another example of the procedure for the specific utterance determination process will be described with reference to Fig. 10. Fig. 10 is a flowchart showing another example of the flow of the specific utterance determination process according to this embodiment. The determination process shown in Fig. 10 may be executed in response to, for example, the PTT button of the intercom (user terminal 10) being turned on.

[0055] First, as shown in FIG. 10, the user terminal 10 receives, for example, a voice input from a speaker and converts the voice into voice packets through a voice recognition process (step S201).

[0056] Next, the user terminal 10 determines whether or not predetermined identification information is included in the voice packet converted in step S201 (step S202).

[0057] If the voice packet contains predetermined identification information (step S203: Yes), the user terminal 10 assigns, for example, a predetermined identifier to the voice packet (step S204). On the other hand, if the voice packet does not contain predetermined identification information (step S203: No), step S204 is skipped and the process proceeds to step S205.

[0058] Next, the user terminal 10 transmits to the information processing device 100, for example, if an identifier was assigned in step S204, the voice packet with the identifier assigned, or if no identifier was assigned, the voice packet converted in step S201 (step S205).

[0059] Next, the information processing device 100 receives, for example, the voice packets transmitted in step S205 from the user terminal 10 used by the speaker (i.e., the speech originator) (step S206).

[0060] Next, the information processing device 100 transmits, for example, the voice packets received in step S206 to the user terminal 10 used by the user who will converse with the speaker (i.e., the speech receiving side) (step S207).

[0061] Next, the user terminal 10 receives, for example, the voice packets transmitted in step S207 from the information processing device 100 (step S208).

[0062] Next, the user terminal 10 outputs, for example, the voice packets received in step S208 (step S209). Note that the voice packets may be output by distinguishing information about the voice packets to which identifiers have been assigned in step S204 from information about other voice packets. After step S209 is executed, the determination process shown in FIG. 10 ends.

[0063] 2. Hardware Configuration Example Next, a hardware configuration example of the information processing device 100 according to the present embodiment will be described. The information devices of the user terminal 10 and the information processing device 100 according to the above-described embodiments are realized, for example, by a computer 1000 configured as shown in FIG. 11 . The following description will be given using the information processing device 100 as an example. FIG. 11 is a block diagram showing a hardware configuration example of the information processing device 100 according to the present embodiment. The computer 1000 has a CPU 1100, a RAM 1200, a ROM 1300, a HDD (Hard Disk Drive) 1400, a communication interface 1500, and an input / output interface 1600. The components of the computer 1000 are connected by a bus 1050.

[0064] The CPU 1100 operates and controls each component based on programs stored in the ROM 1300 or the HDD 1400. For example, the CPU 1100 loads the programs stored in the ROM 1300 or the HDD 1400 into the RAM 1200 and executes processing corresponding to the various programs.

[0065] The ROM 1300 stores boot programs such as a Basic Input Output System (BIOS) that is executed by the CPU 1100 when the computer 1000 is started, and programs that depend on the hardware of the computer 1000 .

[0066] HDD 1400 is a computer-readable recording medium that non-temporarily records programs executed by CPU 1100 and data used by such programs. Specifically, HDD 1400 is a recording medium that records the proposed program according to the present disclosure, which is an example of program data 1450.

[0067] The communication interface 1500 is an interface for connecting the computer 1000 to an external network 1550 (e.g., the Internet). For example, the CPU 1100 receives data from other devices and transmits data generated by the CPU 1100 to other devices via the communication interface 1500.

[0068] The input / output interface 1600 is an interface for connecting the input / output device 1650 and the computer 1000. For example, the CPU 1100 receives data from an input device such as a keyboard or a mouse via the input / output interface 1600. The CPU 1100 also transmits data to an output device such as a display, a speaker, or a printer via the input / output interface 1600. The input / output interface 1600 may also function as a media interface for reading programs recorded on a predetermined recording medium. Examples of media include optical recording media such as a DVD (registered trademark) (Digital Versatile Disc) or a PD (Phase Change Rewritable Disk), magneto-optical recording media such as an MO (Magneto-Optical disk), tape media, magnetic recording media, or semiconductor memory.

[0069] For example, when the computer 1000 functions as the information processing device 100 according to the embodiment, the CPU 1100 of the computer 1000 executes the proposed program loaded onto the RAM 1200, thereby realizing functions such as the control unit 190. The proposed program according to the present disclosure and data in the storage unit 120 are stored in the HDD 1400. The CPU 1100 reads and executes the program data 1450 from the HDD 1400, but as another example, the CPU 1100 may obtain these programs from another device via an external network 1550.

[0070] 3. Summary As described above, the information processing system 1 includes a first terminal including a control unit that converts the voice of a speaker into voice packets and transmits the voice packets to an information processing device; an information processing device that includes a control unit that receives voice packets from the first terminal, determines whether the voice packets include predetermined identification information, and, if the voice packets include the identification information, assigns a predetermined identifier to the voice packets, stores the voice packets, and transmits the voice packets with the identifier assigned to a second terminal; and a second terminal that receives the voice packets with the identifier assigned from the information processing device and includes a control unit that executes processing to output the voice packets with the identifier assigned.

[0071] In this way, if the voice packet into which the speaker's voice is converted contains predetermined identification information to indicate that it is a specific utterance, such as an important utterance, by assigning a predetermined identifier to the voice packet, the information processing system 1 can more easily distinguish the specific utterance from multiple utterances via an intercom, etc.

[0072] The identification information is a predetermined keyword in the speech or a predetermined signal.

[0073] This allows the speaker to indicate in various ways that an utterance is a particular utterance, such as an important utterance.

[0074] Furthermore, the process performed by the information processing device 100 to determine whether or not identification information is included includes a process to determine whether or not a keyword is included in the audio during a predetermined period from the start of or before the end of the audio data in the audio packet.

[0075] This allows the speaker to indicate at various times that an utterance is a particular utterance, such as an important utterance.

[0076] The first terminal and the second terminal are in intercommunication (intercom).

[0077] This allows the information processing device 100 and the like to more easily distinguish a specific utterance even when multiple utterances are made, such as through an intercom.

[0078] The predetermined signal is a signal that is transmitted in response to a predetermined user operation of the PTT button of the intercom.

[0079] This allows the speaker to indicate that the utterance is a specific utterance, such as an important utterance, by operating the intercom.

[0080] Furthermore, the process of converting the speaker's voice into voice packets by the first terminal includes a process of transcribing the speaker's voice using voice recognition and converting it into text as voice data in the voice packets.

[0081] This allows the information processing device 100 and the like to assign a predetermined identifier not only to the speaker's voice itself but also to the voice data converted into text, making it easier to distinguish between multiple utterances.

[0082] Furthermore, the process of assigning an identifier to a voice packet and storing the same by the information processing device 100 includes a process of assigning an identifier to a voice packet according to the content of the identification information and storing the voice packet.

[0083] This allows the information processing device 100 and the like to better classify and distinguish multiple utterances.

[0084] Furthermore, the process of assigning an identifier to a voice packet and storing the same by the information processing device 100 includes a process of assigning a hierarchical identifier to the voice packet according to the content of the identification information and storing the voice packet.

[0085] This allows the information processing device 100 and the like to better classify and distinguish multiple utterances.

[0086] Furthermore, the process of outputting the voice packets to which the identifiers have been assigned by the second terminal includes a process of outputting information relating to the voice packets to which the identifiers have been assigned, distinguishing it from information relating to other voice packets.

[0087] This allows the information processing device 100 and the like to more easily distinguish a specific utterance from multiple utterances.

[0088] In addition, the process by the second terminal of outputting the voice packet to which the identifier has been assigned includes a process of displaying information relating to the voice packet to which the identifier has been assigned in a manner that distinguishes it from information relating to other voice packets by assigning a different color to the information or by highlighting the information.

[0089] This allows the information processing device 100 and the like to more easily distinguish a specific utterance from multiple utterances.

[0090] In addition, the process by the second terminal of outputting a voice packet to which an identifier has been assigned includes a process of displaying information relating to the voice packet to which the identifier has been assigned in a distinguishable manner by hiding information relating to other voice packets.

[0091] This allows the information processing device 100 and the like to more easily distinguish a specific utterance from multiple utterances.

[0092] Furthermore, the process of outputting the voice packet to which the identifier is assigned by the second terminal includes a process of outputting the voice of the speaker of the voice packet to which the identifier is assigned in real time.

[0093] This allows the information processing device 100 and the like to more easily distinguish a specific utterance from among multiple utterances exchanged in real time.

[0094] The second terminal executes a process of outputting a predetermined sound or voice during a predetermined period before or after outputting the voice of the speaker of the voice packet to which the identifier is assigned.

[0095] This allows the information processing device 100 and the like to inform each user conversing with the speaker that the utterance is a specific utterance, such as an important utterance, among multiple utterances exchanged in real time.

[0096] The information processing system 1 also includes a first terminal having a control unit that converts the speaker's voice into a voice packet, determines whether the voice packet contains predetermined identification information, and if the voice packet contains the identification information, assigns a predetermined identifier to the voice packet and transmits the voice packet to an information processing device; an information processing device having a control unit that receives the voice packet with the identifier from the first terminal and transmits the voice packet with the identifier to a second terminal; and a second terminal having a control unit that receives the voice packet with the identifier from the information processing device and outputs the voice packet with the identifier.

[0097] This allows the information processing system 1 to more easily distinguish a specific utterance from multiple utterances via an intercom or the like.

[0098] Although the preferred embodiments of the present disclosure have been described in detail above with reference to the accompanying drawings, the technical scope of the present disclosure is not limited to such examples. It is clear that a person skilled in the art of the present disclosure can conceive of various modified or altered examples within the scope of the technical idea described in the claims, and it is understood that these also naturally fall within the technical scope of the present disclosure.

[0099] Furthermore, the effects described herein are merely descriptive or exemplary and are not limiting. In other words, the technology according to the present disclosure may achieve other effects that will be apparent to those skilled in the art from the description of this specification, in addition to or in place of the above-described effects.

[0100] The present technology can also be configured as follows: (1) An information processing system including: a first terminal including a control unit that performs processing to convert a speaker's voice into voice packets and transmit the voice packets to an information processing device; the information processing device including a control unit that performs processing to receive the voice packets from the first terminal, determine whether the voice packets include predetermined identification information, and if the voice packets include the identification information, assign a predetermined identifier to the voice packets and store them, and transmit the voice packets with the identifier assigned to a second terminal; and the second terminal including a control unit that performs processing to receive the voice packets with the identifier assigned from the information processing device and output the voice packets with the identifier assigned. (2) The information processing system according to (1), wherein the identification information is a predetermined keyword or a predetermined signal during speech. (3) The information processing system according to (2), wherein the process by the information processing device to determine whether the identification information is included includes a process of determining whether the keyword is included in the voice for a predetermined period from the start of or before the end of the voice data in the voice packet. (4) The information processing system according to (2) or (3), wherein the first terminal and the second terminal are intercommunication devices (intercoms). (5) The information processing system according to (4), wherein the predetermined signal is a signal transmitted in response to a predetermined user operation of a PTT (Push To Talk) button of the intercom. (6) The information processing system according to any one of (1) to (5), wherein the process by the first terminal to convert the speaker's voice into the voice packet includes a process of transcribing the speaker's voice using speech recognition and converting the transcribed voice data into text as voice data in the voice packet. (7) The information processing system described in any one of (1) to (6), wherein the process by the information processing device of assigning the identifier to the voice packet and storing it includes a process of assigning the identifier to the voice packet according to the content of the identification information and storing it.(8) The information processing system according to any one of (1) to (7), wherein the process by the information processing device of assigning the identifier to the voice packet and storing the identifier includes a process of assigning the identifier for each layer according to the content of the identification information to the voice packet and storing the voice packet. (9) The information processing system according to any one of (1) to (8), wherein the process by the second terminal of outputting the voice packet to which the identifier has been assigned includes a process of outputting information about the voice packet to which the identifier has been assigned, distinguishing it from information about other voice packets. (10) The information processing system according to any one of (1) to (9), wherein the process by the second terminal of outputting the voice packet to which the identifier has been assigned includes a process of displaying information about the voice packet to which the identifier has been assigned, distinguishing it from information about other voice packets by assigning a different color to the information about the voice packet to which the identifier has been assigned or by highlighting it. (11) The information processing system according to any one of (1) to (10), wherein the process by the second terminal to output the voice packet to which the identifier is assigned includes: distinguishing and displaying information about the voice packet to which the identifier is assigned by hiding information about other voice packets. (12) The information processing system according to any one of (1) to (11), wherein the process by the second terminal to output the voice packet to which the identifier is assigned includes: outputting, in real time, the voice of the speaker of the voice packet to which the identifier is assigned. (13) The information processing system according to any one of (1) to (12), wherein the second terminal executes: outputting a predetermined sound or voice for a predetermined period before or after outputting the voice of the speaker of the voice packet to which the identifier is assigned.(14) An information processing system including: a first terminal including a control unit that executes the following processes: converting a speaker's voice into a voice packet; determining whether the voice packet includes predetermined identification information; if the voice packet includes the identification information, assigning a predetermined identifier to the voice packet; and transmitting the voice packet to an information processing device; the information processing device including a control unit that executes the following processes: receiving the voice packet with the identifier assigned from the first terminal; and transmitting the voice packet with the identifier assigned to a second terminal; and the second terminal including a control unit that executes the following processes: receiving the voice packet with the identifier assigned from the information processing device; and outputting the voice packet with the identifier assigned. (15) An information processing method for transmitting and receiving voice packets converted from a speaker's voice, comprising: a computer converting the speaker's voice into voice packets; the computer determining whether the voice packets contain predetermined identification information; if the voice packets contain the identification information, the computer assigning a predetermined identifier to the voice packets; and if the computer receives the voice packets with the assigned identifier, storing the voice packets in a storage device.

[0101] REFERENCE SIGNS LIST 10 User terminal 11 Communication unit 12 Storage unit 13 Input unit 14 Voice recognition unit 15 PTT processing unit 16 Transmitting / receiving unit 17 Output unit 19 Control unit 50 Network 100 Information processing device 110 Communication unit 120 Storage unit 130 Transmitting / receiving unit 140 Determination unit 150 Assignment unit 190 Control unit 1000 Computer 1050 Bus 1100 CPU 1200 RAM 1300 ROM 1400 HDD 1450 Program data 1500 Communication interface 1550 External network 1600 Input / output interface 1650 Input / output device

Claims

1. A first terminal comprising a control unit that executes a process of converting a speaker's voice into voice packets and transmitting the voice packets to an information processing device, an information processing device comprising a control unit that receives the voice packets from the first terminal, determines whether predetermined identification information is included in the voice packets, and when the identification information is included in the voice packets, assigns and stores a predetermined identifier to the voice packets and transmits the voice packets with the assigned identifier to a second terminal, and the second terminal comprising a control unit that receives the voice packets with the assigned identifier from the information processing device and executes a process of outputting the voice packets with the assigned identifier, the information processing system comprising the above components.

2. The information processing system according to claim 1, wherein the identification information is a predetermined keyword during speaking or a predetermined signal.

3. The information processing system according to claim 2, wherein the process of determining by the information processing device whether the identification information is included includes a process of determining whether the keyword is included in the voice during a predetermined period from the start or before the end of the voice data in the voice packets.

4. The information processing system according to claim 2, wherein the first terminal and the second terminal are intercommunications (Incom).

5. The information processing system according to claim 4, wherein the predetermined signal is a signal transmitted in response to a predetermined user operation of the push-to-talk (PTT) button of the Incom.

6. The information processing system according to claim 1, wherein the process of converting the speaker's voice into voice packets by the first terminal includes a process of converting the speaker's voice into text by voice recognition and converting the text into voice data in the voice packets.

7. The information processing system according to claim 1, wherein the process of assigning and storing the identifier to the voice packets by the information processing device includes a process of assigning and storing the identifier corresponding to the content of the identification information to the voice packets.

8. The process of the information processing apparatus for attaching and storing the identifier to the voice packet includes a process of attaching and storing a hierarchical identifier corresponding to the content of the identification information to the voice packet. The information processing system according to claim 1.

9. The process of the second terminal for outputting the voice packet with the identifier attached includes a process of outputting information related to the voice packet with the identifier attached separately from information related to other voice packets. The information processing system according to claim 1.

10. The process of the second terminal for outputting the voice packet with the identifier attached includes a process of displaying the information related to the voice packet with the identifier attached separately from the information related to other voice packets by assigning different colors or highlighting. The information processing system according to claim 1.

11. The process of the second terminal for outputting the voice packet with the identifier attached includes a process of displaying the information related to the voice packet with the identifier attached separately by making the information related to other voice packets non-displayed. The information processing system according to claim 1.

12. The process of the second terminal for outputting the voice packet with the identifier attached includes a process of outputting the voice of the speaker of the voice packet with the identifier attached in real time. The information processing system according to claim 1.

13. The second terminal executes a process of outputting a predetermined sound or voice during a predetermined period before or after the output of the voice of the speaker of the voice packet with the identifier attached. The information processing system according to claim 12.

14. A first terminal including a control unit that executes a process of converting a speaker's voice into voice packets, determining whether the voice packets include predetermined identification information, assigning a predetermined identifier to the voice packets when the voice packets include the identification information, and transmitting the voice packets to an information processing device; the information processing device including a control unit that executes a process of receiving the voice packets with the identifier assigned thereto from the first terminal and transmitting the voice packets with the identifier assigned thereto to a second terminal; and the second terminal including a control unit that executes a process of receiving the voice packets with the identifier assigned thereto from the information processing device and outputting the voice packets with the identifier assigned thereto. An information processing system.

15. An information processing method for transmitting and receiving voice packets converted from a speaker's voice, the method comprising: a computer converting the speaker's voice into voice packets; the computer determining whether the voice packets include predetermined identification information; the computer assigning a predetermined identifier to the voice packets when the voice packets include the identification information; and the computer storing the voice packets in a storage device when the voice packets with the identifier assigned thereto are received.

Citation Information

Patent Citations

  • How to improve the distinction between dictation and commands

    JP2004510239A

  • System and method for processing message

    JP2005228255A

  • Position sharing service providing device, method thereof, and computer program thereof

    JP2016126789A

  • Method and terminal for displaying instant messaging message

    US20170373994A1

  • Communication system and terminal device

    WO2013089236A1