Programs, systems, and methods for generating conversation records

The system addresses the issue of duplicate speech recording in conversation records by using a server device to identify and correctly attribute speech from multiple user devices, ensuring accurate and reliable meeting minutes.

JP2025085014APending Publication Date: 2025-06-03ALT PTY LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025036475
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

When generating a record of a conversation using multiple user devices, speech recognition can misidentify the speaker, leading to duplicate recording of a user's speech content, which can result in uncertain and unreliable meeting minutes.

Method used

A program, system, and method that utilize a server device with a processor unit to receive voice data from user devices, identify the speaker based on voiceprint or image analysis, and determine if the speaker is associated with the device that transmitted the voice data. If associated, the voice data is recorded correctly; if not, it is discarded to avoid duplicate recording.

Benefits of technology

This solution effectively prevents duplicate recording of speech content, enhancing the accuracy and reliability of conversation records by correctly attributing speech to the appropriate user.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025085014000001_ABST
    Figure 2025085014000001_ABST
Patent Text Reader

Abstract

To provide programs, etc. for generating conversation records.SOLUTION: A program for generating a record of a conversation using a plurality of user devices including a first user device and a second user device is provided. The program, when executed on a server device provided with a processor unit, causes the processor unit to execute processing that includes: receiving data indicating a first voice from the first user device; identifying a speaker of the first voice; determining whether the identified speaker is a user associated with the first user device; when the identified speaker is determined to be the user associated with the first user device, outputting data indicating the first voice data and data indicating that the speaker of the first voice is the identified speaker; and when it is determined that the identified speaker is not the user associated with the first user device, discarding the data indicating the first voice.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a program, a system, and a method for generating a record of a conversation. More specifically, the present invention relates to a program, a system, and a method for generating a record of a conversation using a plurality of user devices including a first user device and a second user device.

Background Art

[0002] There is known a system that uses speech recognition to create minutes of a meeting (for example, Patent Document 1).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] When generating a record of a conversation using a plurality of user devices including a first user device and a second user device by using speech recognition, the voice uttered by the first user to the microphone of the first user device used by the first user is output from the speaker of the second user device used by the second user, and the output voice is input to the second user device by the microphone of the second user device. As a result, the voice of the first user may be misrecognized as if it were uttered by the second user. In this case, in the generated record of the conversation, the speech of the first user is recorded, and the speech of the first user is further recorded as the speech of the second user.

[0005] The inventor of the present invention considered this to be a problem because if the speech content of the first user is recorded repeatedly, the generated minutes of the meeting may become uncertain.

[0006] The present invention has been made in view of the above problems, and an object thereof is to provide a program, a system, and a method for generating a conversation record that can avoid duplicate recording of the user's speech content.

Means for Solving the Problems

[0007] The present invention provides, for example, the following items.

[0008] (Item 1) A program for generating a conversation record using a plurality of user devices including a first user device and a second user device, wherein when the program is executed in a server device including a processor unit, Receiving data indicating a first voice from the first user device; Identifying the speaker of the first voice; Determining whether the identified speaker is a user associated with the first user device; When it is determined that the identified speaker is a user associated with the first user device, outputting data indicating the first voice data and data indicating that the speaker of the first voice is the identified speaker; When it is determined that the identified speaker is not a user associated with the first user device, discarding the data indicating the first voice A program that causes the processor unit to perform a process including the above. (Item 2) Determining whether the identified speaker is a user associated with the first user device includes Determining whether the identified speaker is a first user who holds a first account used in the first user device, and When it is determined that the identified speaker is the first user, the outputting includes outputting data indicating the first voice and data indicating that the speaker of the first voice is the first user. The program according to Item 1. (Item 3) Determining whether the identified speaker is a user associated with the first user device includes determining whether the identified speaker is a guest user who uses a first account used in the first user device, When it is determined that the identified speaker is the guest user, the outputting includes outputting data indicating the first voice and data indicating that the speaker of the first voice is the guest user, according to the program described in item 1 or item 2. (Item 4) Identifying the speaker of the first voice includes identifying the speaker of the first voice based on the voiceprint of the first voice, according to the program described in any one of items 1 to 3. (Item 5) Identifying the speaker of the first voice includes identifying the speaker of the first voice based on an image when the first voice is uttered, according to the program described in any one of items 1 to 3. (Item 6) The outputting includes causing the first user device to display text indicating the first voice and an icon indicating the speaker, according to the program described in any one of items 1 to 5. (Item 7) The outputting includes recording data indicating the first voice and data indicating that the speaker of the voice is the identified speaker, according to the program described in any one of items 1 to 6. (Item 8) A system for generating a record of a conversation using a plurality of user devices including a first user device and a second user device, receiving means for receiving data indicating a first voice from the first user device, identifying means for identifying the speaker of the first voice, determining means for determining whether the identified speaker is a user associated with the first user device, Output means for outputting data indicating the first voice data and data indicating that the speaker of the first voice is the identified speaker when it is determined that the identified speaker is a user associated with the first user device; Discarding means for discarding data indicating the first voice when it is determined that the identified speaker is not a user associated with the first user device A system comprising. (Item 9) A method for generating a record of a conversation using a plurality of user devices including a first user device and a second user device, comprising: Receiving data indicating a first voice from the first user device; Identifying the speaker of the first voice; Determining whether the identified speaker is a user associated with the first user device; Outputting data indicating the first voice data and data indicating that the speaker of the first voice is the identified speaker when it is determined that the identified speaker is a user associated with the first user device; Discarding the data indicating the first voice when it is determined that the identified speaker is not a user associated with the first user device A method including. (Item 10) A program for generating a record of a conversation using a plurality of user devices including a first user device and a second user device, the program, when executed on the first user device including voice input means and a processor unit, Receiving data indicating a first voice input via the voice input means; Identifying the speaker of the first voice; Determining whether the identified speaker is a user associated with the first user device; When it is determined that the identified speaker is a user associated with the first user device, output data indicating the first voice and data indicating that the speaker of the first voice is the identified speaker. When it is determined that the identified speaker is not a user associated with the first user device, discard the data indicating the first voice. A program that causes the processor unit to perform a process including the above. (Item 11) A user device for generating a record of a conversation using a plurality of user devices including a first user device and a second user device, wherein the user device is the first user device. Voice input means for inputting voice. Receiving means for receiving data indicating a first voice input via the voice input means. Identifying means for identifying the speaker of the first voice. Determining means for determining whether or not the identified speaker is a user associated with the first user device. Output means for outputting data indicating the first voice and data indicating that the speaker of the first voice is the identified speaker when it is determined that the identified speaker is a user associated with the first user device. Discarding means for discarding the data indicating the first voice when it is determined that the identified speaker is not a user associated with the first user device. A user device comprising the above. (Item 12) A program for generating a record of a conversation using a plurality of user devices including a first user device and a second user device, the method being executed in the first user device, the method comprising: Receiving data indicating a first voice input via the voice input means of the first user device. Identifying the speaker of the first voice. Determining whether or not the identified speaker is a user associated with the first user device. When it is determined that the identified speaker is the user associated with the first user device, outputting data indicating the first voice and data indicating that the speaker of the first voice is the identified speaker; When it is determined that the identified speaker is not the user associated with the first user device, discarding the data indicating the first voice A method comprising:

Advantages of the Invention

[0009] According to the present invention, it is possible to provide a program, a system, and a method for generating a conversation record that can avoid duplicate recording of the user's speech content.

Brief Description of the Drawings

[0010]

Figure 1A

Figure 1B

Figure 1C

Figure 2A

Figure 2B

Figure 2C

Figure 2D

Figure 3

Figure 4A

Figure 4B

Figure 5

Figure 6

Figure 7

Figure 8

Mode for Carrying Out the Invention

[0011] Hereinafter, embodiments of the present invention will be described with reference to the drawings.

[0012] 1. Minutes generation application FIGS. 1 and 2 are diagrams schematically showing an example of how minutes of a meeting are generated using a minutes-of-meeting generation application developed by the present inventor.

[0013] In the example shown in FIGS. 1A to 1C, users A and B in two different spaces are having a conversation using user device 10 and user device 20. Users A and B each have an account for using the minutes-of-meeting generation application. User A is logged in to his / her account via user device 10. User B is logged in to his / her account via user device 20. FIGS. 1A(a), 1B(a), 1C(a) show the state of the conversation, and FIGS. 1A(b), 1B(b), 1C(b) show an example of the minutes of the meeting displayed on the display unit of user device 10 or user device 20 at that time.

[0014] First, as shown in Fig. 1A(a), when user A says "I agree." to user B, user A speaks "I agree." to the microphone of user device 10. Then, the minutes generation application recognizes the voice by user A (i.e., recognizes "I agree."), and by recognizing that the voice is the voice spoken by user A to user device 10, records "I agree." as the speech by user A.

[0015] Then, as shown in Fig. 1A(b), on the display unit of user device 10 or user device 20, an icon indicating user A and the speech content "I agree." are displayed. Thereby, it is visually understood that user A has said "I agree."

[0016] After user A's speech as shown in Fig. 1A(a), the speech is output from the speaker of user device 20 as shown in Fig. 1B(a). Thereby, user B can hear user A's speech. At this time, if the microphone of user device 20 picks up the voice output from the speaker of user device 20, the minutes generation application recognizes the voice output from the speaker of user device 20 (i.e., recognizes "I agree."), and by misrecognizing that the voice is the voice spoken by user B to user device 20, records "I agree." as the speech by user B.

[0017] Then, as shown in Fig. 1B(b), on the display unit of user device 10 or user device 20, an icon indicating user B and the speech content "I agree." are displayed. However, this recording and display are incorrect. Regardless of user B's intention, it is recorded as if user B also agreed, so the generated minutes are uncertain and of low credibility.

[0018] Therefore, the minutes generation application identifies the speaker of the voice input to the microphone of user device 20, and deletes the incorrect recording and display based on the identification result.

[0019] In this way, as shown in FIG. 1C(b), the icon indicating User B and the statement "I agree." are deleted, leaving only the correct records and generating correct minutes. This can enhance the credibility of the minutes.

[0020] In the examples shown in FIGS. 2A to 2D, User A, User B, and User C who are in two different spaces are having a conversation using User A's user device 10 and User B's user device 20. User A, User B, and User C each have an account for using the minutes generation application. User A and User C are in the same space. User A is logged in to his / her account via user device 10. User B is logged in to his / her account via user device 20. User C participates in the conversation as a guest user using user device 10. FIGS. 2A(a), 2B(a), 2C(a), and 2D(a) show the state of the conversation, and FIGS. 2A(b), 2B(b), 2C(b), and 2D(b) show an example of the minutes displayed on the display unit of user device 10 or user device 20 at that time.

[0021] First, as shown in FIG. 2A(a), when User C speaks "I agree." to User A and User B, User C speaks "I agree." to the microphone of user device 10. Then, the minutes generation application recognizes the voice by User C (i.e., recognizes "I agree."), and by misrecognizing that the voice is the voice spoken by User A to user device 10, records "I agree." as a statement by User A.

[0022] Then, as shown in FIG. 2A(b), on the display unit of user device 10 or user device 20, the icon indicating User A and the statement "I agree." are displayed. However, this record and display are incorrect. Regardless of User A's intention, it is recorded as if User A agreed, so the generated minutes are uncertain and have low credibility.

[0023] Therefore, the minutes generation application identifies the speaker of the voice input to the microphone of the user device 10, and based on the identification result, changes and records it to the correct speaker.

[0024] In this way, as shown in FIG. 2B(b), the icon indicating user A is changed to the icon indicating user C, and the icon indicating user C and the speech content "I agree." are displayed. The speech by the correct speaker (user C) is recorded, and correct minutes are generated.

[0025] After user C speaks as shown in FIG. 2A(a), the speech is output from the speaker of the user device 20 as shown in FIG. 2C(a). As a result, user B can hear the speech of user C. At this time, if the microphone of the user device 20 picks up the voice output from the speaker of the user device 20, the minutes generation application recognizes the voice output from the speaker of the user device 20 (that is, recognizes "I agree."), and by misrecognizing that the voice is the voice spoken by user B to the user device 20, records "I agree." as the speech by user B.

[0026] Then, as shown in FIG. 2C(b), the display unit of the user device 10 or the user device 20 displays the icon indicating user B and the speech content "I agree." However, this record and display are incorrect. Regardless of the intention of user B, it is recorded as if user B also agreed, so the generated minutes are uncertain and of low credibility.

[0027] Therefore, the minutes generation application identifies the speaker of the voice input to the microphone of the user device 20, and based on the identification result, deletes the incorrect record and display.

[0028] In this way, as shown in FIG. 2D, the icon indicating User B and the statement "I agree." are deleted, leaving only the correct records and generating the correct minutes of the meeting. This can enhance the credibility of the minutes of the meeting.

[0029] In the above example, it was described by taking as an example that each user has an account for using the minutes-of-meeting generation application, but the present invention is not limited to this. Even when each user does not have an account, it is possible to generate correct minutes of the meeting by the minutes-of-meeting generation application.

[0030] In the above example, it was described that after the misrecognized statement is recorded, based on the result of speaker identification, the speaker of the recorded statement is changed or the recorded statement is deleted, but the present invention is not limited to this. For example, before recording a statement, the speaker may be identified, and only correct information may be recorded based on the result of speaker identification. At this time, for example, in the example shown in FIG. 1B, the statement "I agree." by User B is neither recorded nor displayed.

[0031] The minutes-of-meeting generation application can be implemented, for example, by the system for generating the conversation record of the present invention described below.

[0032] 2. Configuration of the system for generating the conversation record of the present invention FIG. 3 shows an example of the configuration of a system 1000 for generating a conversation record.

[0033] The system 1000 includes at least one user device 100, a server device 200 connected to at least one user device 100 via a network 400, and a database unit 300 connected to the server device 200.

[0034] The user device 100 can be any terminal device such as a smartphone, tablet, personal computer, smart glasses, etc. The user device 100 can communicate with the server device 200 via the network 400. Here, the type of the network 400 is not limited. For example, the user device 100 may communicate with the server device 200 via the Internet, or may communicate with the server device 200 via a LAN. Although two user devices 100 are depicted in FIG. 2, the number of user devices 100 is not limited to this. The number of user devices 100 can be any number of 2 or more.

[0035] The server device 200 can communicate with at least one user device 100 via the network 400. Also, the server device 200 can communicate with the database unit 300 connected to the server device 200.

[0036] For example, information related to the user's account is stored in the database unit 300 connected to the server device 200. The information related to the user's account includes at least the user's identifier (e.g., username) and the features identifying the user (e.g., voiceprint, or face image, etc.). The information related to the user's account may also include other information. Note that the user's identifier and the features identifying the user do not necessarily have to be associated with the user's account. For example, it is sufficient that the user's identifier and the features identifying the user are associated and stored.

[0037] FIG. 4A shows an example of the configuration of the user device 100.

[0038] The user device 100 includes a communication interface unit 110, an input unit 120, a display unit 130, a memory unit 140, a processor unit 150, a voice input unit 160, and a voice output unit 170.

[0039] The communication interface unit 110 controls communication via the network 400. The processor unit 150 of the user device 100 can receive information from outside the user device 100 and transmit information to the outside of the user device 100 via the communication interface unit 110. The communication interface unit 110 can control communication in any way.

[0040] The input unit 120 enables the user to input information into the user device 100. It does not matter in what manner the input unit 120 enables the user to input information into the user device 100. For example, when the input unit 120 is a touch panel, the user may input information by touching the touch panel. Alternatively, when the input unit 120 is a mouse, the user may input information by operating the mouse. Alternatively, when the input unit 120 is a keyboard, the user may input information by pressing the keys of the keyboard.

[0041] The display unit 130 can be any display for displaying information.

[0042] The memory unit 140 stores programs for executing processes in the user device 100, data required for the execution of those programs, and the like. The memory unit 140 stores, for example, part or all of a program for generating a conversation record (for example, a program that realizes the processes shown in FIGS. 7 to 8 described later). The memory unit 140 may store an application that implements any function. The memory unit 140 may store, for example, a voice recognition application that converts voice data into text. Here, it does not matter how the program is stored in the memory unit 140. For example, the program may be pre-installed in the memory unit 140. Alternatively, the program may be installed in the memory unit 140 by being downloaded via the network 400. The memory unit 140 can be implemented by any storage means.

[0043] The processor unit 150 controls the operation of the entire user device 100. The processor unit 150 reads out the program stored in the memory unit 140 and executes the program. Thereby, it is possible to make the user device 100 function as a device that executes desired steps. The processor unit 150 may be implemented by a single processor or may be implemented by a plurality of processors.

[0044] The voice input unit 160 is any means for receiving voice. The voice input unit 160 is, for example, a microphone.

[0045] The voice output unit 170 is any means for outputting voice. The voice output unit 170 is, for example, a speaker.

[0046] In addition to the above configuration, the user device 100 may include, for example, any camera capable of taking images. The camera may be a camera built into the user device 100 or an external camera attached to the user device 100.

[0047] In the example shown in FIG. 4A, each component of the user device 100 is provided within the user device 100, but the present invention is not limited to this. Any of the components of the user device 100 may be provided outside the user device 100. For example, when each of the input unit 120, the display unit 130, the memory unit 140, the processor unit 150, the voice input unit 160, and the voice output unit 170 is composed of separate hardware components, the respective hardware components may be connected via an arbitrary network. At this time, the type of network is not limited. Each hardware component may be connected via, for example, a LAN, may be wirelessly connected, or may be wired connected. The user device 100 is not limited to a specific hardware configuration. For example, it is also within the scope of the present invention to configure the processor unit 150 with an analog circuit instead of a digital circuit. The configuration of the user device 100 is not limited to that described above as long as its function can be realized.

[0048] FIG. 4B shows an example of the configuration of the server device 200.

[0049] The server device 200 includes a communication interface unit 210, a memory unit 220, and a processor unit 230.

[0050] The communication interface unit 210 controls communication via the network 400. Also, the communication interface unit 210 controls communication with the database unit 300. The processor unit 230 of the server device 200 can receive information from outside the server device 200 via the communication interface unit 210, and can transmit information to the outside of the server device 200. For example, the processor unit 230 of the server device 200 receives data indicating voice from at least one user device 100 via the network 400. The data indicating voice may be, for example, voice data (e.g., an analog signal of a voice waveform, a digital signal of a voice waveform), voice recognition data (e.g., text data obtained by converting voice into text by voice recognition), or a combination of those data. For example, the processor unit 230 of the server device 200 receives image data from at least one user device 100 via the network 400. The image data may be, for example, still image data or moving image data. For example, the processor unit 230 of the server device 200 transmits data indicating voice (e.g., voice data, voice recognition data, or a combination thereof) and data indicating who the speaker of the voice is to at least one user device 100 via the network 400. For example, the processor unit 230 of the server device 200 can receive information regarding the user's account from the database unit 300. The communication interface unit 210 can control communication in any method.

[0051] The memory unit 220 stores programs necessary for the execution of the processing of the server device 200 and data necessary for the execution of such programs. For example, part or all of a program for generating a conversation record (for example, a program that realizes the processing shown in FIGS. 7 to 8 described later) is stored. The memory unit 220 may store, for example, a voice recognition application that converts voice data into text. The memory unit 220 can be implemented by any storage means.

[0052] The processor unit 230 controls the operation of the entire server device 200. The processor unit 230 reads out the programs stored in the memory unit 220 and executes those programs. Thereby, it is possible to make the server device 200 function as a device that executes desired steps. The processor unit 230 may be implemented by a single processor or by a plurality of processors.

[0053] In the example shown in FIG. 4B, each component of the server device 200 is provided within the server device 200, but the present invention is not limited to this. Any of the components of the server device 200 may be provided outside the server device 200. For example, when each of the memory unit 220 and the processor unit 230 is composed of separate hardware components, the hardware components may be connected via any network. At this time, the type of network is not limited. Each hardware component may be connected via, for example, a LAN, may be wirelessly connected, or may be wired connected. The server device 200 is not limited to a specific hardware configuration. For example, it is also within the scope of the present invention to configure the processor unit 230 with an analog circuit instead of a digital circuit. The configuration of the server device 200 is not limited to the above-described one as long as its function can be realized.

[0054] In the example shown in FIGS. 3 and 4B, the database unit 300 is provided outside the server device 200, but the present invention is not limited to this. It is also possible to provide the database unit 300 inside the server device 200. At this time, the database unit 300 may be implemented by the same storage means as the storage means that implements the memory unit 220, or may be implemented by storage means different from the storage means that implements the memory unit 220. In any case, the database unit 300 is configured as a storage unit for the server device 200. The configuration of the database unit 300 is not limited to a specific hardware configuration. For example, the database unit 300 may be composed of a single hardware component, or may be composed of a plurality of hardware components. For example, the database unit 300 may be configured as an external hard disk device of the server device 200, or may be configured as a storage on the cloud connected via a network.

[0055] FIG. 5 shows an example of the configuration of the processor unit 230 of the server device 200.

[0056] The processor unit 230 includes a receiving means 231, an identifying means 232, a determining means 233, an outputting means 234, and a discarding means 235.

[0057] The receiving means 231 is configured to receive data indicating voice. The data indicating voice may be, for example, voice data, voice recognition data, or a combination of voice data and voice recognition data. The receiving means 231 can receive, for example, data indicating voice from at least one user device 100 via the communication interface unit 210. The data indicating voice received by the receiving means 231 may be the data itself transmitted from at least one user device 100, or data generated by the processor unit 230 processing the data transmitted from at least one user device 100 (for example, voice recognition data generated by the processor unit 230 performing voice recognition).

[0058] In one embodiment, the receiving means 231 can receive voice data from at least one user device 100 via the communication interface unit 210. In another embodiment, the receiving means 231 can receive voice recognition data from at least one user device 100 via the communication interface unit 210. In another embodiment, the receiving means 231 can receive voice recognition data generated from the voice data received from at least one user device 100 via the communication interface unit 210.

[0059] Data indicating the received voice is passed to the identification means 232.

[0060] The identification means 232 is configured to identify the speaker of the voice indicated by the received data. The identification means 232 can use any means to identify the speaker of the voice.

[0061] In one embodiment, the identification means 232 can identify the speaker based on, for example, the voiceprint of the voice. For example, the identification means 232 can identify the speaker by comparing the voiceprint of each of a plurality of users registered in the database unit 300 with the voiceprint of the voice. For example, when the voiceprint of the voice matches one of the voiceprints of a plurality of users registered in the database unit 300, the user with that voiceprint is identified as the speaker of the voice, and when the voiceprint of the voice does not match any of the voiceprints of a plurality of users registered in the database unit 300, or when it matches a plurality of the voiceprints of a plurality of users registered in the database unit 300, it can be made unidentifiable. Alternatively, for example, among the voiceprints of a plurality of users registered in the database unit 300, the user with the voiceprint that is most similar to the voiceprint of the voice can be identified as the speaker of the voice.

[0062] The identification means 232 can extract a voiceprint from a voice using a technique known in the field of voice analysis or voice authentication. Further, the identification means 232 can determine the match or similarity of voiceprints using a technique known in the field of voice analysis or voice authentication.

[0063] In another embodiment, the identification means 232 can identify the speaker of the voice based on, for example, an image when the voice is uttered. The image can be a plurality of still images or moving images. For example, the identification means 232 can analyze the image when the voice is uttered and identify the person whose lips are moving as the speaker of the voice. At this time, the identification means 232 also identifies who the person in the image is. For example, the identification means 232 can identify the person in the image by comparing each of the face images of a plurality of users registered in the database unit 300 with the face image of the person in the image. For example, when the face image of the person in the image matches one of the face images of a plurality of users registered in the database unit 300, the user of that face image is identified as the person in the image, and when the face image of the person in the image does not match any of the face images of a plurality of users registered in the database unit 300, or when it matches a plurality of the face images of a plurality of users registered in the database unit 300, it can be set as unidentified. Alternatively, for example, among the face images of a plurality of users registered in the database unit 300, the user of the face image that is most similar to the face image of the person in the image can be identified as the person in the image.

[0064] The identification means 232 can identify the person whose lips are moving using a technique known in the field of image analysis or image authentication. Further, the identification means 232 can determine the match or similarity of face images using a technique known in the field of image analysis or image authentication.

[0065] The determination means 233 is configured to determine whether the speaker identified by the identification means 232 is the user associated with the user device that transmitted the data indicating the voice. For example, in a conversation using a plurality of user devices including a first user device and a second user device, when receiving data indicating voice from the first user device, the determination means 233 determines whether the speaker identified by the identification means 232 is the user associated with the first user device. The user associated with the user device can be the user who uses the user device. For example, when the first user logs in to his / her account via the first user device, the first user becomes the user associated with the first user device. For example, when the second user uses the first user device as a guest user, the second user becomes the user associated with the first user device. Registration as a guest user of the first user device is performed, for example, by transmitting a guest user request to the server device 200 via the first user device.

[0066] The output means 234 is configured to output the data indicating the voice and the data indicating that the speaker of the voice is the speaker identified by the identification means 232 when it is determined that the speaker identified by the identification means 232 is the user associated with the user device that transmitted the data indicating the voice. For example, when receiving data indicating voice from the first user device and when it is determined that the speaker identified by the identification means 232 is the user associated with the first user device, the output means 234 outputs the data indicating the voice and the data indicating that the speaker of the voice is the speaker identified by the identification means 232.

[0067] The mode in which the output means 234 outputs data is not limited. For example, the output means 234 may transmit the data to at least one user device 100 and cause the at least one user device 100 to output the data. At this time, when the data indicating the voice is voice data, the processor unit 230 may convert the voice data into voice recognition data by performing voice recognition processing on the voice data, and then the output means 234 may transmit the voice recognition data to at least one user device 100. Alternatively, after the output means 234 transmits the voice data to at least one user device 100, the at least one user device 100 may convert the voice data into voice recognition data by performing voice recognition processing on the voice data. At least one user device 100 that has received the data can, for example, display the data on the display unit 130. On the display unit 130, for example, text indicating the voice and an icon indicating the speaker may be displayed. This can be a visible record of the conversation, and if the conversation is during a meeting, this can be the minutes of the meeting. For example, at least one user device 100 can record the data in its storage means (for example, an internal storage device or an external storage device). In the storage means, data indicating the voice and data indicating that the speaker of the voice is the speaker identified by the identification means 232 are recorded in pairs. This can be a record of the conversation, and if the conversation is during a meeting, this can be the minutes of the meeting.

[0068] The output means 234 can, for example, record the data in the storage means of the database unit 300 or the server device. In the storage means of the database unit 300 or the server device, data indicating the voice and data indicating that the speaker of the voice is the speaker identified by the identification means 232 are recorded in pairs. This can be a record of the conversation, and if the conversation is during a meeting, this can be the minutes of the meeting.

[0069] The discarding means 235 is configured to discard the data indicating the voice when it is determined that the speaker identified by the identification means 232 is not the user associated with the user device that transmitted the data indicating the voice. For example, when receiving data indicating voice from a first user device, if the speaker identified by the identification means 232 is not the user associated with the first user device, the discarding means 235 discards the data indicating the voice. This can be the case, for example, when the speaker identified by the identification means 232 is the user associated with the second user device, or when the speaker identified by the identification means 232 is not associated with any user device. When the speaker identified by the identification means 232 is not the user associated with the user device that transmitted the data indicating the voice, as described above, the data indicating the voice may indicate the voice emitted by the speaker rather than the voice emitted by the user, so it is unnecessary data for creating a record of the conversation. By the discarding means 235 discarding the data indicating the voice, the processor unit 230 can avoid generating an incorrect conversation record.

[0070] The manner in which the discarding means 235 discards the data is not limited. The discarding means 235 can discard the data in any manner as long as the data is not used for generating the final conversation record. For example, the discarding means 235 may completely erase the data, or may erase the data in a recoverable manner. For example, the discarding means 235 may discard the data immediately after the determination by the determination means 233, or may discard the data after holding the data for a certain period. For example, the discarding means 235 may, as shown in FIGS. 1 and 2, transmit data indicating voice to the user device, and after the data is output from the user device, cause the user device to discard the data. For example, the discarding means 235 may record the data in the database unit 300 or the storage means of the server device and then discard the data.

[0071] In the example shown in FIG. 5 described above, each component of the processor unit 230 is provided within the same processor unit 230, but the present invention is not limited thereto. A configuration in which each component of the processor unit 230 is distributed across a plurality of processor units is also within the scope of the present invention. At this time, the plurality of processor units may be located within the same hardware component, or may be located within separate hardware components that are in the vicinity or remote from each other.

[0072] 3. Processing by the system for generating the conversation record of the present invention FIG. 6 shows an example of a data flow 600 in a system 1000 for generating a record of a conversation according to the present invention. In this example, a case where a first user and a second user are having a conversation using the first user device 100 1 and the second user device 100 2 will be described by taking as an example the case where the first user makes a statement. In FIG. 6, the data flow between the first user device 100 1 and the second user device 100 2 and the server device 200 and the database unit 300 is shown.

[0073] First, the first user makes a statement to the voice input unit 160 (e.g., a microphone) of the first user device 100 1 .

[0074] In step S601, the first voice that has been spoken is input into the first user device 100 1 via the voice input unit 160 of the first user device 100 1 . The input voice is treated as data indicating the first voice. The data indicating the first voice may be, for example, voice data, or may be voice recognition data when the first user device 100 1 performs voice recognition processing.

[0075] In step S602, the data indicating the first voice is transmitted to the server device 200. When the server device 200 receives the data indicating the first voice, the processing of step S607 is performed on the received data indicating the first voice.

[0076] In step S603, the data indicating the first voice is transmitted to the second user device 100 2 The data indicating the first voice may be directly transmitted from the first user device 100 1 to the second user device 100 2 or may be indirectly transmitted from the first user device 100 1 to the second user device 100 2 For example, the data indicating the first voice can be transmitted to the second user device 100 2 via the server device 200. At this time, a copy of the data transmitted in step S602 may be transmitted to the second user device 100 2

[0077] When the second user device 100 2 receives the data indicating the first voice, in step S604, the first voice is output from the voice output unit 170 (for example, a speaker) of the second user device 100 2 Thereby, the second user can listen to the speech of the first user.

[0078] In step S605, the voice output from the voice output unit 170 of the second user device 100 2 is input to the second user device 100 2 via the voice input unit 160 of the second user device 100 2 The input voice is treated as data indicating the second voice. The data indicating the second voice may be, for example, voice data, or may be voice recognition data when the second user device 100 2 performs voice recognition processing.

[0079] In step S606, the data indicating the second voice is transmitted to the server device 200. When the server device 200 receives the data indicating the second voice, the process of step S607 is performed on the received data indicating the second voice.

[0080] ​In step S607, a process for generating a conversation record is performed. Details of step S607 will be described later with reference to FIG. 7.

[0081] By the process of step S607, data indicating the first voice received in step S602 is output together with data indicating that the speaker of the first voice is the first user. On the other hand, by the process of step S607, data indicating the second voice received in step S606 is discarded.

[0082] In step S608, the data indicating the first voice and the data indicating that the speaker of the first voice is the first user, which are output in step S607, are sent to the first user device 100 1 When the data indicating the voice is voice data, the processor unit 230 may convert the voice data into voice recognition data by performing voice recognition processing on the voice data, and then send the voice recognition data to the first user device 100 1 Or, after sending the voice data to the first user device 100 1 the first user device 100 1 may convert the voice data into voice recognition data by performing voice recognition processing on the voice data.

[0083] When the first user device 100 1 receives those data, in step S609, the received data is displayed on the display unit 130 of the first user device 100 1 For example, as shown in FIG. 1, text indicating the first voice and an icon indicating that the speaker of the first voice is the first user may be displayed on the display unit 130. Thereby, a conversation record in which the utterance of the first user is recorded is created, and the first user can view this. Alternatively, instead of this, or in addition to this, in step S609, the received data may be recorded in the storage means of the first user device 100 1 Thereby, a conversation record in which the utterance of the first user is recorded is created.

[0084] In step S610, the data indicating the first voice output in step S607 and the data indicating that the speaker of the first voice is the first user are transmitted to the second user device 100. 2 When the data indicating the voice is voice data, the processor unit 230 may convert the voice data into voice recognition data by performing voice recognition processing on the voice data, and then transmit the voice recognition data to the second user device 100. 2 Or, after transmitting the voice data to the second user device 100, 2 the second user device 100 may convert the voice data into voice recognition data by performing voice recognition processing on the voice data. 2

[0085] When the second user device 100 2 receives those data, in step S611, the received data is displayed on the display unit 130 of the second user device 100. 2 For example, as shown in FIG. 1, text indicating the first voice and an icon indicating that the speaker of the first voice is the first user may be displayed on the display unit 130. 2 Thereby, a conversation record in which the speech of the first user is recorded is created, and the second user can view this.

[0086] Instead of or in addition to steps S608 to S611, steps 612 and S613 may be performed.

[0087] In step S612, the data indicating the first voice output in step S607 and the data indicating that the speaker of the first voice is the first user are transmitted to the database unit 300.

[0088] When the database unit 300 receives those data, in step S613, the received data is recorded in the database unit 300. For example, in the database unit 300, data indicating the first voice and data indicating that the speaker of the first voice is the first user are recorded as a pair. Thereby, a conversation record in which the utterance of the first user is recorded is created.

[0089] In the conversation record created in this way, since the data indicating the second voice received in step S606 is not recorded, the utterance content of the first user is not recorded repeatedly in the created conversation record.

[0090] FIG. 7 is a flowchart showing an example of the process of step S607. The process of step S607 is performed in the processor unit 230 of the server device 200.

[0091] In step S701, the receiving means 231 of the processor unit 230 receives data indicating the first voice. The data indicating the first voice can be received from the first user device 100 1 via the communication interface unit 210. The data indicating the first voice received may be, for example, voice data received from the first user device 100 1 or voice recognition data generated by voice recognition processing by the first user device 100 1 or voice recognition data generated by voice recognition processing by the processor unit 230 for the voice data received from the first user device 100 1 or a combination of voice data and voice recognition data.

[0092] In step S702, the identification means 232 of the processor unit 230 identifies the speaker of the first voice indicated by the data received in step S701.

[0093] In one embodiment, the identification means 232 identifies the speaker of the first voice, for example, based on the voiceprint of the voice. For example, the identification means 232 can identify the speaker of the first voice by comparing the voiceprint of each of a plurality of users registered in the database unit 300 with the voiceprint of the first voice. For example, when the voiceprint of the first voice matches one of the voiceprints of a plurality of users registered in the database unit 300, the user of that voiceprint can be identified as the speaker of the first voice. Alternatively, for example, among the voiceprints of a plurality of users registered in the database unit 300, the user with the voiceprint most similar to the voiceprint of the first voice can be identified as the speaker of the first voice.

[0094] In another embodiment, the identification means 232 can identify the speaker of the first voice, for example, based on an image when the voice is uttered. For example, the identification means 232 analyzes the image when the first voice is uttered and can identify the person whose lips are moving as the speaker of the first voice. At this time, the identification means 232 also identifies who the person in the image is. For example, the identification means 232 can identify the person in the image by comparing the face image of each of a plurality of users registered in the database unit 300 with the face image of the person in the image. For example, when the face image of the person in the image matches one of the face images of a plurality of users registered in the database unit 300, the user of that face image can be identified as the person in the image. Alternatively, for example, among the face images of a plurality of users registered in the database unit 300, the user with the face image most similar to the face image of the person in the image can be identified as the person in the image.

[0095] In step S703, the determination means 233 of the processor unit 230 determines whether the speaker identified in step S702 is a user associated with the first user device 100 1 or not. The user associated with the user device can be the user who uses that user device. For example, the first user is the first user device 100 1When logging in to their own account via [the relevant means], the first user will become the user associated with the first user device 100 1 For example, when the second user uses the first user device 100 as a guest user 1 the second user will become the user associated with the first user device 100 1 The determination means 233 collates, for example, the account of the user using the first user device 100

[0096] or the account of the guest user using the first user device 100 1 with the speaker identified in step S702 to determine whether the speaker identified in step S702 is the user associated with the first user device 100 1 If it is determined that the speaker identified in step S702 is the user associated with the first user device 100 1 the process proceeds to step S704. If it is determined that the speaker identified in step S702 is not the user associated with the first user device, the process proceeds to step S705.

[0097] In step S704, the output means 234 of the processor unit 230 outputs the data indicating the first voice and the data indicating that the speaker of the first voice is the speaker identified in step S702. The manner in which the output means 234 outputs the data is not limited. For example, as described in step S608 or step S610, the data may be transmitted to the first user device 100

[0098] and / or the second user device 100 1 Alternatively, as described in step S612, the data may be transmitted to the database unit 300. 2

[0099] ​In step S705, the deletion means 235 of the processor unit 230 deletes the data indicating the first voice. The manner in which the deletion means 235 deletes the data is not limited. The deletion means 235 can delete the data in any manner as long as the data is not used for generating the final conversation record. For example, the deletion means 235 may completely erase the data, or may erase the data in a recoverable manner. For example, the deletion means 235 may delete the data immediately after the determination by the determination means 233, or may delete the data after holding the data for a certain period. For example, the deletion means 235 may, for example, like in step S608, step S610, step S612, the first user device 100 1 and / or the second user device 100 2 and / or after transmitting the data to the database unit 300, the first user device 100 1 and / or the second user device 100 2 and / or may cause the database unit 300 to delete the data.

[0100] For example, in the case of the first voice received in step S602, the first voice is received from the first user device in step S701, the speaker is identified as the first user in step S702, and in step S703, since it is determined that the first user is the user associated with the first user device, the process proceeds to step S704, and the data indicating the first voice and the data indicating that the speaker of the first voice is the first user are output.

[0101] For example, in the case of the second voice received in step S606, the second voice is received from the second user device in step S701, the speaker is identified as the first user in step S702, and in step S703, since it is determined that the first user is not the user associated with the second user device, the process proceeds to step S705, and the data indicating the second voice is deleted.

[0102] FIG. 8 is a flowchart showing details of step S703 and step S704. Step S703 includes step S7031 and step S7032, and step S704 includes step S7041 and step 7042.

[0103] In step S7031, the determination means 233 of the processor unit 230 determines whether the speaker identified in step S702 is a first user who holds a first account used in the first user device 100 1 (the user device that transmitted data indicating the first voice). That is, in step S7031, it is determined whether the speaker identified in step S702 is a user who logged in to their own account via the first user device 100 1 among the users associated with the first user device 100 1 Whether it is a user who logged in to their own account via the first user device 100

[0104] The determination means 233 can make the determination, for example, by collating the account of the user who uses the first user device 100 1 with the speaker identified in step S702.

[0105] If it is determined that the speaker identified in step S702 is a first user who holds a first account used in the first user device 100 1 the process proceeds to step S7041 of step S704. If it is determined that the speaker identified in step S702 is not a first user who holds a first account used in the first user device 100 1 the process proceeds to step S7032.

[0106] In step S7032, the determination means 233 of the processor unit 230 determines whether the speaker identified in step S702 is the first user device 100 1Determine whether the guest user is using the first account used in the 1 (user device that transmitted the data indicating the first voice). That is, in step S7032, it is determined whether the speaker identified in step S702 is a guest user who uses the first user device 100 1 among the users associated with the first user device 100

[0107] The determination means 233 can make a determination, for example, by collating the account of the guest user who uses the first user device 100 1 with the speaker identified in step S702.

[0108] If it is determined that the speaker identified in step S702 is a guest user who uses the first account used in the first user device 100 1 , the process proceeds to step S7042 of step S704. If it is determined that the speaker identified in step S702 is not a guest user who uses the first account used in the first user device 100 1 , the process proceeds to step S705.

[0109] In step S7041, the output means 234 of the processor unit 230 outputs the data indicating the first voice and the data indicating that the speaker of the first voice is the first user. As described in step S704, the mode in which the output means 234 outputs the data is not limited.

[0110] In step S7042, the output means 234 of the processor unit 230 outputs the data indicating the first voice and the data indicating that the speaker of the first voice is a guest user. As described in step S704, the mode in which the output means 234 outputs the data is not limited.

[0111] In this way, data corresponding to the user associated with the first user device 100 1 can be output.

[0112] In the example shown in FIG. 8, it has been described that step S703 includes step S7031 and step S7032, but the present invention is not limited thereto. For example, step S703 may not include step S7032, or step S703 may not include step S7031. For example, when step S703 does not include step S7032, if it is determined as No in step S7031, the process can proceed to step S705.

[0113] When step S703 does not include step S7032, correspondingly, step S704 may not include step S7042. When step S703 does not include step S7031, correspondingly, step S704 may not include step S7041.

[0114] Taking the case shown in FIG. 1 as an example, an example of the processing in FIGS. 6 to 8 will be described.

[0115] First, user A speaks "I agree." to the microphone of user device 10.

[0116] In step S601, the voice of "I agree." is input to user device 10 via the microphone.

[0117] In step S602, the data indicating the voice of "I agree." is transmitted to server device 200. In this example, it is assumed that the voice data of "I agree." is transmitted. When server device 200 receives the data indicating the voice of "I agree.", the process of step S607 is performed on the voice data of "I agree.".

[0118] In step S603, the voice data of "I agree." is transmitted to user device 20 of user B.

[0119] When the user device 20 receives the voice data of "I agree.", at step S604, the user device 20 outputs the voice of "I agree." from the speaker. As a result, user B can hear the statement of "I agree." made by user A.

[0120] At step S605, the voice of "I agree." output from the speaker of the user device 20 is input to the user device 20 via the microphone of the user device 20. In this example, this voice is referred to as the "second 'I agree.'"

[0121] At step S606, the data indicating the second "I agree." is transmitted to the server device 200. In this example, it is assumed that the voice data of the second "I agree." is transmitted. When the server device 200 receives the voice data of the second "I agree.", the process of step S607 is performed on the voice data of the second "I agree."

[0122] At step S607, the following processing is performed on the voice data of "I agree."

[0123] At step S701, the receiving means 231 of the processor unit 230 receives the voice data of "I agree."

[0124] At step S702, the identifying means 232 of the processor unit 230 identifies the speaker of the voice indicated by the voice data of "I agree." For example, the voiceprint is extracted from the voice data of "I agree.", and the speaker is identified by comparing the voiceprint with the voiceprints included in the information related to the accounts of a plurality of users stored in the database unit 300. Here, it is identified that the speaker is user A.

[0125] At step S703, the determining means 233 of the processor unit 230 determines whether user A is the user associated with the user device 10 that has transmitted the voice data of "I agree."

[0126] Here, in step S7031, the determination means 233 of the processor unit 230 determines whether user A is a user who holds the account used in the user device 10. Since user A is a user who holds the account used in the user device 10, it is determined as Yes, and the process proceeds to step S7041 of step S704.

[0127] In step S7041 of step S704, the output means 234 of the processor unit 230 outputs the voice data of "I agree." and the data indicating that the speaker of the voice of "I agree." is user A.

[0128] In this way, by the process of step S607, the voice data of "I agree." received in step S602 is output together with the data indicating that the speaker of the voice of "I agree." is user A.

[0129] In step S607, the following process is performed on the voice data of the second "I agree."

[0130] In step S701, the receiving means 231 of the processor unit 230 receives the voice data of the second "I agree."

[0131] In step S702, the identification means 232 of the processor unit 230 identifies the speaker of the voice indicated by the voice data of the second "I agree." For example, the voiceprint is extracted from the voice data of the second "I agree.", and the speaker is identified by comparing the voiceprint with the voiceprints included in the information regarding the accounts of a plurality of users stored in the database unit 300. Here, it is identified that the speaker is user A.

[0132] In step S703, the determination means 233 of the processor unit 230 determines whether user A is a user associated with the user device 20 that has sent the voice data of the second "I agree." Since user A is not a user associated with the user device 20, it is determined as No, and the process proceeds to step S705.

[0133] In step S705, the deletion means 235 of the processor unit 230 deletes the voice data of the second "I agree."

[0134] In this way, by the process of step S607, the voice data of the second "I agree." received in step S606 is deleted and not used for creating the conversation record.

[0135] In step S608, the data output in step S607 is transmitted to the user device 10. At this time, the processor unit 230 can perform voice recognition processing on the voice data of "I agree." to convert it into voice recognition data (for example, text data of "I agree.") and transmit it.

[0136] When the user device 10 receives those data, in step S609, the received data is displayed on the display unit 130 of the user device 10. For example, as shown in FIG. 1C, the display unit 130 may display text indicating the voice of "I agree." and an icon indicating that the speaker of the voice of "I agree." is user A. Thereby, a conversation record in which the speech of user A is recorded is created, and user A can view this.

[0137] In step S610, the data output in step S607 is transmitted to the user device 20. At this time, the processor unit 230 can perform voice recognition processing on the voice data of "I agree." to convert it into voice recognition data (for example, text data of "I agree.") and transmit it.

[0138] When the user device 20 receives those data, in step S611, the received data is displayed on the display unit 130 of the user device 20. For example, as shown in FIG. 1C, text indicating the voice of "I agree." and an icon indicating that the speaker of the voice of "I agree." is user A may be displayed on the display unit 130. Thereby, a conversation record in which the speech of user A is recorded is created, and user B can view this.

[0139] In the conversation record created in this way, since the data indicating the second voice of "I agree." received in step S606 is not recorded, the speech content of user A is not recorded repeatedly in the created conversation record.

[0140] Taking the case shown in FIG. 2 as an example, an example of the processing in FIGS. 6 to 8 will be described.

[0141] First, user C speaks "I agree." to the microphone of user device 10.

[0142] In step S601, the voice of "I agree." is input to the user device 10 via the microphone.

[0143] In step S602, the data indicating the voice of "I agree." is transmitted to the server device 200. In this example, it is assumed that the voice data of "I agree." is transmitted. When the server device 200 receives the data indicating the voice of "I agree.", the processing of step S607 is performed on the voice data of "I agree.".

[0144] In step S603, the voice data of "I agree." is transmitted to the user device 20 of user B.

[0145] When the user device 20 receives the voice data of "I agree.", in step S604, the voice of "I agree." is output from the speaker of the user device 20. Thereby, user B can hear the speech of "I agree." by user C.

[0146] In step S605, the voice of "I agree." output from the speaker of the user device 20 is input to the user device 20 via the microphone of the user device 20. In this example, this voice is referred to as the "second 'I agree.'"

[0147] In step S606, the data indicating the second "I agree." is transmitted to the server device 200. In this example, it is assumed that the voice data of the second "I agree." is transmitted. When the server device 200 receives the voice data of the second "I agree.", the process of step S607 is performed on the voice data of the second "I agree.".

[0148] In step S607, the following processing is performed on the voice data of "I agree.".

[0149] In step S701, the receiving means 231 of the processor unit 230 receives the voice data of "I agree.".

[0150] In step S702, the identifying means 232 of the processor unit 230 identifies the speaker of the voice indicated by the voice data of "I agree.". For example, the voiceprint is extracted from the voice data of "I agree.", and the speaker is identified by comparing the voiceprint with the voiceprints included in the information regarding the accounts of a plurality of users stored in the database unit 300. Here, it is identified that the speaker is User C.

[0151] In step S703, the determining means 233 of the processor unit 230 determines whether User C is the user associated with the user device 10 that has transmitted the voice data of "I agree.".

[0152] Here, in step S7031, the determining means 233 of the processor unit 230 determines whether User C is the user who holds the account used in the user device 10. Since User C is not the user who holds the account used in the user device 10, it is determined as No, and the process proceeds to step S7032.

[0153] In step S7032, the determination means 233 of the processor unit 230 determines whether user C is a guest user who uses the account being used in the user device 10. Since user C is a guest user who uses the account being used in the user device 10, it is determined as Yes, and the process proceeds to step S7042 of step S704.

[0154] In step S7042 of step S704, the output means 234 of the processor unit 230 outputs the voice data of "I agree." and the data indicating that the speaker of the voice of "I agree." is user C.

[0155] In this way, by the process of step S607, the voice data of "I agree." received in step S602 is output together with the data indicating that the speaker of the voice of "I agree." is user C.

[0156] In step S607, the following process is performed on the second voice data of "I agree."

[0157] In step S701, the receiving means 231 of the processor unit 230 receives the second voice data of "I agree."

[0158] In step S702, the identification means 232 of the processor unit 230 identifies the speaker of the voice indicated by the second voice data of "I agree." For example, the voiceprint is extracted from the second voice data of "I agree.", and the speaker is identified by comparing the voiceprint with the voiceprints included in the information regarding the accounts of a plurality of users stored in the database unit 300. Here, it is identified that the speaker is user C.

[0159] In step S703, the determination means 233 of the processor unit 230 determines whether user C is a user associated with the user device 20 that has sent the second voice data of "I agree."

[0160] Here, in step S7031, the determination means 233 of the processor unit 230 determines whether user C is a user who holds the account used in the user device 20. Since user C is not a user who holds the account used in the user device 20, it is determined as No, and the process proceeds to step S7032.

[0161] In step S7032, the determination means 233 of the processor unit 230 determines whether user C is a guest user who uses the account used in the user device 20. Since user C is not a guest user who uses the account used in the user device 20, it is determined as No, and the process proceeds to step S705.

[0162] In step S705, the deletion means 235 of the processor unit 230 deletes the voice data of the second "I agree."

[0163] In this way, by the process of step S607, the voice data of the second "I agree." received in step S606 is deleted and not used for creating the conversation record.

[0164] In step S608, the data output in step S607 is transmitted to the user device 10. At this time, the processor unit 230 can perform voice recognition processing on the voice data of "I agree." and convert it into voice recognition data (for example, text data of "I agree.") and transmit it.

[0165] When the user device 10 receives those data, in step S609, the received data is displayed on the display unit 130 of the user device 10. For example, on the display unit 130, as shown in FIG. 2D, text indicating the voice of "I agree." and an icon indicating that the speaker of the voice of "I agree." is user C can be displayed. Thereby, a conversation record in which the speech of user C is recorded is created, and user A and user C can view this.

[0166] In step S610, the data output in step S607 is transmitted to the user device 20. At this time, the processor unit 230 can perform speech recognition processing on the speech data of "I agree." to convert it into speech recognition data (for example, text data of "I agree.") and transmit it.

[0167] When the user device 20 receives those data, in step S611, the received data is displayed on the display unit 130 of the user device 20. For example, as shown in FIG. 2D, the display unit 130 may display text indicating the voice of "I agree." and an icon indicating that the speaker of the voice of "I agree." is user C. Thereby, a record of the conversation in which the speech of user C is recorded is created, and user B can view this.

[0168] In the record of the conversation created in this way, since the data indicating the second voice of "I agree." received in step S606 is not recorded, the speech content of user A is not recorded repeatedly in the created record of the conversation.

[0169] In the example described above, before recording the speech, the speaker is identified, and only the correct information is recorded based on the result of the speaker identification. However, the present invention is not limited to this. For example, as shown in FIGS. 1B, 2A, and 2C, after the misrecognized speech is recorded, the speaker of the recorded speech may be changed or the recorded speech may be deleted based on the result of the speaker identification.

[0170] In the example described above, a conversation using two user devices (the first user device 100 1 and the second user device 100 2 , or the user device 10 and the user device 20) was described as an example. However, the present invention can be applied to conversations using any number of user devices.

[0171] In the example described above with reference to FIGS. 6 to 8, it was explained that the processing is performed in a specific order. However, the order of each processing is not limited to that described, and it may be performed in any logically possible order.

[0172] In the example described above with reference to FIGS. 7 to 8, it was explained that the processing of each step shown in FIGS. 7 to 8 is realized by the processor unit 230 of the server device 200 and the program stored in the memory unit 220. However, the present invention is not limited to this. At least one of the processing of each step shown in FIGS. 7 to 8 may be realized by a hardware configuration such as a control circuit.

[0173] In the example described above, it was explained that the processor unit 230 of the server device 200 performs the processing of step S607. However, the present invention is not limited to this. It is also within the scope of the present invention for the processor unit 150 of the user device 100 to perform the processing of step S607. At this time, the server device 200 may be omitted from the system 1000.

[0174] When the processor unit 150 of the user device 100 performs the processing of step S607, step S607 may include step S701' instead of step S701.

[0175] In step S701', data indicating the voice input via the voice input unit 160 in step S601 is received.

[0176] The subsequent steps are the same as those described above with reference to FIG. 7, and the description is omitted here.

[0177] The present invention is not limited to the above-described embodiments. It is understood that the scope of the present invention should be interpreted only by the claims. Those skilled in the art will understand that they can implement an equivalent range based on the description of the present invention and common technical knowledge from the description of the specific preferred embodiments of the present invention.

Industrial Applicability

[0178] The present invention is useful in that it can provide a program, a system, and a method for generating a record of a conversation that can avoid duplicate recording of the user's utterance content.

Explanation of Signs

[0179] 10, 20, 100 User devices 200 Server device 300 Database section 400 Network 1000 System

Claims

[Claim 1] The invention described in this specification.

Citation Information

Patent Citations

  • Minutes creation method, its device and its program

    JP2008225068A