Information processing device, information processing method, and information processing program

The information processing device addresses the inconvenience of manual editing in conventional speech-to-text systems by automatically masking and revealing privacy-sensitive information, improving user convenience and privacy protection in speech-to-text conversion.

JP7791014B2Active Publication Date: 2025-12-23LY CORP
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2022034738
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-03-07
Publication Date
2025-12-23
Estimated Expiration
2042-03-07

AI Technical Summary

Technical Problem

Conventional speech-to-text technologies lack user convenience in handling privacy-sensitive information, requiring manual editing of displayed content to protect confidentiality.

Method used

An information processing device that converts user utterances into text while masking privacy-sensitive information, using trained models or confidential information tables to identify and obscure such content, and allowing users to selectively reveal masked information.

Benefits of technology

Enhances user convenience by automatically protecting privacy-sensitive information during speech-to-text conversion, enabling secure and efficient handling of confidential data without manual editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007791014000001
    Figure 0007791014000001
  • Figure 0007791014000002
    Figure 0007791014000002
  • Figure 0007791014000003
    Figure 0007791014000003
Patent Text Reader

Abstract

To improve convenience for a user.SOLUTION: An information processing device for the present application includes a reception unit, a generation unit, and a provision unit. The reception unit accepts user's speech. The generation unit generates speech text information, which is text information of the user's speech converted into text by masking parts of the speech accepted by the reception unit that satisfies a confidentiality condition. The provision unit provides the speech text information generated by the generation unit.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing device, an information processing method, and an information processing program. [Background technology]

[0002] Conventionally, there is known a technology for converting a user's speech into text information and providing the converted text information. For example, Patent Document 1 discloses a voice input device including a voice recognition unit that recognizes the user's voice, a character conversion unit that converts the voice recognized by the voice recognition unit into a character string, and a character string display unit that displays the character string converted by the character conversion unit. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent Publication No. 2021-168020 Summary of the Invention [Problem to be solved by the invention]

[0004] However, the above-mentioned conventional technologies have room for improvement in terms of improving user convenience. For example, in the technology described in Patent Document 1, all of the contents of a user's utterance are converted into text and displayed, so if there is content that the user wants to keep secret from the perspective of privacy, for example, the displayed character string needs to be edited, and there is room for improvement.

[0005] The present application has been made in view of the above, and has as its object to provide an information processing device, an information processing method, and an information processing program that can improve user convenience. [Means for solving the problem]

[0006] The information processing device according to the present application includes a receiving unit, a generating unit, and a providing unit. The receiving unit receives a user's utterance. The generating unit generates utterance text information, which is text information obtained by converting the user's utterance into text by masking a portion of the utterance received by the receiving unit that satisfies a confidentiality condition. The providing unit provides the utterance text information generated by the generating unit. [Effects of the Invention]

[0007] According to one aspect of the embodiment, an effect is achieved in that it is possible to improve convenience for users. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 is a diagram illustrating an example of information processing according to an embodiment. [Figure 2] FIG. 2 is a diagram illustrating an example of the configuration of the information processing device according to the embodiment. [Figure 3] FIG. 3 is a diagram illustrating an example of spoken text information displayed on the display unit of the information processing device according to the embodiment. [Figure 4] FIG. 4 is a diagram showing an example in which masking of a selected portion of spoken text information displayed on the display unit of the information processing device according to the embodiment is released. [Figure 5] FIG. 5 is a flowchart showing an example of information processing by the processing unit of the information processing device according to the embodiment. [Figure 6] FIG. 6 is a flowchart showing an example of information processing by the processing unit of the information processing device according to the embodiment. [Figure 7] FIG. 7 is a diagram illustrating another example of the configuration of the information processing device according to the embodiment. [Figure 8] FIG. 8 is a hardware configuration diagram illustrating an example of a computer that realizes the functions of the information processing device according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0009] Hereinafter, modes for implementing an information processing device, an information processing method, and an information processing program according to the present application (hereinafter referred to as "embodiments") will be described in detail with reference to the drawings. Note that the information processing device, the information processing method, and the information processing program according to the present application are not limited to these embodiments. Furthermore, the respective embodiments can be appropriately combined within the scope of not causing any contradiction in the processing content. Furthermore, the same components in the following embodiments will be assigned the same reference numerals, and redundant explanations will be omitted.

[0010] [1. An example of information processing] FIG. 1 is a diagram showing an example of information processing according to an embodiment, and in this embodiment, an information processing method is executed by an information processing device 1.

[0011] The information processing device 1 is a device that can utilize an AI (Artificial Intelligence) assistant function that supports interactive voice operation, and a user U can control peripheral devices and obtain various information by interacting with the information processing device 1. The peripheral devices are various devices such as lighting equipment, refrigerators, washing machines, air conditioners, television sets, dishwashers, dish dryers, induction cookers, and microwave ovens.

[0012] Furthermore, when the user U speaks to the information processing device 1 to acquire various information, the information processing device 1 transmits input information indicating instructions from the user U to the information providing device 2 (see FIG. 2). The information processing device 1 can acquire content (for example, various types of information such as news, traffic information, weather, and music) provided from the information providing device 2 via the network N (see FIG. 2) in accordance with the input information, and display the acquired content on a display unit or output it from a speaker.

[0013] Furthermore, the information processing device 1 has a text conversion function that converts the speech of the user U into text and outputs the converted text information. With this text conversion function, the user U can create, for example, text to be used in emails or SNS (Social Networking Service), text to be posted on electronic bulletin boards or review sites, etc., by speaking to the information processing device 1. The text conversion function will be mainly described below with reference to FIG. 1.

[0014] As shown in Fig. 1, when a user U uses the text function of the information processing device 1, the user U speaks to the information processing device 1 (step S1). In the example shown in Fig. 1, the user U speaks, "I recently moved into an apartment at 1-2-3, A-ku, Tokyo. It's a very comfortable place to live, and Tokkyo Taro and others helped me out a lot when I moved. I'm having a party soon, so please come along. My contact number is 090-0190-xxxx." Note that "090-0190-xxxx" is a phone number, and "xxxx" is a four-digit number combination.

[0015] The information processing device 1 receives an utterance from the user U (step S2). Then, the information processing device 1 generates utterance text information, which is text information obtained by converting the utterance of the user U received in step S2 into text by masking parts that satisfy the confidentiality conditions (step S3). The confidentiality conditions are words that indicate privacy-related content such as an address, name, or telephone number, and can be set or changed by the user U operating the information processing device 1.

[0016] 1, "A-ku 1-2-3," "Takken Taro," and "090-0190-xxxx" are parts (words or word groups, etc.) that satisfy the confidentiality conditions, and the information processing device 1 masks the parts that satisfy the confidentiality conditions. "A-ku 1-2-3" is information that indicates an address, and "Takken Taro" is information that indicates a name.

[0017] The information processing device 1 masks the portions that satisfy the confidentiality conditions in the process of converting the user U's speech into text, but can also mask the portions that satisfy the confidentiality conditions after converting the user U's speech into text.

[0018] The information processing device 1 has a trained model that receives, for example, audio information or text information corresponding to the utterance of user U received in step S2 as input and outputs a confidentiality score indicating the degree of confidentiality for each of the multiple element parts that make up the utterance of user U. The information processing device 1 inputs the audio information or text information corresponding to the utterance of user U received in step S2 to the trained model, and obtains, from the trained model, a confidentiality score that indicates the degree of confidentiality for each of the multiple element parts that make up the utterance of user U.

[0019] The information processing device 1 generates utterance text information by treating, among the multiple element parts, element parts whose confidentiality scores are equal to or greater than a threshold as parts that satisfy the confidentiality conditions. Note that, when the input to the trained model is text information, the information processing device 1 converts the voice information corresponding to the utterance of the user U received in step S2 into text information, and inputs the converted text information to the trained model.

[0020] Furthermore, the information processing device 1 may have a trained model that receives, as input, for example, audio information or text information corresponding to the utterance of user U received in step S2, and outputs, as utterance text information, utterance text information obtained by masking parts of the utterance of user U that satisfy confidentiality conditions. In this case, the information processing device 1 inputs the audio information or text information corresponding to the utterance of user U received in step S2 to the trained model, and can obtain text information output from the trained model as utterance text information.

[0021] Furthermore, instead of or in addition to the trained model, the information processing device 1 may have a confidential information table including information indicating multiple words to be concealed. In this case, the information processing device 1 can generate spoken text information by regarding text information or audio information corresponding to words included in the confidential information table as a portion that satisfies the confidentiality conditions.

[0022] In addition, the character strings included in the confidential information table may be character strings expressed in regular expressions. In this case, the text information or audio information specified by the regular expressions shown in the confidential information table is used as the part that satisfies the confidentiality conditions to generate spoken text information.

[0023] Masking is performed by replacing the portion of the utterance of user U that satisfies the confidentiality conditions with at least one of a specific character, symbol, and pattern. In the example shown in Figure 1, speech text information is generated in which the portion of the utterance of user U that satisfies the confidentiality conditions is replaced with the character "X."

[0024] Specifically, in the example shown in Figure 1, the spoken text information is, "I recently moved into an apartment in XXXXX, Tokyo, and it's very comfortable. XXXX and others were very helpful to me when I moved. I'm having a party soon, so please come along. My contact information is XXX-XXXX-XXXX."

[0025] The information processing device 1 provides the utterance text information generated in step S3 (step S4). For example, the information processing device 1 can provide the utterance text information to the user U by displaying the utterance text information on a display unit. The information processing device 1 can also provide the utterance text information to the user U by outputting the utterance text information as sound from a speaker.

[0026] When providing spoken text information by voice, the information processing device 1 can silence the masked portion or set a preset sound (for example, a beep). The information processing device 1 can also provide the spoken text information to others by transmitting the spoken text information to their terminal devices. This makes it possible to prevent the content of that portion from being revealed to others, even if the utterance of the user U includes a portion that satisfies the confidentiality conditions.

[0027] Furthermore, the information processing device 1 can unmask a masked portion of the utterance text information displayed on the display unit when the user U selects the masked portion. This allows the user U, when wanting to know the masked content, to easily find out the masked content by operating the information processing device 1. Note that the information processing device 1 can also re-mask a portion of the utterance text information when the user U unselects the masked portion.

[0028] In this way, the information processing device 1 according to the embodiment generates utterance text information, which is text information obtained by converting the utterance of the user U into text by masking parts of the utterance of the user U that satisfy confidentiality conditions, and provides the generated utterance text information to the user U. This enables the information processing device 1 to improve convenience for the user U.

[0029] The configuration of an information processing system including the information processing device 1 that performs such processing will be described in detail below.

[0030] [2. Information Processing System Configuration] Next, the configuration of an information processing system including an information processing device 1 according to an embodiment will be described with reference to Fig. 2. Fig. 2 is a diagram showing an example of the configuration of the information processing device 1 according to an embodiment. As shown in Fig. 2, an information processing system 100 according to an embodiment includes an information processing device 1 and an information providing device 2. The information processing device 1 and the information providing device 2 are connected to each other via a network N so as to be able to communicate with each other via a wired or wireless connection. Note that the information processing system 100 shown in Fig. 2 may include a plurality of information processing devices 1 and a plurality of information providing devices 2.

[0031] The information processing device 1 is, for example, a smart speaker, a desktop PC (Personal Computer), a notebook PC, a tablet terminal, a mobile phone, or a PDA (Personal Digital Assistant), etc. Note that the information processing device 1 is not limited to the above examples and may be, for example, a smart watch or a wearable device.

[0032] The information providing device 2 provides online services to the user U. The services provided by the information providing device 2 include, but are not limited to, online services such as a search service, an information providing service, an e-commerce service, an auction service, a music distribution service, and a video distribution service. The information providing services include various services such as a search service provided by a search site, a news distribution service provided by a news site, a traffic information providing service provided by a traffic information site, and a weather information providing service provided by a weather information site.

[0033] The information providing device 2 is an information processing device capable of communicating with various devices via a predetermined network N such as the Internet, and is realized by, for example, a server device or a cloud system. For example, the information providing device 2 is connected to various other devices via the network N so as to be able to communicate with them.

[0034] [3. Information Processing Device 1] As shown in FIG. 2, the information processing device 1 according to the embodiment includes a communication unit 10, a display unit 11, an operation unit 12, a memory unit 13, an audio input unit 14, an audio output unit 15, a position detection unit 16, and a processing unit 17.

[0035] [3.1. Communication Unit 10] The communication unit 10 is realized by, for example, a network interface card (NIC), etc. The communication unit 10 is connected to a network N by wire or wirelessly, and transmits and receives information to and from the information providing device 2 via the network N.

[0036] [3.2. Display section 11] The display unit 11 is, for example, an LCD (Liquid Crystal Display) or an organic EL (Electro Luminescence) display.

[0037] [3.3. Operation unit 12] The operation unit 12 includes, for example, a keyboard including keys for inputting letters, numbers, and spaces, an enter key, and arrow keys, a mouse, a power button, etc. When the display unit 11 is a touch panel display device, the operation unit 12 may include a touch panel.

[0038] [3.4. Storage section 13] The storage unit 13 is realized by, for example, a semiconductor memory element such as a random access memory (RAM) or a flash memory, or a storage device such as a hard disk or an optical disk.

[0039] Various types of information are stored in the storage unit 13. For example, the storage unit 13 stores information transmitted from the information providing device 2 and acquired by the processing unit 17 via the network N and the communication unit 10. The storage unit 13 also stores voice information and text information corresponding to the speech of the user U.

[0040] [3.5. Audio Input Unit 14] The voice input unit 14 converts a voice signal, which is a signal of the voice uttered by the user U, into a digital signal and outputs the converted digital signal, which is a voice digital signal, as voice information to the processing unit 17. The voice input unit 14 includes, for example, a microphone and an AD (Analog to Digital) converter that converts the voice signal, which is an electrical analog signal output from the microphone, into a digital signal.

[0041] [3.6. Audio output unit 15] The audio output unit 15 includes, for example, a DA (Digital to Analog) converter that converts a digital audio signal, which is audio information output from the processing unit 17, into an analog audio signal, and a speaker that converts the analog audio signal output from the DA converter into sound and outputs it.

[0042] [3.7. Position detection unit 16] The position detection unit 16, for example, detects the position of the information processing device 1 and outputs position data that is data on the detected position of the information processing device 1 to the processing unit 17. The position detection unit 16 receives a plurality of positioning signals transmitted from a plurality of positioning satellites in the GNSS (Global Navigation Satellite System), and detects the position of the information processing device 1 based on the received plurality of positioning signals.

[0043] [3.8. Processing section 17] The processing unit 17 is a controller, and is realized, for example, by a CPU (Central Processing Unit) or an MPU (Micro Processing Unit) executing various programs stored in a storage device inside the information processing device 1 using RAM as a working area.

[0044] The processing unit 17 may be realized by an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array). The processing unit 17 includes a receiving unit 20, a generating unit 21, a providing unit 22, and a learning unit 23.

[0045] [3.8.1. Reception Unit 20] The reception unit 20 receives the speech of the user U based on the digital audio information output from the audio input unit 14. For example, when the user U performs a specific operation using the operation unit 12, the reception unit 20 receives the speech of the user U made thereafter.

[0046] The specific operation differs for each of the information acquisition function, information transmission function, device control function, and text conversion function, for example. The information acquisition function is a function of acquiring specific information from the information providing device 2 in response to the utterance of the user U, the information transmission function is a function of transmitting information in response to the utterance of the user U, and the device control function is a function of controlling peripheral devices in response to the utterance of the user U. As described above, the text conversion function is a function of converting the utterance of the user U into text and outputting the text information that is the textual information.

[0047] Furthermore, when the user U utters a specific keyword, the reception unit 20 can also receive subsequent utterances from the user U. Whether the user U has uttered a specific keyword is determined by voice recognition of the voice information signal output from the voice input unit 14. Note that the specific keyword differs for each of the information acquisition function, information transmission function, device control function, and text conversion function, for example.

[0048] The reception unit 20 has a voice recognition function that converts voice information output from the voice input unit 14 in response to speech by the user U into text information. The reception unit 20 may also have a function that analyzes the meaning of the text information converted by the voice recognition function.

[0049] The reception unit 20 outputs voice information or text information corresponding to the user's utterance to the generation unit 21. The voice information corresponding to the user's utterance is voice information output from the voice input unit 14, and the text information corresponding to the user's utterance is information obtained by converting the voice information corresponding to the user's utterance into text using a voice recognition function. The reception unit 20, for example, associates the voice information and text information corresponding to the utterance of the user U with each utterance of the user U and stores them in the storage unit 13.

[0050] The reception unit 20 also receives a setting or change of anonymity conditions that specify whether masking is required. For example, the reception unit 20 receives a setting or change of anonymity conditions when information on the anonymity conditions is input or changed in response to an operation by the user U on the operation unit 12.

[0051] [3.8.2. Generation unit 21] In the text conversion function, the generation unit 21 generates utterance text information, which is text information obtained by converting the utterance of the user U received by the reception unit 20 into text by masking parts that satisfy confidentiality conditions.

[0052] The confidentiality conditions are, for example, words indicating privacy-related content such as addresses, names, and telephone numbers, or words indicating content that violates public order and morals. Words indicating content that violates public order and morals include, for example, discriminatory or derogatory words, obscene words, and words that affirm or encourage crimes. As described above, the confidentiality conditions can be generated or changed by the user U.

[0053] The generation unit 21 has a trained model that takes, for example, audio information or text information corresponding to the user U's utterance as input and outputs a confidentiality score indicating the degree of confidentiality for each of the multiple element parts that make up the user U's utterance.

[0054] The generation unit 21 inputs the audio information or text information corresponding to the utterance of the user U received by the reception unit 20 into the trained model, and acquires from the trained model an anonymity score indicating the degree of anonymity for each of the multiple element parts constituting the utterance of the user U. The generation unit 21 generates utterance text information by regarding, among the multiple element parts, element parts whose anonymity score is equal to or greater than a threshold as parts that satisfy the anonymity conditions.

[0055] In addition, the generation unit 21 may have a trained model that takes, for example, audio information or text information corresponding to the utterance of user U as input and outputs, as utterance text information, utterance text information in which parts of the utterance of user U that satisfy confidentiality conditions are masked.

[0056] In this case, the generation unit 21 inputs the voice information or text information corresponding to the utterance of the user U received by the reception unit 20 into the trained model, and can obtain the text information output from the trained model as the utterance text information.

[0057] The generation unit 21 may have a plurality of trained models according to the use of the spoken text information, for example. The use of the spoken text information may be for email, SNS, word-of-mouth posting, or electronic bulletin board use, for example. The generation unit 21 can determine the use of the spoken text information according to, for example, an operation by the user U on the operation unit 12, a designation by the user U through speech, or the type of application displayed on the display unit 11.

[0058] The generation unit 21 may have, for example, a plurality of trained models according to the context of the user U. The context includes the current location of the user U or the motion state of the user U. The current location or the motion state of the user U is determined based on the location detected by the location detection unit 16, for example.

[0059] The trained model is generated by machine learning using a neural network such as a convolutional neural network or a recurrent neural network, but is not limited to such examples. For example, instead of a neural network, the trained model may be generated using machine learning using other learning algorithms, such as a learning algorithm for a regression method such as linear regression, multiple regression, or logistic regression.

[0060] Furthermore, instead of or in addition to the trained model, the generation unit 21 may have a confidential information table including information indicating a plurality of words to be concealed. In this case, the information processing device 1 can generate spoken text information by regarding text information or audio information corresponding to words included in the confidential information table as a portion that satisfies the confidentiality conditions.

[0061] The character strings included in the confidential information table may be character strings expressed by regular expressions. In this case, the text information or audio information specified by the regular expressions shown in the confidential information table is used as the part that satisfies the confidentiality conditions to generate the spoken text information. For example, in the case of a telephone number, a character string expressed by a regular expression is "[0-9]{3}-[0-9]{4}-[0-9]{4}".

[0062] Furthermore, instead of or in addition to the trained model, the generation unit 21 may have a confidentiality information table including information on confidentiality for each different word. The information on confidentiality is information indicating the confidentiality level, which is set to, for example, two or more levels. For example, the generation unit 21 can lower the confidentiality level to be masked as the proportion of words with set confidentiality levels included in the utterance of the user U increases.

[0063] For example, suppose there are three confidentiality levels, from level 1 to level 3, with the confidentiality levels increasing in order from level 1 to level 3. That is, the lowest confidentiality level is level 1, the next lowest confidentiality level is level 2, and the highest confidentiality level is level 3.

[0064] In this case, if the rate at which words with a confidentiality level set are included in the utterance of user U is less than the first threshold, the generation unit 21 targets words with a confidentiality level of level 3 as masking targets. Furthermore, if the rate at which words with a confidentiality level set are included in the utterance of user U is equal to or greater than the first threshold and less than the second threshold, the generation unit 21 targets words with a confidentiality level of level 2 or higher as masking targets. Furthermore, if the rate at which words with a confidentiality level set are included in the utterance of user U is equal to or greater than the second threshold, the generation unit 21 targets words with a confidentiality level of level 1 or higher as masking targets.

[0065] Furthermore, instead of or in addition to the proportion of words set to a confidentiality level included in the utterance of the user U, the generation unit 21 can also determine the confidentiality level to be masked according to the purpose of the utterance text information and the context of the user U. The user U can also specify the confidentiality level to be masked by operating the operation unit 12 or by speaking to the information processing device 1. For example, the generation unit 21 can also mask words with a confidentiality level equal to or higher than the level specified by the reception unit 20.

[0066] Masking is performed by replacing the portion of user U's speech that satisfies the confidentiality conditions with at least one of specific characters, symbols, and patterns. Specific characters include, for example, "X," "·," "-," and "?", and symbols include, for example, "□," "◆," "※," and "◯." Patterns include, for example, checkerboard patterns and geometric patterns. Masking may also be performed by replacing the portion of user U's speech that satisfies the confidentiality conditions with, for example, a space.

[0067] The generation unit 21 can also generate utterance text information by using both the trained model and the secret information table. The generation unit 21 also stores a combination of voice information and text information corresponding to the utterance of the user U in the storage unit 13 for each utterance of the user U.

[0068] In addition, instead of masking with characters or symbols with the same number of characters as the part that satisfies the confidentiality conditions, the generation unit 21 can also mask with characters or symbols with a different number of characters than the part that satisfies the confidentiality conditions, for example, when the confidentiality level is above a threshold value.

[0069] [3.8.3.Providing Department 22] The providing unit 22 provides the utterance text information generated by the generating unit 21. For example, the providing unit 22 provides the utterance text information to the user U by causing the display unit 11 to display the utterance text information.

[0070] Fig. 3 is a diagram showing an example of spoken text information displayed on the display unit 11 of the information processing device 1 according to the embodiment. The example shown in Fig. 3 is an example of spoken text information displayed on the display unit 11 of the information processing device 1 when a user U utters, "When I went shopping at the AAA supermarket near my house, I met and talked with Tokkyo Hanako, and she told me that her son, Jiro, joined BBB Shoji this year."

[0071] The spoken text information shown in Figure 3 is, "When I went shopping at the XXXXXXX store near my house, I met and talked with Mr. XXXX, who told me that his son, XX, joined the expensive XXXXX this year." In the spoken text information shown in Figure 3, the parts that satisfy the confidentiality conditions, "AAA Supermarket," "Hanako Patent," and "BBB Trading," are masked with the character string "X."

[0072] Masking is performed by replacing the portion of the utterance of user U that satisfies the confidentiality conditions with at least one of a specific character, a symbol, and a pattern. In the example shown in Fig. 3, speech text information is generated in which the portion of the utterance of user U that satisfies the confidentiality conditions is replaced with the character "X."

[0073] The providing unit 22 can also convert the spoken text information into audio information by speech synthesis and output the converted audio information to the audio output unit 15. This allows the providing unit 22 to provide the spoken text information to the user U by voice. When converting the spoken text information into audio information by speech synthesis, the providing unit 22 can convert a masked part of the spoken text information into a specific sound (for example, a "beep" sound), and can also silence the masked part of the spoken text information.

[0074] In addition, based on the dialogue with user U, the providing unit 22 can transmit information corresponding to user U's utterance to the information providing device 2, obtain information corresponding to user U's utterance from the information providing device 2, and control peripheral devices corresponding to user U's utterance.

[0075] Furthermore, the providing unit 22 removes the masking when a masked portion of the utterance text information displayed on the display unit 11 is selected. The user U can select the masked portion by operating the operation unit 12, for example.

[0076] 4 is a diagram showing an example in which masking of a selected portion of spoken text information displayed on the display unit 11 of the information processing device 1 according to the embodiment is released. In the example shown in FIG. 4, "XXXXX" is selected by the user U, and the masking of "XXXXX" selected by the user U is released, so that "BBB Trading" is displayed.

[0077] The providing unit 22 can generate, as the utterance information, information including, for example, text information obtained by converting the utterance of the user U into text without masking, and information indicating the position of the part of the utterance of the user U that satisfies the confidentiality conditions. In this case, the providing unit 22 can provide the utterance text information to the user U or transmit it to an external device via the communication unit 10 based on the generated utterance information.

[0078] [3.8.4. Learning Section 23] The learning unit 23 can generate and update a trained model. For example, the learning unit 23 generates and updates a trained model using training data that includes speech information and text information corresponding to utterances of the user U that satisfy confidentiality conditions.

[0079] The learning data is generated, for example, based on the past speech history of the user U. For example, the learning unit 23 generates the learning data based on the speech of the user U that is accepted by the accepting unit 20 as words that satisfy the confidentiality conditions.

[0080] In addition, the learning unit 23 can also generate learning data by displaying the user U's past speech history on the display unit 11 and having the user U select words that satisfy the confidentiality conditions by operating the operation unit 12, etc.

[0081] For example, the learning unit 23 acquires the past speech history of the user U from the storage unit 13. The past speech history of the user U includes voice information and text information corresponding to the past utterances of the user U, and the learning unit 23 generates, as learning data, a combination of voice information and text information corresponding to words that satisfy the confidentiality conditions selected by the user U.

[0082] The learning unit 23 can also generate, as training data, a combination of information indicating the positions of words that satisfy the confidentiality conditions selected by the user U and audio information or text information corresponding to past utterances of the user U. The trained model generated or updated using such training data is a model that outputs a confidentiality score indicating the degree of confidentiality for each position (position from the first word) of a word included in the audio information or text information corresponding to the utterances of the user U.

[0083] In this case, the generation unit 21 inputs audio information or text information corresponding to the utterance of the user U into the trained model, and generates the utterance text information by masking positions where the confidentiality score output from the trained model is equal to or greater than a threshold. Note that the generation unit 21 can set the threshold higher, for example, as the above-mentioned confidentiality level increases.

[0084] The learning unit 23 can generate multiple trained models according to the application of the utterance text information, and can also generate multiple trained models according to the context of the user U. Note that the learning unit 23 only needs to be able to generate the trained models described above, and the learning process by the learning unit 23 is not limited to the process described above.

[0085] [4. Processing Procedure] Next, a procedure of information processing by the processing unit 17 of the information processing device 1 according to the embodiment will be described. Figures 5 and 6 are flowcharts showing an example of information processing by the processing unit 17 of the information processing device 1 according to the embodiment.

[0086] First, a description will be given of Fig. 5. Fig. 5 shows an example of the utterance text information generation process performed by the information processing device 1. The processing unit 17 of the information processing device 1 receives an utterance from the user U (step S10).

[0087] Next, the processing unit 17 generates utterance text information, which is text information obtained by converting the user U's utterance received in step S10 into text by masking parts that satisfy the confidentiality conditions (step S11). In the processing of step S11, the processing unit 17 inputs, for example, audio information or character information corresponding to the user U's utterance into the trained model, and obtains, from the trained model, a confidentiality score indicating the degree of confidentiality for each of the multiple element parts that make up the user U's utterance. The processing unit 17 generates utterance text information by regarding, among the multiple element parts, element parts whose confidentiality scores are equal to or greater than a threshold as parts that satisfy the confidentiality conditions.

[0088] Then, the processing unit 17 provides the utterance text information generated in step S11 (step S12), and ends the processing shown in Fig. 5. For example, in the processing of step S12, the processing unit 17 can, for example, display the utterance text information on the display unit 11 or output the utterance text information as sound from the audio output unit 15.

[0089] Next, Fig. 6 will be described. Fig. 6 shows an example of a learning process performed by the information processing device 1. The processing unit 17 of the information processing device 1 generates learning data (step S20). In the process of step S20, for example, the processing unit 17 generates the learning data based on an utterance of the user U that is accepted by the accepting unit 20 as a word that satisfies the confidentiality condition.

[0090] Next, a trained model is generated or updated using the training data generated in step S20 (step S21), and the processing shown in FIG. 6 ends.

[0091] [5. Modifications] The above-mentioned information processing device 1 has been described as a device such as a terminal device operated by a user U, but the information processing device 1 may also be a server device, etc. Fig. 7 is a diagram showing another example of the configuration of the information processing device 1 according to the embodiment.

[0092] 7 converts the speech of a user U into voice information or text information, and transmits the user speech information including the converted voice information or text information to the information processing device 1 via the network N. The reception unit 20 of the information processing device 1 shown in FIG. 7 receives the user speech information transmitted from the terminal device 3 via the network N and the communication unit 10, thereby receiving the speech of the user U.

[0093] The generation unit 21 of the information processing device 1 generates utterance text information based on the audio information or text information included in the user utterance information. The provision unit 22 provides the utterance text information by transmitting the utterance text information generated by the generation unit 21 to the terminal device 3 via the communication unit 10 and the network N. The terminal device 3 receives the utterance text information transmitted from the information processing device 1, and displays the received utterance text information or outputs it as sound.

[0094] Note that some of the functions of the processing unit 17 in the information processing device 1 shown in Fig. 2 may be realized by the terminal device 3 shown in Fig. 7. Also, some of the functions of the processing unit 17 in the information processing device 1 shown in Fig. 2 may be realized by the information providing device 2.

[0095] 2 has a trained model and a secret information table for each user U, and the trained model in this case can also be called an on-device model. Also, the information processing device 1 shown in FIG. 7 may have a trained model and a secret information table that are common to all users U, or may have a trained model and a secret information table for each user U.

[0096] In the above example, the information processing device 1 shown in Fig. 2 has been described as having an AI assistant function that supports interactive voice operations, but the information processing device 1 may be a device that does not have the AI ​​assistant function. For example, the information processing device 1 may be a voice recorder, etc. The text conversion function of the information processing device 1 can also be used, for example, when converting meeting minutes into text.

[0097] [6. Hardware Configuration] The information processing device 1 according to the embodiment described above is realized by, for example, a computer 80 configured as shown in Fig. 8. Fig. 8 is a hardware configuration diagram showing an example of the computer 80 that realizes the functions of the information processing device 1 according to the embodiment. The computer 80 has a CPU 81, a RAM 82, a ROM (Read Only Memory) 83, an HDD (Hard Disk Drive) 84, a communication interface (I / F) 85, an input / output interface (I / F) 86, and a media interface (I / F) 87.

[0098] The CPU 81 operates and controls each part based on programs stored in the ROM 83 or the HDD 84. The ROM 83 stores a boot program executed by the CPU 81 when the computer 80 starts up, programs that depend on the hardware of the computer 80, and the like.

[0099] The HDD 84 stores programs executed by the CPU 81, data used by such programs, etc. The communication interface 85 receives data from other devices via the network N (see FIG. 2) and sends it to the CPU 81, and transmits data generated by the CPU 81 to other devices via the network N.

[0100] The CPU 81 controls output devices such as a display and a printer, and input devices such as a keyboard and a mouse, via the input / output interface 86. The CPU 81 acquires data from the input devices via the input / output interface 86. The CPU 81 also outputs generated data to the output devices via the input / output interface 86.

[0101] The media interface 87 reads a program or data stored in a recording medium 88 and provides it to the CPU 81 via the RAM 82. The CPU 81 loads the program or data from the recording medium 88 onto the RAM 82 via the media interface 87 and executes the loaded program. The recording medium 88 is, for example, an optical recording medium such as a DVD (Digital Versatile Disc) or a PD (Phase Change Rewritable Disc), a magneto-optical recording medium such as an MO (Magneto-Optical disk), a tape medium, a magnetic recording medium, or a semiconductor memory.

[0102] For example, when the computer 80 functions as the information processing device 1 according to the embodiment, the CPU 81 of the computer 80 executes programs loaded onto the RAM 82 to realize the functions of the processing unit 17. In addition, the HDD 84 stores data in the storage unit 13. The CPU 81 of the computer 80 reads and executes these programs from a recording medium 88, but as another example, these programs may be acquired from another device via the network N.

[0103] [7. Other] Furthermore, among the processes described in the above embodiments and modifications, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically using known methods. In addition, the information including the processing procedures, specific names, various data, and parameters shown in the above documents and drawings can be changed as desired unless otherwise specified. For example, the various information shown in each drawing is not limited to the information shown in the drawings.

[0104] Furthermore, the components of each device shown in the figure are conceptual functional components and do not necessarily have to be physically configured as shown in the figure. In other words, the specific form of distribution and integration of each device is not limited to that shown in the figure, and all or part of them can be functionally or physically distributed and integrated in any unit depending on various loads, usage conditions, etc.

[0105] Furthermore, the above-described embodiments and modifications can be combined as appropriate within the scope of not causing any contradiction in the processing content.

[0106] [8. Effects] As described above, the information processing device 1 according to the embodiment includes the receiving unit 20, the generating unit 21, and the providing unit 22. The receiving unit 20 receives an utterance from the user U. The generating unit 21 generates utterance text information, which is text information obtained by converting the utterance of the user U into text by masking parts of the utterance received by the receiving unit 20 that satisfy confidentiality conditions. The providing unit 22 provides the utterance text information generated by the generating unit 21. This enables the information processing device 1 to improve convenience for the user U.

[0107] The generation unit 21 also has a trained model that receives as input audio information or text information corresponding to the utterance of the user U received by the reception unit 20, and outputs an anonymity score indicating the degree of anonymity for each of a plurality of element parts constituting the utterance of the user U. The generation unit 21 uses the trained model to generate utterance text information by regarding, among the plurality of element parts, element parts whose anonymity scores are equal to or greater than a threshold as parts that satisfy the anonymity conditions. This allows the information processing device 1 to appropriately detect parts that satisfy the anonymity conditions.

[0108] The generation unit 21 has a trained model that receives as input audio information or text information corresponding to the utterance of the user U received by the reception unit 20, and outputs, as utterance text information, text information obtained by masking parts of the utterance of the user U that satisfy the confidentiality conditions, and generates utterance text information using the trained model. This allows the information processing device 1 to appropriately detect parts that satisfy the confidentiality conditions.

[0109] Furthermore, the generation unit 21 generates the spoken text information using a confidential information table including information indicating a plurality of words to be confidentialized, thereby enabling the information processing device 1 to appropriately detect a portion that satisfies the confidentiality conditions.

[0110] The information processing device 1 also includes a learning unit 23 that updates the trained model based on the past speech history of the user U. This enables the information processing device 1 to improve the accuracy of detecting portions that satisfy the confidentiality conditions using the trained model.

[0111] Furthermore, the generation unit 21 masks the part of the utterance of the user U that satisfies the confidentiality condition by converting the part into at least one of a specific character, a symbol, and a pattern. This allows the information processing device 1 to appropriately present the masked part to the user U.

[0112] Furthermore, the providing unit 22 provides the utterance text information to the user U by displaying the utterance text information on the display unit 11. This allows the information processing device 1 to improve the convenience for the user U.

[0113] Furthermore, the providing unit 22 removes the masking when a masked portion of the utterance text information displayed on the display unit 11 is selected. This enables the information processing device 1 to improve the convenience for the user U.

[0114] The above describes the embodiments of the present application in detail based on the drawings, but this is merely an example, and the present invention can be implemented in other forms that include the embodiments described in the Disclosure of the Invention section and that have been modified and improved in various ways based on the knowledge of those skilled in the art.

[0115] Furthermore, the above-mentioned "section, module, unit" can be read as "means" or "circuit," etc. For example, an acquisition unit can be read as an acquisition means or an acquisition circuit. [Explanation of symbols]

[0116] 1. Information processing equipment 2 Information provision device 3 Terminal Devices 10. Communications Department 11 Display section 12 Control section 13 Storage section 14 Audio input section 15 Audio output section 16 Position detection unit 17 Processing section 20 Reception 21 Generation part 22 Providing Department 23 Learning Department 100 Information Processing Systems

Claims

1. a reception unit that receives utterances from a user; a generation unit that generates utterance text information, which is text information obtained by converting the user's utterance into text by masking a portion of the utterance that satisfies a confidentiality condition among the utterance received by the reception unit; a providing unit that provides the utterance text information generated by the generating unit, The generation unit Using a confidential information table including information indicating a plurality of words to be concealed and information indicating a confidentiality level for each of the words, the greater the proportion of words to which a confidentiality level is set that are included in the utterance, the lower the confidentiality level to be the target of the masking, and the speech text information is generated.

1. An information processing device comprising:

2. A reception unit that receives utterances from a user; a generation unit that generates utterance text information, which is text information obtained by converting the user's utterance into text by masking a portion of the utterance that satisfies a confidentiality condition among the utterance received by the reception unit; a providing unit that provides the utterance text information to the user by displaying the utterance text information generated by the generating unit on a display unit, The providing unit When a masked portion of the speech text information displayed on the display unit is selected, the masking is released.

1. An information processing device comprising:

3. The generation unit The system has a trained model that receives as input audio information or text information corresponding to the utterance received by the reception unit and outputs a confidentiality score indicating the degree of confidentiality for each of a plurality of element parts constituting the user's utterance, and generates the utterance text information by using the trained model, with element parts of the plurality of element parts whose confidentiality score is equal to or greater than a threshold being treated as parts that satisfy the confidentiality condition.

2. The information processing apparatus according to claim 1, wherein:

4. The generation unit a trained model that receives as input voice information or text information corresponding to the utterance received by the reception unit, and outputs as the utterance text information text information obtained by masking a portion of the user's utterance that satisfies the confidentiality condition, and generates the utterance text information using the trained model; 2. The information processing apparatus according to claim 1, wherein:

5. The generation unit The spoken text information is generated using a confidential information table including information indicating a plurality of words to be confidentialized.

3. The information processing apparatus according to claim 2, wherein:

6. A learning unit that updates the trained model based on the user's past speech history.

5. The information processing apparatus according to claim 3, wherein the information processing apparatus is a computer.

7. The generation unit The part of the user's speech that satisfies the confidentiality condition is masked by converting the part into at least one of a specific character, a symbol, and a pattern.

6. The information processing device according to claim 1, wherein:

8. 1. A computer-implemented information processing method, comprising: a reception step of receiving an utterance from a user; a generating step of generating utterance text information, which is text information obtained by converting the user's utterance into text by masking a portion of the utterance that satisfies a confidentiality condition among the utterance received by the receiving step; a providing step of providing the utterance text information generated by the generating step, The generating step includes: Using a confidential information table including information indicating a plurality of words to be concealed and information indicating a confidentiality level for each of the words, the greater the proportion of words to which a confidentiality level is set that are included in the utterance, the lower the confidentiality level to be the target of the masking, and the speech text information is generated. An information processing method comprising:

9. A computer-implemented information processing method, comprising: a reception step of receiving an utterance from a user; a generating step of generating utterance text information, which is text information obtained by converting the user's utterance into text by masking a portion of the utterance that satisfies a confidentiality condition among the utterance received by the receiving step; a providing step of providing the utterance text information to the user by displaying the utterance text information generated by the generating step on a display unit, The providing step includes: When a masked portion of the speech text information displayed on the display unit is selected, the masking is released. An information processing method comprising:

10. A reception procedure for receiving a user's utterance; a generation step of generating utterance text information, which is text information obtained by converting the user's utterance into text by masking a portion of the utterance that satisfies a confidentiality condition among the utterance received by the reception step; a providing step of providing the utterance text information generated by the generating step; The generating procedure includes: Using a confidential information table including information indicating a plurality of words to be concealed and information indicating a confidentiality level for each of the words, the greater the proportion of words to which a confidentiality level is set that are included in the utterance, the lower the confidentiality level to be the target of the masking, and the speech text information is generated. An information processing program characterized by:

11. A reception procedure for receiving a user's utterance; a generation step of generating utterance text information, which is text information obtained by converting the user's utterance into text by masking a portion of the utterance that satisfies a confidentiality condition among the utterance received by the reception step; a providing step of providing the utterance text information to the user by displaying the utterance text information generated by the generating step on a display unit; The providing step comprises: When a masked portion of the speech text information displayed on the display unit is selected, the masking is released. An information processing program characterized by:

Citation Information

Patent Citations

  • Display mode alteration program and display control device

    JP2006244048A

  • Consultation fee settlement terminal device and user terminal device

    JP2007140856A

  • System and program for managing personal information

    JP2008186473A

  • Document management device and program

    JP2016012210A

  • Information processing device, information processing method and program

    JP2020149628A