Agent device program

The agent device program addresses privacy issues in multi-user voice assistant systems by estimating user numbers and managing personal information access, ensuring privacy and convenience for all users.

JP7767479B2Active Publication Date: 2025-11-11HONDA MOTOR CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024014365
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-02-01
Publication Date
2025-11-11
Estimated Expiration
2044-02-01

AI Technical Summary

Technical Problem

Conventional voice assistant services do not adequately address privacy concerns when multiple users, such as passengers and drivers, use the same system, leading to potential exposure of personal information.

Method used

A program for an agent device that collects user information, estimates the number of users, and determines the appropriate use and output of personal information based on user privacy settings, ensuring privacy protection in multi-user environments.

Benefits of technology

The program effectively protects user privacy by preventing personal information from being revealed to other users, allowing for convenient assistance tailored to individual needs while maintaining confidentiality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007767479000001
    Figure 0007767479000001
  • Figure 0007767479000002
    Figure 0007767479000002
  • Figure 0007767479000003
    Figure 0007767479000003
Patent Text Reader

Abstract

To protect the privacy of a user of an agent device.SOLUTION: A program for an agent device causes a computer to execute: input processing (S10) of receiving instruction information by a voice or a touch operation from a user and user information on the user; recording processing (S20) of recording information including personal information on the user in a storage unit; estimation processing (S30) of estimating the number of users currently using the agent device from the user information; determination processing (S40) of determining whether the personal information in the storage unit can be used or not on the basis of whether the estimated number of users is single or plural; decision processing (S50) of deciding output information to the instruction information by using the information in the storage unit on the basis of a result of the determination; and output processing (S60) of outputting a voice and / or an image as the decided output information.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a program for controlling an agent device. [Background technology]

[0002] As part of this type of technology, voice assistant services installed on smartphones and in-vehicle agent devices are becoming increasingly popular with the advancement of AI (Artificial Intelligence) technology. These voice assistant services allow users to enjoy a variety of services, such as operating devices connected to a communication network, such as air conditioners and lights, searching for information on the Internet, playing music and news, and reading out schedules and messages. Generally, voice assistant services, also known as personal voice assistant services, are intended to be used by a single user, and security concerns arise when multiple users use the service. In response to such concerns, for example, Patent Document 1 discloses a technology for setting access privileges for a new user through a spoken introduction from a trusted user. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent No. 6949149 Summary of the Invention [Problem to be solved by the invention]

[0004] In the conventional technology, access privileges are set for multiple users, but it is not assumed that multiple users who have been set will use the service at the same time, and for example, an incoming message from one user may be seen by other users. Such concerns are more pronounced in vehicles where there may be passengers in addition to the driver. [Means for solving the problem]

[0005] The program for an agent device according to one aspect of the present invention includes an input process for receiving instruction information from a user by voice or touch operation and user information relating to the user; At least one of the following will be collected from the device used by the user: information about SNS, SMS, emails, content subscription or viewing, internet information previously collected, and air conditioner setting information, which are used as programs other than the agent device program. Information including personal information as The computer is caused to perform a recording process for recording in a memory unit, an estimation process for estimating the number of users currently using the system based on the user information, a judgment process for determining whether or not personal information in the memory unit can be used based on whether the estimated number of users is single or multiple, a decision process for using the information in the memory unit based on the judgment result to determine output information for instruction information, and an output process for outputting audio and / or images as the determined output information. [Effects of the Invention]

[0006] According to the present invention, when a plurality of users use the system, it is possible to protect the privacy of the users. [Brief explanation of the drawings]

[0007] [Figure 1] FIG. 1 is a schematic diagram illustrating a configuration of an agent system according to an embodiment. [Figure 2A] FIG. 2 is a diagram illustrating an example of the configuration of a main part of a server device. [Figure 2B] FIG. 2B is a diagram illustrating a functional configuration of the assist unit in FIG. 2A; [Figure 3] FIG. 10 is a diagram illustrating ranked personal information. [Figure 4] 10 is a flowchart illustrating an example of a flow of processing executed by a server device. DETAILED DESCRIPTION OF THE INVENTION

[0008] <Summary> The agent system according to the embodiment uses technologies such as voice recognition and natural language processing to interpret the content of speech made by service users who use the agent service, and responds to questions and carries out requests (which may also be called voice instructions) input by voice from the service users, and is also called a voice assistant. In the embodiment, as an example, a communication terminal such as a smartphone that executes an application program (called an app) for using the agent service is linked with a server device that executes a program for the agent service to provide the agent service to service users who use the communication terminal.

[0009] When an agent service is used in an environment where multiple service users (hereinafter simply referred to as users) exist, for example, if one user requests the agent system to notify the user of an incoming message on his or her communication terminal, the agent system that receives the request will read out the message, and the contents of the message will become known to the other users. To avoid such a situation, the agent system according to the embodiment takes into consideration protecting the privacy of users when using the agent service in an environment where multiple users exist. The configuration of such an agent system will be described in more detail with reference to the drawings.

[0010] <Agent System> FIG. 1 is a schematic diagram illustrating the configuration of an agent system (hereinafter referred to as a voice assistance system) 400 according to an embodiment. As shown in FIG. 1, the voice assistance system 400 is configured so that a server device 200 and a plurality of devices 100 can communicate with each other via a network 300. The server device 200 is a server device for the voice assistance system 400. The devices 100 correspond to communication terminals used by each user. Although FIG. 1 illustrates three devices 100, 100a, 100b, and 100c, there are actually a large number of devices 100 corresponding to the number of users.

[0011] The server device 200 executes a program for the voice assistance system 400 to provide voice assistance in response to questions and requests from each device 100. By executing an app for the voice assistance system 400, each device 100 transmits questions and requests input by the user via voice to the server device 200, and performs operations based on the answers and instructions transmitted from the server device 200. The device 100 may be, for example, a smartphone, a phablet, a tablet, a smartwatch, a laptop PC, a desktop PC, an Internet TV, a Homehub, a PDA, a mobile phone, and various home appliances, or may be a voice assistant device installed in a vehicle.

[0012] The network 300 has a function of connecting the server device 200 and the plurality of devices 100 so that they can communicate with each other, and is, for example, the Internet or a wired or wireless LAN (Local Area Network).

[0013] In the embodiment, for example, each device 100 records the speech content (corresponding to the above-mentioned voice instruction) of each user, and transmits the recorded data as user speech information to the server device 200. The server device 200 performs voice recognition on the speech information transmitted from each device 100, interprets the user's instruction, and performs voice assistance via the corresponding device 100.

[0014] <Server device> 2A is a diagram illustrating an example of the configuration of the main parts of server device 200. Server device 200 has a CPU that performs various calculations, a storage device that stores various data, programs, etc., and includes, as functional components, a communication unit 210, a voice recognition unit 220, an assist unit 230, and a storage unit 240.

[0015] The communication unit 210 communicates data with a plurality of devices 100 via the network 300 . The speech recognition unit 220 converts the speech information received from the multiple devices 100 via the communication unit 210 into text data by performing speech recognition on each of the speech information. As an example, the speech recognition unit 220 performs acoustic analysis on the recorded data and converts it into text using a speech recognition dictionary such as an acoustic model, a language model, and a pronunciation dictionary.

[0016] The assist unit 230 interprets (in other words, performs semantic analysis of) the instruction content from the user based on the converted text data, and executes an assist process according to the interpreted instruction content. The storage unit 240 stores, for example, information about users who use the voice assistant service (which may be called personal information) for each user. The personal information includes registration information and usage information.

[0017] The registration information is information that associates a user name (which may be an account serving as a user ID (identification)) with device information of the device 100 used by the user. The device information may include, for example, a device name (which may be a device ID), a model name, an IP (Internet Protocol) address, and its specifications (for example, an output sound pressure level, frequency characteristics, crossover frequency, input impedance, allowable input, etc. of a speaker serving as an output device, a screen size, resolution, etc. of a display serving as an output device).

[0018] The usage information is information acquired from the device 100 used by the user when using the voice assistant service. For example, the usage information includes information about the social networking service (SNS), short message service (SMS), and email used on the device 100 used by the user (including the names of specific accounts followed by the user on SNS), keywords used in past searches, a history of destinations set in the past, information about content subscribed to or viewed through subscriptions, information on the Internet acquired in the past, and setting information for electronic devices such as air conditioners, electronic devices installed in vehicles, etc. It may also include the number of times the user has used the voice assistant service (which may be the frequency of use per specified period) and information stored on the user's device 100 (e.g., the user's address, password, etc.), or information inferred by the server device 200 from information related to the SNS etc. used by the user (e.g., the user's family composition, etc.).

[0019] In the embodiment, when a user starts using a voice assistant service (at the time of registration), registration information is recorded in the storage unit 240 of the server device 200 based on the user's permission. Also, when a registered user uses the voice assistant service, usage information indicating the usage history is recorded in the storage unit 240 of the server device 200. Generally, the usage information increases each time the service is used, so it can be said that the usage information stored in the storage unit 240 is updated.

[0020] <Use of personal information> A user's personal information is highly private and confidential information. When a user receives voice assistant services alone, using the user's personal information may result in more convenient assistance for the user than if the user's personal information were not used at all. However, when multiple users are present and one of them receives voice assistant services, it may be inconvenient if the user's personal information becomes known to the other users.

[0021] For example, it would be convenient for a user to have keywords used in a previous search automatically set for the next search, but the user may not want others to know about them. The same applies to destinations set in previous route searches. As another example, it is convenient for a user if the name of the content to which the user subscribes is automatically set the next time the user searches, but the user may not want others to know it. As another example, although it is convenient for a user to have information on the Internet that the user has previously acquired automatically displayed or read aloud, the user may not want others to know about it.

[0022] In consideration of the above, in this embodiment, the personal information of a user is ranked, for example, into five levels according to the level of confidentiality. The highest level of confidentiality is called Level 1, and the lowest level of confidentiality is called Level 5.

[0023] <Personal information ranking> 3 is a diagram illustrating an example of ranked personal information. When performing the assist process, the assist unit 230 differentiates the rank of personal information that can be used when using the personal information between a case where a single user receives voice assistant services and a case where one of multiple users receives voice assistant services in an environment where multiple users exist. For example, when a user receives voice assistant services alone, the user's personal information is available from level 2 to level 5. In contrast, when multiple users are present in an environment where only one of them receives voice assistant services, the user's personal information is not available. This makes it possible to prevent a user's personal information from being known to other users in an environment where multiple users are present. Furthermore, when a user receives voice assistant services alone, as described above, it becomes possible to have the voice assistant perform assistance processing that is convenient for that user. In addition, when a user receives voice assistant services alone, the level of personal information that can be used is not limited to the above-mentioned levels 2 to 5, and may be changed as appropriate.

[0024] <Assist processing> Fig. 2B is a diagram illustrating an example of the functional configuration of the assisting unit 230 of Fig. 2A. The assisting unit 230 includes a semantic analyzing unit 231, an executing unit 232, a personal information acquiring unit 233, and a recording control unit 234. The semantic analysis unit 231 performs semantic analysis of the content of the instruction from the user based on the text data obtained by the voice recognition unit 220 .

[0025] The execution unit 232 executes an assist process according to the instruction content semantically analyzed by the semantic analysis unit 231. For example, if the instruction content is a question, an answer to the question is determined and generated by a search engine (not shown) in the server device 200 or a search server (not shown) connected to the network 300. The execution unit 232 transmits the determined or generated answer as text information or audio information to the device 100 of the user who asked the question via the communication unit 210.

[0026] Furthermore, when the instruction content semantically analyzed by semantic analysis unit 231 is a setting instruction for an electronic device (not shown) connected to network 300 or an electronic device provided in a vehicle (not shown), execution unit 232 determines an operation to output a setting instruction signal to the electronic device to be set. Execution unit 232 transmits the setting instruction signal to the electronic device to be set via communication unit 210.

[0027] Furthermore, if the instruction content semantically analyzed by the semantic analysis unit 231 is an instruction to output and play specific music content through speakers (or headphones, etc.), the execution unit 232 determines the operation of playing audio through speakers (or headphones, etc.) of the user's device 100 based on music content information stored in the user's device 100 or a music server (not shown) connected to the network 300. The execution unit 232 transmits a signal indicating the storage location of the corresponding music content and a playback instruction to the user's device 100 via the communication unit 210. The audio playback may be either streaming playback or download playback.

[0028] As described above, when the user starts using the voice assistant service (at the time of registration), the personal information acquisition unit 233 acquires registration information from the user's device 100 based on the user's permission. Furthermore, as described above, when a registered user uses the voice assistant service, the personal information acquisition unit 233 acquires new usage information from the device 100 of the user.

[0029] The recording control unit 234 controls the operation of recording the above-mentioned registration information and the operation of recording the above-mentioned usage information in the storage unit 240. In addition, the recording control unit 234 controls the operation of reading the registration information stored in the storage unit 240 and the operation of reading the usage information stored in the storage unit 240, as necessary.

[0030] In the above configuration, the functions of the voice recognition unit 220 and the assistance unit 230 can be realized by a CPU (not shown) of the server device 200 and a program for the voice assistance system 400. The program for the voice assistance system 400 is a program for performing voice recognition on recorded data of user utterances transmitted from multiple devices 100, and providing the user with the above-mentioned voice assistance service.

[0031] <Explanation of the flowchart> 4 is a flowchart illustrating an example of the flow of processing executed by the server device 200. The CPU of the server device 200 repeatedly executes the processing illustrated in FIG.

[0032] 4, server device 200 performs input processing and proceeds to step S20. More specifically, server device 200 receives, via communication unit 210, location information indicating the current location of device 100 used by the user and instruction information (corresponding to the above-mentioned speech information) indicating a question or request input by the user to device 100 through speech. The user may input a question or request by touching an operating member of device 100, operating a button, operating a key, or the like (collectively referred to as touch operations). In this case, the instruction information corresponds to text information generated by device 100 based on the touch operation.

[0033] In step S20, the server device 200 performs a recording process and proceeds to step S30. More specifically, the server device 200 records the instruction information transmitted from the device 100 in a predetermined area of ​​the storage unit 240 (which may be referred to as an instruction information storage area).

[0034] In step S30, the server device 200 performs an estimation process and proceeds to step S40. More specifically, the server device 200 estimates whether there is one or more users based on the location information of each device 100 transmitted from the multiple devices 100 to the server device 200. If the distance between the multiple devices 100 whose location information was input in step S10 is within a predetermined distance (e.g., 2 m), the server device 200 estimates that there is one user. On the other hand, if the distance between the multiple devices 100 exceeds the predetermined distance, or if the number of devices 100 whose location information was input in step S10 is one, the server device 200 estimates that there is one user.

[0035] In step S40, server device 200 performs a determination process and proceeds to step S50. More specifically, based on whether the number of users estimated in step S30 is one or multiple, server device 200 determines whether the personal information in storage unit 240 can be used if the number is one, and determines whether the personal information in storage unit 240 cannot be used if the number is multiple.

[0036] In step S50, the server device 200 performs a determination process and proceeds to step S60. More specifically, the assist unit 230 of the server device 200 determines output information for the instruction information using personal information stored in the storage unit 240 based on the result of the determination in step S40. Specifically, if the estimated number of users is one and the instruction from the user is a question, the assist unit 230 determines and generates an answer to the question using personal information stored in the storage unit 240. Furthermore, if the estimated number of users is one and the instruction from the user is a setting instruction for an electronic device or the like connected to the network 300, the assist unit 230 determines an output operation of a setting instruction signal for the electronic device to be set using personal information stored in the storage unit 240. Furthermore, if the estimated number of users is one and the instruction is an instruction to output and play specific music content through headphones or the like, the assist unit 230 determines an audio playback operation using personal information stored in the storage unit 240. On the other hand, if the estimated number of users is plural, the assisting unit 230 of the server device 200 determines output information for each piece of instruction information without using the personal information in the storage unit 240.

[0037] In step S60, the server device 200 performs output processing and ends the processing in Fig. 4. More specifically, the server device 200 outputs the output information determined in step S50 to the corresponding devices 100 via the communication unit 210.

[0038] According to the embodiment described above, the following advantageous effects are achieved. (1) The program for the voice assistance system 400 executed by the server device 200 constituting the voice assistance system 400 as an agent device causes the computer (server device 200) to execute the following steps: an input process (S10) for accepting instruction information by voice or touch operation from the user and user information about the user; a recording process (S20) for recording information including personal information about the user in the memory unit 240; an estimation process (S30) for estimating the number of users currently using the system from the user information; a determination process (S40) for determining whether the personal information in the memory unit 240 can be used based on whether the estimated number of users is single or multiple; a decision process (S50) for using the information in the memory unit 240 based on the result of the determination to determine output information for the instruction information; and an output process (S60) for outputting audio and / or images as the determined output information. This configuration makes it possible to protect the privacy of users when multiple users use the voice assistant service by preventing one user's personal information from being revealed to other users.

[0039] (2) In the program (1) above, if the number of users estimated by the estimation process (S30) is one, the judgment process (S40) determines whether personal information about the user corresponding to the instruction information received in the input process (S10) can be used, and the decision process (S50) determines the output information for the instruction information using information including personal information in the memory unit 240. With this configuration, when there is a single user, the assist process is performed using the personal information of the user, so that it is possible to perform the assist process more convenient for a single user than when the personal information is not used.

[0040] (3) In the program (1) above, if the number of users estimated by the estimation process (S30) is more than one, the judgment process (S40) determines whether the personal information can be used or can be used within a specified range, and the decision process (S50) determines the output information for the instruction information using information excluding personal information in the memory unit 240, or using personal information within a specified range and information excluding personal information in the memory unit 240. With this configuration, when there are multiple users, the assist process is performed by disabling the use of personal information of the users or by limiting the range of use, so that it is possible to perform an assist process that is convenient for the user while avoiding a situation in which important personal information of one user is known to other users.

[0041] (4) In the program of (3) above, the user information includes relationship level information indicating the strength of the relationship with other users, and the personal information includes importance level information indicating the level of importance of each piece of information, and the determination process (S50) changes a predetermined range of the personal information in the memory unit 240 based on the relationship level information and the importance level information. With this configuration, it becomes possible to change the scope of use of personal information restricted when there are multiple users, depending on the importance of the personal information and the strength of the relationship with other users.

[0042] (5) In the program (4) above, the relationship level information indicates the relationship level increasing in the order of acquaintance, friend, and family, and the determination process (S50) widens the specified range of personal information the higher the relationship level of the multiple users, and narrows the specified range the lower the relationship level of the multiple users. With this configuration, it becomes possible to appropriately change the scope of use of personal information that is restricted when there are multiple users.

[0043] (6) In the program of (4) above, the importance level information indicates a higher importance level as the confidentiality of the personal information increases, and the determination process (S50) changes the specified range of personal information by changing the threshold of the importance level of the personal information used to determine the output information. With this configuration, it becomes possible to appropriately change the scope of use of personal information that is restricted when there are multiple users.

[0044] (7) In the programs (1) to (6) above, the user is an occupant of the vehicle in which the voice assist system 400 is installed, and the determination process (S40) determines whether the personal information in the memory unit 240 can be used based on whether the number of occupants estimated in the estimation process (S30) is single or multiple. With this configuration, it is possible to prevent a situation in which personal information of one user is known to other users in a vehicle in which passengers other than the driver are riding, and to protect the privacy of users.

[0045] The above embodiment can be modified in various ways, and modifications will be described below. (Variation 1) In the above description, an example has been described in which, when the instruction content from the user is a question, output information in response to the instruction information is output as audio information (or text information) to the user's device 100. In Modification 1, instead of outputting audio information as the output information, or together with audio information as the output information, image information (including video) may be output.

[0046] When the device 100 acquires audio information from the server device 200, the device 100 outputs and plays the audio information through a speaker (or headphones connected via a wired or wireless connection) provided in the device 100. When the device 100 acquires image information from the server device 200, the device 100 outputs and plays the image information through a display provided in the device 100.

[0047] (Variation 2) In the above explanation, an example has been shown in which the server device 200 is configured as a single server device, but the server device 200 may be configured using a virtual server function on a cloud, or may be configured as a distributed system across multiple devices.

[0048] The above description is merely an example, and the present invention is not limited to the above-described embodiment and modifications as long as the features of the present invention are not impaired. One or more of the above-described embodiment and modifications can be arbitrarily combined, and modifications can also be combined with each other. [Explanation of symbols]

[0049] 100, 100a, 100b, 100c device, 200 server device, 210 communication unit, 220 voice recognition unit, 230 assist unit, 231 semantic analysis unit, 232 execution unit, 233 personal information acquisition unit, 234 recording control unit, 240 memory unit, 300 network, 400 voice assist system

Claims

1. A program for an agent device, an input process for receiving instruction information from a user by voice or touch operation and user information relating to the user; a recording process for recording in a storage unit, as information including personal information, at least one of information regarding SNS, SMS, and emails used as a program different from the agent device program, information regarding content subscription or viewing, information on the Internet previously acquired, and air conditioner setting information, which are acquired from a device used by the user; an estimation process for estimating the number of users currently using the service based on the user information; a determination process for determining whether the personal information in the storage unit can be used based on whether the estimated number of users is one or more; a determination process for determining output information for the instruction information by using the information in the storage unit based on the result of the determination; an output process for outputting audio and / or images as the determined output information; A program for an agent device, characterized by causing a computer to execute the above.

2. 2. The program according to claim 1, If the number of users estimated by the estimation process is one, the determination process determines whether the personal information related to the user corresponding to the instruction information received in the input process is usable; the determination process determines the output information in response to the instruction information using information including the personal information stored in the storage unit. A program characterized by:

3. 2. The program according to claim 1, When the number of users estimated by the estimation process is plural, The determination process determines whether or not the personal information can be used or whether or not the personal information can be used within a predetermined range; the determination process determines the output information for the instruction information using the information stored in the storage unit excluding the personal information, or determines the output information using the personal information within the predetermined range stored in the storage unit and the information excluding the personal information. A program characterized by:

4. 4. The program according to claim 3, the user information includes relationship level information indicating the strength of a relationship with another user; the personal information includes importance level information indicating the level of importance for each program different from the agent device program, the determination process changes the predetermined range of the personal information in the storage unit based on the relationship level information and the importance level information; A program characterized by:

5. 5. The program according to claim 4, the relationship level information indicates relationship levels increasing in the order of acquaintance, friend, and family; the determination process widens the predetermined range of the personal information as the relationship level of the plurality of users increases, and narrows the predetermined range as the relationship level of the plurality of users decreases; A program characterized by:

6. 5. The program according to claim 4, The importance level information indicates a higher importance level as the confidentiality of the personal information is higher, the determination process varies the predetermined range of the personal information by changing a threshold value of the importance level of the personal information used to determine the output information; A program characterized by:

7. 7. The program according to claim 1, the user is a passenger in a vehicle in which the agent device is installed, the determination process determines whether the personal information in the storage unit can be used based on whether the number of occupants estimated in the estimation process is one or more. A program characterized by:

Citation Information

Patent Citations

  • Agent device, system, control method for agent device, and program

    JP2020147214A

  • Information processing device and information processing method

    JP2023016213A

  • Speech-based privilege management for voice assistant systems

    JP6949149B2

  • Data disclosure device, data disclosure method, and program

    WO2020115863A1