Automated response devices and programs
The automatic response device addresses the challenge of providing user-specific responses by using characteristic information to select and output tailored responses, enhancing personalization.
Patent Information
- Application Number
- JP2025009593
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2045-01-23
AI Technical Summary
Conventional automatic response technologies struggle to provide responses that are appropriate for individual users due to their reliance on predetermined settings.
An automatic response device that acquires characteristic information from user interactions, selects a response scenario based on this information, and outputs tailored responses using voice or text output methods.
Enables personalized responses by reflecting user characteristics, such as expertise, personality, and online behavior, ensuring appropriate responses are provided.
Smart Images

Figure 0007809228000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to automatic response technology. [Background technology]
[0002] BACKGROUND ART Automatic response technologies that automatically respond to user utterances are known (see, for example, Non-Patent Documents 1 and 2). [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] "ForeSight Voice Mining", [online], NTT Technocross Corporation, [Retrieved October 10, 2024], Internet<https: / / www.ntt-tx.co.jp / products / foresight_vm / > [Non-patent document 2] "CTBASE Call Center Solution", [online], NTT Technocross Corporation, [Retrieved October 10, 2024], Internet<https: / / www.ntt-tx.co.jp / products / ctbase / > Summary of the Invention [Problem to be solved by the invention]
[0004] However, conventional automatic response technologies respond based on predetermined settings, making it difficult to provide responses that are appropriate for individual users.
[0005] In view of this, an automatic response technique is provided that can provide a response suited to each individual user. [Means for solving the problem]
[0006] In the first form, the automatic response device acquires characteristic information including information based on the content of information sent by the user to a third party, selects a response scenario based on the characteristic information, and outputs information representing the response content based on the response scenario in response to the content of the user's speech.
[0007] In the second form, the automatic response device acquires characteristic information of the user from information about the user in accordance with acquisition criteria that can be switched depending on the content of the user's utterance, selects a response scenario based on the characteristic information, and outputs information representing the response content based on the response scenario in response to the content of the utterance.
[0008] In the third form, the automatic answering device acquires characteristic information including information based on the content of information sent by the user to a third party, sets the voice response content to the spoken content in accordance with voice setting standards that can be switched based on the characteristic information of the user, and outputs information representing the voice response content at that voice setting.
[0009] In the fourth form, the automatic answering device acquires the user's characteristic information, sets the voice response content for the user's speech content in accordance with voice setting standards that can be switched depending on the user's speech content, and outputs information representing the voice response content using the voice setting. [Effects of the Invention]
[0010] This makes it possible to provide a response that is appropriate for each individual user. [Brief explanation of the drawings]
[0011] [Figure 1] FIG. 1 is a block diagram illustrating an automatic response system according to an embodiment. [Figure 2] FIG. 2 is a block diagram illustrating the automatic answering device of the embodiment. [Figure 3] FIG. 3 is a diagram illustrating information stored in the storage unit of the embodiment. [Figure 4] FIG. 4 is a flow diagram illustrating the automatic response method of the embodiment. [Figure 5] FIG. 5 is a block diagram illustrating a hardware configuration of the automatic answering device according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0012] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. [First embodiment] In the first embodiment, an automatic answering device acquires characteristic information including information based on the content of information transmitted by a user to a third party, selects a response scenario based on the characteristic information, and outputs information representing a response based on the response scenario in response to the user's utterance. Here, it can be assumed that the user's personality is reflected in the content of information transmitted by a user to a third party. Therefore, by selecting a response scenario for the user's utterance based on the characteristic information including information based on the content of the transmitted information, a response appropriate for each individual user can be provided.
[0013] <Configuration> As illustrated in FIG. 1, the automatic answering system 1 of this embodiment has an automatic answering device 11 and a terminal device 12. The automatic answering device 11 and the terminal device 12 are configured to be able to communicate with each other via a network. A user 1000 speaks using the terminal device 12, and the content of the speech is transmitted to the automatic answering device 11 via the network. The automatic answering device 11 responds to the speech content, and the response content is transmitted to the terminal device 12 via the network. The terminal device 12 outputs the response content, and the user 1000 perceives the response content. The network may be a wide area network such as the Internet, or a local area network such as an intranet or an office network.
[0014] <Automatic Answering Machine 11> As illustrated in FIG. 2, the automatic answering device 11 includes a communication interface 111, a calculation unit 112, and a storage unit 113. The calculation unit 112 includes a control unit 112a, a memory 112b, a scenario setting unit 112c, a user registration unit 112d, a reception unit 112e, an acquisition unit 112f, a scenario selection unit 112g, a voice setting unit 112h, and a response unit 112i. The communication interface 111 has a function of communicating with a network. The communication interface 111 is, for example, an Ethernet interface, a Wi-Fi (registered trademark) interface, an optical fiber interface, a cellular network interface, or the like. The calculation unit 112 is configured by a computer including a processor (hardware processor) such as a central processing unit (CPU) and memories such as random-access memory (RAM) and read-only memory (ROM) executing a predetermined program. This computer may include one processor and memory, or multiple processors and memories. This program may be installed on the computer or may be pre-recorded in a ROM or the like. Furthermore, at least a part of the calculation unit 112 may be configured with an electronic circuit that realizes processing functions without using a program, rather than an electronic circuit that realizes functional configuration by loading a program like a CPU. The electronic circuit that makes up the calculation unit 112 may include multiple CPUs. The storage unit 113 has a function of recording information. The storage unit 113 is, for example, a hard disk, a magneto-optical disc (MO), a semiconductor memory, etc.
[0015] <Terminal Device 12> The terminal device 12 has a communication interface with the function of communicating with a network, a calculation unit with the function of performing information processing, a memory unit with the function of recording information, a user interface for inputting information from the user 1000 and outputting information to the user 1000, etc. Information input from the user 1000 via the user interface is, for example, voice input, character input, selection input, electroencephalogram input, etc. Information output to the user 1000 via the user interface is, for example, voice output, acoustic output, character output, image output, video output, etc. Specific examples of the user interface are a microphone, a mouse, a keyboard, a touch panel, a speaker, a display, an electroencephalograph, etc.
[0016] <Pre-processing> As a pre-processing step, multiple response scenarios are set. A response scenario represents a series of responses that are automatically given in response to a user's utterance. A response scenario may represent one or multiple responses to the utterance, may represent multiple responses that branch depending on the user's utterance, may represent a series of responses when utterances and responses are repeated, may represent knowledge (information) to be given to a generation AI that generates the response, or may represent the settings (e.g., role, personality, character) of the generation AI that generates the response.
[0017] In the pre-processing, multiple response scenarios are set according to the user's characteristics. Even if the user's utterance content is substantially the same (e.g., identical, nearly identical, or similar), the response content based on these multiple response scenarios will differ from one another. This is because even if the user's utterance content is substantially the same, the appropriate response content will differ depending on the user's characteristics. In particular, in this embodiment, multiple response scenarios are set according to the information transmitted by the user to a third party. The information transmitted by the user to a third party often reflects the user's characteristics (e.g., expertise, personality, age, gender, language, nationality, region, occupation, family composition, hobbies, special skills, preferences, interests, etc.). Therefore, even if the utterance content is substantially the same, the appropriate response content will differ depending on the information transmitted by the user to a third party. In other words, in this embodiment, multiple response scenarios are set according to the user's characteristics estimated from the transmitted information. For example, multiple response scenarios may be generated according to the level of expertise, multiple response scenarios according to differences in personality, or multiple response scenarios according to differences in age. The information transmitted to a third party may be information transmitted to an unspecified number of people, information transmitted to a specified number of people, or information transmitted to a specified few people. For example, the information transmitted is content transmitted to a third party online. The information transmitted may be at least one of a post, a writing, a comment, a reply, a question, an answer, an inquiry, a response, an application, or an acceptance, which is transmitted to a third party. The information transmitted may be at least one of a chronological sequence of multiple pieces of information, content added in real time, content changed in real time, or content updated in real time. For example, the information transmitted is information transmitted on a social networking service (SNS), an app, a distribution service site, a Q&A site, an online shopping site, etc.In addition, multiple response scenarios may be generated according to user attributes (e.g., age, gender, nationality, language, region, occupation, family structure, highest level of education, field of expertise, income, means of commuting, disability, listening ability, hobbies, special skills, preferences, interests, religion, frequency of Internet use, type of search engine or browser used, type of point service, electronic money, or credit card used, purchase history, travel frequency, health condition, electronic devices owned, investment details, etc.). Also, multiple response scenarios may be generated according to the user's online access information (e.g., URLs and site names of viewed pages, search words, names of accessed services, access logs, and other browser-related information). In other words, multiple response scenarios may be generated according to the user's attributes and characteristics inferred from the user's online access information.
[0018] A scenario setting unit 112c of the automatic response device 11 (FIG. 2) receives multiple pieces of response scenario information 1133 related to multiple response scenarios, and the scenario setting unit 112c stores the received multiple pieces of response scenario information 1133 in the storage unit 113 (FIG. 3). Each piece of response scenario information 1133 corresponds to one of the response scenarios. For example, the response scenario information 1133 may directly represent a response content based on one of the response scenarios, or may represent information for identifying the response content, or may represent information for generating the response content. For example, the response scenario information 1133 may represent a template for the response content, information indicating the flow or branching of the response content (e.g., a flow chart or a chart), a collection of information for generating the response content (e.g., keywords), a collection of links for generating the response content, knowledge (information), commands, or instructions for causing a generation AI to generate the response content, or settings (e.g., role, personality, character) of the generation AI that generates the response content.
[0019] Here, multiple pieces of response scenario information 1133 are stored in the storage unit 113 so that response scenario information 1133 related to a response scenario corresponding to the user's characteristics can be identified using characteristic information representing the user's characteristics. For example, in the storage unit 113, characteristic information representing the user's characteristics may be associated with response scenario information 1133 related to a response scenario corresponding to the user's characteristics, or the characteristic information representing the user's characteristics may be included in the response scenario information 1133 related to a response scenario corresponding to the user's characteristics. In this embodiment, the characteristic information representing the user's characteristics includes information based on the information transmission content of the user to a third party. The information based on the information transmission content may include, for example, the information transmission content itself, some information (e.g., keywords) included in the information transmission content, or information representing an analysis result of the information transmission content. For example, if the information transmission content includes at least one of a plurality of pieces of information transmission content in chronological order, content added in real time, content changed in real time, or content updated in real time, the analysis result of the information transmission content may be information based on future information transmission content (second information transmission content) estimated from the information transmission content. In the case of such information transmissions, it may be possible to infer future information transmissions from the information transmissions. For example, if the information transmissions involve the user's movement, the region represented by the information transmissions (e.g., place names or latitude and longitude) can be estimated based on the change in the region represented by the information transmissions. For example, if the information transmissions involve daily actions toward a goal, the level of achievement represented by the information transmissions (e.g., score or number of followers) can be estimated based on the change in the level of achievement represented by the information transmissions. Furthermore, for example, the analysis result of the information transmissions may be the type of medium (media) through which the information transmissions were transmitted. This is because, depending on the type of medium, it may be possible to infer the attributes of the user who transmits the information (e.g., age, gender, nationality, language, hobbies, special skills, preferences, interests, etc.). Examples of media include internet media, social media, mass media, and sales promotion media. Examples of media include platforms, websites, SNS, apps, distribution service sites, Q&A sites, and online shopping sites.Furthermore, for example, the content of information posted by a user to a third party may include multiple pieces of content posted through multiple different media, and the analysis result of the content of the information posted may be the ratio of the multiple pieces of content posted through the multiple different media.Furthermore, the analysis result of the content of the information posted may be the user's attributes estimated from the content of the information posted.For example, the user's attributes can be estimated from a theme analysis of the content of the information posted, a language style analysis, an engagement analysis, an analysis of follower characteristics, a time period, a frequency, etc.Furthermore, the characteristic information representing the user's characteristics may include information representing other user attributes or may include the user's online access information.
[0020] <User registration process> Next, the user registration process of this embodiment will be described. A user 1000 of the automatic response system 1 of this embodiment must first register as a user. The user 1000 who registers as a user inputs user information required for a user registration request into the terminal device 12. The user information includes information that represents the characteristics of the user 1000. The user information of this embodiment includes information for identifying the user 1000 (e.g., name, username, account name, telephone number, email address, address, etc.), as well as information regarding information transmission by the user 1000 to a third party. Information transmission to a third party may be information transmission to an unspecified number of people, information transmission to a specified number of people, or information transmission to a specified few people. Information transmission to a third party is, for example, information transmission to a third party online. Examples of information transmission to a third party include posts, writings, comments, replies, questions, answers, inquiries, responses, applications, acceptances, etc. The information regarding the transmission of information by user 1000 to a third party may, for example, represent the medium used by user 1000 (e.g., platform, site, SNS, app, distribution service site, Q&A site, online shopping site, etc.), or may include information for identifying user 1000 in these (e.g., name, username, account name, phone number, email address, address, etc.). In addition, the user information may include information representing the attributes of user 1000, or may include online access information of user 1000.
[0021] The terminal device 12 transmits a user registration request including the input user information to the automatic response device 11 via the network. The user registration request is received by the communication interface 111 of the automatic response device 11 (FIG. 2) and sent to the user registration unit 112d. The user registration unit 112d issues a user ID to the user 1000, associates the user ID of the user 1000 with the user information, and stores the associated information in the memory unit 113. That is, as illustrated in FIG. 3, the user registration unit 112d associates the user ID 1131a of the user 1000 with the user information 1131b, and stores this pair in the memory unit 113 as user registration information 1131. The issued user ID is also sent to the communication interface 111, which then transmits the user ID to the terminal device 12 via the network. The terminal device 12 receives the user ID and stores it in its memory unit (not shown).
[0022] <Characteristics information setting process> Next, the characteristic information setting process of this embodiment will be described. The acquisition unit 112f of the automatic response device 11 (FIG. 2) performs a characteristic information setting process to acquire characteristic information representing the characteristics of the registered user 1000 at a predetermined opportunity and store the information in the storage unit 113. The characteristic information setting process may be performed when the above-mentioned user registration process is performed, or may be performed at a predetermined opportunity after the user registration process. The characteristic information setting process may be performed once, multiple times, or after updating the characteristic information.
[0023] The characteristic information acquired in this embodiment includes at least information based on the content of information transmitted by the user 1000 to a third party. As described above, the information based on the content of information transmitted may include, for example, the content of information transmitted itself, a portion of information included in the content of information transmitted, or information representing the analysis result of the content of information transmitted. For example, the analysis result of the content of information transmitted may be information based on future content of information transmitted (second content of information transmitted) estimated from the content of information transmitted, the type of medium through which the content of information transmitted, the ratio of multiple contents transmitted through multiple different media, or user attributes estimated from the content of information transmitted. For example, the acquisition unit 112f may acquire characteristic information including information based on the content of information transmitted by the user 1000 to a third party using the user information 1131b stored in the storage unit 113. For example, the acquisition unit 112f may use "information for identifying the user 1000" (e.g., account name) included in the user information 1131b to acquire the content of information transmitted by the user 1000 to a third party from the Internet (e.g., post content), and acquire characteristic information based on the content. Also, for example, the acquisition unit 112f may use "information regarding information transmitted by the user 1000 to a third party" (e.g., information indicating a specific SNS and a username) included in the user information 1131b to acquire the content of information transmitted by the user 1000 to a third party from the Internet (e.g., post content), and acquire characteristic information based on the content.
[0024] Additionally, the characteristic information acquired in this embodiment may include user information of the user 1000. For example, the acquiring unit 112f may acquire characteristic information including at least a part of the user information 1131b stored in the storage unit 113 (FIG. 3). For example, the acquiring unit 112f may acquire characteristic information including the type of medium (e.g., SNS) used by the user 1000, or may acquire characteristic information including the attributes of the user 1000 (e.g., age, gender, field of specialty, etc.), or may acquire characteristic information including online access information of the user 1000 (e.g., URLs of viewed pages, site names, etc.).
[0025] The acquiring unit 112f associates the user ID of the user 1000 with the acquired characteristic information and stores the associated information in the storage unit 113. That is, as illustrated in FIG. 3, the acquiring unit 112f associates the user ID 1132a with the characteristic information 1132b and stores the pair of the associated user IDs in the storage unit 113 as characteristic registration information 1132. For example, the user ID 1132a of the user 1000 is the same as the user ID 1131a in the user registration information 1131. When the characteristic information setting process is executed multiple times, the characteristic information 1132b is updated with the newly acquired characteristic information. For example, the newly acquired characteristic information may be added to the characteristic information 1132b, the characteristic information 1132b may be overwritten with the newly acquired characteristic information, or only a portion of the newly acquired characteristic information that was not present in the characteristic information 1132b before the update may be added to the characteristic information 1132b.
[0026] <Automatic response processing> Next, the automatic response process of this embodiment will be described with reference to FIG. When the user 1000 wishes to start the automatic response process, the user inputs information for requesting the start of the automatic response process to the terminal device 12. As a result, the terminal device 12 transmits start request information including the user ID of the user 1000 to the automatic response device 11 via the network. The start request information is received by the communication interface 111 of the automatic response device 11 (FIG. 2) and sent to the acquisition unit 112f.
[0027] Step S112f-1: The acquisition unit 112f, which receives the start request information, acquires characteristic information representing the characteristics of the user 1000. As described above, the characteristic information in this embodiment includes information based on the content of information transmitted by the user 1000 to a third party. The information based on the content of information transmitted may, for example, include the content of information transmitted itself, a portion of information included in the content of information transmitted, or information representing the analysis result of the content of information transmitted. The analysis result of the content of information transmitted may, for example, be information based on future content of information transmitted (second content of information transmitted) estimated from the content of information transmitted, the type of medium by which the content of information transmitted, the ratio of multiple contents transmitted through multiple different media, or the attributes of the user estimated from the content of information transmitted. For example, the acquisition unit 112f acquires the characteristic information of the user 1000 from the characteristic registration information 1132 in the storage unit 113 (FIG. 3). For example, the user ID of the user 1000 is used to search for the user ID 1132a in the characteristic registration information 1132 in the storage unit 113, and the characteristic information 1132b associated with the user ID 1132a that matches the user ID is acquired. The acquisition unit 112f may update the characteristic information 1132b using the characteristic information setting process described above, and then acquire the characteristic information of the user 1000 from the updated characteristic information 1132b, or may acquire the characteristic information of the user 1000 directly using the characteristic information setting process. The acquisition unit 112f may also acquire the characteristic information of the user 1000 based on, for example, certain acquisition criteria. For example, the acquisition unit 112f may acquire only information based on the content of the information transmission as characteristic information, or may acquire information based on the content of the information transmission and user information as characteristic information, or may acquire only a specific part of any of these as characteristic information, based on the certain acquisition criteria. For example, of the characteristic information 1132b, only characteristic information newer than a predetermined date and time may be acquired, or only characteristic information older than a predetermined date and time may be acquired, or only characteristic information in a predetermined field may be acquired. For example, the acquisition criteria in this embodiment may be set in advance, or may be set by the user 1000 using the terminal device 12. The acquired characteristic information is sent to the scenario selection unit 112g.
[0028] Step S112g:The scenario selection unit 112g receives the characteristic information acquired by the acquisition unit 112f. Using the received characteristic information, the scenario selection unit 112g extracts response scenario information 1133 corresponding to the characteristic information from among multiple pieces of response scenario information 1133 stored in the storage unit 113 (FIG. 3). For example, the scenario selection unit 112g may extract response scenario information 1133 associated with the received characteristic information, or may extract response scenario information 1133 including the received characteristic information. As described above, each piece of response scenario information 1133 corresponds to one of the response scenarios. Therefore, the scenario selection unit 112g selects a response scenario based on the received characteristic information. For example, if the characteristic information includes the content of the information to be transmitted, the scenario selection unit 112g selects a response scenario based on the content of the information to be transmitted. For example, if the characteristic information includes part of the information (e.g., keywords) included in the content of the information to be transmitted, the scenario selection unit 112g selects a response scenario based on the part of the information to be transmitted. For example, if the characteristic information includes information representing the analysis result of the information transmission content, the scenario selection unit 112g selects a response scenario based on the information representing the analysis result. For example, if the analysis result of the information transmission content is information based on future information transmission content (second information transmission content) estimated from the information transmission content, the scenario selection unit 112g estimates future information transmission content from the information transmission content and selects a response scenario based on the future information transmission content. If the analysis result of the information transmission content is the type of medium through which the information transmission content was transmitted, the scenario selection unit 112g selects a response scenario based on the type of medium through which the information transmission content was transmitted. If the information transmission content includes multiple contents transmitted through multiple different media and the analysis result of the information transmission content is the ratio of the multiple contents transmitted through the multiple different media, the scenario selection unit 112g selects a response scenario based on the ratio of the multiple contents transmitted through the multiple different media. If the analysis result of the information transmission content is user attributes estimated from the information transmission content, the scenario selection unit 112g selects a response scenario based on the user attributes.The extracted response scenario information 1133 is sent to the response unit 112i.
[0029] Step S112h-1: The voice setting unit 112h sets the voice of the response content to the utterance content of the user 1000 in accordance with certain voice setting criteria. The voice setting criteria in this embodiment may be set in advance or may be set by the user 1000 using the terminal device 12. The content of the voice setting is sent to the response unit 112i.
[0030] Step S112e: The response unit 112i sends information requesting the user 1000 to speak (utterance request information) to the communication interface 111. The communication interface 111 transmits the received utterance request information to the terminal device 12 (FIG. 1) via the network. The terminal device 12, which has received the utterance request information, accepts the utterance from the user 1000. The utterance of the user 1000 may be input by voice, text, selective input, electroencephalogram input, or other methods. The terminal device 12, to which the utterance of the user 1000 has been input, transmits information representing the content of the input utterance (utterance content information) to the automatic response device 11 via the network. The communication interface 111 of the automatic response device 11 (FIG. 2) receives the utterance content information and sends it to the response unit 112i. As a result, the response unit 112i receives the content of the utterance of the user 1000.
[0031] Step S112g: The response unit 112i uses the sent response scenario information 1133 to acquire a response content to the received utterance content. That is, the response unit 112i acquires a response content based on the response scenario corresponding to the response scenario information 1133 in response to the utterance content of the user 1000. For example, if the response scenario information 1133 directly represents the response content, the response unit 112i acquires the response content represented by the response scenario information 1133. For example, the response unit 112i acquires the response content from the template represented by the response scenario information 1133. For example, if the response scenario information 1133 represents information for identifying the response content, the response unit 112i identifies the response content using the response scenario information 1133. For example, the response unit 112i identifies the response content using information, a set of information, or a set of link destinations indicating the flow or branching of the response content represented by the response scenario information 1133. For example, if the response scenario information 1133 represents information for generating a response content, the response unit 112i generates the response content using the response scenario information 1133. For example, the response unit 112i provides the knowledge (information), commands, instructions, and settings represented by the response scenario information 1133 to the generation AI, causes the generation AI to generate a response content, and receives the generated response content. The response unit 112i outputs information (response content information) representing the response content by voice using the transmitted voice settings. That is, the response unit 112i outputs information representing the response content based on the received response scenario in response to the received utterance content. The response content information is sent to the communication interface 111. The communication interface 111 transmits the response content information to the terminal device 12 via the network. The terminal device 12 receives the response content information and outputs the response content represented by the response content information. The terminal device 12 of this embodiment may output the response content by voice, may use a combination of voice output and text output, or may output using another output method. In this way, the user 1000 perceives the response content.
[0032] Step S112a: The control unit 112a of the automatic response device 11 (FIG. 2) determines whether the response satisfies a termination condition. The termination condition may be, for example, the end of a series of response contents represented by the response scenario, the end of a predetermined number of responses since the start of the automatic response process, or the end of communication with the terminal device 12. Here, if the response does not satisfy the termination condition, the process returns to step S112e. On the other hand, if the response satisfies the termination condition, the automatic response process is terminated.
[0033] <Features of this form> In this embodiment, the acquisition unit 112f of the automatic response device 11 acquires characteristic information including information based on the content of information transmitted by the user 1000 to a third party, the scenario selection unit 112g selects a response scenario based on the characteristic information, and the response unit 112i outputs information representing the response content based on the response scenario in response to the utterance content of the user 1000. Here, the transmitted information reflects the characteristics of the user 1000, and the characteristics of the user 1000 are also reflected in the characteristic information. Therefore, by selecting a response scenario based on the characteristic information, it is possible to provide a response content that is appropriate for each individual user 1000.
[0034] [Modification 1 of the First Embodiment] In the first embodiment, before the response unit 112i receives the utterance content (before step S112e), the acquisition unit 112f acquires the characteristic information (step S112f-1), the scenario selection unit 112g selects a response scenario (step S112g-1), and the voice setting unit 112h sets the voice (step S112h-1). However, these processes may be executed after the response unit 112i receives the utterance content. For example, as illustrated in FIG. 4, after step S112e, the acquisition unit 112f acquires the characteristic information (step S112f-2) in the same manner as in step S112f-1, then the scenario selection unit 112g selects a response scenario (step S112g-2) in the same manner as in step S112g-1, and then the voice setting unit 112h sets the voice (step S112h-2) in the same manner as in step S112h-1, and then the process of step S112g is executed. Alternatively, the processes of steps S112f-1, S112g-1, and S112h-1 may not be performed before step S112e, and the processes of steps S112f-2, S112g-2, and S112h-2 may be performed after step S112e, and then the process of step S112g may be performed.
[0035] [Modification 2 of the First Embodiment] In the first embodiment and its first modification, an example has been shown in which the terminal device 12 outputs the response content by voice. However, the terminal device 12 may output the response content in text or the like instead of outputting the response content by voice. In this case, the voice setting by the voice setting unit 112h (steps S112h-1 and S112h-2) can be omitted. Furthermore, the response unit 112i outputs information representing the response content in text or the like (step S112g).
[0036] [Second embodiment] In the second embodiment, an automatic answering device acquires user characteristic information from information about the user in accordance with acquisition criteria that can be switched depending on the user's utterance content, selects a response scenario based on the characteristic information, and outputs information representing the response content based on the response scenario in response to the utterance content. Here, the utterance content of the user 1000 reflects the characteristics of the user 1000. Therefore, by switching the acquisition criteria for characteristic information depending on the user's utterance content, it is possible to provide a response that is appropriate for each user. Furthermore, by switching the acquisition criteria for characteristic information depending on the user's utterance content, it is possible to provide a response that is appropriate for each user based not only on the static characteristics of each user but also on dynamic characteristics based on the utterance content. The following will mainly explain the differences from the matters described so far, and the same reference numbers will be used to simplify the explanation of matters already described.
[0037] <Configuration> 1, the automatic answering system 2 of this embodiment has an automatic answering device 21 and a terminal device 12. That is, the automatic answering system 2 is obtained by replacing the automatic answering device 11 of the first embodiment with the automatic answering device 21.
[0038] <Automatic Answering Machine 21> 2, the automatic answering device 21 has a communication interface 111, a calculation unit 212, and a storage unit 113. The calculation unit 212 has a control unit 112a, a memory 112b, a scenario setting unit 112c, a user registration unit 112d, a reception unit 112e, an acquisition unit 212f, a scenario selection unit 112g, a voice setting unit 112h, and a response unit 112i. The calculation unit 212 is configured, for example, by a computer including a processor, memory, etc., executing a predetermined program. The electronic circuit that configures the calculation unit 212 may include multiple CPUs.
[0039] <Pre-processing> The pre-processing in this embodiment is the same as that in the first embodiment. However, the multiple response scenarios set in this embodiment may or may not correspond to the content of information transmitted by the user to a third party, as long as they correspond to the characteristics of the user.
[0040] <User registration process> The user registration process in this embodiment is the same as that in the first embodiment.
[0041] <Characteristics information setting process> The characteristic information setting process of this embodiment is the same as that of the first embodiment. However, the characteristic information of this embodiment may or may not include information based on the content of information transmitted by the user 1000 to third parties. Furthermore, the characteristic information of this embodiment may include information representing the attributes of the user, or may include online access information of the user.
[0042] <Automatic response processing> Next, the automatic response process of this embodiment will be described with reference to FIG. First, instead of the acquisition unit 112f, the acquisition unit 212f of the automatic answering device 21 (FIG. 2) executes the process of step S112f-1 described in the first embodiment, the scenario selection unit 112g executes the process of step S112g-1 described in the first embodiment, and the voice setting unit 112h executes the process of step S112h-1 described in the first embodiment. However, the characteristic information of this embodiment may or may not include information based on the content of information transmitted by the user 1000 to third parties. Furthermore, the characteristic information of this embodiment may include information representing the user's attributes or may include the user's online access information. Thereafter, the response unit 112i executes the process of step S112e described in the first embodiment and receives the content of the utterance from the user 1000. The content of the utterance is sent to the acquisition unit 212f. Thereafter, the process of step S212f-2 described below is executed.
[0043] Step S212f-2: The acquisition unit 212f acquires user characteristic information from information about the user 1000 in accordance with acquisition criteria that can be switched depending on the content of the user's utterance received. For example, the acquisition criteria represent at least one of the type of characteristic information, the content range of the characteristic information, the time range of the characteristic information, the regional range of the characteristic information, the language of the characteristic information, or the amount of the characteristic information, and the acquisition unit 212f acquires characteristic information that conforms to the acquisition criteria that can be switched depending on the content of the utterance. For example, depending on the content of the utterance, the type of characteristic information to be acquired may be switched, the content range of the characteristic information to be acquired may be switched, the regional range of the characteristic information to be acquired, the language of the characteristic information to be acquired, or the amount of the characteristic information to be acquired may be switched. Here, the type of characteristic information may be, for example, information based on the content of the information transmission, information representing the user's attributes, access information, or any of these further subdivided types. The same applies to the content range of the characteristic information. The temporal range of the characteristic information may be, for example, the range of the time when the information content was transmitted, the range of the time when the user's attributes were reached, or the range of the time of the access information. The regional range of the characteristic information may be, for example, the range of the region where the information content was transmitted or the region included in the information content, the range of the region represented by the user's attributes, or the range of the region where the server on which the site represented by the access information is implemented is located. The acquisition criteria may be switched depending on, for example, at least one of the type of speech content, the field of the speech content, the time represented by the speech content, the region represented by the speech content, the language of the speech content, the importance of the speech content, the urgency of the speech content, the difficulty of the speech content, or the expertise of the speech content. For example, the type of characteristic information to be acquired may be switched depending on the field of the speech content. For example, the acquisition criteria may be associated with an utterance content label representing the speech content. For example, the acquisition criteria may be associated with a speech content label that represents at least one of the type of speech content, the field of the speech content, the time period represented by the speech content, the region represented by the speech content, the language of the speech content, the importance of the speech content, the urgency of the speech content, the difficulty of the speech content, or the specialization of the speech content.As a result, the acquisition unit 212f can estimate an utterance content label from the utterance content, switch to an acquisition criterion corresponding to the estimated utterance content label (an acquisition criterion according to the utterance content), and acquire characteristic information conforming to the acquisition criterion from information about the user 1000. Note that the estimation of the utterance content label may be performed, for example, using a table in which information included in the utterance content is associated with the utterance content label, or may be performed using a machine learning model that estimates the utterance content label from the utterance content. Furthermore, the information about the user 1000 may be, for example, characteristic information acquired in the characteristic information setting process, or other characteristic information of the user 1000. The acquisition unit 212f sends the acquired characteristic information to the scenario selection unit 112g.
[0044] Thereafter, the scenario selection unit 112g selects a response scenario based on the characteristic information sent from the acquisition unit 212f (step S112g-2), the voice setting unit 112h performs voice setting (step S112h-2), and then the response unit 112i outputs information representing the response content based on the response scenario in response to the utterance content (step S112g). The other processing is the same as that of the first embodiment or the second modification of the first embodiment, except that the automatic response unit 11 is replaced with the automatic response unit 21.
[0045] <Features of this form> In this embodiment, the acquisition unit 212f of the automatic response device 21 acquires user characteristic information from information about the user 1000 in accordance with acquisition criteria that can be switched depending on the content of the user's utterance. The scenario selection unit 112g selects a response scenario based on the characteristic information and outputs information representing a response based on the response scenario in response to the utterance. The user's utterance content reflects the user's 1000 characteristics. Therefore, by selecting a response scenario based on the characteristic information acquired in accordance with acquisition criteria that can be switched depending on the content of the user's utterance, a response tailored to each individual user 1000 can be provided. Furthermore, by switching the acquisition criteria for characteristic information depending on the user's utterance, a response tailored to each individual user can be provided based not only on the static characteristics of each individual user but also on dynamic characteristics based on the utterance content. Furthermore, as shown in the first embodiment, if the characteristic information includes the content of information transmitted by the user to a third party, a response tailored to each individual user 1000 can be provided.
[0046] [Third embodiment] In the third embodiment, an automatic answering device acquires characteristic information including information based on the content of information transmitted by a user to a third party, sets a voice response to the user's speech in accordance with a voice setting standard that can be switched based on the characteristic information, and outputs information representing the voice response content in the voice setting. This enables a voice response according to the user's characteristics, and a response that is appropriate for each individual user.
[0047] <Configuration> 1, the automatic answering system 3 of this embodiment has an automatic answering device 31 and a terminal device 12. That is, the automatic answering system 3 is obtained by replacing the automatic answering device 11 of the first embodiment with the automatic answering device 31.
[0048] <Automatic Answering Machine 21> 2, the automatic answering device 31 has a communication interface 111, a calculation unit 312, and a storage unit 113. The calculation unit 312 has a control unit 112a, a memory 112b, a scenario setting unit 112c, a user registration unit 112d, a reception unit 112e, an acquisition unit 112f, a scenario selection unit 112g, a voice setting unit 312h, and a response unit 112i. The calculation unit 312 is configured, for example, by a computer including a processor, memory, etc., executing a predetermined program. The electronic circuit that configures the calculation unit 312 may include multiple CPUs.
[0049] <Pre-processing> The pre-processing in this embodiment is the same as that in the first embodiment.
[0050] <User registration process> The user registration process in this embodiment is the same as that in the first embodiment.
[0051] <Characteristics information setting process> The characteristic information setting process of this embodiment is the same as that of the first embodiment.
[0052] <Automatic response processing> Next, the automatic response process of this embodiment will be described with reference to FIG. First, instead of the acquisition unit 112f, the acquisition unit 112f of the automatic answering device 31 (FIG. 2) executes the process of step S112f-1 described in the first embodiment, and the scenario selection unit 112g executes the process of step S112g-1 described in the first embodiment. However, the characteristic information acquired by the acquisition unit 112f (step S112f-1) is also sent to the voice setting unit 312h. Thereafter, the process of step S312h-1 described below is executed.
[0053] Step S312h-1: The voice setting unit 312h sets the voice of the response content to the user's utterance content in accordance with the voice setting criteria that are switched based on the characteristic information sent from the acquisition unit 112f. For example, the voice setting criteria represent at least one of the gender of the voice, the expected age of the voice, the speaking speed of the voice, the pitch of the voice, the accent of the voice, the clarity of the voice, the volume of the voice, the intonation of the voice, the quality of the voice, the fluency of the voice, the continuity of the voice, the presence or absence of hesitation in the voice, the rhythm of the voice, the language of the voice, the dialect of the voice, the reaction speed of the voice, the level of fatigue of the voice, and voice changes, and the voice setting unit 312h sets the voice in accordance with the voice setting criteria that are switched based on the characteristic information. For example, depending on the characteristic information, the gender of the voice may be changed, the assumed age of the voice may be changed, the speaking speed of the voice may be changed, the pitch of the voice may be changed, the accent of the voice may be changed, the clarity of the voice may be changed, the volume of the voice may be changed, the intonation of the voice may be changed, the quality of the voice may be changed, the fluency of the voice may be changed, the continuity of the voice may be changed, the presence or absence of hesitation in the voice may be changed, the rhythm of the voice may be changed, the language of the voice may be changed, the dialect of the voice may be changed, the reaction speed of the voice may be changed, the fatigue level of the voice may be changed, or voice changes may be changed. For example, the speaking speed of the voice may be changed depending on the listening ability of the user 1000. For example, the voice setting criteria may be associated with a characteristic information label representing the characteristic information. The characteristic information of this embodiment is the same as that of the first embodiment. As a result, the voice setting unit 312h can estimate the attribute information label from the attribute information, switch to a voice setting standard corresponding to the estimated attribute information label, and set the voice of the response content to the user's utterance content in accordance with the voice setting standard. Note that the attribute information label may be estimated using, for example, a table in which information included in the attribute information is associated with the attribute information label, or may be estimated using a machine learning model that estimates the attribute information label from the attribute information. The voice setting content is sent to the response unit 112i.
[0054] Thereafter, the response unit 112i executes the process of step S112e described in the first embodiment and receives the content of the utterance from the user 1000. Next, the response unit 112i executes the process of step S112g. That is, in step S112g, information (response content information) representing the content of the response voiced in the sent voice settings is output. The other processes are the same as those in the first embodiment, except that the automatic response unit 11 is replaced with the automatic response unit 31.
[0055] <Features of this form> In this embodiment, the acquisition unit 112f of the automatic answering device 31 acquires characteristic information including information based on the content of information transmitted by the user 1000 to a third party, the voice setting unit 312h sets the voice of the response content to the utterance content of the user 1000 in accordance with the voice setting standard that is switched based on the characteristic information, and the response unit 112i outputs information representing the voice response content in the voice setting. This makes it possible to respond by voice according to the characteristics of the user, and to provide a response that is appropriate for each individual user.
[0056] [Modification 1 of the third embodiment] In the third embodiment, the audio setting unit 312h performed audio setting (step S312h-1) before the response unit 112i received the speech content (before step S112e). However, this processing may be executed after the response unit 112i received the speech content. For example, as illustrated in FIG. 4, after step S112e, the audio setting unit 312h may perform audio setting (step S312h-2) in the same manner as in step S312h-1, and then the processing of step S112g may be executed. Alternatively, before step S112e, the processing of step S312h-1 may not be executed, and the processing of step S312h-2 may be executed after step S112e, and then the processing of step S112g may be executed.
[0057] [Modification 2 of the third embodiment] In the third embodiment, before the response unit 112i receives the utterance content (before step S112e), the acquisition unit 112f acquires the characteristic information (step S112f-1), and the scenario selection unit 112g selects a response scenario (step S112g-1). However, these processes may be performed after the response unit 112i receives the utterance content. For example, as illustrated in FIG. 4, after step S112e, the acquisition unit 112f may acquire the characteristic information (step S112f-2) in the same manner as in step S112f-1, and then the scenario selection unit 112g may select a response scenario (step S112g-2) in the same manner as in step S112g-1. Alternatively, the processes of steps S112f-2 and S112g-2 may be performed after step S112e without performing the processes of steps S112f-1 and S112g-1 before step S112e.
[0058] [Fourth embodiment] In the fourth embodiment, an automatic answering device acquires user characteristic information, sets the voice of a response to the user's utterance in accordance with a voice setting standard that can be switched depending on the utterance content, and outputs information representing the voice response content under the voice setting. Here, the utterance content of the user 1000 reflects the characteristics of the user 1000. Therefore, by switching the voice setting standard depending on the user's utterance content, it is possible to provide a response that is appropriate for each individual user. Furthermore, by switching the voice setting standard depending on the user's utterance content, it is possible to provide a response that is appropriate for each individual user based not only on the static characteristics of each individual user but also on dynamic characteristics based on the utterance content.
[0059] <Configuration> 1, the automatic answering system 4 of this embodiment has an automatic answering device 41 and a terminal device 12. That is, the automatic answering system 4 is obtained by replacing the automatic answering device 11 of the first embodiment with the automatic answering device 41.
[0060] <Automatic Answering Machine 41> 2, the automatic answering device 41 has a communication interface 111, a calculation unit 412, and a storage unit 113. The calculation unit 412 has a control unit 112a, a memory 112b, a scenario setting unit 112c, a user registration unit 112d, a reception unit 112e, an acquisition unit 112f, a scenario selection unit 112g, a voice setting unit 412h, and a response unit 112i. The calculation unit 412 is configured, for example, by a computer including a processor, memory, etc., executing a predetermined program. The electronic circuit that configures the calculation unit 412 may include multiple CPUs.
[0061] <Pre-processing> The pre-processing in this embodiment is the same as that in the first embodiment. However, the multiple response scenarios set in this embodiment may or may not correspond to the content of information transmitted by the user to a third party, as long as they correspond to the characteristics of the user.
[0062] <User registration process> The user registration process in this embodiment is the same as that in the first embodiment.
[0063] <Characteristics information setting process> The characteristic information setting process of this embodiment is the same as that of the first embodiment. However, the characteristic information of this embodiment may or may not include information based on the content of information transmitted by the user 1000 to third parties. Furthermore, the characteristic information of this embodiment may include information representing the attributes of the user, or may include online access information of the user.
[0064] <Automatic response processing> Next, the automatic response process of this embodiment will be described with reference to FIG. First, the acquisition unit 112f of the automatic answering device 41 (FIG. 2) executes the process of step S112f-1 described in the first embodiment, the scenario selection unit 112g executes the process of step S112g-1 described in the first embodiment, and the voice setting unit 112h executes the process of step S112h-1 described in the first embodiment. However, the characteristic information of this embodiment may or may not include information based on the content of information transmitted by the user 1000 to third parties. Furthermore, the characteristic information of this embodiment may include information representing the user's attributes or may include the user's online access information. Thereafter, the response unit 112i executes the process of step S112e described in the first embodiment and receives the content of the user's utterance. The content of the utterance is sent to the voice setting unit 412h. Thereafter, the following process of step S412h-2 is executed.
[0065] Step S412h-2: The voice setting unit 412h sets the voice of the response content to the user's utterance content in accordance with the voice setting criteria that can be switched depending on the utterance content sent from the response unit 112i. For example, the voice setting criteria represent at least one of the gender of the voice, the expected age of the voice, the speaking speed of the voice, the pitch of the voice, the accent of the voice, the clarity of the voice, the volume of the voice, the intonation of the voice, the quality of the voice, the fluency of the voice, the continuity of the voice, the presence or absence of hesitation in the voice, the rhythm of the voice, the language of the voice, the dialect of the voice, the reaction speed of the voice, the level of fatigue of the voice, and voice changes, and the voice setting unit 412h sets the voice in accordance with the voice setting criteria that can be switched depending on the utterance content. For example, depending on the content of the speech, the gender of the voice may be switched, the assumed age of the voice may be switched, the speaking speed of the voice may be switched, the pitch of the voice may be switched, the accent of the voice may be switched, the clarity of the voice may be switched, the volume of the voice may be switched, the intonation of the voice may be switched, the quality of the voice may be switched, the fluency of the voice may be switched, the continuity of the voice may be switched, the presence or absence of hesitation in the voice may be switched, the rhythm of the voice may be switched, the language of the voice may be switched, the dialect of the voice may be switched, the reaction speed of the voice may be switched, the fatigue level of the voice may be switched, or changes in the voice may be switched. The voice setting criteria may be switched depending on, for example, at least one of the type of speech content, the field of the speech content, the time period represented by the speech content, the region represented by the speech content, the language of the speech content, the importance of the speech content, the urgency of the speech content, the difficulty of the speech content, or the expertise of the speech content. For example, the speech rate may be switched depending on the type of speech content. For example, the voice setting criteria may be associated with a speech content label representing the speech content. For example, the voice setting criteria may be associated with a speech content label representing at least one of the type of speech content, the field of the speech content, the time period represented by the speech content, the region represented by the speech content, the language of the speech content, the importance of the speech content, the urgency of the speech content, the difficulty of the speech content, or the expertise of the speech content.As a result, the voice setting unit 412h can estimate the speech content label from the speech content, switch to a voice setting standard (voice setting standard according to the speech content) corresponding to the estimated speech content label, and set the voice of the response content to the speech content of the user 1000. For example, the speech content label may represent a type of speech content according to the level of understanding of the user 1000, and when the speech content label (index) representing the level of understanding of the user 1000 estimated from the speech content is a first value, the speech rate of the voice that is the voice setting standard may be set to a first rate, and when the speech content label (index) is a second value, the speech rate of the voice that is the voice setting standard may be set to a second rate that is slower than the first rate. However, the level of understanding of the user 1000 represented by the speech content label (index) of the second value is lower than the level of understanding of the user 1000 represented by the speech content label (index) of the first value. In this case, if the user 1000 has a low level of understanding, the response will be made in a voice with a slow speaking speed, and if the user 1000 has a high level of understanding, the response will be made in a voice with a fast speaking speed. Note that the estimation of the utterance content label may be performed, for example, using a table in which information included in the utterance content is associated with the utterance content label, or may be performed using a machine learning model that estimates the utterance content label from the utterance content. The content of the voice setting is sent to the response unit 112i.
[0066] Thereafter, the response unit 112i executes the process of step S112g. That is, in step S112g, information (response content information) representing the response content by voice in the sent voice setting is output. The other processes are the same as those in the first embodiment, except that the automatic answering device 11 is replaced with the automatic answering device 41.
[0067] <Features of this form> In this embodiment, the acquisition unit 112f of the automatic answering device 41 acquires characteristic information of the user 1000, the voice setting unit 412h sets the voice of the response content to the utterance content of the user 1000 in accordance with the voice setting standard that can be switched depending on the utterance content, and the response unit 112i outputs information representing the voice response content in the voice setting. This makes it possible to respond by voice according to the utterance content of the user, and to provide a response that is appropriate for each user. Furthermore, by switching the voice setting standard depending on the utterance content of the user, it is possible to provide a response that is appropriate for each user based not only on the static characteristics of each user but also on the dynamic characteristics based on the utterance content.
[0068] [Modification 1 of the Fourth Embodiment] In the fourth embodiment, the voice setting unit 412h sets the voice of the response to the user's utterance content in accordance with the voice setting standard, which is switched depending on the utterance content sent from the response unit 112i. However, characteristic information of the user 1000 may also be sent from the acquisition unit 112f to the voice setting unit 412h, and the voice setting unit 412h may set the voice in accordance with the voice setting standard, which is switched depending on the combination of the sent utterance content and characteristic information. The characteristic information may include the content of information transmitted by the user to a third party. For example, the voice setting standard may be associated with a combination of an utterance content label representing the utterance content and a characteristic information label representing the characteristic information. The voice setting unit 412h may estimate the utterance content label from the utterance content, estimate the characteristic information label from the characteristic information, switch to the voice setting standard corresponding to the combination of the estimated characteristic information label and utterance content label, and set the voice of the response to the user's utterance content in accordance with the voice setting standard.
[0069] [Modification 2 of the Fourth Embodiment] In the fourth embodiment, before the response unit 112i receives the utterance content (before step S112e), the acquisition unit 112f acquires the characteristic information (step S112f-1), and the scenario selection unit 112g selects a response scenario (step S112g-1). However, these processes may be performed after the response unit 112i receives the utterance content. For example, as illustrated in FIG. 4, after step S112e, the acquisition unit 112f may acquire the characteristic information (step S112f-2) in the same manner as in step S112f-1, and then the scenario selection unit 112g may select a response scenario (step S112g-2) in the same manner as in step S112g-1. Alternatively, the processes of steps S112f-2 and S112g-2 may be performed after step S112e without performing the processes of steps S112f-1 and S112g-1 before step S112e. Alternatively, the processes of steps S112f-2 and S112g-2 may be executed after step S112e without executing the processes of steps S112f-1 and S112g-1 before step S112e. Alternatively, the response unit 112i may send the utterance content, and the acquisition unit 112f may send the characteristic information of the user 1000 to the voice setting unit 412h, and the voice setting unit 412h may set the voice in accordance with a voice setting standard that is switched depending on the combination of the utterance content and characteristic information sent. The characteristic information may include the content of information to be transmitted by the user to a third party.
[0070] [Modification 3 of the Fourth Embodiment] In this embodiment, as described in the second embodiment, characteristic information of user 1000 may be acquired from information about user 1000 in accordance with acquisition criteria that can be switched depending on the content of the utterance of user 1000. That is, the processing of step S112f-2 in variant example 2 of the fourth embodiment may be replaced with the processing of step S212f-2 in the second embodiment.
[0071] [Hardware configuration] The functions performed by the components described herein may be implemented in circuitry or processing circuitry, including general-purpose processors, application-specific processors, integrated circuits, ASICs (Application Specific Integrated Circuits), a CPU (a Central Processing Unit), conventional circuits, and / or combinations thereof, programmed to perform the described functions. A processor includes transistors and other circuits and is considered to be circuitry or processing circuitry. A processor may also be a programmed processor that executes programs stored in memory.
[0072] In this specification, a circuitry, unit, or means is hardware that is programmed to realize or performs the described functions, which may be any hardware disclosed herein or any hardware known to be programmed to realize or perform the described functions.
[0073] If the hardware is a processor considered to be a type of circuitry, the circuitry, means, or unit is a combination of the hardware and software used to configure the hardware and / or processor.
[0074] For example, the automatic answering devices 11, 21, 31, and 41 in each embodiment are devices configured by a general-purpose or dedicated computer having a processor (hardware processor) such as a central processing unit (CPU) and memories such as random-access memory (RAM) and read-only memory (ROM) executing a predetermined program. That is, the automatic answering devices 11, 21, 31, and 41 in each embodiment have, for example, processing circuitry configured to implement each of the units possessed by the automatic answering devices. This computer may have one processor and memory, or multiple processors and memories. This program may be installed on the computer or may be pre-recorded in a ROM or the like. Furthermore, some or all of the processing units may be configured using electronic circuits that independently realize processing functions, rather than electronic circuits that realize functional configurations by loading programs, such as a CPU. Furthermore, the electronic circuits constituting one device may include multiple CPUs.
[0075] FIG. 5 is a block diagram illustrating the hardware configuration of the automatic answering device 11, 21, 31, or 41 in each embodiment. As illustrated in FIG. 5, the automatic answering device 11, 21, 31, or 41 in this example includes a central processing unit (CPU) 10a, an input unit 10b, an output unit 10c, a random access memory (RAM) 10d, a read-only memory (ROM) 10e, an auxiliary storage device 10f, a communication unit 10h, and a bus 10g. The CPU 10a in this example includes a control unit 10aa, a calculation unit 10ab, and a register 10ac, and executes various calculation processes according to various programs loaded into the register 10ac. The input unit 10b is an input terminal, keyboard, mouse, touch panel, or the like, through which data is input. The output unit 10c is an output terminal, speaker, display, or the like, through which data is output. The communication unit 10h is a LAN card or the like, controlled by the CPU 10a that has loaded a predetermined program. The RAM 10d is a static random access memory (SRAM), a dynamic random access memory (DRAM), or the like, and has a program area 10da where a predetermined program is stored and a data area 10db where various data are stored. The auxiliary storage device 10f is a hard disk, a magneto-optical disc (MO), a semiconductor memory, or the like, and has a program area 10fa where a predetermined program is stored and a data area 10fb where various data are stored. The bus 10g connects the CPU 10a, the input unit 10b, the output unit 10c, the RAM 10d, the ROM 10e, the communication unit 10h, and the auxiliary storage device 10f so that information can be exchanged. The CPU 10a writes the program stored in the program area 10fa of the auxiliary storage device 10f to the program area 10da of the RAM 10d in accordance with the loaded OS (Operating System) program. Similarly, the CPU 10a writes various data stored in the data area 10fb of the auxiliary storage device 10f to the data area 10db of the RAM 10d.The addresses in RAM 10d where the programs and data are written are stored in register 10ac of CPU 10a. Control unit 10aa of CPU 10a sequentially reads out these addresses stored in register 10ac, reads out the programs and data from the areas in RAM 10d indicated by the read addresses, causes calculation unit 10ab to sequentially execute the calculations indicated by the programs, and stores the calculation results in register 10ac. This configuration realizes the functional configuration of automatic answering devices 11, 21, 31, and 41.
[0076] The program describing this processing can be recorded on a computer-readable recording medium. Examples of computer-readable recording media are non-transitory recording media. Examples of such recording media include magnetic recording devices, optical disks, magneto-optical recording media, and semiconductor memories.
[0077] The program may be distributed, for example, by selling, transferring, lending, etc. a portable recording medium such as a DVD or CD-ROM on which the program is recorded. Furthermore, the program may be stored in a storage device of a server computer, and then transferred from the server computer to another computer via a network, thereby distributing the program.
[0078] A computer that executes such a program may first temporarily store the program recorded on a portable recording medium or transferred from a server computer in its own storage device. Then, when executing a process, the computer reads the program stored on its own recording medium and executes the process in accordance with the read program. Alternatively, the computer may read the program directly from a portable recording medium and execute the process in accordance with the program. Furthermore, the computer may execute the process in accordance with the program each time a program is transferred from a server computer to the computer. The server computer may not transfer the program to the computer, but may instead execute the process through a so-called ASP (Application Service Provider) service, which realizes the processing function by issuing an execution instruction and obtaining the results. Furthermore, the server computer may execute the process on a terminal using a so-called SaaS (Software as a Service) service, which allows users to use part of the server computer along with the program. In this embodiment, the program includes information used for computer processing that is equivalent to a program (such as data that is not a direct instruction to the computer but has properties that define computer processing).
[0079] Furthermore, in this embodiment, the device is configured by executing a predetermined program on a computer, but at least a part of the processing contents may be realized by hardware.
[0080] [Other variations] The present invention is not limited to the above-described embodiment. For example, the characteristic information setting process may be omitted, and the characteristic information may be directly acquired in steps S112f-1, 112f-2, and 212f-2 of the automatic response process.
[0081] In the third embodiment, the fourth embodiment, and their variations, instead of selecting a response scenario based on characteristic information, the response scenario may be selected by other means, or the response scenario may be fixed.
[0082] In addition, the above-described embodiments may be combined. Furthermore, the above-described various processes may not only be executed in chronological order as described, but may also be executed in parallel or individually depending on the processing capacity of the device executing the processes or as needed. Needless to say, appropriate modifications are possible within the scope of the present invention.
[0083] [Note] The above contents can be summarized as follows: [Appendix 1] an acquisition unit that acquires characteristic information including information based on the content of information transmitted by the user to a third party; a scenario selection unit that selects a response scenario based on the characteristic information; a response unit that outputs information representing a response based on the response scenario in response to the user's utterance; An automatic answering device having: [Appendix 1.1] 1. The automated answering device of claim 1, The information transmission content includes content transmitted to the third party online. [Appendix 1.2] 1. The automated answering device of claim 1, The information transmission content includes at least one of the following to the third party: posted content, written content, comment content, reply content, question content, answer content, inquiry content, response content, application content, or acceptance content. [Appendix 1.3] 1. The automated answering device of claim 1, The information transmission content includes at least one of a plurality of pieces of information transmitted in time series, content added in real time, content changed in real time, and content updated in real time, according to the automatic answering device. [Appendix 1.4] 1. An automated answering device according to claim 1.3, The scenario selection unit estimates future second information transmission content from the information transmission content, and selects the response scenario based on the second information transmission content. [Appendix 1.5] 1. The automated answering device of claim 1, The scenario selection unit selects the response scenario based on the type of medium through which the information content is transmitted. [Appendix 1.6] 1. The automated answering device of claim 1, The information transmission content includes content transmitted through multiple media, The scenario selection unit selects the response scenario based on a combination of the media through which the information content is transmitted. [Appendix 1.7] 1. The automated answering device of claim 1, The information transmission content includes a plurality of contents transmitted through a plurality of different media, The scenario selection unit selects the response scenario based on a ratio of the plurality of contents transmitted through the plurality of different media.
[0084] [Appendix 2] an acquisition unit that acquires characteristic information of the user from information about the user in accordance with acquisition criteria that can be switched depending on the content of the user's utterance; a scenario selection unit that selects a response scenario based on the characteristic information; a response unit that outputs information representing a response content based on the response scenario in response to the utterance content; An automatic answering device having: [Appendix 2.1] 1. The automated answering device of claim 2, The acquisition criteria are switched according to at least one of the type of speech content, the field of the speech content, the time period represented by the speech content, the region represented by the speech content, the language of the speech content, the importance of the speech content, the urgency of the speech content, the difficulty of the speech content, or the specialization of the speech content. [Appendix 2.2] 1. The automated answering device of claim 2, the acquisition criteria represent at least one of a type of the characteristic information, a content range of the characteristic information, a time range of the characteristic information, a geographical range of the characteristic information, a language of the characteristic information, or an amount of the characteristic information; The acquisition unit acquires the characteristic information that meets the acquisition criteria. [Appendix 2.3] 1. The automated answering device of claim 2, The characteristic information includes the content of information transmitted by the user to a third party.
[0085] [Appendix 3] an acquisition unit that acquires characteristic information including information based on the content of information transmitted by the user to a third party; a voice setting unit that sets a voice for a response to a speech content of the user in accordance with a voice setting standard that is switched based on the characteristic information; a response unit that outputs information representing the content of the response by voice in the voice setting; An automatic answering device having: [Appendix 3.1] 3. The automated answering device of claim 3, The information transmission content includes content transmitted to the third party online. [Appendix 3.2] 3. The automated answering device of claim 3, The information transmission content includes at least one of the following to the third party: posted content, written content, comment content, reply content, question content, answer content, inquiry content, response content, application content, or acceptance content. [Appendix 3.3] 3. The automated answering device of claim 3, The information transmission content includes at least one of a plurality of pieces of information transmitted in time series, content added in real time, content changed in real time, and content updated in real time, according to the automatic answering device. [Appendix 3.4] 3. The automated answering device of claim 3, the voice setting criteria represent at least any of the gender of the voice, the assumed age of the voice, the speaking rate of the voice, the pitch of the voice, the accent of the voice, the clarity of the voice, the volume of the voice, the intonation of the voice, the quality of the voice, the fluency of the voice, the continuity of the voice, the presence or absence of hesitation in the voice, the rhythm of the voice, the language of the voice, the dialect of the voice, the reaction speed of the voice, the fatigue level of the voice, or changes in the voice; The voice setting unit performs the voice setting in accordance with the voice setting criteria. [Appendix 3.5] 3. The automated answering device of claim 3, a scenario selection unit that selects a response scenario based on the characteristic information; The response unit outputs information representing the response content based on the response scenario using voice in the voice settings.
[0086] [Appendix 4] an acquisition unit for acquiring user characteristic information; a voice setting unit that sets a voice of a response content to the user's utterance content in accordance with a voice setting standard that is switched depending on the utterance content; a response unit that outputs information representing the content of the response by voice in the voice setting; An automatic answering device having: [Appendix 4.1] 5. The automated answering device of claim 4, An automatic answering device in which the voice setting criteria are switched according to at least one of the type of speech content, the field of the speech content, the time period represented by the speech content, the region represented by the speech content, the language of the speech content, the importance of the speech content, the urgency of the speech content, the difficulty of the speech content, or the expertise of the speech content. [Appendix 4.2] 5. The automated answering device of claim 4, the voice setting criteria represent at least any of the gender of the voice, the assumed age of the voice, the speaking rate of the voice, the pitch of the voice, the accent of the voice, the clarity of the voice, the volume of the voice, the intonation of the voice, the quality of the voice, the fluency of the voice, the continuity of the voice, the presence or absence of hesitation in the voice, the rhythm of the voice, the language of the voice, the dialect of the voice, the reaction speed of the voice, the fatigue level of the voice, or changes in the voice; The voice setting unit performs the voice setting in accordance with the voice setting criteria. [Appendix 4.3] 5. The automated answering device of claim 4, The voice setting standard sets the speech rate of the voice to a first rate when the utterance content is estimated to be a first value for an index representing the user's understanding level, and sets the speech rate of the voice to a second rate slower than the first rate when the utterance content is estimated to be a second value for an index representing the user's understanding level, The automatic answering device, wherein the level of understanding of the user represented by the index of the second value is lower than the level of understanding of the user represented by the index of the first value. [Appendix 4.4] 5. The automated answering device of claim 4, The characteristic information includes the content of information transmitted by the user to a third party, The voice setting unit performs the voice setting in accordance with the voice setting standard that is switched depending on a combination of the speech content and the characteristic information. [Appendix 4.5] 5. The automated answering device of claim 4, The acquisition unit acquires the characteristic information of the user from information about the user in accordance with acquisition criteria that can be switched depending on the content of the user's utterance.
[0087] [Appendix 5] A program for causing a computer to function as an automatic answering device as set forth in any of appendices 1 to 4. [Appendix 6] An automatic answering method for an automatic answering device according to any one of appendix 1 to appendix 4. [Explanation of symbols]
[0088] 11-41 Automatic answering machine 112f,212f Acquisition part 112g Scenario Selection 112h, 312h, 412h Audio setting section 112i Response Part 12 Terminal equipment
Claims
1. an acquisition unit that acquires information based on the content of information transmitted by the user to a third party; a scenario selection unit that estimates a future second information transmission content from a transition of information based on the information transmission content, and selects a response scenario based on the second information transmission content; a response unit that outputs information representing a response based on the response scenario in response to the user's utterance; An automatic answering device having:
2. an acquisition unit that acquires information based on information transmission content including a plurality of contents transmitted by a user to a third party through a plurality of different media; a scenario selection unit that selects a response scenario based on a ratio of the plurality of contents transmitted through the plurality of different media; a response unit that outputs information representing a response based on the response scenario in response to the user's utterance; An automatic answering device having:
3. an acquisition unit that estimates an utterance content label representing the content of a user's utterance, switches an acquisition criterion for acquiring the user's characteristic information to an acquisition criterion associated with the estimated utterance content label, and acquires, as the user's characteristic information, information that matches the switched acquisition criterion from information based on the content of information transmitted by the user to a third party; a scenario selection unit that selects a response scenario based on the characteristic information; a response unit that outputs information representing a response content based on the response scenario in response to the utterance content; An automatic answering device having:
4. A program for causing a computer to function as the automatic answering device according to any one of claims 1 to 3.
Citation Information
Patent Citations
Information processing apparatus and program
JP2019091387A
Program and information processing method
JP2024129098A
Information output method and information output device
WO2022224670A1