Automatic response system, automatic response method, and program

The automatic response system addresses the challenge of providing tailored responses in facilities by integrating question and situation acquisition with facility data to generate personalized outputs, improving user guidance and reducing employee workload.

WO2025220150A1PCT designated stage Publication Date: 2025-10-23MITSUBISHI ELECTRIC CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/015272
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-17
Publication Date
2025-10-23

AI Technical Summary

Technical Problem

Existing systems fail to provide appropriate responses to diverse user needs and situations in facilities due to employee availability and workload, leading to unanswered questions.

Method used

An automatic response system comprising a question acquisition unit, situation acquisition unit, answer generation unit, and output unit that utilizes facility data, user situation, and additional information to generate and output tailored responses.

Benefits of technology

The system provides more appropriate and personalized responses to facility users, addressing diverse needs and situations, enhancing user guidance and reducing employee workload.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024015272_23102025_PF_FP_ABST
    Figure JP2024015272_23102025_PF_FP_ABST
Patent Text Reader

Abstract

An automatic response system (10) according to the present disclosure is provided with: a question acquisition unit (12) that acquires a question pertaining to the use of a facility from a facility user who is a user of the facility; a situation acquisition unit (13) that acquires additional information indicating the situation of a questioner who is a facility user having asked a question; an answer generation unit (14) that generates an answer to the question using the question, the additional information, and facility data that is data pertaining to the facility; and an output unit (15) that outputs the answer.
Need to check novelty before this filing date? Find Prior Art

Description

Automatic response system, automatic response method and program

[0001] The present disclosure relates to an automatic response system, an automatic response method, and a program for generating answers to questions.

[0002] In recent years, it has become a problem that employees' time is taken up by answering questions from users and guiding users in train stations, commercial facilities, etc. In addition, when a user wants to ask a question, the employee may be busy assisting another user, which can cause the user to be unable to ask a question when they want to. For this reason, automation of user guidance is being considered.

[0003] Patent Document 1 discloses a route guidance system for use in shopping malls, event venues, etc. In the route guidance system described in Patent Document 1, when a user tells a robot their destination, a display device displays the route, and the robot explains the route to the destination to the user using voice and body movements.

[0004] Japanese Patent Application Laid-Open No. 2007-265329

[0005] The technology described in Patent Document 1 can guide users to a specified destination on behalf of an employee. However, users have diverse needs and situations. Therefore, the technology described in Patent Document 1 may not be able to provide information appropriate for the user.

[0006] The present disclosure has been made in consideration of the above, and aims to provide an automatic response system that can provide more appropriate responses to facility users.

[0007] In order to solve the above-mentioned problems and achieve the objectives, the automatic response system disclosed herein comprises a question acquisition unit that acquires questions about facility usage from facility users who are users of the facility; a situation acquisition unit that acquires additional information that indicates the situation of the facility user who asked the question; an answer generation unit that generates an answer to the question using the question, the additional information, and facility data that is data about the facility; and an output unit that outputs the answer.

[0008] The automatic response system according to the present disclosure has the effect of being able to provide more appropriate responses to facility users.

[0009] 5 is a diagram showing an example of the configuration of an automatic response system according to the first embodiment; FIG. 5 shows an example of the configuration of a response generation unit according to the first embodiment; FIG. 6 is a diagram for explaining the automatic response according to the first embodiment; FIG. 7 is a diagram for explaining the automatic response according to the first embodiment; FIG. 8 is a diagram for explaining the automatic response according to the first embodiment; 1 is a diagram showing a specific example of a response that takes into account the situation in the automatic response system of embodiment 1. FIG. 2 is a diagram showing a specific example of a response that takes into account the situation in the automatic response system of embodiment 1. FIG. 3 is a diagram showing another specific example of a response that takes into account the situation in the automatic response system of embodiment 1. FIG. 4 is a diagram showing an example of the configuration of a computer system that realizes each of the information processing units of embodiment 1. FIG. 5 is a diagram showing an example of the configuration of an automatic response system according to embodiment 2. FIG. 6 is a diagram showing a specific example of a response that takes into account the situation in the automatic response system of embodiment 2. FIG. 7 is a diagram showing an example of the configuration of an automatic response system according to embodiment 3.

[0010] An automatic response system, an automatic response method, and a program according to an embodiment will be described in detail below with reference to the accompanying drawings.

[0011] Embodiment 1. FIG. 1 is a diagram illustrating an example of the configuration of an automatic answering system according to an embodiment. The automatic answering system 10 of this embodiment accepts questions from facility users regarding facility use and outputs information related to facility guidance as a response to the questions. The facility may be, for example, a train station, a hospital, a commercial facility, a hotel, a building, an event venue, a conference center, a theme park, a railway facility including multiple stations (including the interior of a railway car), a school, a business office, a factory, etc., but is not limited to these as long as it is available to multiple users. The facility may be a facility available to an unspecified number of people, such as a train station, a facility available to specific people, such as a school or business office, or a facility available to both specific and unspecified people.

[0012] 1, the automatic response system 10 includes a data management unit 11, a question acquisition unit 12, a situation acquisition unit 13, an answer generation unit 14, and an output unit 15. The data management unit 11 and the answer generation unit 14 constitute an information processing unit 20, which is an information processing device. The automatic response system 10 is installed, for example, within a facility, but at least some of the units constituting the automatic response system 10 may be installed outside the facility.

[0013] The data management unit 11 manages facility data, which is data related to facilities. The facility data is data used by the answer generation unit 14 when generating answers, and may include not only data related to the facility itself but also data related to the surrounding area of ​​the facility and data related to the facility outside the facility. The facility data includes, but is not limited to, at least one of map information (map information) within the facility, facility status information related to the status of the facility, map information around the facility, facility information transmitted by at least one of the facility manager and facility employees, and weather information for the area including the facility. The map information includes information that enables determining routes to potential locations that users will visit, such as stores and conference rooms within the facility. For example, the map information may include information represented by a graph in which links such as corridors within the facility are used and the potential locations are nodes. The map information may also include attributes of the corridors, such as whether the corridors include stairs, slopes, elevators, or escalators. Alternatively, the map information within the facility may include image data showing predetermined routes for each potential location that users will visit. In this case, the map information may include image data for each attribute of the corridors, such as whether the corridors include slopes, elevators, or escalators.

[0014] The facility status information includes, but is not limited to, congestion information indicating the degree of congestion at the facility, information such as road closures within the facility, information indicating whether equipment within the facility is broken, and road conditions indicating the condition of surrounding roads. Facility information includes information that a facility manager or employee wants to convey to users, such as, for example, "It is raining, so the facility is slippery" or "The venue for event X starting at 10:00 today has been changed from X to Y." Weather information includes, but is not limited to, at least one of weather (sunny, rainy, strong winds, etc.), temperature, and humidity.

[0015] Furthermore, if the facility is located within a station, the facility data may include a timetable for trains and other railroad vehicles, and if a store is located within the facility, the facility data may include at least one of information about the products handled by the store, information about the food and drink served at the store, information indicating the store's main target demographic, information about the store's recommended products, store-generated information transmitted by store employees, etc. Store-generated information is information that store employees want to convey to users, such as "Today, Shop A is having a rainy day sale," "Shop B is currently running a limited-time sale," or "Restaurant K has a wide selection of cold drinks available that are perfect for hot days."

[0016] The facility data may include static data that does not change in a short period of time, i.e., does not change for a certain period of time, such as map information within the facility, information indicating the store's main target demographic, etc. The facility data may also include dynamic data that may change in a short period of time, i.e., may change over a certain period of time, such as facility status information, facility information, weather information, etc. The facility data may also include both static data and dynamic data.

[0017] The data management unit 11 includes a data acquisition unit 111, a pre-processing unit 112, and a data storage unit 113. The data acquisition unit 111 acquires facility data and outputs the acquired facility data to the pre-processing unit 112.

[0018] Specifically, the data acquisition unit 111 may acquire static facility data by, for example, accepting data input from a facility manager, operator, store employee, etc., or by receiving the data from another device (not shown). For example, the data acquisition unit 111 may receive map information about the area around the facility from a device that manages the map information. Alternatively, a terminal device that can be operated by a facility manager, operator, store employee, etc. may accept input of facility data from the facility manager, operator, store employee, etc., and the data acquisition unit 111 may receive the facility data from the terminal device.

[0019] Furthermore, the data acquisition unit 111 may acquire dynamic data from facility data by receiving sensor information acquired by a sensor (not shown in FIG. 1 ), such as a surveillance camera or a microphone, and calculating the dynamic data based on the received sensor information. For example, the dynamic data may include data acquired by a sensor installed inside or around the facility. For example, the data acquisition unit 111 may calculate congestion information based on video, which is sensor information acquired by a surveillance camera. Alternatively, the data acquisition unit 111 may output the sensor information to the preprocessing unit 112, and the preprocessing unit 112 may calculate the dynamic data from the sensor information. Alternatively, another device (not shown) may calculate the dynamic data based on the sensor information, and the data acquisition unit 111 may receive the dynamic data from the other device.

[0020] Furthermore, for example, for facility transmission information and store transmission information, which are dynamic data, the data acquisition unit 111 may acquire the facility transmission information and store transmission information as text data by accepting input from a sender, or may accept input of the facility transmission information and store transmission information by voice. Furthermore, information broadcast within the facility may be acquired as facility transmission information and store transmission information from a microphone used for in-facility broadcasting at the facility. Alternatively, microphones may be installed near the location where in-facility broadcasting is performed, and information broadcast within the facility may be acquired from the microphone as facility transmission information and store transmission information.

[0021] The preprocessing unit 112 performs preprocessing on the facility data received from the data acquisition unit 111. For example, the preprocessing unit 112 stores the preprocessed facility data in the data storage unit 113 as a database. The preprocessing is, for example, a process of converting the facility data into a data format accessible by the answer generation unit 14. For example, when the answer generation unit 14 reads vectorized data, the preprocessing unit 112 vectorizes the facility data and stores the vectorized facility data in the data storage unit 113. The vectorization includes, for example, a process of converting words in natural language processing into vectors, which are distributed representations according to the meanings of the words.

[0022] The content of the preprocessing performed by the preprocessing unit 112 is not limited to the above example, and may be determined depending on the processing method used in the answer generation unit 14. Depending on the processing method used in the answer generation unit 14, the preprocessing unit 112 may not be provided, and the facility data may be stored directly in the data storage unit 113. The data storage unit 113 stores the facility data.

[0023] The question acquisition unit 12 acquires questions about facility use from facility users (hereinafter also referred to as facility users). Specifically, the question acquisition unit 12 accepts questions about the facility from facility users and outputs the accepted questions to the answer generation unit 14. The question acquisition unit 12 may acquire the questions as voice data, or may acquire the questions as text data by accepting input through operation by the facility users. That is, the question acquisition unit 12 may include a microphone, an input means for accepting input through user operation, such as a touch panel, or both a microphone and an input means. The question acquisition unit 12 may also accept image input. In this case, for example, the image input is accepted by capturing an image with a camera. For example, a facility user may present a photo of a product, a store, or the like and input a question, such as where the product in the photo is located, or where the store is located, by voice or text. The camera may be a camera used as the situation acquisition unit 13 described below. When the question acquisition unit 12 receives a question using an image, it outputs the question together with the image data to the answer generation unit 14 .

[0024] The situation acquisition unit 13 acquires additional information indicating the situation of the facility user who asked the question and outputs the acquired additional information to the answer generation unit 14. For example, the situation acquisition unit 13 includes a sensor such as a camera or an infrared sensor and acquires detection results from the sensor as additional information (auxiliary information). For example, the additional information includes video data of the questioner. Furthermore, the situation acquisition unit 13 may acquire the additional information indicating the situation of the facility user as text data by accepting input through operation by the facility user, or may acquire the additional information indicating the situation of the facility user as audio data. When the situation acquisition unit 13 acquires the additional information as text data or audio data, for example, a question inquiring about the situation is presented to the facility user by the output unit 15, and the situation acquisition unit 13 acquires the answer to the question as the additional information. When the additional information is acquired as text data or audio data, the question acquisition unit 12 may function as the situation acquisition unit 13. Furthermore, the microphone used by the situation acquisition unit 13 to acquire audio data may be shared with the question acquisition unit 12.

[0025] As described above, the additional information may be, for example, situation information indicating the situation of the facility user asking the question. However, the additional information is not limited to this. The additional information may be situation information indicating the situation of the facility user asking the question and attribute information indicating the attributes of the facility user asking the question. For example, video data acquired by a camera can be used to understand both the situation and attributes of the facility user. Furthermore, the situation acquisition unit 13 may present the facility user with a question asking about the attribute in addition to the situation via the output unit 15, and the situation acquisition unit 13 may acquire the answer to the question as the additional information. In this way, the additional information may include the situation information and the attribute information. Alternatively, the situation information and the attribute information may be acquired by different means, such as acquiring video data as situation information and acquiring the attribute information as a response to the question asking about the attribute. Note that the method for acquiring the attributes of the facility user is not limited to acquiring the attributes from video data acquired by a camera. For example, as described below, the attributes of the individual may be identified using an individual identification result.

[0026] The answer generation unit 14 generates an answer regarding facility guidance using the question received from the question acquisition unit 12, the incidental information received from the situation acquisition unit 13, and the facility data stored in the data storage unit 113, and outputs the generated answer to the output unit 15. Specifically, the answer generation unit 14 understands the question and generates an answer using the facility data, taking into account the situation indicated by the incidental information. The answer generation unit 14 can generate an answer using data in multiple formats, including, for example, audio, text, and image, as input, and can output the answer in multiple formats. That is, for example, the answer generation unit 14 can generate a multimodal output in response to a multimodal input such as audio data, text data, and video data, but is not limited thereto. Hereinafter, generating a multimodal output in response to a multimodal input is also referred to as multimodal processing. Details of the answer generation unit 14 will be described later.

[0027] The output unit 15 presents the answer received from the answer generation unit 14 to the facility user who asked the question by outputting the answer. The answer may be output as audio, text data, an image, video, or a combination of two or more of these. Therefore, the output unit 15 includes, for example, at least one of a speaker and a display device such as a display or a touch panel. If the answer generation unit 14 is capable of multimodal processing, the output unit 15 outputs the answer according to the format of the answer received from the answer generation unit 14. For example, if the answer received from the answer generation unit 14 is audio data, the output unit 15 outputs the answer as audio. If the answer received from the answer generation unit 14 is at least one of video data and text data, the output unit 15 outputs the answer by display. If the output unit 15 includes a touch panel, the touch panel may be used as an input means for the question acquisition unit 12 or the situation acquisition unit 13.

[0028] Next, the answer generation unit 14 will be described in detail. FIG. 2 is a diagram illustrating an example of the configuration of the answer generation unit 14 according to the present embodiment. For example, as illustrated in FIG. 2 , the answer generation unit 14 includes a situation recognition unit 141 and a generation unit 142. The situation recognition unit 141 recognizes a situation based on the additional information received from the situation acquisition unit 13 and outputs the recognized situation to the generation unit 142. Note that the situation recognition unit 141 recognizes the facility user who asked the question. However, if the facility user is accompanied by someone, such as a family or a couple, hereinafter, not only the facility user who actually asked the question but also the accompanying person of the facility user will be treated as the facility user who asked the question. Hereinafter, the facility user who asked the question will also be referred to as the questioner. The situation recognized by the situation recognition unit 141 includes, but is not limited to, at least one of the following: what the questioner is wearing (such as the questioner's clothing and accessories), whether the questioner is using an assistive device such as a wheelchair or a cane (whether or not the assistive device is being used), the number of questioners, the questioner's emotions, the size of luggage the questioner is carrying, and the direction of movement of the questioner.

[0029] Furthermore, as described above, when the additional information is information that indicates the situation and attributes of the questioner, the situation recognition unit 141 uses the additional information to recognize the situation and attributes of the questioner, and outputs the recognized situation and attributes to the generation unit 142. The attributes of the questioner may include, for example, at least one of gender, age, generation, presence or absence of a disability, and occupation.

[0030] The generation unit 142 generates an answer regarding facility guidance using the situation (or situation and attributes) received from the situation recognition unit 141, the question received from the question acquisition unit 12, and facility data stored in the data storage unit 113, and outputs the generated answer to the output unit 15.

[0031] In this way, the answer generation unit 14 generates an answer that reflects the situation of the questioner. For example, if a questioner using a wheelchair asks how to get to a destination, the answer generation unit 14 can generate an answer that shows a route that does not have stairs and uses a ramp or elevator.

[0032] 3 and 4 are diagrams illustrating the automatic response of this embodiment. In the example shown in FIG. 3, the automatic response system 10 includes a camera as the situation acquisition unit 13, and a touch panel and a speaker as the output unit 15. In the example shown in FIG. 3, a keyboard (software keyboard) for accepting questions is displayed below the touch panel, and the portion where this keyboard is displayed also functions as the question acquisition unit 12. In the example shown in FIG. 3, the automatic response system 10 further includes a microphone as the question acquisition unit 12. As shown in the upper part of FIG. 3, a questioner 30 carrying a large baggage utters, "Where is the station platform?" to ask the automatic response system 10 for the location of the station platform.

[0033] The answer generation unit 14 of the automatic response system 10 recognizes the situation that the questioner 30 is carrying large luggage from video data received from a camera, which is an example of a sensor used as the situation acquisition unit 13, and generates an answer using the recognized situation, the voice data of the question received from the question acquisition unit 12, and the facility data. Note that the situation recognized in the example shown in FIG. 3 is not limited to "carrying large luggage," but may also be "both hands are full" or "difficulty walking." In the example shown in FIG. 3, since the questioner 30 is carrying large luggage, the answer generation unit 14 generates a route to the station platform, i.e., a route using an elevator, as shown in the lower part of FIG. 3, based on map information within the facility in the facility data, rather than a route using stairs or escalators. Note that there are no particular restrictions on the route generation method, and methods such as searching for the shortest route within a range that satisfies the conditions can be used. In the example shown in FIG. 3, the answer generation unit 14 generates a diagram (image) showing the route and text data describing the diagram as an answer, and outputs the generated answer to the output unit 15, which then outputs the answer. 3, text and a diagram showing the route are displayed as the answer on the touch panel of the output unit 15, but the output method is not limited to this. For example, a speaker of the output unit 15 may also output the answer as voice, or the speaker may output the answer as voice instead of text.

[0034] In the example shown in Fig. 4, the automatic response system 10 is composed of a camera, which is an example of a sensor used as the situation acquisition unit 13, and a main body 40. The main body 40 includes an information processing unit 20, a question acquisition unit 12, and an output unit 15, which are not shown in Fig. 4. In this way, the sensor used as the situation acquisition unit 13 may be provided separately from the main body 40. Furthermore, the automatic response system 10 may include multiple sensors as the situation acquisition unit 13; for example, sensors may be provided both in the main body 40 and outside the main body 40.

[0035] In the example shown in FIG. 4 , a questioner 30 uses a wheelchair and asks the automatic answering system 10, "Where is the exit?" In the example shown in FIG. 4 , a microphone, which is the question acquisition unit 12, receives a question. In the example shown in FIG. 4 , since the questioner 30 uses a wheelchair, the answer generation unit 14 generates an answer that includes a slope without steps as a route to the exit, and the output unit 15 displays the answer. Furthermore, when a question about the location of the exit is asked, the answer generation unit 14 assumes that the questioner will be leaving the facility, and if it is raining based on the facility data, the answer generation unit 14 may add a sentence that corresponds to the outside weather, such as "It's raining outside, so please be careful on your way home."

[0036] 4, the request is for an exit. However, by including map information about the area around the facility in the facility data, for example, when a questioner asks, "Please tell me the route to the department store" if the facility is located inside a train station, the automatic response system 10 may be able to provide guidance on the route from the station to the department store. In this way, the answer generation unit 14 may provide information about the area around the facility as information related to facility guidance. Furthermore, when the facility is located inside a train station, the facility data may include information indicating the train lines that stop at the station, the stations on the lines, the nearest stores and tourist attractions to the station, etc., so that a response to a question such as, "Which train line should I take to get to VV?" may include information about which train line to take and which station to get off at.

[0037] In the example shown in FIG. 3 , the automatic response system 10 is realized as a single device, and in the example shown in FIG. 4 , it is realized by a camera, which is an example of a sensor used as the situation acquisition unit 13, and the main body 40. However, the hardware configuration of the automatic response system 10 is not limited to these examples. For example, the information processing device, which is the information processing unit 20 shown in FIG. 1 , and the input / output device including the question acquisition unit 12, the situation acquisition unit 13, and the output unit 15 may be provided in separate locations. In this case, the input / output device and the information processing device may each include a transceiver unit for communication, and information may be exchanged via the transceiver unit. Furthermore, FIGS. 3 and 4 are merely examples, and the situations considered by the automatic response system 10 and the answers presented by the automatic response system 10 are not limited to these examples.

[0038] Returning to the description of FIG. 2 , the answer generation unit 14 may be, for example, a generative AI (artificial intelligence). Examples of the generative AI include multimodal AI such as Gemini (registered trademark) and GPT (Generative Pre-trained Transformer)-4V, image generation AI such as Stable Diffusion, large-scale language model such as GPT-3, and voice generation AI such as Murf. AI. When the answer generation unit 14 uses multimodal AI, the answer generation unit 14 includes the functions of a situation recognition unit 141 and a generation unit 142. When the answer generation unit 14 uses multimodal AI, before accepting a question from a facility user, such as when starting up the automatic response system 10, an instruction such as "Please consider the situation of the facility user who asked the question and generate video data, text data, and audio data as necessary to provide an answer" is input to the answer generation unit 14. The answer generation unit 14 is also instructed to use facility data as appropriate. As a result, the answer generator 14 takes into consideration the situation of the questioner 30, references the facility data, and generates an answer regarding guidance to the facility using at least one of video data, text data, and audio data.

[0039] Furthermore, specific information to be considered as a situation may be input to the answer generating unit 14, or data indicating a definition of a situation may be included in the facility data. For example, an instruction such as "The situation of the facility user who asked the question includes whether the questioner uses an assistive device such as a wheelchair or a cane, the number of people asking the question, the emotions of the questioner, the clothing of the questioner, what the questioner is wearing, the size of luggage the questioner is carrying, and the direction of movement of the questioner" may be input to the answer generating unit 14.

[0040] When the answer generation unit 14 uses multimodal AI and receives the above-described instructions, the format of the answer is determined by the answer generation unit 14. On the other hand, if it is desired to define specific rules, such as an instruction to provide an answer in the same format as the question, rather than leaving the determination of the answer format to the answer generation unit 14, instructions indicating the rules may be input to the answer generation unit 14, or data indicating the rules may be included in the facility data. For example, typical examples, norms, guidelines, etc. of answers corresponding to questions may be stored in the data storage unit 113 as facility data or separately from the facility data, and the answer generation unit 14 may refer to the rules when creating an answer. Alternatively, pre-learning may be performed to learn these rules. Note that while facility data is used to generate an answer, a portion of the facility data may be used as part of the answer. For example, if the facility data includes images of products sold by a store, the answer generation unit 14 may include the images in the answer.

[0041] For example, in order for the answer generation unit 14 to generate the answers exemplified in FIGS. 3 and 4 , the answer generation unit 14 may input the situation of the questioner 30 to be considered and the question to the multimodal AI in advance, check whether the desired answer is obtained, and if the desired answer is not obtained, instruct the multimodal AI to provide a correct answer corresponding to the situation and the question. The correct answer corresponding to the situation and the question may also be included in the facility data. For example, as shown in FIG. 3 , when the questioner 30 carrying large luggage asks about the route to the destination, if the answer generation unit 14 generates an answer that includes stairs, the answer generation unit 14 may be given an instruction such as, "If the questioner is carrying large luggage, use the elevator instead of the stairs." This instruction may also be included in the facility data. Alternatively, the instruction may be broken down into information such as, "It is difficult for the questioner to walk if he or she is carrying large luggage or has both hands full," and information such as, "If it is difficult to walk, use the elevator instead of the stairs." In addition, when attributes are taken into consideration in addition to situations, instructions for generating an appropriate answer for each situation and attribute may be given to the answer generation unit 14, or the instructions may be included in the facility data.

[0042] Furthermore, when a large-scale language model is used in the answer generation unit 14, the question acquisition unit 12 may accept the question as text data, and the output unit 15 may output the answer as text data. Alternatively, the question acquisition unit 12 may accept the question as voice, and the question acquisition unit 12 or the answer generation unit 14 may convert the voice data of the question into text data. For example, a large-scale language model may function as the generation unit 142, and a situation recognition model that recognizes a situation from video data may be used as the situation recognition unit 141. The situation recognition model may be, for example, a trained model that has been trained to infer a situation from video data by machine learning. Examples of machine learning used to generate the trained model include supervised learning such as a neural network, but are not limited to unsupervised learning, reinforcement learning, and the like. The situation recognition model may also be a model that recognizes attributes.

[0043] Furthermore, the situation recognition model may be a combination of multiple models, such as a combination of a general emotion recognition model that recognizes emotions from video data and an image recognition model that recognizes the presence or absence of assistive devices, the size of luggage, clothing, etc. from video data. The generation unit 142 generates an answer as text data using, for example, text data indicating the situation (or the situation and attributes) received from the situation recognition unit 141 and the question. Alternatively, a speech generation AI may be further used as the generation unit 142, and the answer generated by the large-scale language model may be converted into speech data by the speech generation AI, or an image generation AI may be further used as the generation unit 142, and the answer generated by the large-scale language model may be converted into image data by the image generation AI.

[0044] In this way, a multimodal output may be realized by combining and using multiple models. In this case, for example, output may be performed in all available output formats, or rules for selecting the output format may be determined in advance. For example, a rule may be established in which the status (or status and attributes) of the questioner 30 is associated with the output format, or a rule may be established in which output is to be in the same format as the format of the question asked by the questioner 30. Furthermore, a rule may be established in which the type of question content (e.g., a question about a route to a destination, a question asking about recommended shops, etc.) is associated with the output format. The rules for determining the output format are not limited to these examples.

[0045] When a generative AI, a large-scale language model, or the like is used in the answer generation unit 14, the facility data is vectorized by the preprocessing unit 112 and stored in the data storage unit 113, allowing the answer generation unit 14 to use the facility data as is. Therefore, the facility data can be reflected in the answers of the answer generation unit 14 without requiring additional learning.

[0046] Furthermore, a general automatic conversation program such as a chatbot may be used for the generation unit 142. In this case, a situation recognition model may be used as the situation recognition unit 141, similar to when a large-scale language model is used for the generation unit 142, and text data indicating the situation (or the situation and attributes) recognized by the situation recognition unit 141 may be input to the generation unit 142.

[0047] The automatic conversation program may be a rule-based (scenario-based) program in which response rules are defined in advance, or a machine learning program. When a rule-based automatic conversation program is used, for example, answers corresponding to the content of questions for each situation (or situation and attribute) are defined in advance and set in the automatic conversation program. The machine learning in a machine learning automatic conversation program may be supervised learning, unsupervised learning, or reinforcement learning. For example, when supervised learning is used in the automatic conversation program, a trained model is generated using multiple training datasets including questions for each situation (or situation and attribute) and corresponding answers, which are correct answer data. The generation unit 142 generates an answer by inputting the situation (or situation and attribute) and question corresponding to the questioner 30 into the trained model. Note that even when an automatic conversation program is used in the generation unit 142, the answer generated by the automatic conversation program may be converted into voice data by a voice generation AI or into image data by an image generation AI. In this case, output may be performed in all available output formats, or rules for selecting the output format may be defined in advance, and the output format may be determined according to the rules.

[0048] The questioner 30 may not only be an unspecified person visiting the facility, but also a specific person, such as an employee, a facility manager, or a student if the facility is a school. That is, facility users may include a specific person. For example, if the facility is intended for a specific group of people, only specific people may be considered as questioners 30. In this case, for example, a facial photograph of the specific person may be included in the facility data, the situation acquisition unit 13 may acquire video data from a camera as additional information, and the answer generation unit 14 may identify the questioner 30 based on the facial photograph and generate an answer based on the identification result. This video data is an example of personal authentication information for identifying an individual. For example, for specific people who may potentially use the facility, their identification information (hereinafter also referred to as personal identification information) may be associated with a facial photograph and stored in the data storage unit 113 as facility data. Furthermore, for each piece of personal identification information, individual information about the person corresponding to the personal identification information may be stored in the data storage unit 113 as facility data. Individual information is information regarding the use of a facility by a specific individual, such as, but not limited to, one or more of the following: schedule, accessible locations, and range of information that can be obtained (authority to view information).

[0049] For example, if the facility is a school, the individual information of the facility data includes information indicating the classes each student is taking, and the facility data also includes class location information indicating the time and location (e.g., classroom) of each class. Thus, when a student (questioner 30) asks a question such as, "Where is the next class?", the answer generation unit 14 can determine the next class the student will take based on the individual information of the student and the time and location of the class using the class location information. The answer generation unit 14 generates an answer, reflecting the determined information, such as, "The next class is Class Y, which will be held in Classroom X in Building 3." In this case, the answer may include an image showing the route from the current location to Classroom X in Building 3. When generating an answer, as described above, for example, the situation may be further taken into consideration.

[0050] Furthermore, if the facility is a school, the facility data may include individual information for not only students but also school staff. For example, the individual information may include information indicating the staff's class schedule, the time and location of the staff's exam supervision, and so on. In this case, as in the student example, for example, in response to a question from a staff member such as, "When and where will I be supervising my next exam?", a response such as, "My next exam supervision will be in classroom Z in building 2, starting at 11:00 a.m." is generated. The content of the response described above is merely illustrative, and the content of the response is not limited to the above example. Note that while an example of identifying individuals using facial information has been described here, to identify individuals, for example, a passcode assigned to each individual may be included in the facility data, and the automated response system 10 may identify individuals by accepting input of the passcode when a question is asked. Alternatively, the automated response system 10 may be equipped with a device that performs biometric authentication other than facial recognition, and individuals may be identified using that device. In other words, the personal authentication information may be a passcode, biometric authentication information, or the like.

[0051] Furthermore, for each attribute of a specific person using the facility, such as an employee or student, at least one of information and rules for generating a response to the questioner 30 for that attribute may be defined as attribute-related information, and the attribute-related information may be included in the facility data. The attribute-related information is information about facility use for each attribute, including schedules, accessible locations, the scope of information that can be obtained (information viewing authority), and contact information. If the facility is a school, the attributes may include, for example, status (student, professor, assistant professor, teacher, office worker, etc.) and affiliation (faculty, department, selection). If the facility is a business establishment, the attributes may include job title, affiliation (department, faculty, division), qualifications held, etc. Thus, for example, the attributes may include at least one of status, affiliation, and job title. For example, by including personal attributes in the individual information in the facility data, the response generation unit 14 may identify the individual using a method similar to the above-described individual identification, grasp the attributes based on the individual information, and generate a response according to the attributes.

[0052] For example, if the attribute-related information specifies classes to be taken for each department, as in the example described above, when a student asks, "Where is my next class?", the answer generation unit 14 identifies the department to which the student belongs as an attribute of the student (asker 30). The answer generation unit 14 also identifies classes corresponding to the identified attributes based on the attribute-related information and generates an answer indicating the location of the class based on the class location information. For example, if accessible locations are specified for each attribute, when the asker 30 asks how to get to a destination, the answer generation unit 14 calculates a route that travels within the accessible locations and generates the calculated route as an answer. For example, if the attribute includes information viewing authority, the answer for the asker 30 is generated using information within the range of viewing permission. For example, if the attribute-related information includes a message for each attribute, the answer may be generated by adding the message to a direct answer corresponding to the question, regardless of the content of the question. For example, the message may include a notice of a change in class.

[0053] Furthermore, if the facility is open to the general public, such as a train station, a commercial facility, or a university, the questioner 30 may be both an unspecified person, such as a customer of the facility or a visitor to the university, and a specified person, such as an employee or a student of the facility. The answer generation unit 14 may change the content of the answer depending on whether the questioner 30 is an unspecified person or a specified person. In this case, the attributes of the questioner 30 may include, for example, whether the questioner 30 is a specified person, and rules for generating answers for each attribute, such as accessible locations, the scope of information that can be obtained (information viewing authority), and contact information, may be defined in the attribute-related information, as in the above example. The attribute of whether the questioner 30 is a specified person is determined in the same way as in the case of identifying the individual specific person described above, for example, by including facial information or the like of the specified person in the facility data in advance. For example, the attribute-related information may include information distinguishing between facility data used to generate answers for specified people and facility data used to generate answers for unspecified people. Furthermore, for specified people such as employees, attributes such as affiliation and position may be further included as attributes, as in the above example, and answers may be generated according to these attributes.

[0054] For example, if the facility is a commercial facility and the questioner 30 asks, "Where is the warehouse where product P is stored?", the answer generation unit 14 generates, as an answer, information indicating the warehouse where product P is stored based on the facility data if the questioner 30 is an employee, and generates an answer indicating that the question cannot be answered if the questioner 30 is not an employee. At this time, for example, the answer generation unit 14 further generates an answer depending on the situation, as described above.

[0055] Furthermore, if the answer generation unit 14 cannot recognize the content of the question from the questioner 30 as language, if the question from the questioner 30 is insufficient for generating an answer, or if information about the situation of the questioner 30 is insufficient for generating an answer, the answer generation unit 14 may cause the output unit 15 to output information to the questioner 30 asking for the missing information. That is, for example, if the answer generation unit 14 cannot recognize the question or if there is insufficient information for generating an answer, the answer generation unit 14 of the automatic response system 10 may generate an answer to ask the questioner 30 again. For example, if the content of the question from the questioner 30 cannot be recognized as language, the answer generation unit 14 may output at least one of text data and audio data such as "Please repeat the question again" to the output unit 15, and the output unit 15 may output at least one of text and audio. Furthermore, if the facility is located within a station, the answer generation unit 14 may output at least one of text data and voice data for identifying the platform, such as "Is this the platform for the K line or the L line?" to the output unit 15, and the output unit 15 may output at least one of text and voice. Furthermore, for example, if it is recognized that there are multiple questioners 30 but the number of questioners 30 is unclear because they are overlapping, the answer generation unit 14 may output at least one of text data and voice data for identifying the number of people, such as "How many people are there?", and the output unit 15 may output at least one of text and voice.

[0056] Next, the operation of the automatic answering system 10 of this embodiment will be described. Fig. 5 is a flowchart showing an example of an automatic answering process procedure in the automatic answering system 10 of this embodiment. The automatic answering system 10 acquires facility data (step S1). In detail, the data acquisition unit 111 acquires the facility data and outputs it to the preprocessing unit 112. The facility data may be static data as described above, dynamic data, or both.

[0057] The automatic response system 10 stores the facility data (step S2). Specifically, the preprocessing unit 112 performs preprocessing on the facility data, and stores the preprocessed facility data in the data storage unit 113. The preprocessing is, for example, vectorization as described above, but is not limited to this, and may be conversion into a data format that can be accessed by the answer generation unit 14 or a data format that can be accessed quickly by the answer generation unit 14. Also, as described above, preprocessing does not necessarily have to be performed.

[0058] The automatic answering system 10 acquires the question content (step S3). Specifically, the question acquiring unit 12 acquires the question from the questioner 30 as voice or text data, thereby acquiring the question content, and outputs the acquired question to the answer generating unit 14.

[0059] The automatic response system 10 acquires the situation (step S4). Specifically, the situation acquisition unit 13 acquires additional information indicating the situation of the questioner 30 and outputs the additional information to the answer generation unit 14. For example, the situation acquisition unit 13 may acquire sensor information acquired by a sensor such as a camera as the additional information, or may acquire the additional information by inputting voice or text data from the questioner 30.

[0060] The automatic response system 10 generates an answer (step S5). In detail, the answer generation unit 14 grasps the situation (the situation of the questioner 30) from the additional information, generates an answer to the question based on the situation, and outputs the generated answer to the output unit 15. As described above, the automatic response system 10 may further grasp attributes (the attributes of the questioner 30) from the additional information, and generate an answer to the question based on the situation and attributes.

[0061] The automatic response system 10 outputs the answer (step S6) and ends the automatic response process. Specifically, in step S6, the output unit 15 outputs the answer. The output unit 15 outputs the answer by voice or by display, depending on the format of the answer generated by the answer generation unit 14 (whether the answer is voice data or text data or image data).

[0062] FIG. 6 is a flowchart showing an example of the answer generation process procedure in the answer generation unit 14 shown in step S5 of FIG. 5 . The answer generation unit 14 recognizes a situation and a question (step S11). Specifically, the situation recognition unit 141 recognizes a situation based on additional information and outputs the situation to the generation unit 142. The generation unit 142 recognizes the situation acquired from the situation recognition unit 141 and the question received from the question acquisition unit 12 by converting them into a format required for answer generation processing. Furthermore, for example, when a multimodal AI or a large-scale language model is used in the answer generation unit 14, the generation unit 142 understands the meaning of the situation and the question. Specifically, if the question is "Where is the platform?", the answer generation unit 14 understands that "the person is asking about the route to the platform," and if the recognized situation is "I have large luggage," the answer generation unit 14 understands that "I would prefer a route that avoids stairs as much as possible."

[0063] The answer generator 14 refers to the facility data (step S12). Specifically, the generator 142 reads out facility data necessary for answering the question from the data storage unit 113. For example, in the example where the question is "Where is the platform?" and the situation is "I have large luggage," map information within the station is read out as facility data.

[0064] The answer generation unit 14 generates an answer (step S13) and ends the answer generation process. Specifically, in step S13, the generation unit 142 uses the referenced facility data to generate an answer to the question that takes the situation into consideration, and outputs the generated answer to the output unit 15. For example, in the example where the question is "Where is the platform?" and the situation is "I have large luggage," the generation unit 142 searches for a route to the platform that uses an elevator and does not use stairs, using map information within the station, and generates a diagram showing the route obtained by the search and an explanation for the diagram.

[0065] Through the above processing, the automated answering system 10 of this embodiment can generate an answer that takes into account the situation of the questioner 30. This allows the automated answering system 10 to provide a more appropriate answer to the facility user than when simply answering a specific question from the facility user. Furthermore, the automated answering system 10 may generate an answer based on attributes in addition to the situation, and in this case, it is also possible to provide a more appropriate answer according to the attributes to the facility user. Furthermore, by including dynamic data (dynamic facility data) as facility data, it is possible to provide the facility user with appropriate answers and information based on the current status of the facility.

[0066] FIG. 7 is a diagram illustrating an example of dynamic facility data according to the present embodiment. In the example illustrated in FIG. 7 , the facility is a railway facility (including a railway vehicle) including a station premises or multiple stations. In the example illustrated in FIG. 7 , a questioner 30 asks, "Which car is empty?" To generate an answer to this question, it is necessary to know the degree of congestion in each car. In the example illustrated in FIG. 7 , the data acquisition unit 111 acquires, as dynamic facility data, video data acquired by cameras 50 installed in each railway vehicle. Furthermore, a timetable indicating when each train formation will arrive at the station is also stored in the data storage unit 113 as facility data. Note that the dynamic facility data is updated to the latest data, for example, when new data is acquired. However, past dynamic facility data may also be stored in the data storage unit 113.

[0067] The preprocessing unit 112 may store video data of the interior of the vehicles in the data storage unit 113, or may determine the degree of congestion of each vehicle based on the video data and store the determined result as facility data in the data storage unit 113. The degree of congestion may be indicated, for example, in two stages of whether or not there is a mixture, in three stages of crowded, normal, and empty, by a congestion rate, or in some other way. When the video data is stored in the data storage unit 113, the answer generation unit 14 recognizes the degree of congestion based on the video data and uses it to generate an answer.

[0068] For example, in the example shown in FIG. 7 , the answer generation unit 14 identifies the next arriving train car based on a timetable and determines which cars are available by using the degree of congestion calculated from video data captured inside each car of the train car. This allows the automatic response system 10 to generate an appropriate answer that takes into account the actual situation, such as "Car 2 is available," in response to the question, "Which car is available?" Also, in the example shown in FIG. 7 , the automatic response system 10 is installed on a station platform. The answer generation unit 14 generates an image showing the location of car 2 and the current location as an answer, and the output unit 15 displays the generated image along with the text answer, "Car 2 is available." Note that, as described above, the situation (or the situation and attributes) may also be taken into consideration when generating this answer. For example, in the example shown in Figure 7, if the questioner 30 does not have large luggage, an answer may be generated to guide the questioner to the emptiest available vehicle, and if the questioner 30 has large luggage, an answer may be generated to guide the questioner to the nearest available vehicle (even if it is not the emptiest vehicle).

[0069] Furthermore, the dynamic facility data is not limited to the example shown in Figure 7, but may be information indicating the degree of congestion in the toilets within the facility, information indicating the degree of congestion in the store, or facility-originated information, store-originated information, etc., as described above.

[0070] Next, a specific example of a response that takes into account the situation in the automatic response system 10 of this embodiment will be described. As described above, the situation of the questioner 30 may include, for example, the number of questioners 30. For example, when the questioner 30 asks, "Please tell me a restaurant you recommend," if there is only one questioner 30, the automatic response system 10 selects a quiet restaurant based on the facility data, generates a response indicating that the selected restaurant is recommended, and outputs the response. On the other hand, if the questioner 30 is with a family, the automatic response system 10 selects a family-oriented restaurant, such as a family restaurant, generates a response indicating that the selected restaurant is recommended, and outputs the response.

[0071] 8 to 14 are diagrams showing specific examples of responses that take the situation into consideration in the automatic response system 10 of this embodiment. In the example shown in FIG. 8, a middle-aged or elderly woman wearing elegant clothing and accessories asks, "Do you know of any clothing stores?" The automatic response system 10 recognizes that the questioner 30 is wearing accessories and elegant clothing as her situation, and based on this recognition result, the question, and information in the facility data indicating the products sold by each store and the target audience of each store, determines that Shop A and Shop B are recommended clothing stores for middle-aged or elderly people. The automatic response system 10 then generates information providing directions to Shop A and Shop B as a response to the question and displays it on the output unit 15. In this way, a response that takes the situation of the questioner 30 into consideration is generated.

[0072] In the example shown in FIG. 8 , the automatic response system 10 further generates a response including basis information indicating the basis of the response based on the situation used to generate the response. In the example shown in FIG. 8 , the automatic response system 10 generates basis information of "mother generation" based on the situation that the customer is a middle-aged or elderly woman wearing elegant clothing and accessories as the basis for deciding to guide them to Shop A and Shop B. As a result, in the example shown in FIG. 8 , not only the suggested store names "Shop A" and "Shop B" but also the text "How about a fashionable clothing store popular with the mother generation?" are displayed. Furthermore, touching the "Shop A" and "Shop B" portions on the display screen shown in FIG. 8 may display detailed information about each store. Note that the specific content of the sentence (text) in the response is not limited to the example shown in FIG. 8 .

[0073] In the example shown in FIG. 9 , a young woman wearing a uniform asks, "Do you know of any clothing stores?" The automatic response system 10 recognizes that the situation of the questioner 30 is that of a young woman wearing a uniform, and based on this recognition result, the question, and information in the facility data indicating the products sold by each store and the target audience of each store, determines that Shop C and Shop D are recommended clothing stores for female high school students (high school girls). The automatic response system 10 then generates information directing to Shop C and Shop D as an answer to the question and displays it on the output unit 15. In the example shown in FIG. 9 , similar to the example shown in FIG. 8 , evidence information is also displayed. Specifically, the text "high school girl" is generated as evidence information, and an answer including the evidence information is displayed.

[0074] In the example shown in FIG. 10 , a middle-aged man wearing a classic hat and classic clothing asks, "Do you know where a clothing store is?" The automatic response system 10 recognizes that the situation of the questioner 30 is a middle-aged man wearing a classic hat and classic clothing. Based on this recognition result, the question, and information in the facility data indicating the products sold by each store and the target audience of each store, the automatic response system 10 determines that an e-shop is a recommended clothing store for middle-aged men who like classic fashion. The automatic response system 10 then generates information providing directions to the e-shop as an answer to the question and displays it on the output unit 15. In the example shown in FIG. 10 , similar to the examples shown in FIGS. 8 and 9 , evidence information is also displayed. Specifically, the text "dandy" is generated as evidence information, and an answer including the evidence information is displayed. Note that in the examples shown in FIGS. 8 , 9 , and 10 , the gender of the questioner 30 is also taken into consideration. However, the gender of the questioner 30 is also an attribute, and these examples can be said to generate an answer taking into consideration both the situation, such as clothing, and the attribute.

[0075] 11 and 12 show examples in which the facial expression or emotion of the questioner 30 is taken into consideration as the situation of the questioner 30. In both of the examples shown in Fig. 11 and Fig. 12, the questioner 30 is asking where his / her home is.

[0076] In the example shown in Fig. 11, the questioner 30 asks "Where is the platform?" in a calm and normal manner. The automatic answering system 10 recognizes that the questioner 30 is normal, i.e., calm, from the questioner's facial expression, movements, tone of voice, etc., and generates an answer while engaging in a dialogue with the questioner 30. In the example shown in Fig. 11, the automatic answering system 10 asks the questioner 30, "Where do you want to take the train?", and the questioner 30 replies, "I want to go to the XX direction," and based on the answer from the questioner 30, the automatic answering system 10 generates and outputs the answer, "If you are going to the XX direction, please proceed to platform 3."

[0077] On the other hand, in the example shown in FIG. 12 , the questioner 30 appears flustered and asks, "Where is the platform?" The automated response system 10 recognizes that the questioner 30 is in a panic and is in a state of stress based on the questioner's facial expression, movements, tone of voice, and the like, and outputs a response including multiple types of information at once without engaging in a dialogue with the questioner 30. In the example shown in FIG. 12 , the automated response system 10 simultaneously displays the following information as responses: "Towards XX, Platform 3," "Towards YY, Platform 4," and "Towards ZZ, Platform 6." As illustrated in FIGS. 11 and 12 , the content of the response, the method of interaction with the questioner 30, and the like may be determined depending on the facial expression or emotions of the questioner 30.

[0078] 13 and 14 show examples in which the situation of the questioner 30 is considered based on the facial expression or emotion of the questioner 30 and the number of people. In the example shown in FIG. 13, the questioner 30 is a group of four people, arm around each other's shoulders, chatting and appearing to be having fun, and asking, "Can you recommend a bar?" Based on the facial expression, movements, and tone of voice of the questioner 30, the automatic answering system 10 recognizes that the questioner 30 is in a high-energy state (a state of being in a good mood) and recognizes that the questioner 30 is part of a group, and generates a response that provides directions to lively establishments "Bar A" and "Bar B." In the example shown in FIG. 13, the response also includes evidence information, as in the examples shown in FIGS. 8, 9, and 10. Information such as "For all of you who are excited, how about a lively bar with all-you-can-drink options?" is also presented to the output unit 15.

[0079] In the example shown in Fig. 14, the questioner 30 is a man and a woman holding hands, and asks, "Can you recommend a bar?" The automatic answering system 10 recognizes that the questioner 30 is a couple based on the questioner's facial expression, movements, tone of voice, etc., and generates a response recommending quiet bars "Bar C" and "Bar D." In the example shown in Fig. 14, the response also includes evidence information, as in the examples shown in Figs. 8, 9, and 10, and the output unit 15 also displays information such as "How about a quiet bar for a couple?"

[0080] The specific examples shown above are merely illustrative, and the responses that take the situation into consideration in the automatic answering system 10 are not limited to these examples. Furthermore, in FIGS. 8, 9, 10, and 14, responses that take gender into consideration are generated, but this is not limiting. The situation may not be recognized based on gender (male or female), and the response may not include any gender-related content. In other words, the response generator 14 may generate a gender-independent response. This allows the automatic answering system 10 to output responses that take LGBTQ (Lesbian, Gay, Bisexual, Transgender, Queer / Questioning) into consideration.

[0081] The automatic response system 10 may also generate LGBTQ-friendly responses that do not distinguish between genders or that take both genders into consideration. For example, when providing restroom guidance, the system may provide a response that includes both genders, such as by providing the locations of men's and women's restrooms regardless of the user's gender. Rules for generating LGBTQ-friendly responses may be stored in the data storage unit 113 as facility data or separately from the facility data, and the response generation unit 14 may refer to these rules when creating a response. Rules may also be established for ethical standards to be reflected in responses, not limited to LGBTQ considerations, and these rules may be stored in the data storage unit 113 as facility data or separately from the facility data, and the response generation unit 14 may refer to these rules when creating a response. When a generation AI is used in the response generation unit 14, these rules may be specified in advance to the generation AI.

[0082] Next, the hardware configuration of the automatic answering system 10 of this embodiment will be described. The information processing units 20 in the automatic answering system 10 of this embodiment function as the information processing units 20 by executing a computer program on the computer system, which is a computer program that describes the processing of each of the information processing units 20. FIG. 15 is a diagram showing an example configuration of a computer system that realizes each of the information processing units 20 of this embodiment. As shown in FIG. 15, this computer system includes a control unit 101, an input unit 102, a memory unit 103, a display unit 104, a communication unit 105, and an output unit 106, which are connected via a system bus 107.

[0083] In FIG. 15 , the control unit 101 is a processor such as a CPU (Central Processing Unit) and executes a program describing the processing performed by the information processing unit 20 of this embodiment. The input unit 102 is composed of, for example, a keyboard, buttons, a mouse, etc., and is used by a user of the computer system to input various information. The memory unit 103 includes various memories such as RAM (Random Access Memory) and ROM (Read Only Memory) and a storage device such as a hard disk, and stores programs to be executed by the control unit 101, necessary data obtained during processing, etc. The memory unit 103 is also used as a temporary storage area for programs. The control unit 101 and the memory unit 103 constitute, for example, a processing circuit. The processing circuit may be a single circuit or multiple circuits. The display unit 104 is composed of a display, an LCD (Liquid Crystal Display), etc., and displays various screens to the user of the computer system. Note that a touch panel in which the input unit 102 and the display unit 104 are integrated may also be used. The communication unit 105 is a receiver and transmitter that perform communication processing. The output unit 106 is a speaker or the like. Note that Fig. 15 is just an example, and the configuration of the computer system that realizes each of the information processing units 20 is not limited to the example shown in Fig. 15. For example, the output unit 106 may not be provided.

[0084] Here, an example of the operation of the computer system until the program of this embodiment is ready to be executed will be described. In the computer system having the above configuration, for example, the program is installed in storage unit 103 from a CD-ROM or DVD-ROM inserted in a CD (Compact Disc)-ROM drive or DVD (Digital Versatile Disc)-ROM drive (not shown). Then, when the program is executed, the program read from storage unit 103 is stored in the main storage area of ​​storage unit 103. In this state, control unit 101 executes the processing as each of information processing units 20 of this embodiment in accordance with the program stored in storage unit 103.

[0085] In the above description, a program describing the processing in each of the information processing units 20 is provided using a CD-ROM or DVD-ROM as a recording medium, but this is not limited to this. Depending on the configuration of the computer system, the capacity of the program to be provided, etc., it is also possible to use a program provided via a transmission medium such as the Internet via the communication unit 105.

[0086] The program of this embodiment causes a computer system to execute, for example, the steps of acquiring a question regarding facility usage from a facility user who is a user of the facility, acquiring additional information indicating the status of the facility user who asked the question, generating an answer to the question using the additional information, the question, and facility data that is data regarding the facility, and outputting the answer.

[0087] The preprocessing unit 112 and the answer generation unit 14 shown in FIG. 1 are realized by the control unit 101 shown in FIG. 15 executing a program stored in the storage unit 103 shown in FIG. 15. The storage unit 103 is also used to realize the preprocessing unit 112 and the answer generation unit 14. The data acquisition unit 111 shown in FIG. 1 is realized by at least one of the communication unit 105 and the input unit 102 shown in FIG. 15. Some functions of the data acquisition unit 111 may be realized by the control unit 101 and the storage unit 103. The data storage unit 113 shown in FIG. 1 is part of the storage unit 103 shown in FIG. 15. The information processing unit 20 may be realized by multiple computer systems. For example, the information processing unit 20 may be realized by a cloud computer system.

[0088] Furthermore, as described above, the question acquisition unit 12 of this embodiment is realized by at least one of an input means such as a touch panel or a keyboard, and a microphone. As described above, the situation acquisition unit 13 of this embodiment is realized by at least one of a sensor such as a camera, an input means such as a touch panel or a keyboard, and a microphone. The output unit 15 is realized by at least one of a display such as a touch panel, and a speaker. As described above, if the display that realizes the output unit 15 includes a display and is a touch panel, the touch panel may also function as the question acquisition unit 12. Furthermore, the touch panel may also function as the situation acquisition unit 13. When the question acquisition unit 12 and the situation acquisition unit 13 perform processing, the control unit 101 and the memory unit 103 in the computer system described above may be used to realize them. When the question acquisition unit 12, the situation acquisition unit 13, the information processing unit 20, and the output unit 15 are integrated, the entire automatic response system 10 may be regarded as the computer system illustrated in FIG. 15 . In this case, for example, the computer system may include a microphone as the input unit 102, and may further include a sensor as the situation acquisition unit 13. In this case, for example, the question acquisition unit 12 is realized by the input unit 102, the situation acquisition unit 13 is realized by at least one of the sensor and the input unit 102, and the output unit 15 is realized by at least one of the display unit 104 and the output unit 106.

[0089] As described above, the automated answering system 10 of this embodiment generates an answer to a question from the questioner 30 using the situation of the questioner 30. This allows a more appropriate answer to be provided to the facility user than simply answering a specific question from the facility user. Furthermore, the automated answering system 10 may generate an answer based on attributes in addition to the situation, and in this case, it is also possible to provide a facility user with a more appropriate answer according to the attributes. Furthermore, by including dynamic facility data as facility data, it is possible to provide the facility user with appropriate answers and information based on the current status of the facility.

[0090] Second Embodiment Fig. 16 is a diagram showing an example of the configuration of an automatic response system according to a second embodiment. The automatic response system 10a of this embodiment includes an automatic response device 20a and a terminal device 60. The automatic response device 20a is similar to the information processing unit 20 of the first embodiment, except that a transmitting / receiving unit 16 is added. The terminal device 60 includes a question acquisition unit 12, a situation acquisition unit 13, an output unit 15, and a transmitting / receiving unit 17. Components having the same functions as those of the first embodiment are assigned the same reference numerals as those of the first embodiment, and redundant explanations will be omitted. Below, differences from the first embodiment will be mainly explained.

[0091] In the first embodiment, an example has been described in which a facility user present in a facility uses the automatic answering system 10 installed in the facility. In this embodiment, a user who is about to use the facility is also treated as a facility user. The terminal device 60 is a device that can be operated by the facility user, such as a smartphone, tablet, or personal computer. The terminal device 60 may also be a mobile terminal that can be carried by the facility user. The facility user uses the automatic answering system 10a by using the terminal device 60, for example, at home or on the way to the facility. Note that the location where the facility user uses the automatic answering system 10a is not limited to this and may be any location, even within the facility.

[0092] The question acquisition unit 12 of the terminal device 60 acquires a question about the facility from a facility user by at least one of voice and text, similar to the question acquisition unit 12 of the first embodiment. Note that, similar to the first embodiment, the question may include an image. The question acquisition unit 12 outputs the acquired question to the transceiver unit 17. The question acquisition unit 12 is, for example, at least one of an input means such as a touch panel or keyboard of a smartphone, tablet, personal computer, or the like, and a microphone. When an image is included in the question, for example, the facility user specifies image data stored in the terminal device 60, and the input means accepts the specification. Note that, when a question is input as voice, the question acquisition unit 12 may convert the voice data into text data and output it to the transceiver unit 17, or may output the voice data directly to the transceiver unit 17.

[0093] Similar to the situation acquisition unit 13 in the first embodiment, the situation acquisition unit 13 acquires additional information indicating the situation (the situation of the questioner) of the questioner 30, who is a facility user asking a question. Similar to the first embodiment, the additional information may further include information indicating the attributes of the facility user. The situation acquisition unit 13 may acquire the additional information in the form of at least one of audio and text, or may acquire image data of the questioner 30 as the additional information. Furthermore, for example, the output unit 15 may output a screen to the questioner 30 asking questions such as the number of people using the facility and whether or not they will be using assistive devices, and acquire answers entered on the screen as the additional information. The situation acquisition unit 13 may be realized, for example, by a camera built into the terminal device 60, such as a smartphone, tablet, or personal computer, and may output video data acquired by the camera to the transceiver 17 as the additional information. This camera may be, for example, an internal camera that captures the questioner from inside the terminal device 60.

[0094] The output unit 15 outputs the answer by at least one of audio output and display, similar to the output unit 15 in the first embodiment. The output unit 15 is, for example, at least one of a display of the terminal device 60 and a microphone of the terminal device 60, but is not limited thereto. The display of the terminal device 60 may be a touch panel. In this case, the touch panel may also function as at least one of the question acquisition unit 12 and the situation acquisition unit 13.

[0095] The transmitting / receiving unit 17 communicates with the automatic answering device 20a to exchange information with the automatic answering device 20a. For example, the transmitting / receiving unit 17 transmits the question received from the question acquiring unit 12 and the additional information received from the situation acquiring unit 13 to the automatic answering device 20a. The transmitting / receiving unit 17 also outputs the answer received from the automatic answering device 20a to the output unit 15.

[0096] The operations of the question acquisition unit 12, the situation acquisition unit 13, the output unit 15, and the transmission / reception unit 17 may be performed by installing application software that provides a facility guidance service in the terminal device 60, or by the terminal device 60 accessing the automatic answering device 20a. For example, the automatic answering device 20a may have a function as a web server, and the above operations may be realized by the terminal device 60 accessing the web server.

[0097] The transmitter / receiver 16 of the automatic answering device 20a communicates with the terminal device 60 to exchange information with the terminal device 60. For example, the transmitter / receiver 16 outputs a question and additional information received from the terminal device 60 to the answer generation unit 14. The transmitter / receiver 16 also outputs an answer received from the answer generation unit 14 to the terminal device 60. The answer generation unit 14, like the answer generation unit 14 of the first embodiment, generates an answer using the question, additional information, and facility data, and outputs the generated answer to the transmitter / receiver 16. The automatic answering device 20a of this embodiment is realized by a computer system, like the information processing unit 20 of the first embodiment. The automatic answering device 20a may be installed inside or outside the facility.

[0098] In this embodiment, facility users can ask questions using the terminal device 60, allowing them to use the automatic response system 10a regardless of their location. Therefore, for example, before using a facility, they can obtain information about the facility in advance, either at home or on the way to the facility. This allows facility users to use the facility efficiently based on the information they obtain in advance. Furthermore, by including map information for generating a route to the facility in the facility data, when a facility user asks how to get to the facility, it is possible to present a route from the facility user's current location to the facility as a response. Furthermore, for example, facility data may also include data on reported lost items. Thus, if a facility user suspects that they may have left something behind at the facility after using the facility, they can inquire about the lost item by indicating where and what they left behind, and the automatic response system 10a can respond as to whether the corresponding lost item has been reported.

[0099] 17 and 18 are diagrams showing specific examples of responses that take the situation into consideration in the automatic response system 10a of this embodiment. In both of the examples shown in Fig. 17 and Fig. 18, it is assumed that it is raining around the facility and that data indicating the weather is included in the facility data.

[0100] In the example shown in FIG. 17 , the facility is a department store, and a questioner 30, who is planning to visit the department store, uses a terminal device 60 to ask a question before arriving at the department store. In the example shown in FIG. 17 , it is raining, and the questioner 30 is holding an umbrella. The terminal device 60 accepts the question, which is voice data, "Please tell me the way to the department store," and transmits the question and accompanying information, which is video data of the questioner 30, to the automatic response device 20 a. Using the question, accompanying information, and facility data, the automatic response device 20 a determines that the questioner 30 is holding an umbrella even though it is raining, and therefore should be guided along a normal route. The automatic response device 20 a generates a normal route to the department store (e.g., the shortest route) as an answer and transmits the answer to the terminal device 60. The terminal device 60 displays the route to the department store as an answer. 17, the questioner 30 is holding an umbrella, but this is not limiting. Similarly, if the questioner 30 is holding a closed umbrella, the automatic response device 20a may generate a response to guide the questioner 30 along a normal route even if it is raining. In other words, the automatic response device 20a may generate a response to guide the questioner 30 along a normal route as long as the questioner 30 is holding an umbrella, regardless of the state of the umbrella.

[0101] In the example shown in FIG. 18 , the facility is a conference hall where a briefing session is scheduled. In the example shown in FIG. 18 , the questioner 30 is dressed formally and does not have an umbrella. The terminal device 60 receives a question in the form of voice data, such as "Please tell me the way to the briefing session venue." The terminal device 60 transmits the question and accompanying information, which is video data of the questioner 30, to the automatic answering device 20 a. Using the question, accompanying information, and facility data, the automatic answering device 20 a determines that it is raining, the questioner 30 is not carrying an umbrella, and is dressed formally, and therefore should guide the questioner 30 along a covered route that will keep him / her dry. The automatic answering device 20 a generates an answer that indicates a route that will keep him / her dry and transmits the answer to the terminal device 60. The terminal device 60 displays the route to the conference hall as the answer. As illustrated in FIGS. 17 and 18 , the automatic answering device 20 a can provide an answer appropriate for the questioner 30 by determining a route to the facility based on the situation of the questioner 30.

[0102] As described above, in this embodiment, a facility user asks a question using terminal device 60, and terminal device 60 outputs a response. Therefore, facility users can use automatic response system 10a even when they are outside the facility.

[0103] In addition, a transmission / reception unit 16 may be added to the automatic response system 10 shown in embodiment 1, and the transmission / reception unit 16 may communicate with the terminal device 60, thereby making it possible to respond to both questions from facility users within the facility described in embodiment 1 and questions using the terminal device 60 described in this embodiment.

[0104] Third Embodiment Fig. 19 is a diagram showing an example of the configuration of an automatic response system according to a third embodiment. The automatic response system 10b of this embodiment is similar to the automatic response system 10 of the first embodiment, except that a rating acquisition unit 18 is added and the data storage unit 113 further stores rating information. Components having the same functions as those of the first embodiment are assigned the same reference numerals as those of the first embodiment, and redundant explanations will be omitted. Below, differences from the first embodiment will be mainly explained.

[0105] In this embodiment, after the exchange with the questioner 30 is completed, the evaluation acquisition unit 18 acquires an evaluation result indicating an evaluation of the answer from the questioner 30 and stores the acquired evaluation result in the data storage unit 113 along with the corresponding question and answer. Here, an example in which evaluation information is stored in the data storage unit 113 is described. However, this is not limiting. An evaluation information storage unit that stores evaluation information separately from the data storage unit 113 may be provided. For example, after outputting the answer, the output unit 15 outputs a question inquiring about the evaluation of the answer from the automatic response system 10b, such as "Did you get the information you wanted?" or "Was the answer appropriate?" Furthermore, if the evaluation from the questioner 30 includes negative words such as "unsatisfactory" or "not good," the evaluation acquisition unit 18 may cause the output unit 15 to output a question further asking the questioner 30 what the problem was and accept input from the questioner 30. The evaluation acquisition unit 18 may acquire the evaluation result by voice or as text data. Alternatively, the output unit 15 may present options indicating the evaluation result, and the evaluation acquisition unit 18 may acquire the selection made by the questioner 30 as the evaluation result.

[0106] The evaluation acquisition unit 18 may also evaluate the content of the answer based on the behavior of the questioner 30. For example, after the automatic response system 10b outputs a route to a destination to the questioner 30 who has asked for directions, it may analyze the behavior of the questioner 30 using video data captured of the questioner 30, and if the questioner 30 is still lost, it may evaluate the answer as being difficult to understand.

[0107] The evaluation information may be used for re-learning or additional learning by the answer generation unit 14 together with the corresponding questions and answers, for example, or may be referred to when the answer generation unit 14 generates an answer. Alternatively, the operator, manager, etc. of the automatic response system 10b may check the evaluation information, identify areas for improvement using the results of the check, and instruct the identified areas for improvement to the answer generation unit 14. Alternatively, the operator, manager, etc. of the automatic response system 10b may check the evaluation information, use the results of the check, to decide what type of re-learning or additional learning to perform for the answer generation unit 14, and the re-learning or additional learning may be performed using the determined results.

[0108] The evaluation acquisition unit 18 is, for example, at least one of an input means by operation and a microphone, similar to the question acquisition unit 12. The input means, the microphone, and other hardware may be shared with the question acquisition unit 12.

[0109] As described above, in this embodiment, the evaluation from the questioner 30 is acquired and reflected in the answer, thereby improving the accuracy of the answer. Note that the evaluation acquisition unit 18 may be added to the automatic answering system 10a of the second embodiment so that the evaluation result is reflected, or the evaluation acquisition unit 18 may be added to an automatic answering system that combines the first and second embodiments so that the evaluation result is reflected.

[0110] The configurations shown in the above embodiments are merely examples, and may be combined with other known technologies, or different embodiments may be combined with each other. It is also possible to omit or modify parts of the configurations as long as they do not deviate from the gist of the invention.

[0111] 10, 10a, 10b Automatic response system, 11 Data management unit, 12 Question acquisition unit, 13 Situation acquisition unit, 14 Answer generation unit, 15 Output unit, 16, 17 Transmitting / receiving unit, 18 Evaluation acquisition unit, 20 Information processing unit, 20a Automatic response device, 30 Questioner, 40 Main unit, 50 Camera, 60 Terminal device, 111 Data acquisition unit, 112 Preprocessing unit, 113 Data storage unit, 141 Situation recognition unit, 142 Generation unit.

Claims

1. An automatic response system comprising: a question acquisition unit that acquires a question regarding facility usage from a facility user who is a user of the facility; a situation acquisition unit that acquires additional information indicating the situation of the facility user who asked the question; an answer generation unit that generates an answer to the question using the question, the additional information, and facility data that is data regarding the facility; and an output unit that outputs the answer.

2. The automatic response system described in claim 1, characterized in that the answer generation unit is capable of generating the answer using multiple types of data including voice, text, and images as input, and is capable of outputting the answer in multiple formats.

3. The automatic response system described in claim 1 or 2, characterized in that the situation includes at least one of what the questioner is wearing, whether the questioner is using an assistive device, the number of questioners, the questioner's emotions, and the size of the luggage the questioner is carrying.

4. The automatic answering system according to any one of claims 1 to 3, wherein the additional information includes video data of the questioner.

5. An automatic response system as described in any one of claims 1 to 4, characterized in that the facility data includes static data that does not change for a certain period of time and dynamic data that may change even during the certain period of time.

6. The automated response system of claim 5, wherein the dynamic data includes data acquired by sensors installed within and / or around the facility.

7. The automatic response system according to claim 5 or 6, wherein the dynamic data includes facility-originated information issued by at least one of the facility manager and the facility employee.

8. An automatic answering system according to any one of claims 5 to 7, wherein the dynamic data includes store-originated information issued by an employee of a store within the facility.

9. The automated response system according to any one of claims 5 to 8, wherein the dynamic data includes weather information.

10. An automatic response system as described in any one of claims 1 to 9, characterized in that the answer generation unit comprises: a situation recognition unit that recognizes the situation based on the additional information; and a generation unit that generates the answer to the question using the situation, the question, and facility data.

11. The automatic response system described in claim 10, characterized in that the additional information further indicates the attributes of the questioner, the situation recognition unit recognizes the situation and the attributes based on the additional information, and the generation unit generates the answer to the question using the question, the situation, the attributes, and facility data.

12. The automated response system according to claim 11, wherein the attributes include at least one of gender, age, generation, presence or absence of a disability, and occupation.

13. The automatic response system described in claim 11 or 12, characterized in that the facility users include predetermined specific people, the facility data includes individual information that is information regarding the use of the facility by each of the specific people, the situation acquisition unit acquires personal authentication information for identifying the individual, and the answer generation unit identifies the individual using the personal authentication information and generates the answer using the corresponding individual information based on the identification result.

14. The automatic response system described in claim 12 or 13, characterized in that the attributes include at least one of status, affiliation, and position; the facility data includes attribute-related information, which is information regarding use of the facility for each of the attributes, and individual information, which is information regarding use of the facility for each specific person; the individual information includes the attributes of the corresponding individual; the situation acquisition unit acquires personal authentication information for identifying the individual; and the answer generation unit identifies the individual using the personal authentication information, recognizes the attributes of the identified individual based on the identification result and the individual information, and generates the answer using the attribute-related information corresponding to the attributes.

15. An automatic response system as described in any one of claims 1 to 14, characterized in that the answer generation unit generates the answer including basis information indicating the basis for the answer based on the situation used to generate the answer.

16. An automatic response system as described in any one of claims 1 to 15, characterized in that it comprises: a pre-processing unit that converts the facility data into data in a format accessible to the answer generation unit; and a data storage unit that stores the facility data after the conversion by the pre-processing unit, wherein the answer generation unit generates the answer using the facility data stored in the data storage unit.

17. An automatic answering system according to any one of claims 1 to 16, characterized in that the answer generation unit generates answers that are not dependent on gender.

18. An automatic response system as described in any one of claims 1 to 17, characterized in that the answer generation unit generates an answer that asks the questioner to ask again when it cannot recognize the question or when there is insufficient information to generate the answer.

19. An automatic answering system as described in any one of claims 1 to 18, characterized in that it comprises an evaluation acquisition unit that acquires an evaluation result of the answer from the questioner after the answer is output.

20. An automatic answering system as described in any one of claims 1 to 19, comprising: a terminal device operable by the questioner and including the question acquisition unit, the situation acquisition unit, and the output unit; and an automatic answering device including the answer generation unit, wherein the terminal device transmits the additional information and the question to the automatic answering device, the automatic answering device generates the answer using the facility data and the additional information and the question received from the terminal device, and transmits the generated answer to the terminal device, and the terminal device outputs the answer.

21. An automatic response method in an automatic response system, comprising the steps of: acquiring a question regarding facility usage from a facility user who is a user of the facility; acquiring additional information indicating the status of the facility user who asked the question; generating an answer to the question using the question, the additional information, and facility data that is data regarding the facility; and outputting the answer.

22. A program that causes a computer system to execute the following steps: acquiring a question regarding facility usage from a facility user who is a user of the facility; acquiring additional information indicating the status of the facility user who asked the question; generating an answer to the question using the question, the additional information, and facility data that is data regarding the facility; and outputting the answer.

Citation Information

Patent Citations

  • Guide display system, guide display method and guide display program

    JP2017220181A

  • Object Management System

    JP7411303B1