Program, information processing device, and method
The voice data providing system addresses passive lifestyles in elderly care settings by generating personalized audio content based on user interactions, enhancing self-efficacy and promoting independent living while reducing loneliness and preventing dementia.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-11-21
- Publication Date
- 2026-03-04
AI Technical Summary
Elderly individuals in nursing homes or hospitals often experience passive lifestyles due to one-sided care, leading to reduced social participation, loneliness, and decreased self-efficacy, which can increase the need for care and potentially contribute to conditions like dementia.
A voice data providing system that interacts with users, primarily the elderly, by asking questions, receiving answers, and generating personalized audio data in a radio program format, which is then played back to enhance self-efficacy and facilitate independent living, with the option to share user status data with external parties if necessary.
The system improves self-efficacy and enables independent living by providing personalized audio content, reducing feelings of loneliness and potentially preventing conditions like dementia through increased social participation and external collaboration.
Smart Images

Figure 0007824485000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a program, an information processing device, and a method. [Background technology]
[0002] With the increasing elderly population, various technologies are being offered to diagnose the cognitive functions of the elderly and prevent dementia.
[0003] Furthermore, against the backdrop of an increase in single-person elderly households or households consisting of only elderly people, various technologies have been provided for monitoring the elderly and detecting the onset of diseases such as dementia.
[0004] Patent Document 1 discloses technology such as an elderly care system that monitors the elderly and provides a dementia prevention method called "reminiscence therapy" to elderly people and other people who are targets of dementia prevention measures, allowing them to listen to reminiscence information related to the "reminiscence therapy" and accepting answers to questions. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Patent Publication No. 2021-131832 Summary of the Invention [Problem to be solved by the invention]
[0006] Regardless of whether they have dementia or not, when elderly people enter nursing homes or hospitals, they are essentially subject to the facility's management and receive one-sided care, leading to a passive lifestyle. While serviced housing for the elderly, which offers meals and other services, is available for single-person or elderly-only households, these services also lead to a passive lifestyle. The same is true for elderly people living with family or using day care services at nursing homes. While these services and care are convenient for elderly people, they tend to lose the ability to live independently and make their own decisions. This can lead to feelings of loneliness due to reduced social participation, which in turn reduces self-efficacy and ultimately increases the need for care.
[0007] Therefore, this disclosure describes a technology that provides predetermined audio data, allows people to listen to that audio data, improves self-efficacy, and enables collaboration with external parties when necessary. [Means for solving the problem]
[0008] According to one embodiment of the present disclosure, there is provided a program for providing predetermined voice data to a user when executed by a computer including a processor and a memory, wherein the memory stores a database for storing personal data relating to the user. The program causes the processor to repeatedly execute the following steps: providing first question data to the user at a first timing for each predetermined period of time for asking the user questions about the user; accepting input of answers to the questions about the user from the user and acquiring first answer data indicating the content of the answer; extracting personal data about the user from the acquired first answer data and storing it in a database; and generating first question data for asking the user further questions, if necessary, based on the first answer data; generating audio data edited in a predetermined format to be provided to the user, based on the personal data about the user; providing the generated audio data to the user at a second timing for each predetermined period of time that is different from the first timing; acquiring playback status data indicating the playback status of the user for the provided audio data; generating user status data indicating the user's status, based on the personal data about the user and the acquired playback status data about the user; and outputting the personal data about the user and the generated user status data to an external output destination, if necessary, depending on the content of the generated user status data. [Effects of the Invention]
[0009] According to the present disclosure, questions about a user are asked, answers to the questions are received, and audio data edited in a predetermined format is generated based on personal data extracted from the answer data and provided to the user. Furthermore, depending on the content of user status data based on the playback status of the audio data, the personal data and user status data are output to an external output destination when necessary. Therefore, by having the user listen to the audio data, it is possible to improve self-efficacy and to collaborate with external parties when necessary. This helps the user maintain an independent lifestyle based on their own judgment, making it possible to prevent dementia and other conditions. [Brief explanation of the drawings]
[0010] [Figure 1] 1 is a block diagram showing the overall configuration of a voice data providing system 1 according to an embodiment of the present disclosure. [Figure 2] 2 is a block diagram showing a functional configuration of the terminal device 10 of FIG. 1. FIG. [Figure 3] 2 is a block diagram showing the functional configuration of the server 20 of FIG. 1. FIG. [Figure 4] FIG. 4 is a diagram showing an example of the data structure of a user database 2021 in FIG. 3. [Figure 5] FIG. 4 is a diagram showing an example of the data structure of a personal database 2022 in FIG. 3. [Figure 6] 10 is a flowchart showing an example of the flow of a voice data generation process performed by the voice data providing system 1. [Figure 7] 10 is a flowchart showing an example of the flow of a voice data providing process performed by the voice data providing system 1. [Figure 8] 10 is a diagram showing an example of a screen displaying questions and answers displayed on the terminal device 10. FIG. [Figure 9] FIG. 10 is a diagram showing an example of a screen for providing voice data displayed on the terminal device 10. [Figure 10] FIG. 2 is a block diagram showing the basic hardware configuration of a computer 90. DETAILED DESCRIPTION OF THE INVENTION
[0011] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. In all drawings illustrating the embodiments, common components are designated by the same reference numerals, and repeated explanations will be omitted. Note that the following embodiments do not unduly limit the content of the present disclosure described in the claims. Furthermore, not all of the components shown in the embodiments are necessarily essential components of the present disclosure. Furthermore, each drawing is a schematic diagram and is not necessarily a precise illustration.
[0012] In the following description, a "processor" refers to one or more processors. The at least one processor is typically a microprocessor such as a CPU (Central Processing Unit), but may also be another type of processor such as a GPU (Graphics Processing Unit). The at least one processor may be single-core or multi-core.
[0013] Furthermore, the at least one processor may be a processor in the broad sense, such as a hardware circuit (for example, a field-programmable gate array (FPGA) or an application-specific integrated circuit (ASIC)) that performs part or all of the processing.
[0014] In the following explanation, information that produces an output in response to an input may be described using the expression "xxx database," but this information may be data of any structure, or may be a learning model such as a neural network that produces an output in response to an input. Therefore, "xxx database" may be referred to as "xxx information."
[0015] Furthermore, in the following description, the configuration of the tables that make up each database is an example, and one table may be divided into two or more tables, or all or part of two or more tables may be one table.
[0016] In addition, in the following explanation, processing may be described using the "program" as the subject, but since a program is executed by a processor to perform specified processing while appropriately using a memory unit and / or an interface unit, etc., the subject of the processing may also be the processor (or a device such as a controller that has that processor).
[0017] The program may be installed in a device such as a computer, or may be stored in, for example, a program distribution server or a computer-readable (e.g., non-transitory) recording medium. Also, in the following description, two or more programs may be realized as one program, or one program may be realized as two or more programs.
[0018] Furthermore, in the following description, identification numbers are used as identification information for various objects, but other types of identification information (for example, identifiers including alphabetic characters or symbols) may also be used.
[0019] In addition, in the following description, when describing elements of the same type without distinguishing between them, reference symbols (or common symbols among the reference symbols) may be used, and when describing elements of the same type with distinction between them, the identification numbers (or reference symbols) of the elements may be used.
[0020] In the following description, the control lines and information lines are those that are considered necessary for the description, and do not necessarily represent all the control lines and information lines in the product. All components may be interconnected.
[0021] <Summary> The following describes a voice data providing system according to the present disclosure. The voice data providing system according to the present disclosure asks questions to users, primarily elderly people, and receives answers. Based on personal data extracted from the answer data, the system generates and provides edited voice data in a predetermined format, such as a radio program format. The voice data providing system reflects the user's personal data, such as yesterday's events and today's schedule, in voice data in a radio program format, and provides the generated voice data to the user. The voice data providing system according to the present disclosure is a system provided as a web service, for example, by a cloud server or the like, as a so-called SaaS (Software as a Service), and is configured to be accessible by users through predetermined authentication.
[0022] As mentioned above, when elderly people move into senior housing with support services, nursing homes, or hospitals, they basically follow the management of the facility and live a passive life with one-sided care. While this provides great convenience for the elderly receiving services and care, it also tends to prevent them from living independently based on their own decisions. This leads to a decrease in social participation for the elderly, which can lead to feelings of loneliness, a decrease in self-efficacy, and, as a result, an increase in the level of care required.
[0023] Therefore, the voice data providing system according to the present disclosure uses, for example, AI to ask questions about the user and accept input of answers in an interactive format. Such dialogue is performed at a first timing every predetermined period (for example, every day). During such interaction, the user's personal data, such as yesterday's events and today's schedule, is extracted, and this is reflected in voice data in a format similar to a radio program, which is then generated and provided to the user at a second timing every predetermined period (for example, every day). Then, depending on the content of the user status data based on the playback status of the voice data, the personal data and user status data are output (as intervention data) to an external output destination, such as a relative, a medical institution, or a care facility, if necessary.
[0024] Furthermore, the voice data providing system according to the present disclosure outputs the user's personal data, for example, schedule data, as a reminder at a third timing different from the timing at which the voice data is provided.
[0025] Furthermore, the audio data providing system according to the present disclosure accepts input such as a request from the user or an external person (for example, a relative) for audio data in a format such as a radio program, and reflects the request in the audio data.
[0026] This configuration allows users to listen to audio data, improving their sense of self-efficacy, and also allows for external collaboration when necessary. This helps users maintain an independent lifestyle based on their own judgment, making it possible to prevent dementia and other conditions.
[0027] First Embodiment The following describes an audio data providing system 1 according to an embodiment of the present disclosure. In the following description, for example, when a terminal device 10 accesses a server 20, the server 20 responds with information for generating a screen on the terminal device 10. The terminal device 10 generates and displays a screen based on the information received from the server 20.
[0028] <1 Overall configuration of voice data provision system 1> FIG. 1 is a block diagram showing the overall configuration of a voice data providing system 1 according to a first embodiment of the present disclosure. As shown in FIG. 1, the voice data providing system 1 includes a plurality of terminal devices (terminal device 10A and terminal device 10B are shown in FIG. 1; hereinafter, they may be collectively referred to as "terminal device 10"), a server 20, and an external server 30. The terminal devices 10, the server 20, and the external server 30 are connected to each other via a network 80 so that they can communicate with each other. The network 80 may be a wired or wireless network. Examples of the network 80 include 4G and 5G mobile communication systems, LTE (Long Term Evolution), and wireless networks (e.g., Wi-Fi (registered trademark)) that can connect to the Internet via a predetermined access point. In the case of a wireless connection, the network 80 may use communication protocols such as Z-Wave (registered trademark), ZigBee (registered trademark), and Bluetooth (registered trademark). In addition, in the case of a wired connection, the network may also include a network that is directly connected using a USB (Universal Serial Bus) cable or the like.
[0029] In this embodiment, the server 20 is a web server (including a cloud server), and exchanges information with the terminal device 10 via web pages. A web page browser for viewing web pages is installed on the terminal device 10, but a dedicated application for providing the services of the server 20 may also be installed so that the pages can be viewed using the dedicated application. In this embodiment, the voice data providing system 1 is described as being configured such that the terminal device 10 and the server 20 are connected via a network 80, but it may also be configured on-premise using various types of individual computer devices or the like.
[0030] The terminal device 10 is a device operated by each user. Here, a user is a person who uses the terminal device 10 to answer questions, which are functions of the voice data providing system 1, and receives voice data, such as elderly people who are the target users of the voice data providing system 1. The terminal device 10 is realized by a desktop PC (Personal Computer), a laptop PC (notebook PC), or the like. Alternatively, the terminal device 10 may be a tablet compatible with a mobile communication system, a mobile terminal such as a smartphone, or the like.
[0031] The terminal device 10 is communicatively connected to the server 20 via a network 80. The terminal device 10 is connected to the network 80 by communicating with communication devices such as a wireless base station 81 conforming to communication standards such as 4G, 5G, and LTE (Long Term Evolution), and a wireless LAN router 82 conforming to a wireless LAN (Local Area Network) standard such as IEEE (Institute of Electrical and Electronics Engineers) 802.11. As shown as a terminal device 10B in FIG. 1 , the terminal device 10 includes a communication IF (Interface) 12, an input device 13, an output device 14, a memory 15, a storage unit 16, and a processor 19.
[0032] The communication IF 12 is an interface for inputting and outputting signals so that the terminal device 10 can communicate with external devices. The input device 13 is an input device (e.g., a keyboard, a touch panel, a touch pad, a pointing device such as a mouse, etc.) for receiving input operations from a user. The output device 14 is an output device (e.g., a display, a speaker, etc.) for presenting information to a user. The memory 15 is for temporarily storing programs and data processed by the programs, etc., and is a volatile memory such as a DRAM (Dynamic Random Access Memory). The storage unit 16 is a storage device for saving data, such as a flash memory or an HDD (Hard Disc Drive). The processor 19 is hardware for executing an instruction set written in a program, and is composed of an arithmetic unit, registers, peripheral circuits, etc.
[0033] Server 20 is a device that provides predetermined voice data to users, primarily targeting elderly people. It generates and provides edited voice data in a predetermined format, such as a radio program format. Server 20 uses, for example, AI to interactively ask questions about the user and accept answers, extracts the user's personal data, and generates and provides voice data in a radio program format. Server 20 also outputs the personal data and user status data to an external output destination as needed, depending on the content of the user status data based on the playback status of the voice data, etc.
[0034] The server 20 is a computer connected to a network 80. The server 20 includes a communication IF 22, an input / output IF 23, a memory 25, a storage 26, and a processor 29.
[0035] The communication IF 22 is an interface for inputting and outputting signals so that the server 20 can communicate with external devices. The input / output IF 23 functions as an interface with an input device for receiving input operations from a user and an output device for presenting information to the user. The memory 25 is for temporarily storing programs and data processed by the programs, etc., and is a volatile memory such as a DRAM (Dynamic Random Access Memory). The storage 26 is a storage device for saving data, such as a flash memory or an HDD (Hard Disc Drive). The processor 29 is hardware for executing an instruction set written in a program, and is composed of an arithmetic unit, registers, peripheral circuits, etc.
[0036] The external server 30 is, for example, a server device that provides a service (generative AI service) using a generative AI system. The external server 30 provides a prompt including text data to a large-scale language model, and outputs text data that is a structured document based on the text data. The external server 30 may be a large-scale language model trained with a large amount of text data, or a model obtained by transfer learning or fine-tuning the large-scale language model. Examples of large-scale language models include GPT-3, GPT-4, and GPT-5 developed by OpenAI, and Gemini developed by Google.
[0037] Note that the voice data providing system 1 according to the embodiment of the present disclosure uses a service provided by a generation AI system when analyzing response data in a user's natural language to a question, extracting personal data, and generating text data or voice data that will be the basis for the voice data. However, as will be described later, the server 20 may be provided with a language model 2023, and an AI service may be used by the language model 2023. In this case, the external server 30 may not be provided. Conversely, when an AI service is used by the external server 30, the language model 2023, which will be described later, may not be provided.
[0038] <1.1 Configuration of the terminal device 10> FIG. 2 is a block diagram showing the functional configuration of the terminal device 10 constituting the voice data providing system 1 of the first embodiment. As shown in FIG. 2, the terminal device 10 includes a plurality of antennas (antenna 111, antenna 112), wireless communication units (first wireless communication unit 121, second wireless communication unit 122) corresponding to the respective antennas, an operation reception unit 130 (including a touch-sensitive device 131 and a display 132), a position information sensor 140, a camera 150, a storage unit 160, and a control unit 170. The terminal device 10 also has functions and configurations (e.g., a battery for storing power, a power supply circuit for controlling the supply of power from the battery to each circuit, etc.) that are not specifically shown in FIG. 2. As shown in FIG. 2, the blocks included in the terminal device 10 are electrically connected by a bus or the like.
[0039] The antenna 111 emits a signal emitted by the terminal device 10 as a radio wave. The antenna 111 also receives a radio wave from space and provides the received signal to the first radio communication unit 121.
[0040] The antenna 112 emits a signal emitted by the terminal device 10 as a radio wave. The antenna 112 also receives a radio wave from space and provides the received signal to the second radio communication unit 122.
[0041] The first wireless communication unit 121 performs modulation / demodulation processing and the like for transmitting and receiving signals via the antenna 111 so that the terminal device 10 can communicate with other wireless devices. The second wireless communication unit 122 performs modulation / demodulation processing and the like for transmitting and receiving signals via the antenna 112 so that the terminal device 10 can communicate with other wireless devices. The first wireless communication unit 121 and the second wireless communication unit 122 are communication modules including a tuner, an RSSI (Received Signal Strength Indicator) calculation circuit, a CRC (Cyclic Redundancy Check) calculation circuit, a high-frequency circuit, etc. The first wireless communication unit 121 and the second wireless communication unit 122 perform modulation / demodulation and frequency conversion of wireless signals transmitted and received by the terminal device 10, and provide the received signals to the control unit 170.
[0042] The operation reception unit 130 has a mechanism for receiving input operations from the user. Specifically, the operation reception unit 130 is configured as a touch screen and includes a touch-sensitive device 131 and a display 132. The touch-sensitive device 131 receives input operations from the user of the terminal device 10. The touch-sensitive device 131 detects the user's touch position on the touch panel, for example, by using a capacitive touch panel. The touch-sensitive device 131 outputs a signal indicating the user's touch position detected by the touch panel to the control unit 170 as an input operation. The terminal device 10 may also be provided with a physical keyboard (not shown) that can be used for input, and may receive the user's input operations via the keyboard.
[0043] Display 132 displays data such as images, videos, and text under the control of control unit 170. Display 132 is realized by, for example, an LCD (Liquid Crystal Display) or an organic EL (Electro-Luminescence) display.
[0044] The position information sensor 140 is a sensor that detects the position of the terminal device 10, and is, for example, a GPS (Global Positioning System) module. The GPS module is a receiving device used in a satellite positioning system. In the satellite positioning system, signals are received from at least three or four satellites, and the current position of the terminal device 10 equipped with the GPS module is detected based on the received signals. The position information sensor 140 may be a transmitting / receiving device based on a communication standard used in a short-range communication system between information devices. Specifically, the position information sensor 140 uses the 2.4 GHz band, such as a Bluetooth (registered trademark) module, to receive beacon signals from other information devices equipped with a Bluetooth (registered trademark) module.
[0045] Camera 150 is a device that receives light with a light receiving element and outputs the received light as a captured image in accordance with the control of control unit 170. Camera 150 may be configured, for example, by an imaging device such as a digital camera or a video camera.
[0046] Storage unit 160 is configured with memory 15, such as a flash memory, and storage unit 16, and stores data and programs used by terminal device 10. In one aspect, storage unit 160 stores user information 161.
[0047] The user information 161 is information about a user who uses the terminal device 10 to answer questions and receive voice data, which is a function of the voice data providing system 1. The user information may include user identification data (such as a user ID) that identifies the user, the user's name, gender, date of birth, etc.
[0048] The control unit 170 is configured by, for example, the processor 19, and controls the operation of the terminal device 10 by reading a program stored in the storage unit 160 and executing instructions included in the program. The control unit 170 is, for example, an application that is pre-installed in the terminal device 10. The control unit 170 operates in accordance with the program to fulfill the functions of an input operation reception unit 171, a transmission / reception unit 172, a notification control unit 173, and a data processing unit 174.
[0049] The input operation receiving unit 171 performs processing to receive input operations by the user to an input device such as the touch-sensitive device 131 .
[0050] The transmitting / receiving unit 172 performs processing for the terminal device 10 to transmit and receive data to and from external devices such as the server 20 in accordance with a communication protocol.
[0051] The notification control unit 173 performs processing to present information to the user. The notification control unit 173 performs processing to display a display image on the display 132, etc.
[0052] The data processing unit 174 performs calculations on data that the terminal device 10 has received as input in accordance with a program, and outputs the calculation results to a memory or the like.
[0053] <1.2 Functional configuration of server 20> 3 is a diagram showing the functional configuration of the server 20 constituting the voice data providing system 1 of the embodiment 1. As shown in FIG. 3, the server 20 performs the functions of a communication unit 201, a storage unit 202, and a control unit 203.
[0054] The communication unit 201 performs processing for the server 20 to communicate with external devices.
[0055] The storage unit 202 stores data and programs used by the server 20. The storage unit 202 stores a user database 2021, a personal database 2022, a language model 2023, and the like.
[0056] The user database 2021 is a database for storing and holding various data relating to users, such as elderly people who are the target users of the voice data providing system 1, who use the voice data providing system 1 to answer questions and receive voice data. The user database 2021 stores, for example, user identification data (such as a user ID) that identifies the user, the user's name, gender, address (which may be current or past), career history, hobbies, date of birth, etc. These data are entered with the user's consent, and some data may not be included. Details will be described later.
[0057] The personal database 2022 is a database for storing and holding personal data related to users who use the voice data providing system 1. The personal database 2022 stores, for example, answer data to questions and personal data extracted from the answer data, specifically data such as yesterday's events and today's plans, linked to user identification data (such as a user ID) that identifies the user. Furthermore, the user's nickname or name used in a radio program, which will be described later, may be stored as personal data. Furthermore, the user's vital data, etc. may be stored as personal data. Details will be described later.
[0058] The language model 2023 is a language analysis model for providing a service (generative AI service) using a generative AI system that analyzes data of a user's natural language responses to questions, extracts personal data, and generates text data that will be the basis for voice data or voice data. When the language model 2023 receives input of a prompt including data of the response to the question, it outputs an intent analysis result and specific text data. The language model 2023 may be based on a large-scale language model with advanced natural language processing capabilities, such as GPT-3, GPT-4, or GPT-5 developed by OpenAI, or Gemini developed by Google.
[0059] The control unit 203 performs the functions shown in various modules, such as a reception control module 2031, a transmission control module 2032, a question data providing module 2033, an answer receiving module 2034, a question data generation module 2035, a voice data generation module 2036, a voice data providing module 2037, a playback status data acquisition module 2038, a user status data generation module 2039, a personal data providing module 2040, a question provision and acquisition module 2041, and an external output module 2042, by the processor 29 of the server 20 performing processing according to the program.
[0060] The reception control module 2031 controls the process by which the server 20 receives signals from external devices in accordance with a communication protocol.
[0061] The transmission control module 2032 controls the process in which the server 20 transmits signals to external devices in accordance with a communication protocol.
[0062] The question data providing module 2033 controls a process of providing the user with first question data for asking a question about the user at a first timing every predetermined period. The question data providing module 2033 transmits, for example, question data (first question data) indicating a predetermined question to the user, specifically text data of a question sentence, to the terminal device 10 used by the user via the communication unit 201 and displays the question on the display 132 of the terminal device 10. The question data providing module 2033 provides, as question data, text data of a dialogue-style question such as a question asking about today's events or tomorrow's plans, to be reflected in the content of the voice data provided by the voice data providing system 1. The question data providing module 2033 provides the question data to the user at a first timing (for example, a set time in the evening) every predetermined period (for example, every day).
[0063] For example, when a user uses the voice data providing system 1 for the first time, the question data providing module 2033 may provide text data of an interactive question in natural language that is set as a template in advance as question data. As will be described later, the question data providing module 2033 may provide text data of a question generated by the question data generation module 2035 after the answer receiving module 2034 receives an answer input from the user. Furthermore, when a user has used the voice data providing system 1 many times and personal data of the user has been accumulated in the personal database 2022, the question data providing module 2033 may provide text data of a question generated by the question data generation module 2035 (at the start of the relevant timing).
[0064] The question data providing module 2033 may send the question data via a communication tool such as LINE (registered trademark) or a chat tool, or via an application provided by the voice data providing system 1, or by email.
[0065] The answer receiving module 2034 receives from the user an answer to a question about the user provided by the question data providing module 2033, and controls processing for acquiring first answer data indicating the content of the answer. The user, for example, operates the terminal device 10 to input an answer to a question about the user and transmit it to the server 20. The answer receiving module 2034 receives the answer data (first answer data) indicating the content of the answer transmitted from the terminal device 10 by receiving it via the communication unit 201. The answer receiving module 2034 receives the user's answer to a question about the user provided by the question data providing module 2033, such as a question asking about today's events or tomorrow's plans.
[0066] The answer receiving module 2034 receives the text data entered as answer data when the user inputs an answer to a question as text. The answer receiving module 2034 may also receive image data such as photographs, stamps (image data that visualizes the response content), etc. In addition to text data, the answer receiving module 2034 may also receive voice data uttered by the user using the microphone function of the terminal device 10 (not shown).
[0067] The question data generation module 2035 controls the process of extracting personal data related to the user from the first answer data acquired by the answer receiving module 2034 and storing the extracted personal data in the personal database 2022. For example, the question data generation module 2035 performs natural language processing, which is an existing technology, on the text data that is the acquired answer data, analyzes the answer content, and extracts the answer content corresponding to the question content as personal data. Furthermore, the question data generation module 2035 performs speech recognition processing, which is an existing technology, on the audio data that is the acquired answer data, analyzes the answer content, and extracts the answer content corresponding to the question content as personal data. Specifically, if the question data is about an event that happened yesterday and the answer data is about going shopping, the question data generation module 2035 extracts "I went shopping yesterday" as personal data. The question data generation module 2035 then stores this personal data in the personal database 2022.
[0068] The question data generation module 2035 may, for example, extract the user's future schedule data, specifically schedule data such as "I will visit a friend's house tomorrow at 1:00 p.m." as personal data related to the user, and store it in the personal database 2022.
[0069] Furthermore, the question data generation module 2035 controls the process of generating first question data for asking the user further questions when necessary, based on the first answer data. For example, the question data generation module 2035 determines, based on the answer received from the user, whether the personal database 2022 has accumulated enough personal data about the user to enable the voice data generation module 2036 (described later) to generate voice data. If it is determined that this is possible, the voice data generation module 2036 performs processing. If it is determined that this is not possible (the above-mentioned "when necessary"), question data (first question data) for asking the user further questions is generated. When generating question data, the question data generation module 2035 performs, for example, natural language processing, which is an existing technology, to analyze the personal data and the answer content, and generate text data of a question in a dialogue format in natural language regarding the content acquired as personal data. Specifically, when the answer data indicates that the user went shopping, the question data generation module 2035 generates question data such as "What did you go out to buy?" or "Where did you go?" as further question data. The question data generation module 2035 repeats these processes until it is determined that voice data can be generated.
[0070] In the process of extracting personal data and generating question data, the question data generation module 2035 may use, for example, a service (a generation AI service) provided by a generation AI system by the language model 2023 or the external server 30. The question data generation module 2035 may generate a prompt for extracting personal data or generating question data from the personal data and the content of the answer, provide this to the AI service by the language model 2023 or the external server 30, and obtain the output data as personal data or question data. Note that the question data generation module 2035 may use only either the language model 2023 or the external server 30, or may use both depending on the processing content.
[0071] The audio data generation module 2036 controls the process of generating audio data edited in a predetermined format to be provided to the user based on personal data about the user stored in the personal database 2022. The audio data generation module 2036 generates audio data in a format similar to that of a radio program, for example, based on the personal data about the user. The audio data generation module 2036 generates audio data in any data format, such as AAC, ATRAC, mp3, or mp4. When generating audio data, the audio data generation module 2036 may refer to data stored in the user database 2021 (such as gender, career history, hobbies, and date of birth) or to answer data to past questions. Furthermore, the audio data generation module 2036 may refer to data related to the date on which music data, weather data, and audio data are provided (such as calendar data and events that occurred on that day) from web services such as other external servers.
[0072] The audio data in a radio program-like format generated by the audio data generation module 2036 may be, for example, a radio program in natural language personalized for each user. In this radio program, a personality (host) speaks about the user's personal data (the user's answers to questions), addresses the user by their name (not shown) included in the personal data, and includes content encouraging the user to reflect on the day's events and remember their upcoming plans (schedule). This audio data (radio program) uses positive language to enhance the self-efficacy of elderly users who are the target users of the audio data providing system 1, and is designed to instill positive emotions in the users. The audio data (radio program) also includes features like a typical radio program, such as a weather section, a quiz section, a voting section, an exercise section, and music playback. Among these sections, sections that obtain user responses are configured to grasp the user's interests and cognitive response characteristics from changes in the user's response tendencies and selection patterns. Content personalized for each user may also include a section for the user's own reminiscence or a section for brain training tailored to the user's age, etc. Therefore, the user's personal data is not converted directly into voice data.
[0073] Additionally, the audio data generated by the audio data generation module 2036 may include music and other data according to the user's preferences stored as personal data, and may be in multiple languages or dialects.
[0074] In one aspect, the audio data generation module 2036 may generate audio data based on past playback status data acquired by a playback status data acquisition module 2038 (described later) and the content of past user status data generated by a user status data generation module 2039. For example, if the user does not play (listen to) the audio data to the end for some reason, the audio data generation module 2036 may re-incorporate the content that was not played back, or may generate audio data without including the same content if it is determined that the content does not suit the user's preferences.
[0075] In a certain aspect, the voice data generation module 2036 may generate voice data based on the content of response data from the user to question data for the user, which is provided by the question providing and acquiring module 2041, which will be described later. For example, the question providing and acquiring module 2041 may ask the user questions such as impressions or questionnaires about the provided voice data (radio program), and when the user answers the questions, the voice data generation module 2036 may incorporate the content of the answers as feedback and generate future voice data.
[0076] Furthermore, in a certain aspect, the voice data generation module 2036 may generate voice data based on the contents of the reflection data input from the user or an external party and accepted by the question provision and acquisition module 2041, which will be described later. For example, the voice data generation module 2036 may incorporate the contents of a request accepted by the question provision and acquisition module 2041 from the user or from an external party (for example, a relative of the user) into future voice data.
[0077] In the process of generating text data or voice data that is the source of voice data, the voice data generation module 2036 may use, for example, the language model 2023 or a service (a generation AI service) provided by a generation AI system by the external server 30. The question data generation module 2035 may generate a prompt for generating text data or voice data from personal data, provide it to the language model 2023 or the AI service by the external server 30, and obtain output data as text data or voice data. Note that the voice data generation module 2036 may use only either the language model 2023 or the external server 30, or may use both depending on the processing content.
[0078] The voice data providing module 2037 controls a process of providing the voice data generated by the voice data generation module 2036 to the user at a second timing different from the first timing for each predetermined period. For example, the voice data providing module 2037 transmits the voice data generated by the voice data generation module 2036 to the terminal device 10 used by the user via the communication unit 201, and displays a notification on the display 132 of the terminal device 10 that the voice data has been received. The voice data providing module 2037 provides the voice data to the user at a second timing (for example, a set time the next morning) for each predetermined period (for example, every day).
[0079] The voice data providing module 2037 may transmit the voice data via a communication tool or chat tool such as LINE, via an application provided by the voice data providing system 1, or by email. The voice data providing module 2037 may also transmit and provide the voice data to a device such as an artificial intelligence speaker (smart speaker). The voice data providing module 2037 may transmit the voice data using the same means (channel) as when the question data providing module 2033 transmitted the question data, or may transmit the voice data using a different means.
[0080] Furthermore, the voice data providing module 2037 may transmit the voice data at a third timing different from the first timing and the second timing, for example, several hours after the second timing (again). At this time, the voice data provided at the second timing may be transmitted, or a portion of the voice data may be transmitted.
[0081] The playback status data acquisition module 2038 controls the process of acquiring playback status data indicating the playback status of the user for the audio data provided by the audio data providing module 2037. The playback status data acquisition module 2038 transmits an instruction signal to the terminal device 10, for example, to monitor the playback status of the audio data by the user and transmit the playback status data resulting from the monitoring. The terminal device 10 transmits the playback status data to the server 20 in accordance with the instruction signal, and the playback status data acquisition module 2038 accepts the playback status data transmitted from the terminal device 10 by receiving it via the communication unit 201. The playback status data acquisition module 2038 instructs the terminal device 10 to transmit the playback status data, for example, after a predetermined time has elapsed (for example, 12 hours) since the audio data providing module 2037 transmitted the audio data.
[0082] The playback status data acquisition module 2038, for example, monitors the playback status of audio data by the user, and acquires, as playback status data, one or more of the following data: data indicating whether the user has played the audio data, data indicating how far the user has played the audio data (in chronological order), and the time at which the user played the audio data.
[0083] The user status data generation module 2039 controls the process of generating user status data indicating the status of a user based on personal data about the user stored in the personal database 2022 and the playback status data of the user acquired by the playback status data acquisition module 2038. The user status data generation module 2039, for example, analyzes the personal data and playback status data to determine whether or not there is an abnormality with respect to the user and outputs the result.
[0084] Specifically, the user state data generation module 2039 may determine that there is something wrong with the user if the analysis result of the playback status data shows that the user has not played back audio data for a certain period of time. Also, the user state data generation module 2039 may determine that there is something wrong with the user if the analysis result of the personal data shows that the user has not responded to questions for a certain period of time. Furthermore, the user state data generation module 2039 may combine these data to analyze the user's living situation and determine whether there is anything wrong with the user.
[0085] The personal data providing module 2040 controls a process of providing personal data about the user, which is stored in the personal database 2022, to the user at a third timing different from the first timing and the second timing. For example, the personal data providing module 2040 provides the personal data to the user at the third timing, for example, several hours after the second timing.
[0086] For example, the personal data providing module 2040 may provide the user with future schedule data as personal data at a third timing, for example, a timing corresponding to the schedule.
[0087] The question providing and acquiring module 2041 controls a process of providing the user with second question data for asking a question about the voice data provided to the user by the voice data providing module 2037 at the second timing, at a third timing different from the first timing and the second timing. The question providing and acquiring module 2041 transmits, for example, question data (second question data) for asking the user a question about the voice data provided by the voice data providing module 2037, specifically, text data of a question sentence, to the terminal device 10 used by the user via the communication unit 201 and displays the question on the display 132 of the terminal device 10. The question providing and acquiring module 2041 provides, as the question data, text data of an interactive question such as an impression of the voice data provided by the voice data providing module 2037, a question such as a questionnaire, for example, whether today's voice data (radio program) was good / bad, areas for improvement, content that the user would like to see included, etc. The question providing and acquiring module 2041 provides the question data to the user at the third timing (for example, immediately after the second timing).
[0088] The question providing and acquiring module 2041 also receives input of an answer to a question related to the voice data from the user and controls processing for acquiring second answer data indicating the content of the answer. The user, for example, operates the terminal device 10 to input an answer to a question related to the voice data and transmit it to the server 20. The question providing and acquiring module 2041 receives the answer data (second answer data) indicating the content of the answer transmitted from the terminal device 10 by receiving it via the communication unit 201. The question providing and acquiring module 2041 acquires, as the answer data, text data of answers to questions such as impressions of the voice data provided by the voice data providing module 2037, whether today's voice data (radio program) was good / bad, areas for improvement, and content that the user would like to see included. This answer data is incorporated into the content of future voice data by the voice data generation module 2036 as described above.
[0089] Furthermore, the question providing and acquiring module 2041 receives input of content to be reflected in the voice data from the user or an external output destination, and controls the process of acquiring reflection data indicating the input content. The user or an external party (e.g., the user's relative) operates the terminal device 10, for example, to input content to be reflected in the voice data and transmit it to the server 20. The question providing and acquiring module 2041 receives the reflection data indicating the content to be reflected transmitted from the terminal device 10 via the communication unit 201. The question providing and acquiring module 2041 acquires, as reflection data, requests from the user, specifically text data such as the voice quality, gender, and speaking speed of the program personality, music to be played during the voice data program, and topics to be covered. The question providing and acquiring module 2041 may also be configured to acquire, as reflection data, things to be looked up, things to be searched for using a search service, etc., and reflect the results in the voice data. The question providing and acquiring module 2041 acquires, as reflection data, requests from the user's relatives, specifically text data such as questions to be asked of the user and confirmation items. This reflected data is incorporated into the content of future voice data by the voice data generation module 2036 as described above.
[0090] The external output module 2042 controls the process of outputting personal data about the user and the generated user status data to an external output destination when necessary, depending on the content of the user status data generated by the user status data generation module 2039. For example, the external output module 2042 transmits the user status data to an external output destination, such as an output destination registered as the user's relative, via the communication unit 201. The external output module 2042 also transmits the user status data to an external output destination, such as a medical institution or a welfare facility, or both, via the communication unit 201. The medical institution and welfare facility may be a facility registered as the user's regular doctor, or may be a facility close to the user's residential area.
[0091] The external output module 2042 may output, for example, personal data and user status data to a pre-registered external output destination having a predetermined relationship with the user, such as an output destination registered as a relative of the user. The external output module 2042 may transmit, for example, the playback status of audio data included in the playback status data as user status data to notify the relative. This is intended to provide a monitoring function for relatives and to notify abnormalities, and by notifying the relative that the elderly person in question is in good health, it is possible to provide a sense of security to the relative.
[0092] The external output module 2042 may, for example, select an external output destination to which the personal data and user status data are output according to the content of the user status data, and output the data to the selected external output destination. The external output module 2042 may, for example, determine the user's abnormal state, select an external output destination (a medical institution or a welfare facility) corresponding to the abnormality, and output the data to the selected medical institution or welfare facility. This is for the purpose of notifying the medical institution or welfare facility of the abnormality. At this time, the external output module 2042 may extract intervention data to be output to the external output destination (a medical institution or a welfare facility) from the personal data and the user status data, and output the extracted intervention data to the external output destination. Furthermore, the external output module 2042 may perform a time-series analysis of the personal data and the user status data, generate intervention data from the results of the time-series analysis, and output the generated intervention data to the external output destination (a medical institution or a welfare facility).
[0093] Note that the question provision and acquisition module 2041 may be configured to generate voice data (radio programs) that better meet the needs of the user by, for example, performing additional learning on the language model 2023 regarding the content of a request received from the user. In this case, it becomes possible to generate voice data (radio programs) that correspond to the user's emotions and preferences obtained from the user's answer data through a dialogue with the user, thereby making it possible to provide a more personalized radio program.
[0094] <2 Data Structure> FIG. 4 is a diagram showing an example of the data structure of the user database 2021 in FIG.
[0095] As shown in FIG. 4, each record in the user database 2021 includes an item "user ID," an item "user name," an item "gender," an item "date of birth," and the like.
[0096] The item "user ID" is data for identifying each user who uses the voice data providing system 1 to answer questions and receive voice data.
[0097] The item "user name" is the name or the like of a user who uses the voice data providing system 1 to answer questions and receive voice data.
[0098] The item "gender" is the gender of the user who uses the voice data providing system 1 to answer questions and receive voice data.
[0099] The item "Date of Birth" is the date of birth of a user who uses the voice data providing system 1 to answer questions and receive voice data.
[0100] When a new user is registered in the voice data providing system 1, the server 20 adds a record to the user database 2021.
[0101] FIG. 5 is a diagram showing an example of the data structure of the personal database 2022 in FIG.
[0102] As shown in FIG. 5, each record in the personal database 2022 includes an item "user ID," an item "answer data details," and the like.
[0103] The item “user ID” is information that identifies each user who uses the voice data providing system 1 to answer questions and receive voice data, and corresponds to the item “user ID” in the user database 2021.
[0104] The item "Answer data details" is personal data stored by the question data generation module 2035 when a user answers a question using the voice data providing system 1, and specifically includes the item "Question sending date" and items "Question 1," "Answer 1," "Question 2," "Answer 2," etc.
[0105] The item "question sending date" is data indicating the date on which the question data providing module 2033 sent the question data.
[0106] The items "Question 1", "Question 2", etc. are the contents of the question data sent by the question data providing module 2033. The item "Question 1" stores, for example, the question data that was first sent by the question data providing module 2033 as the question of the day (for example, question data that was set in advance as a template), and "Question 2" and subsequent items store, for example, question data generated by the question data generation module 2035 through a dialogue with the user.
[0107] The items “Answer 1”, “Answer 2”, etc. are contents extracted from the answer data accepted by the answer acceptance module 2034. The items “Answer 1”, “Answer 2”, etc. store personal data extracted by the question data generation module 2035 from the text data of the answer data entered by the user, for example.
[0108] When a new user is registered in the voice data providing system 1, the server 20 adds a record to the personal database 2022. When the question data is transmitted, the question data providing module 2033 of the server 20 adds a record to the item "answer data details" of the personal database 2022. When the question data generation module 2035 of the server 20 extracts personal data from the answer data, the server 20 stores the answer data in the items "answer 1," "answer 2," etc. of the personal database 2022.
[0109] <3 operations> Hereinafter, with reference to FIGS. 6 and 7, a description will be given of a voice data generation process and a voice data provision process (method) performed by the voice data provision system 1 according to the embodiment of the present disclosure.
[0110] FIG. 6 is a flowchart showing an example of the flow of the voice data generation process performed by the voice data providing system 1.
[0111] In step S101, the question data providing module 2033 of the server 20 provides the user with first question data for asking questions about the user at a predetermined period (for example, every day) and at a first timing (for example, at a set time in the evening). In step S101, as the question data, text data of a dialogue-style question such as a question to be reflected in the content of the voice data provided by the voice data providing system 1, such as asking about today's events or tomorrow's plans, is transmitted to the terminal device 10 used by the user via the communication unit 201 and displayed on the display 132 of the terminal device 10.
[0112] In step S102, the answer receiving module 2034 of the server 20 receives an input of an answer to the question about the user provided in step S101, and acquires first answer data indicating the content of the answer. In step S102, for example, the answer from the user to the question about the user provided in step S101, such as an interactive question asking about today's events or tomorrow's plans, is received.
[0113] In step S103, the question data generation module 2035 of the server 20 extracts personal data related to the user from the first answer data received in step S102 and stores the extracted data in the personal database 2022, and determines whether the personal data of the user has been sufficiently accumulated in the personal database 2022 and whether voice data can be generated in step S105. If it is determined that voice data cannot be generated and further questions are necessary ("Y" in step S103), the process proceeds to step S104, and if it is determined that voice data can be generated and further questions are unnecessary ("N" in step S103), the process proceeds to step S105.
[0114] In step S104, the question data generation module 2035 of the server 20 generates first question data for asking the user a further question based on the first answer data received in step S102. In step S104, for example, natural language processing is performed to analyze the personal data and the answer content, and text data of an interactive question in natural language is generated regarding the content to be acquired as personal data.
[0115] In step S105, the voice data generation module 2036 of the server 20 generates voice data edited in a predetermined format to be provided to the user, based on the personal data about the user stored in the personal database 2022. In step S105, for example, natural language processing or the like is performed based on the personal data about the user, and voice data in a format similar to a radio program in natural language personalized for each user is generated.
[0116] As described above, the voice data providing system 1 extracts personal data of the user by interactively asking questions about the user and accepting input of answers at a first timing (for example, a set time in the evening) at predetermined intervals (for example, every day) using, for example, the language model 2023 or the external server 30. Voice data in a radio program format in natural language personalized for each user is then generated.
[0117] FIG. 7 is a flowchart showing an example of the flow of the voice data providing process performed by the voice data providing system 1.
[0118] In step S201, the voice data providing module 2037 of the server 20 provides the voice data generated in step S105 to the user at a second timing (for example, a set time the next morning) at predetermined intervals (for example, every day). In step S201, for example, the voice data generated in step S105 is transmitted to the terminal device 10 used by the user via the communication unit 201, and a notification that the voice data has been received is displayed on the display 132 of the terminal device 10.
[0119] In step S202, the playback status data acquisition module 2038 of the server 20 acquires playback status data indicating the playback status of the user for the audio data provided in step S201. In step S202, for example, the playback status acquisition module 2038 monitors the playback status of the audio data played by the user and causes the terminal device 10 to transmit the playback status data as a result. Also in step S202, the playback status data transmitted from the terminal device 10 is received via the communication unit 201 and accepted.
[0120] In step S203, the user status data generation module 2039 of the server 20 generates user status data indicating the status of the user based on the personal data stored in the personal database 2022 and the playback status data acquired in step S202. In step S203, for example, the personal data and playback status data are analyzed to determine whether or not there is an abnormality with respect to the user, and the result is output.
[0121] In step S204, the external output module 2042 of the server 20 outputs the personal data about the user and the generated user status data to an external output destination if necessary, depending on the content of the user status data generated in step S203. In step S204, for example, the user status data is transmitted via the communication unit 201 to an external output destination, such as an output destination registered as the user's relatives, or to a medical institution, welfare facility, or both.
[0122] As described above, the voice data providing system 1 provides users with voice data in a radio program format in natural language that is personalized for each user. This voice data (radio program) uses positive language to improve the self-efficacy of elderly users, who are the target users of the voice data providing system 1, and its content is designed to give the user a positive feeling. Furthermore, user status data indicating the user's condition is generated based on the playback status of the provided voice data, and depending on the content, personal data and user status data are output to an external output destination, such as the user's relatives, medical institution, or welfare facility, if necessary. Thus, by utilizing the voice data providing system according to the present disclosure, users can improve their self-efficacy by listening to the voice data, and external collaboration is possible if necessary.
[0123] <4 Screen example> Hereinafter, with reference to FIGS. 8 and 9, examples of screens for questions and answers displayed on the terminal device 10 by the voice data providing system 1 and examples of screens for providing voice data will be described.
[0124] Fig. 8 is a diagram showing an example of a screen of questions and answers displayed on the terminal device 10. The example screen of Fig. 8 shows an example of a screen in which questions and answers are asked in an interactive format by the question data providing module 2033 and the answer receiving module 2034 of the server 20. This corresponds to step S101 in Fig. 6 (after steps S102 to S104 have been performed).
[0125] 8, a communication screen 1311 provided by a communication tool such as LINE is displayed on the display 132 of the terminal device 10. This communication screen 1311 displays question data display fields 1312, 1313, 1315, and 1316 provided by the question data providing module 2033, and answer data display fields 1314 and 1317 entered by the user and accepted by the answer accepting module 2034. The communication screen 1311 also has an input field 1318 where the user enters an answer, and a send button 1319 for sending the answer entered in the input field 1318.
[0126] 8, when a question is sent in natural language, the user inputs an answer to the question in an input field 1318 and presses a send button 1319 by tapping on the screen, etc. As a result, the input answer is displayed as in an answer data display field 1314. By repeating such dialogue, personal data related to the user is extracted by the question data generation module 2035, and voice data based on the personal data is generated by the voice data generation module 2036.
[0127] Fig. 9 is a diagram showing an example of a screen for providing voice data displayed on the terminal device 10. The screen example of Fig. 9 shows an example of a screen in which voice data has been provided by the voice data providing module 2037 of the server 20. This corresponds to step S201 in Fig. 7.
[0128] 9, a communication screen 1321 provided by a communication tool such as LINE is displayed on the display 132 of the terminal device 10. This communication screen 1321 displays a message display field 1322 indicating that voice data will be provided, which is provided by the voice data providing module 2037, and a voice data display field 1323 to be provided. The voice data display field 1323 is provided with a play button 1324 that can play the voice data. The communication screen 1321 also has an input field 1325 similar to the example shown in FIG. 8, and a send button 1326 for sending text entered in the input field 1325.
[0129] 9, when the audio data is provided, the user presses the play button 1324 by tapping on the screen, etc. This makes it possible to play (listen to) the audio data provided by the audio data providing module 2037.
[0130] <Summary> As described above, according to the present embodiment, a process of asking a question about the user and receiving an answer at a first timing (e.g., a set time in the evening) every predetermined period (e.g., every day) is performed interactively using, for example, the language model 2023 or the external server 30, to extract personal data of the user. Then, voice data in a radio program format is generated in natural language personalized for each user, and the voice data is provided to the user every predetermined period (e.g., every day) at a second timing (e.g., a set time the next morning). This voice data (radio program) is positively addressed to improve the self-efficacy of elderly people who are the target users of the voice data providing system 1, and its content is designed to give the user a positive feeling. Furthermore, user status data indicating the user's status is generated based on the playback status of the provided voice data, and the personal data and user status data are output, if necessary, to an external output destination, such as the user's relatives, a medical institution, or a welfare facility, depending on the content. Therefore, by using the voice data providing system according to the present disclosure, it is possible to improve the self-efficacy of users by having them listen to the voice data, and to collaborate with external parties if necessary. This will help users maintain an independent lifestyle based on their own judgment, making it possible to prevent dementia and other conditions.
[0131] Furthermore, according to this embodiment, the user's personal data, for example, schedule data, is output as a reminder at a third timing different from the timing at which the voice data is provided. This allows the user to receive reminders about their planned schedule, which can be useful for memorizing.
[0132] Furthermore, according to this embodiment, requests or other inputs from the user or an outside party (e.g., a relative) are accepted for audio data in a format similar to a radio program, and are reflected in the audio data. This makes it possible to provide a more personalized radio program, thereby improving the user's sense of self-efficacy in their daily lives.
[0133] <Basic computer hardware configuration> 10 is a block diagram showing the basic hardware configuration of a computer 90. The computer 90 includes at least a processor 901, a main memory device 902, an auxiliary memory device 903, and a communication IF 991 (interface), which are electrically connected to one another by a communication bus 921.
[0134] The processor 901 is hardware for executing an instruction set written in a program, and is composed of an arithmetic unit, registers, peripheral circuits, and the like.
[0135] The main memory device 902 is used to temporarily store programs, data to be processed by the programs, etc. For example, it is a volatile memory such as a DRAM (Dynamic Random Access Memory).
[0136] The auxiliary storage device 903 is a storage device for saving data and programs, such as a flash memory, a hard disk drive (HDD), a magneto-optical disk, a CD-ROM, a DVD-ROM, or a semiconductor memory.
[0137] The communication IF 991 is an interface for inputting and outputting signals for communicating with other computers via a network using wired or wireless communication standards. The network is composed of the Internet, a LAN, various mobile communication systems constructed by wireless base stations, etc. For example, the network includes 3G, 4G, and 5G mobile communication systems, LTE (Long Term Evolution), and wireless networks (e.g., Wi-Fi (registered trademark)) that can connect to the Internet via a predetermined access point. In the case of a wireless connection, communication protocols include, for example, Z-Wave (registered trademark), ZigBee (registered trademark), and Bluetooth (registered trademark). In the case of a wired connection, the network also includes a direct connection using a USB (Universal Serial Bus) cable, etc.
[0138] It should be noted that the computer 90 can be virtually realized by distributing all or part of each hardware configuration across multiple computers 90 and interconnecting them via a network. In this way, the computer 90 is a concept that includes not only a computer 90 housed in a single housing or case, but also a virtualized computer system.
[0139] <Basic functional configuration of computer 90> The following describes the functional configuration of a computer realized by the basic hardware configuration (FIG. 10) of the computer 90. The computer includes at least the functional units of a control unit, a storage unit, and a communication unit.
[0140] The functional units of the computer 90 can also be realized by distributing all or part of the functional units among multiple computers 90 interconnected via a network. The computer 90 is a concept that includes not only a single computer 90 but also a virtualized computer system.
[0141] The control unit is realized by the processor 901 reading out various programs stored in the auxiliary storage device 903, expanding them in the main storage device 902, and executing processing in accordance with the programs. The control unit can realize functional units that perform various types of information processing depending on the type of program. In this way, the computer is realized as an information processing device that performs information processing.
[0142] The storage unit is realized by a main storage device 902 and an auxiliary storage device 903. The storage unit stores data, various programs, and various databases. Furthermore, the processor 901 can allocate a storage area corresponding to the storage unit in the main storage device 902 or the auxiliary storage device 903 in accordance with the programs. Furthermore, the control unit can cause the processor 901 to execute processes for adding, updating, and deleting data stored in the storage unit in accordance with the various programs.
[0143] A database refers to a relational database, which manages data sets called masters and tables in a tabular format structurally defined by rows and columns, by relating them to each other. In a database, a table is called a table, a master, a column in a table is called a column, and a row in a table is called a record. In a relational database, relationships between tables and masters can be set and associated. Typically, each table and each master has a column set as a primary key to uniquely identify a record, but setting a primary key to a column is not essential. The control unit can cause the processor 901 to add, delete, or update records in specific tables and masters stored in the storage unit according to various programs. Furthermore, by storing data, various programs, and various databases in the storage unit, it can be considered that the information processing device and information processing system according to the present disclosure have been manufactured.
[0144] Note that the databases and masters in this disclosure may include any data structure in which information is structurally defined (such as a list, dictionary, associative array, or object). The data structure also includes data that can be considered as a data structure by combining data with functions, classes, methods, etc. written in any programming language.
[0145] The communication unit is realized by the communication IF 991. The communication unit realizes a function of communicating with other computers 90 via a network. The communication unit can receive information transmitted from other computers 90 and input the information to the control unit. The control unit can cause the processor 901 to execute information processing on the received information in accordance with various programs. In addition, the communication unit can transmit information output from the control unit to other computers 90.
[0146] Furthermore, the above-described configurations, functions, processing units, processing means, etc. may be partially or entirely implemented in hardware, for example, by designing them as integrated circuits. The present invention can also be realized by software program code that implements the functions of the embodiments. In this case, a storage medium on which the program code is recorded is provided to a computer, and a processor included in the computer reads the program code stored in the storage medium. In this case, the program code itself read from the storage medium implements the functions of the above-described embodiments, and the program code itself and the storage medium on which it is stored constitute the present invention. Examples of storage media for providing such program code include flexible disks, CD-ROMs, DVD-ROMs, hard disks, SSDs, optical disks, magneto-optical disks, CD-Rs, magnetic tape, non-volatile memory cards, and ROMs.
[0147] Furthermore, the program code that realizes the functions described in this embodiment can be implemented in a wide range of program or script languages, such as assembler, C / C++, perl, Shell, PHP, and Java (registered trademark).
[0148] Furthermore, the program code of the software that realizes the functions of the embodiments may be distributed via a network and stored in a storage means such as a computer's hard disk or memory, or in a storage medium such as a CD-RW or CD-R, and the processor of the computer may read and execute the program code stored in the storage means or storage medium.
[0149] The functions performed by the components described herein may be implemented in circuitry or processing circuitry, including a general-purpose processor, an application-specific processor, an integrated circuit, an application-specific integrated circuit (ASIC), a central processing unit (CPU), conventional circuitry, and / or combinations thereof, programmed to perform the described functions. A processor includes transistors and other circuits and is considered to be circuitry or processing circuitry. A processor may also be a programmed processor that executes programs stored in memory. In this specification, a circuitry, unit, or means is hardware that is programmed to realize or performs the described functions, which may be any hardware disclosed herein or any hardware known to be programmed to realize or perform the described functions. If the hardware is a processor considered to be a type of circuitry, the circuitry, means, or unit is a combination of the hardware and software used to configure the hardware and / or processor.
[0150] Although the embodiments of the present disclosure have been described above, they can be implemented in various other forms, and various omissions, substitutions, and modifications can be made. These embodiments, modifications, and omissions, substitutions, and modifications are included in the technical scope of the claims and their equivalents.
[0151] For example, while the embodiment of the present disclosure is configured to play audio data, it may also be configured to display related image data or the audio being played as subtitles on the display 132 when playing the audio data. Furthermore, the audio data may be configured to be playable or skippable for each segment. While the above has been described, these may be implemented in various other forms, and may be implemented with various omissions, substitutions, and modifications. These embodiments and modifications, as well as those incorporating omissions, substitutions, and modifications, are within the technical scope of the claims and their equivalents.
[0152] <Additional Notes> The matters explained in the above embodiments will be supplemented below.
[0153] (Supplementary Note 1) A program for providing predetermined voice data to a user when executed by a computer including a processor 29 and a memory 25, wherein the memory 25 stores a database (2022) for storing personal data relating to the user, and the program causes the processor 29 to repeatedly perform the following steps: providing first question data for asking the user a question relating to the user at a first timing for each predetermined period (S101); accepting an input of an answer to the question relating to the user from the user and acquiring first answer data indicating the content of the answer (S102); extracting personal data relating to the user from the acquired first answer data and storing it in the database; and generating first question data for asking the user a further question, if necessary, based on the first answer data (S104). a step (S103) of executing a program; a step (S105) of generating, based on personal data about the user, audio data edited in a predetermined format to be provided to the user; a step (S201) of providing the generated audio data to the user at a second timing different from the first timing for each predetermined period; a step (S202) of acquiring playback status data indicating the playback status of the user for the provided audio data; a step (S203) of generating user status data indicating the status of the user based on the personal data about the user and the acquired playback status data about the user; and a step (S204) of outputting the personal data about the user and the generated user status data to an external output destination if necessary depending on the content of the generated user status data.
[0154] (Appendix 2) A program described in (Appendix 1), in which, in the output step, personal data about the user and the generated user status data are output to an external output destination that has been registered in advance and has a predetermined relationship with the user.
[0155] (Appendix 3) A program described in (Appendix 1), in which, in the output step, an external output destination to which personal data about the user and the generated user status data are output is selected according to the content of the generated user status data, and the program outputs the data to the selected external output destination.
[0156] (Appendix 4) A program described in (Appendix 2) or (Appendix 3), in which, in the output step, personal data about the user and the generated user status data are output to either a medical institution, a welfare facility, or both as external output destinations.
[0157] (Appendix 5) A program described in (Appendix 4), in which, in the output step, intervention data to be output to an external output destination is extracted from personal data about the user and the generated user status data depending on the content of the generated user status data, and the extracted intervention data is output to the external output destination.
[0158] (Appendix 6) A program described in (Appendix 5), in which, in the output step, a time-course analysis is performed on personal data about the user and the generated user status data, intervention data is generated from the results of the time-course analysis, and the generated intervention data is output to an external output destination.
[0159] (Appendix 7) The program described in (Appendix 1), wherein in the step of acquiring playback status data for the user, one or more of the following data are acquired: data indicating whether the user has played audio data; data indicating how far the user has played the audio data; and the time at which the user played the audio data.
[0160] (Supplementary Note 8) The program according to (Supplementary Note 1), wherein in the step of extracting personal data relating to the user, future schedule data of the user is extracted.
[0161] (Appendix 9) The program described in (Appendix 1), further comprising the step of providing personal data about the user to the user at a third timing different from the first timing and the second timing.
[0162] (Appendix 10) The program described in (Appendix 1) further causes the program to execute a step of providing all or part of the audio data provided to the user at the second timing to the user at a third timing different from the first timing and the second timing.
[0163] (Supplementary Note 11) The program according to (Supplementary Note 1), wherein in the step of generating audio data, the audio data is generated based on the content of past playback status data or user state data.
[0164] (Appendix 12) The program further executes the steps of: providing the user with second question data at a third timing different from the first timing and the second timing, the second question data being used to ask a question about the voice data provided to the user at the second timing; and receiving input from the user of an answer to the question about the voice data and acquiring second answer data indicating the content of the answer. In the step of generating voice data, the program described in (Appendix 1) generates voice data based on the content of the second answer data.
[0165] (Appendix 13) A program described in any of (Appendix 1) to (Appendix 3), further comprising: a step of receiving input of content to be reflected in the voice data from a user or an external output destination; acquiring reflection data indicating the input content; and, in a step of generating voice data, generating the voice data based on the content of the reflection data.
[0166] (Appendix 14) The program described in (Appendix 1), wherein in the step of generating voice data, the voice data is generated using a generation AI service.
[0167] (Supplementary Note 15) An information processing device that provides predetermined voice data to a user includes a control unit 203 and a memory 25 (storage unit 202), wherein the memory 25 stores a database (2022) for storing personal data related to the user, and the control unit 203 repeatedly executes the following steps: a step (S101) of providing the user with first question data for asking the user a question related to the user at a first timing for each predetermined period; a step (S102) of accepting an input of an answer to the question related to the user from the user and acquiring first answer data indicating the content of the answer; and a step (S104) of extracting personal data related to the user from the acquired first answer data and storing it in the database, and generating first question data for asking the user a further question if necessary based on the first answer data. an information processing device that executes the steps of: generating (S103) audio data edited in a predetermined format based on personal data about the user to be provided to the user; providing (S105) the generated audio data to the user at a second timing different from the first timing for each predetermined period of time; acquiring (S202) playback status data indicating a playback status of the user for the provided audio data; generating user status data indicating a status of the user based on the personal data about the user and the acquired playback status data about the user; and outputting (S204) the personal data about the user and the generated user status data to an external output destination when necessary, depending on the content of the generated user status data.
[0168] (Supplementary Note 16) A method for providing predetermined voice data to a user, executed by a computer including a processor 29 and a memory 25, wherein the memory 25 stores a database (2022) for storing personal data related to the user, and the method includes the steps of: a step (S101) in which the processor 29 provides the user with first question data for asking the user a question related to the user at a first timing for each predetermined period; a step (S102) in which the processor 29 receives input of an answer to the question related to the user from the user and acquires first answer data indicating the content of the answer; and a step (S104) in which the processor 29 extracts personal data related to the user from the acquired first answer data, stores the extracted personal data in the database, and generates first question data for asking the user a further question, if necessary, based on the first answer data. a step (S103) of executing a program; a step (S105) of generating, based on personal data about the user, audio data edited in a predetermined format to be provided to the user; a step (S201) of providing the generated audio data to the user at a second timing different from the first timing for each predetermined period; a step (S202) of acquiring playback status data indicating the playback status of the user for the provided audio data; a step (S203) of generating user status data indicating the status of the user based on the personal data about the user and the acquired playback status data about the user; and a step (S204) of outputting the personal data about the user and the generated user status data to an external output destination if necessary depending on the content of the generated user status data. [Explanation of symbols]
[0169] 1: Voice data provision system 10: Terminal device 10A: Terminal equipment 10B: Terminal device 13: Input device 14: Output device 15: Memory 16: Storage section 19: Processor 20: Server 25: Memory 26: Storage 29: Processor 80: Network 81: Wireless base station 82: Wireless LAN router 90: Computer 111: Antenna 112: Antenna 121: First wireless communication unit 122: Second wireless communication unit 130: Operation reception unit (touch screen) 131: Touch-sensitive devices 132: Display 140: Location information sensor 150: Camera 160: Storage section 161: User information 170: Control unit 171: Input operation reception unit 172: Transmitter / receiver 173: Notification control section 174: Data processing section 201: Communications Department 202: Storage section 203: Control unit 901: Processor 902: Main memory 903 :Auxiliary storage device 921: Communication bus 2021: User Database 2022: Personal Database 2023: Language Model 2031: Receiving control module 2032: Transmission control module 2033: Question data provision module 2034: Answer reception module 2035: Question data generation module 2036: Voice data generation module 2037: Audio data provision module 2038: Playback status data acquisition module 2039: User state data generation module 2040: Personal data provision module 2041: Question provision and acquisition module 2042: External output module
Claims
1. A program for causing a computer having a processor and a memory to execute the program and providing predetermined voice data to a user, the memory stores a database for storing personal data relating to the user; The program causes the processor to: providing the user with first question data for asking the user a question about the user at a first timing for each predetermined period; receiving an input of an answer to a question about the user from the user and acquiring first answer data indicating the content of the answer; a step of repeatedly executing a step of extracting personal data related to the user from the acquired first response data, storing the extracted personal data in the database, and generating first question data based on the first response data for asking the user further questions if necessary; generating audio data edited in a predetermined format for provision to the user based on personal data relating to the user; providing the generated voice data to the user at a second timing different from the first timing for each predetermined period of time; acquiring playback status data indicating a playback status of the provided audio data by the user; generating user status data indicating a status of the user based on personal data about the user and the acquired playback status data of the user; and outputting, if necessary, personal data relating to the user and the generated user status data to an external output destination, depending on the content of the generated user status data.
2. 2. The program according to claim 1, wherein in the outputting step, personal data relating to the user and the generated user status data are output to an external output destination that has been registered in advance and has a predetermined relationship with the user.
3. The program of claim 1, wherein in the output step, an external output destination to which the personal data about the user and the generated user status data are output is selected according to the content of the generated user status data, and the data is output to the selected external output destination.
4. 4. The program according to claim 2, wherein in the outputting step, the personal data about the user and the generated user status data are output to an external output destination, such as a medical institution or a welfare facility, or both.
5. 5. The program according to claim 4, wherein in the output step, intervention data to be output to an external output destination is extracted from personal data about the user and the generated user status data according to the content of the generated user status data, and the extracted intervention data is output to the external output destination.
6. The program according to claim 5, wherein in the output step, a time-course analysis is performed on personal data about the user and the generated user status data, the intervention data is generated from the results of the time-course analysis, and the generated intervention data is output to an external output destination.
7. The program of claim 1, wherein in the step of acquiring playback status data for the user, one or more of the following data are acquired: data indicating whether the user has played the audio data, data indicating how far the user has played the audio data, and the time at which the user played the audio data.
8. 2. The program according to claim 1, wherein in the step of extracting personal data relating to the user, future schedule data of the user is extracted.
9. The program further comprises: The program according to claim 1 , further comprising: providing personal data relating to the user to the user at a third timing different from the first timing and the second timing.
10. The program further comprises: The program according to claim 1, further comprising a step of providing all or part of the voice data provided to the user at the second timing to the user at a third timing different from the first timing and the second timing.
11. 2. The program according to claim 1, wherein in the step of generating the audio data, the audio data is generated based on the content of the past playback status data or the past user state data.
12. The program further comprises: providing second question data for asking a question about the voice data provided to the user at the second timing to the user at a third timing different from the first timing and the second timing; receiving an input of an answer to a question regarding the voice data from the user, and acquiring second answer data indicating the content of the answer; 2. The program according to claim 1, wherein in the step of generating the voice data, the voice data is generated based on the content of the second response data.
13. The program further comprises: receiving an input of content to be reflected in the voice data from the user or the external output destination, and acquiring reflection data indicating the input content; 4. The program according to claim 1, wherein in the step of generating the audio data, the audio data is generated based on the content of the reflection data.
14. The program according to claim 1 , wherein the step of generating the voice data uses a generation AI service to generate the voice data.
15. An information processing device that includes a control unit and a memory and provides predetermined voice data to a user, the memory stores a database for storing personal data relating to the user; The control unit providing the user with first question data for asking the user a question about the user at a first timing for each predetermined period; receiving an input of an answer to a question about the user from the user and acquiring first answer data indicating the content of the answer; a step of repeatedly executing a step of extracting personal data related to the user from the acquired first response data, storing the extracted personal data in the database, and generating first question data based on the first response data for asking the user further questions if necessary; generating audio data edited in a predetermined format for provision to the user based on personal data relating to the user; providing the generated voice data to the user at a second timing different from the first timing for each predetermined period of time; acquiring playback status data indicating a playback status of the provided audio data by the user; generating user status data indicating a status of the user based on personal data about the user and the acquired playback status data of the user; and outputting, if necessary, personal data relating to the user and the generated user status data to an external output destination, according to the content of the generated user status data.
16. 1. A method for providing predetermined audio data to a user, the method being executed by a computer having a processor and a memory, the method comprising: the memory stores a database for storing personal data relating to the user; The method further comprises the processor: providing the user with first question data for asking the user a question about the user at a first timing for each predetermined period; receiving an input of an answer to a question about the user from the user and acquiring first answer data indicating the content of the answer; a step of repeatedly executing a step of extracting personal data related to the user from the acquired first response data, storing the extracted personal data in the database, and generating first question data based on the first response data for asking the user further questions if necessary; generating audio data edited in a predetermined format for provision to the user based on personal data relating to the user; providing the generated voice data to the user at a second timing different from the first timing for each predetermined period of time; acquiring playback status data indicating a playback status of the provided audio data by the user; generating user status data indicating a status of the user based on personal data about the user and the acquired playback status data of the user; and outputting personal data about the user and the generated user status data to an external output destination, if necessary, depending on the content of the generated user status data.
Citation Information
Patent Citations
User state confirmation system, user state confirmation method, communication terminal device, user state reporting method, and computer program
JP2014197264A
Recollection information transmitting and elderly person watching system
JP2021131832A