Control device and digital family buddhist altar control program

The integration of generative AI into digital Buddhist altars allows for a virtual reality conversation space by using a control device that generates synthetic voices and maintains thematic relevance, addressing the limitations of pre-recorded responses and enhancing user interaction.

JP2026003441APending Publication Date: 2026-01-13MITSUI CHEMICALS INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024101395
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-24
Publication Date
2026-01-13

AI Technical Summary

Technical Problem

Conventional digital Buddhist altars lack the ability to create a virtual reality conversation space between users and the deceased, as they rely on pre-recorded information and cannot provide accurate responses to user questions.

Method used

Introduce generative AI technology to facilitate a virtual reality conversation space by using a control device equipped with a voice conversation function that elicits information through questioning, accumulates recollections, and generates synthetic voices based on the deceased's autobiography, with features like silent time measurement and thematic regression to maintain conversation relevance.

Benefits of technology

Establishes a virtual reality conversation space between users and the deceased, enabling dynamic and relevant interactions by generating personalized responses based on accumulated recollections, thus enhancing user engagement and emotional connection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026003441000001_ABST
    Figure 2026003441000001_ABST
Patent Text Reader

Abstract

To establish a virtual reality conversation space between users using a digital family Buddhist altar and the deceased by introducing a generation AI technology to the digital family Buddhist altar and associating it with information on a personal history created by the deceased before his or her death.SOLUTION: In the digital family Buddhist altar system, the user 12 who has created the personal history becomes the deceased, and when the user (a relative or the like) speaks to the user 12 of the deceased in front of the family Buddhist altar, the contents matching the contents and based on the personal history created in life are outputted by the voice of the deceased. As a result, the interactive-type generating AI executing unit 52A can extract the contents of the personal history or generate a new sentence based on the registered contents of the personal history instead of a conversation with fixed contents, and can respond to the user's question or the like with the voice of the deceased user 12. For this reason, it is possible to deepen the relationship between the user and the deceased user 12 in a situation where they are actually having a conversation.SELECTED DRAWING: Figure 14
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a control device and a digital Buddhist altar control program. [Background technology]

[0002] There are known voice conversation systems that use terminal devices such as interactive robots and smartphones to realize free conversation with humans. For example, Patent Document 1 listed below describes a conversation voice providing system that uses synthesized human voice as the voice used in the conversation with humans.

[0003] Furthermore, Patent Document 2 discloses a system that allows a person to reminisce about fond memories, events, or things from the past and then talk about the memories in a way that suits each person. Voice recognition technologies are described in Patent Documents 3 and 4.

[0004] Incidentally, a digital Buddhist altar is one of the devices that utilizes a conversational voice provision system that uses synthesized human voice. An example of a homepage is "https: / / www.cohaco.rip / ".

[0005] The basic components of a digital Buddhist altar include a touch sensor, a motion sensor, an LCD display, a speaker, and a microphone, and it has the ability to display personal photos, videos, and messages, as well as the ability to link with a smartphone app to leave video messages from the deceased. It is believed that messages are pre-recorded and played back.

[0006] More specifically, it has the ability to recognize the user's words, and when they call out to it, it will display a familiar face (such as a photo of the deceased), speak messages left behind by the deceased, and send notifications of important days such as the anniversary of their death.

[0007] However, conventional digital Buddhist altars cannot provide the functions of a conversational voice providing system, in particular the ability to create a virtual space such as a virtual conversation between the user and the deceased.

[0008] In other words, even if the digital Buddhist altar outputs the voice of the deceased, it is outputting pre-recorded information, and does not necessarily output accurate answers to the user's questions.

[0009] Therefore, it is possible to add a function using generative AI technology that responds to questions from users using the voice of the deceased. [Prior art documents] [Patent documents]

[0010] [Patent Document 1] Japanese Patent Application Laid-Open No. 2016-110151 [Patent Document 2] Japanese Patent Publication No. 2020-57350 [Patent Document 3] Japanese Patent Application Publication No. 2019-105735 [Patent Document 4] Japanese Patent Application Publication No. 8-115093 Summary of the Invention [Problem to be solved by the invention]

[0011] However, in a conversation using generative AI, if there is no underlying information (or if there is little information), it will not be able to generate an answer to the question. Therefore, even if generative AI technology is simply introduced, it will not be possible to enhance the conversation function between the user and the deceased as a digital Buddhist altar, and it will not be possible to establish a virtual reality conversation space between the user and the deceased.

[0012] Taking the above facts into consideration, the present invention aims to provide a control device and a digital Buddhist altar control program that can establish a virtual reality conversation space between the user of the digital Buddhist altar and the deceased by introducing generative AI technology into the digital Buddhist altar and associating it with information about the deceased's autobiography created during their lifetime. [Means for solving the problem]

[0013] The control device of the present invention is executed using a predetermined scenario program and an interactive generation AI execution unit that assists the progress of the scenario program, and is equipped with a voice conversation function that accepts answers from the user by providing the user with questions to elicit information related to a specific theme that prompts the user to reminisce about their personal history, and a recollection control unit that accumulates multiple pieces of recollection information of the user obtained by repeating the voice conversation function, and the interactive generation AI execution unit has a user synthetic voice generation function that generates synthetic voice of the user, and in response to questions from a user who wishes to have a dialogue with the user, generates a response to the question based on the accumulated recollection information of the user, and when outputting the response aloud, generates and outputs a synthetic voice similar to that of the user.

[0014] In the control device of the present invention, the recollection control unit creates a summary of the user's personal history by era or by specific theme based on the recollection information.

[0015] The control device of the present invention is characterized by further comprising a measurement unit that measures the silent time from when the question is provided until the user starts to answer, and a judgment unit that judges when the user has finished answering based on the silent time measured by the measurement unit.

[0016] The control device of the present invention is characterized in that it further has a regression instruction unit that, if the user's speech content deviates from an original theme based on a scenario program while the voice conversation function is being executed, returns the user to the original theme through free conversation by the interactive generation AI execution unit.

[0017] A digital Buddhist altar control program according to the present invention is characterized in that it causes a computer to operate as the control device described above. [Effects of the Invention]

[0018] According to the present invention, by introducing generative AI technology into a digital Buddhist altar and associating it with information about the deceased's autobiography created during their lifetime, it is possible to establish a virtual reality conversation space between the user of the digital Buddhist altar and the deceased. [Brief explanation of the drawings]

[0019] [Figure 1] 1 is a diagram illustrating a configuration of an example of a voice conversation system according to a first embodiment. [Figure 2] 2A is a block diagram showing an example of the hardware configuration of a terminal device of the voice conversation system shown in FIG. 1, and FIG. 2B is a block diagram showing an example of the hardware configuration of a voice conversation generation device according to the first embodiment. [Figure 3] 2 is a functional block diagram showing an example of the functions of each component of the autobiography creation system shown in FIG. 1. FIG. [Figure 4] 4 is a control flowchart showing a main routine of processing based on the personal history creating scenario program according to the first embodiment. [Figure 5] FIG. 5 is a transition diagram listing options on the screen displayed on the display unit 22 during execution of the flowchart of FIG. 4. [Figure 6] 1A and 1B are front views of screens displayed on the display unit, in which (A) shows the overall top screen A, (B) shows the free conversation screen A1, and (C) shows the setting screen A2. [Figure 7] 10A and 10B are front views of the screens displayed on the display unit, where (A) shows the personal history menu screen B and (B) shows the personal history setting screen C. [Figure 8] 1A and 1B are front views of screens displayed on the display unit, where (A) shows an episode selection screen D1, (B) shows an episode selection screen D2, and (C) shows an interactive mode screen E. FIG. [Figure 9] This is a front view of the screens displayed on the display unit, where (A) is the screen F for viewing and editing episodes, (B) is the screen F1 showing the contents of the conversation file, (C) is the conversation file content editing screen F11, (D) is the screen F2 listing the target parties during recollection, and (E) is the screen F21 listing the conversations in the episode. [Figure 10] This is a front view of the screen displayed on the display unit, where (A) shows a conversation file list screen G1 with the send list tag selected, and (B) shows a conversation file list screen G2 with the public list tag selected. [Figure 11] 4 is a control flowchart for determining whether an answer has ended, using the user confrontation information acquisition unit and the answer end time determination unit of FIG. 3, where (A) is a flowchart showing the flow of the user-specific predetermined time setting interrupt routine, and (B) is a control flowchart showing the answer end determination interrupt routine. [Figure 12] 4 is a control flowchart showing the flow of a theme deviation determination interruption routine in the theme return instruction unit of FIG. 3; [Figure 13] 1. FIG. 11 is a schematic diagram of a digital Buddhist altar system according to a second embodiment, in which a digital Buddhist altar user terminal device is added to the autobiography creating system shown in FIG. [Figure 14] FIG. 10 is a functional block diagram showing an example of the functions of each component of the digital Buddhist altar system according to the second embodiment. [Figure 15] 10 is a control flowchart showing the flow of processing in the digital Buddhist altar system according to the second embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0020] Fig. 1 is a schematic diagram illustrating an example of an autobiography creation system 10 according to a first embodiment. As shown in Fig. 1, the autobiography creation system 10 according to the first embodiment includes a terminal device 14 that is operated by a user 12, who is the target of autobiography creation, to create the autobiography, and a voice conversation generation device 16 as a control device. These are configured to be able to communicate with each other via a network NW such as the Internet.

[0021] The terminal device 14 is a terminal device that has a function of acquiring voice information including the voice uttered by the user 12 .

[0022] The terminal device 14 can be configured as a device capable of at least acquiring surrounding sounds and performing data communication, such as a smartphone, a personal computer, a tablet terminal, a mobile phone, an interactive robot, or a smart speaker.

[0023] Fig. 2(A) is a block diagram showing an example of the hardware configuration of the terminal device 14 shown in Fig. 1. As shown in Fig. 2, the terminal device 14 mainly includes a microphone 18, a speaker 20, a display unit 22, an operation unit 24, and a control device 26. Hereinafter, the microphone 18, the speaker 20, the display unit 22, and the operation unit 24 will be collectively referred to as a UI (user interface) 28.

[0024] The microphone 18 is a device, typically a microphone, that can collect sounds generated around the terminal device 14. The speaker 20 is a device that can notify various pieces of information by voice to the user 12 or the like around the terminal device 14. The specific configurations of the microphone 18 and the speaker 20 are not particularly limited, and may be, for example, devices built into the terminal device 14 or externally connected devices.

[0025] The display unit 22 is an example of a user interface that can provide various information to a user who uses the terminal device 14. This display unit 22 can be a well-known display device such as a liquid crystal display (LCD) or an organic electro-luminescent display (OLED).

[0026] The operation unit 24 is an example of a user interface that allows an operator (i.e., the user 12) who operates the terminal device 14 to perform a predetermined input operation. In the first embodiment, the operation unit 24 may be a keyboard or a pointing device when the terminal device 14 is a personal computer. In addition, in the case where the terminal device 14 is a smartphone or the like, the operation unit 24 may be a touch panel that also serves as the display unit 22.

[0027] The control device 26 is a device capable of performing various controls of the terminal device 14, and may be configured, for example, by a well-known computer. The control device 26 configured by a well-known computer may include, for example, at least one processor 26A, a ROM (Read Only Memory) 26B and a RAM (Random Access Memory) 26C as examples of memory, a storage 26D, a communication interface 26E, and an input / output interface 26F. These components may be connected to each other so as to be able to communicate with each other via an internal bus 26G.

[0028] The processor 26A may be configured with, for example, a CPU (Central Processing Unit) and has the function of executing various programs and controlling each part. Specifically, the processor 26A reads various control programs stored in, for example, the ROM 26B or the storage 26D, and executes the programs using the RAM 26C as a work area. The processor 26A controls the operation of each component of the terminal device 14, specifically the microphone 18, the speaker 20, etc., and performs various arithmetic processing in accordance with the programs.

[0029] The ROM 26B stores various programs and various data.

[0030] The RAM 26C temporarily stores programs or data as a working area.

[0031] The storage 26D can be configured with a recording medium such as a hard disk drive (HDD), a solid state drive (SSD), a flash memory, etc. The storage 26D may at least temporarily store various programs including an operating system and various data necessary to operate the first terminal device 14.

[0032] The communication interface 26E can transmit and receive predetermined data via wired communication using a communication cable or wireless communication based on wireless communication standards such as Wi-Fi (registered trademark) or Bluetooth (registered trademark). This communication interface 26E allows the control device 26 of the terminal device 14 to communicate with other electronic devices, such as the voice conversation generation device 16, via the network NW.

[0033] The input / output interface 26F may be connected to various components of the terminal device 14 and may be capable of transmitting and receiving various information, such as control signals for the components and audio signals collected by the microphone 18. Although not shown, the input / output interface 26F may further include a connector to which an external recording medium can be connected, and various drives such as a DVD drive.

[0034] Fig. 2(B) is a block diagram showing an example of the hardware configuration of the voice conversation generation device 16 according to the first embodiment shown in Fig. 1. The voice conversation generation device 16 according to the first embodiment is a device used by a user of the terminal device 14 to create his or her personal history.

[0035] The voice conversation generation device 16 may be configured, for example, with one or more server computers. As shown in Fig. 2(B), the voice conversation generation device 16 includes at least one processor 30, a ROM 32 and a RAM 34 as examples of memory, a storage 36, and a timer 38.

[0036] These various components are connected to each other via an internal bus 42 so as to be able to communicate with each other. Furthermore, among these components, those except for the timer 38 may perform substantially the same functions as the corresponding various components of the control device 26 of the terminal device 14 described above. Therefore, the details of these components will be explained in the above description and will not be described here. Note that the voice conversation generation device 16 according to the first embodiment may further include other components in addition to the components described above, such as a user interface such as a display unit or an operation unit.

[0037] Timer 38 is a timing means for measuring the passage of time, and in the first embodiment, is capable of measuring the time elapsed since the synthetic speech was generated. Note that if there is no need for voice conversation generation device 16 to measure the time since the synthetic speech was generated, timer 38 may be omitted.

[0038] Figure 3 is a functional block diagram showing an example of an autobiography creation control system executed by the voice conversation generation device 16 shown in Figure 1. Each functional block does not limit the hardware configuration, and part or all of it is executed by a software program (scenario program) based on the above-mentioned hardware configuration (see Figure 2(B)).

[0039] 3, the voice conversation generation device 16 is mainly composed of a control unit 50 and a dialogue-based generation AI execution unit 52. The control unit 50 operates based on a scenario program. The dialogue-based generation AI execution unit 52 executes processing in response to instructions from the control unit 50 as part of its operation based on the scenario program.

[0040] (summary generation) The control unit 50 is connected to a personal history creation receiving unit 54, a voice information transmitting unit 56, a voice information receiving unit 58, and a user confrontation information acquiring unit 70.

[0041] When the autobiography creation receiving unit 54 receives an instruction to create an autobiography from the terminal device 14, it sends a notification that an instruction to create an autobiography has been received to the control unit 50. The control unit 50 selects a theme for creating the autobiography and sends the selected theme to the interactive generation AI execution unit 52.

[0042] The voice information transmitting unit 56 transmits a question by voice to the terminal device 14. The information to be transmitted is information generated by the dialogue generation AI executing unit 52 and converted into voice information, and the control unit 50 sends it to the terminal device 14 via the voice information transmitting unit 56.

[0043] The voice information receiving unit 58 receives voice (answer to a question) from the terminal device and sends it to the control unit 50. The control unit 50 sends the received voice information to the dialogue-based generation AI executing unit 52.

[0044] A scenario database 60 is connected to the control unit 50 and the interactive generation AI execution unit 52. When the interactive generation AI execution unit 52 acquires a theme, it creates a scenario based on the theme based on the scenario stored in the scenario database 60, and performs the following processes: voice analysis, accumulation of information after voice analysis (recollection information), question generation, answer reception, free conversation, and summary generation.

[0045] Free conversation includes backchanneling, empathy, and conformity, and refers to conversation that is considerate of the conversation partner user 12 (which may include visual actions). In the following, "backchanneling" is used as an example of free conversation, but when applying other free conversation, "backchanneling" can be replaced with other free conversation.

[0046] The recollection information is accumulated (stored) in the recollection information / summary information storage unit 62 , but may also be stored in an internal storage area of ​​the dialogue-based generation AI execution unit 52 .

[0047] The recollection information can be provided in a manner that allows it to be read, edited, and made public based on predetermined keywords (date, person being recollected (mother, father, etc.)) in response to an instruction from the user 12.

[0048] As shown in Table 1, the scenario database 60 stores scenarios for each theme (such as a theme according to age group).

[0049] [Table 1] The dialogue-type generation AI execution unit 52 reads out a scenario based on the theme, and generates questions for the user 12, and further questions, responses, etc. in response to the answers from the user 12.

[0050] That is, the control unit 50 uses the interactive generation AI execution unit 52 to ask questions to the user 12 by voice, receive answers (episodes) from the user 12's own recollections by voice, and collect episodes that fit the theme in accordance with the scenario while sequentially responding to the answers from the user 12 (recollection method). A summary of the collected episodes (recollection information) is generated for each theme.

[0051] When the interactive generation AI execution unit 52 generates a summary for each theme, it sends it to the recollection information / summary information storage unit 62 .

[0052] (Create your own history) The recollection information and summary information storage unit 62 is connected to the personal history creation unit 64. In response to instructions from the terminal device 14, the personal history creation unit 64 creates a personal history, for example, by era or by theme, using summaries of multiple eras and topics. The created personal history is, for example, sent to a printing process, printed on paper media in predetermined units (for example, one A4-sized sheet), and mailed to the user 12. Of course, this does not deny the possibility of sending the personal history to the terminal device 14 as data, but for the user 12, receiving the personal history on paper media allows them to enjoy the pleasure of physically filing the personal history each time it is sent.

[0053] The recollection information / summary information storage unit 62 is connected to an update unit 66, and it is possible to edit the generated recollection information and summary via the update unit 66 in response to an instruction from the terminal device 14.

[0054] For example, as a variation of the autobiography mentioned above, based on keywords that appear throughout the years (for example, a close friend "Y" from childhood to adulthood), it is possible to extract recollections of the close friend Y from multiple summaries by age group and re-edit them into a single summary.In addition, information about the close friend Y can be added as a topic to part of the autobiography.

[0055] For example, there are times in conversations when episodes featuring "Y" appear across different generations (see, for example, Episode 1 and Episode 2).

[0056] (Episode 1) "Back then, in the fifth and sixth grades, we did things like horse-riding battles and group gymnastics, which were pretty exciting. I had a friend named Y. He was the leader of the horse-riding battles, and I used to carry him from the front of the class."

[0057] (Episode 2) I used to play with "Y" a lot in elementary school. We were together until junior high school, and I think we were the best of friends in school. We often played pranks together, and both of us got scolded by our dad.

[0058] The name of a person ("Y") that appears in both Episode 1 and Episode 2 above is tagged, and information about that person is stored as an episode.

[0059] Then, by editing episodes centered around that person and visualizing them on one sheet as memories of that person, you can accumulate all kinds of people and visualize and realize that your life has been surrounded by many different people.

[0060] (Judgment of completion of answer) The voice conversation generation device 16 includes a user confrontation information acquisition unit 70. In the first embodiment, the user confrontation information acquisition unit 70 acquires the time from when the user 12 receives a question to when the user starts answering (silent time) as a target for determining when the user 12 has finished answering. Note that sounds that are not considered to be answers (for example, growling, coughing, sneezing, mumbling, etc.) are treated as silent time.

[0061] The user confrontation information acquisition unit 70 is connected to the answer end time determination unit 72, and sends the acquired silent time to this answer end time determination unit 72. When the received silent time has exceeded a predetermined time ts, the answer end time determination unit 72 determines that the answer has ended, and notifies the control unit 50.

[0062] More specifically, some users 12 may recall a question and then start answering (start speaking), resulting in silent periods. During these silent periods, the dialogue-based generation AI execution unit 52 continues to receive voice information and performs voice analysis, so there are cases where silent periods (or voices that are not considered to be answers) are determined to be answers.

[0063] Therefore, the control unit 50 is configured so that the interactive generation AI execution unit 52 specifies the end time of the answer so as not to misinterpret the answer from the user 12. Note that since the length of the silent period (so-called "pause") may differ depending on the user 12, it is preferable to set the above-mentioned predetermined time ts for each user 12.

[0064] In addition, in the first embodiment, the time when the answer ends is determined using the silent time, but the time when the answer ends may also be determined by detecting the line of sight and facial expression of the user 12 and using the line of sight and facial expression, or by using a combination of at least two or more of the silent time, line of sight, and display.

[0065] (Thematic regression control) A theme return instruction unit 68 is connected to each of the control unit 50 and the interactive generation AI execution unit 52 .

[0066] The theme regression instruction unit 68 compares the initial theme received from the control unit 50 with the current theme received from the dialogue-based generation AI execution unit 52, and if the current theme deviates from the initial theme, it instructs the dialogue-based generation AI execution unit 52 to send questions or the like to encourage theme regression. Whether or not there has been a deviation can be determined, for example, by comparing keywords or the like predetermined for each theme. Alternatively, the determination may be made by the dialogue-based generation AI execution unit 52.

[0067] When the user 12 is reminiscing, it is ideal to speak along a given theme, but due to a chain of memories, the user may speak off-topic.

[0068] Therefore, when an answer that deviates from the original topic is received, the interactive generation AI execution unit 52 presents the user 12 with a question that will make it easier for them to return to the original topic in order to collect information that is in line with the original topic. This allows information related to the original topic to be collected in a very natural way (especially without upsetting the user 12) while providing a response, rather than a mechanical response such as "That's not the topic."

[0069] The operation of the first embodiment will be described below.

[0070] FIG. 4 is a control flowchart showing a main routine of processing based on the personal history creating scenario program according to the first embodiment.

[0071] 5 is a transition diagram listing options on the screen displayed on the display unit 22 during execution of the flowchart of FIG. 4, and FIG. 6 is a front view showing an example of a screen displayed on the display unit 22.

[0072] In explaining the flowchart of FIG. 4, the processing flow will be explained in reference to FIGS.

[0073] As shown in Fig. 4, in step 100, the overall top screen (screen A shown in Fig. 6(A)) is displayed. Screen A displays menu screen options of "Autobiography Mode," "Conversation," "Settings," and "Exit."

[0074] The "personal history mode" is a mode for transitioning to one for creating a personal history, and Figure 5 mainly shows the transition state in this "personal history mode."

[0075] The "Have a conversation" mode offers options for conversation settings, and unlike the autobiography mode, it has functions such as conversation with the AI ​​and reminiscence therapy, allowing the AI ​​to have free conversations with the user or conversations based on a scenario.

[0076] The "Settings" mode has a basic information screen and options for necessary conversation information, and transitions to various settings such as initial settings and mid-session changes.

[0077] "Exit" means that the application is terminated.

[0078] In the next step 102, it is determined whether "End" has been selected, and if the determination is affirmative, the routine ends. If the determination is negative in step 102, the process branches depending on the selected mode (step 104).

[0079] If "Have a conversation" is selected in step 104, the process proceeds to step 106, where a conversation screen (see screen A1 shown in FIG. 6(B)) is displayed, and then the process proceeds to step 108, where "Have a conversation" mode processing is executed, and the process returns to step 100.

[0080] If "Settings" is selected in step 104, the process proceeds to step 110, where a settings screen (see screen A2 shown in FIG. 6(C)) is displayed, and then the process proceeds to step 112, where "Settings" mode processing is executed, and the process returns to step 100.

[0081] If "Personal History Mode" is selected in step 104, the process proceeds to step 114, where a personal history menu screen (see screen B shown in FIG. 7(A)) is displayed, and then the process proceeds to step 116, where the process branches depending on the selected menu. Screen B displays the following menu screen options: "Tell an Episode," "View / Edit Episode," "Send / Publish Episode," "Set Personal History," and "Back."

[0082] ("Setting your personal history" mode) If "Set personal history" is selected in step 116, the process proceeds to step 118, where the personal history setting screen (see screen C shown in Figure 7(B)) is displayed, and then the process proceeds to step 120, where the personal history setting mode processing is executed, and the process returns to step 100.

[0083] ("Tell a story" mode)

[0084] If "Tell an episode" is selected in step 116, the process proceeds to step 122, where an episode selection screen (see screens D1 and D2 shown in FIGS. 8(A) and 8(B)) is displayed. These screens D1 and D2 are displayed in chronological order.

[0085] First, screen D1 displays options for eras (the eras listed in Table 1, including birth, childhood, elementary school, junior high school, high school, and working life), and when one is selected, the screen switches to screen D2, which displays options for topics (themes listed in Table 1, including sports day, classroom, and Valentine's Day).

[0086] As an example, in Figure 5, when the elementary school period is selected on screen D1 and the sports day is selected on screen D2, the process moves from step 122 to step 124, where an instruction to start the generated AI is given, and then the process moves to step 126, where the interactive mode screen (see screen E in Figure 8(C)) is displayed and the interactive mode is executed, and the process returns to step 100.

[0087] The screen E has a conversation display section where the status of the voice conversation is displayed in text in real time, so that the user 12 can recognize the content of the conversation not only by hearing but also by sight.

[0088] ("Watch / Edit Episode" mode) If "View / Edit Episode" is selected in step 116, the process proceeds to step 128, where a screen for viewing / editing episodes (see screen F shown in FIG. 9(A)) is displayed. On this screen F, the options displayed are "Date," "People," "Back," and "Other Episodes." A list of conversation files categorized by the date tag of the conversation and the person being reminisced about in the conversation (person tag) is displayed.

[0089] Screen F displays a list of conversation files with the date tag selected. Screen F displays each conversation file, "Watch other episodes," and "Go back" as options.

[0090] When one conversation file is selected on this screen F, a screen showing the contents of the conversation file (see screen F1 shown in FIG. 9(B)) is displayed. On screen F1, "Edit" and "Back" are displayed as options.

[0091] When "Edit" is selected on screen F1, a screen is displayed in which the contents of the displayed conversation file can be edited (see screen F11 shown in Figure 9(C)), and user 12 can edit the contents as needed (proceed from step 128 to step 130).

[0092] On the other hand, when a person tag is selected on screen F with the date tag selected, a list of the people who were the subject of the recollection is displayed (see screen F2 in Figure 9(D)). Screen F2 displays the options "Dad," "Mom," "Akemi," "Masahiro," "Satoshi," "Other Episodes," and "Back."

[0093] On this screen F2, for example, if "Mom" is selected, a conversation list of episodes related to the mother is displayed (see screen F21 in FIG. 9(E)). On screen F21, the options displayed are each conversation file, "Watch other episodes," and "Go back." The subsequent processing is the same as the flow from screen F to screen F1 to screen F11, so a description thereof will be omitted here (proceed from step 128 to step 130).

[0094] ("Send / Publish Episode" mode) If "Send / Publish Episode" is selected in step 116, the process proceeds to step 132, where a screen for sending / publishing episodes (see screens G1 and G2 shown in Figures 10(A) and 10(B)) is displayed. On this screen G1, the options displayed are "Send List," "Publish List," "Back," and "Send." A list of conversation files categorized into send list tags and publish list tags is displayed.

[0095] Screen G1 displays a list of conversation files with the send list tag selected, and screen G2 displays a list of conversation files with the publish list tag selected.

[0096] If an option is selected in step 132 , the send / publish mode is executed in step 134 and the process returns to step 100 .

[0097] By publishing an episode, users other than the user 12 can view the conversation file. However, even if the user 12 does not publish the conversation file, the conversation file may be viewable upon authentication by the user 12.

[0098] The above is the basic processing flow of the autobiography creation scenario program. By combining the generation AI function of the interactive generation AI execution unit with the scenario program for creating an autobiography using the reminiscence method, user 12 can reminisce in small chunks at their own pace, without having to make time in advance, compared to reminiscing through conversation in a physical face-to-face setting.

[0099] (Comparison with existing technology) Tables 2 and 3 compare the issues with existing technologies (requesting a personal history creation service or creating one yourself) with the effects that can be solved by the first embodiment (combination of a scenario program and generation AI) regarding the procedures for creating a personal history.

[0100] First, Table 2 lists the challenges associated with existing techniques (requesting autobiography writing services or writing your own).

[0101] [Table 2]

[0102] Next, Table 3 lists the effects of the first embodiment.

[0103] [Table 3]

[0104] (Judgment of completion of answer) There are cases where the user 12 thinks about a question and has difficulty finding an answer. In such cases, the dialogue-based generation AI execution unit 52 analyzes the silent state, such as growling, coughing, sneezing, or muttering, and may arrive at an answer that is different from the user 12's intention.

[0105] Therefore, in the first embodiment, the silent state of the user 12 is detected, and the dialogue-based generation AI execution unit 52 is instructed as to whether or not to end the answer depending on the so-called "pause" specific to the user 12.

[0106] 11A and 11B are flowcharts showing a control process for determining whether or not a response has ended, using the user confrontation information acquisition unit 70 and the response end time determination unit 72 shown in FIG.

[0107] Here, before the answer end determination interrupt routine of Fig. 11(B) is executed, a user-specific predetermined time setting interrupt routine (see Fig. 11(A)) is executed to set, on a regular or irregular basis, a predetermined time ts to be compared with a silent time, which is a criterion for determining whether an answer has ended, for each user 12. First, the flow of the user-specific predetermined time setting interrupt routine of Fig. 11(A) will be described.

[0108] As shown in FIG. 11(A), in step 150, the silent time t from the transmission of the question to the start of the answer voice is measured.

[0109] Next, in step 152, the measured silent time t is stored, and the process proceeds to step 154.

[0110] In step 154, it is determined whether it is time to update the predetermined time ts. Although details are omitted, n silent periods t (1 to n) are stored until it is time to update the predetermined time ts.

[0111] If the determination in step 154 ​​is negative, the process returns to step 150 and the above steps are repeated. If the determination in step 154 ​​is positive, the process proceeds to step 156, where the history of stored silent periods t(1 to n) is also synchronized, the predetermined time ts corresponding to the so-called "pause" specific to the user 12 is updated, and the process returns to step 150. The stored silent periods t(1 to n) may be reset when updated, or the number of stored periods may be increased by a fixed amount or without limit.

[0112] FIG. 11(B) is a control flowchart showing the answer completion determination interrupt routine.

[0113] In step 160, the current predetermined time ts is read, and then the process proceeds to step 162 to determine whether or not the system is waiting for an answer. If the determination in step 162 is negative, it is determined that it is not time to determine whether or not the system is finished answering, and the process returns to step 160. If the determination in step 162 is positive, it is determined that it is time to determine whether or not the system is at the start of a silent period.

[0114] If the determination in step 164 is negative, it is determined that the user is currently answering, and the process returns to step 160. If the determination in step 164 is positive, it is determined that the utterance for answering has stopped, the predetermined time ts specific to the user 12 is read, and the process then proceeds to step 166, where the timer 38 is reset and started. Note that the reset may be performed at other times, such as when the timer is stopped (described later), but since it is important that the timer is reset reliably before starting, the timer is reset immediately before starting.

[0115] In the next step 170, it is determined whether the timer measurement value has passed the predetermined time ts read out in step 166.

[0116] If the determination in step 170 is negative, the process proceeds to step 172 to determine whether or not a reply voice has been detected. If the determination in step 172 is negative, it is determined that the silent state continues, and the process returns to step 170.

[0117] If the determination in step 172 is affirmative, it is determined that the silent state has ended, and the process proceeds to step 174, where the timer 38 is stopped, and the process returns to step 160. As mentioned above, the timer may be reset immediately after the timer is stopped.

[0118] On the other hand, if the determination at step 170 is affirmative, the process proceeds to step 176 to stop the timer 38, and then to step 178 to instruct the dialogue-based generation AI execution unit 52 to end the answer, and the process returns to step 160.

[0119] Here, depending on the user 12, there may be a so-called "pause" before the answer is given, and during this "pause" sounds other than the speech (e.g., grunting, coughing, sneezing, mumbling, etc.) may cause misinterpretation in the voice analysis by the dialogue generation AI execution unit 52.

[0120] Therefore, by performing an answer determination process (see Figures 11(A) and (B)) using the user confrontation information acquisition unit 70 and answer end time determination unit 72 of Figure 3, it is possible to avoid misinterpretation of voice when the user 12 receives a question and answers it.

[0121] In the first embodiment, as shown in FIG. 11(A), the predetermined time for determining the end of the response based on the silent period (so-called "pause") is learned periodically or irregularly for each user 12, but the predetermined time may be fixed uniformly or fixed for each user 12.

[0122] In addition, in the first embodiment, the silent time was measured and compared with a predetermined time ts to determine the end time of the answer, but the predetermined time ts may also be set using the history of the ratio of silent time to speaking time of user 12 to determine the end time of the answer.

[0123] For example, there are users 12 who speak quickly, those who speak slowly, those who have short silent periods, and those who have long silent periods, and so on. Therefore, if the predetermined time ts is set based on the history of the ratio of silent periods to speaking periods for each user, it is possible to set a predetermined time ts that is more suitable for each user than a uniquely determined predetermined time ts.

[0124] More specifically, for a user 12 who tends to speak quickly and have short speech periods and short silence periods, the predetermined time ts can be set short. On the other hand, for a user 12 who tends to speak slowly and have long speech periods and long silence periods, the predetermined time ts can be set long.

[0125] Furthermore, in the first embodiment, the silent time is measured and compared with a predetermined time ts to determine the end time of the answer, but the end time of the answer may also be determined by detecting the line of sight or facial expression of the user 12 and based on the direction of the line of sight, the joy, anger, sadness, or happiness of the facial expression, etc. The end time of the answer may also be determined by combining at least two or more of "pause," "line of sight," and "facial expression." When two or more factors are combined, a weight (importance) may be set for each of them.

[0126] (Thematic regression control) Incidentally, when the user 12 is having a conversation based on reminiscence, the user 12 may have a conversation that deviates from the original topic. A conversation that deviates from the original topic may ultimately result in wasted time in creating a summary based on the original topic, so it is preferable to have a conversation based on the original topic as much as possible.

[0127] Therefore, in the first embodiment, when the content of the conversation of the user 12 deviates from the original topic, a question or pose is made to bring the conversation back to the original topic.

[0128] FIG. 12 is a control flowchart showing the flow of a theme deviation determination interrupt routine in the theme return instruction unit 68 of FIG.

[0129] In step 180, the set theme (initial theme) is read out, and then the process proceeds to step 182, where a request is made to the dialogue generation AI execution unit 52 to analyze the theme being dialogued.

[0130] In the next step 184, it is determined whether the initial theme and the theme being dialogued match. If the determination in step 184 is negative, it is determined that the themes match, and the process returns to step 180. If the determination in step 184 is positive, the process proceeds to step 186, where a request is made to the dialogue-based generation AI execution unit 52 for a dialogue that will return to the set theme (the initial theme), and the process returns to step 180.

[0131] While answering questions, user 12 may give answers that deviate from the original topic as memories gradually return to them. This is not a problem from the perspective of recollection for user 12, but from the perspective of the original purpose of creating a summary of the original topic in a short period of time, deviation from the topic poses a problem.

[0132] Therefore, by carrying out a dialogue that brings the user 12 back to the original topic so as not to diminish the user's desire to reminisce, the user 12 can quickly return to the original topic without losing his or her desire to reminisce.

[0133] (Second embodiment) A second embodiment of the present invention will be described below. In the second embodiment, the same components as those in the first embodiment are denoted by the same reference numerals, and a description of the components will be omitted.

[0134] The feature of the second embodiment is that it is a technology related to a digital Buddhist altar that, when the user 12 passes away, makes use of the summaries by era or theme that the user 12 used to create his / her autobiography while he / she was alive, and when the user 12 has passed away, the user's relatives (hereinafter referred to as "users") ask the user 12 questions in front of the Buddhist altar, the user 12 will answer in his / her voice.

[0135] As background, there is technology that records the voice of the deceased and outputs multiple standard voice data.

[0136] FIG. 13 is a schematic diagram of a digital Buddhist altar system according to a second embodiment, in which a digital Buddhist altar user terminal device 74 is added to the autobiography creating system shown in FIG.

[0137] As shown in FIG. 14, a voice conversation generation device 16A according to the second embodiment includes a digital Buddhist altar control unit 76.

[0138] As described above, the recollection information and summary information storage unit 62 stores summaries for each era or each theme when the personal history was created between the user 12 and the terminal device 14 while the user was alive.

[0139] The interactive generation AI execution unit 52A according to the second embodiment mainly includes a specific user summary information reading function, a voice analysis function, a response generation function, and a specific user synthetic voice generation function.

[0140] The digital Buddhist altar correspondence control unit 76 includes a user identification unit 78 and is connected to the digital Buddhist altar user terminal device 74. The user identification unit 78 instructs the interactive generation AI execution unit 52A to converse with the user 12 (deceased person) specified by the user ID or the like from the digital Buddhist altar correspondence control unit 76.

[0141] Upon receiving this instruction, the interactive generation AI execution unit 52A reads out summary information of the specific user 12 (deceased person), performs voice analysis, and generates a synthetic voice of the specific user.

[0142] When a voice is output from the digital Buddhist altar user terminal device 74, this voice is acquired by the voice information acquisition unit 80. The voice information acquisition unit 80 is connected to the request / response analysis unit 82, which analyzes the requested response based on the voice information and sends the analysis result to the interactive generation AI execution unit 52A.

[0143] The interactive generation AI execution unit 52A generates an answer and outputs a synthetic voice of the specific user 12. The output synthetic voice is acquired by the answer acquisition unit 84 and sent to the digital Buddhist altar user terminal device 74 via the answer voice output unit 86.

[0144] In other words, when the user speaks from the digital Buddhist altar user terminal device 74, an accurate response is given in the voice of a specific user 12, which creates a greater level of emotion than when simply listening to conventional standard voice data (it is possible to create a situation where an actual conversation is taking place), thereby deepening the relationship between the user and the deceased user 12.

[0145] The operation of the second embodiment will be described below with reference to the flowchart of FIG.

[0146] In step 200, it is determined whether or not the user ID has been acquired. If the determination in step 200 is negative, the process proceeds to step 214.

[0147] If the determination in step 200 is affirmative, the process proceeds to step 202, where the user 12 (deceased user 12) is identified, and the process proceeds to step 204.

[0148] In step 204 , summary information of the identified user 12 (specific user 12 ) is read, and the process proceeds to step 206 , where a synthesized voice of the specific user 12 is generated, and the process proceeds to step 208 .

[0149] In step 208, it is determined whether a question or the like has been received from the user. If the determination in step 208 is negative, the process proceeds to step 214. If the determination in step 208 is positive, the process proceeds to step 210, where the voice is analyzed and an answer is generated from the summary information, and the process proceeds to step 212.

[0150] In step 212, a response is output in the voice of the specific user 12, and the process proceeds to step 214.

[0151] In step 214, it is determined whether or not the use of the digital Buddhist altar user terminal device 74 has ended, and if the determination is negative, it is determined that the conversation is still continuing, and the routine returns to step 208. On the other hand, if the determination is positive in step 214, it is determined that the conversation has ended, and this routine ends.

[0152] In the second embodiment, the user's question is answered in the voice of a specific user 12, but the system may also be configured to present related information and return a question in response to a user's casual muttering. For example, if the user says, "I went to my grandchild's sports day the other day," the digital Buddhist altar system may read information about the sports day from the summary information and return a question such as, "When I was a kid, the horse-riding battle was a big hit at sports days, but I wonder what kind of events are popular now?" This allows for a conversation that is not one-sided.

[0153] As explained above, in the digital Buddhist altar system according to the second embodiment, when a user 12 who created a personal history dies and the user (a relative, etc.) speaks to the deceased user 12 in front of the Buddhist altar, the content that matches the content and is based on the personal history that the user created while the user was alive is output in the voice of the deceased.

[0154] As a result, instead of a fixed conversation, the interactive generation AI execution unit 52A can extract content from the autobiography or generate new sentences based on the registered content of the autobiography, and answer questions from the user in the voice of the deceased user 12. This allows the relationship between the user and the deceased user 12 to deepen in a situation where they are actually having a conversation.

[0155] (Addendum) [Appendix 1] The system is executed using a predetermined scenario program and an interactive generation AI execution unit that assists the progress of the scenario program, and is provided with a voice conversation function that receives answers from the user by providing the user with questions to elicit information related to a specific theme that prompts the user to reminisce about their personal history, and has a recollection control unit that accumulates multiple pieces of recollection information of the user obtained by repeating the voice conversation function, The interactive generation AI execution unit A control device having a user synthetic voice generation function for generating synthetic voice of the user, which generates a response to a question posed by a user wishing to converse with the user based on the accumulated recollection information of the user, and when outputting the response aloud, generates and outputs a synthetic voice similar to that of the user.

[0156] [Appendix 2] The recollection control unit 2. The control device according to claim 1, wherein a summary of the user's personal history is created for each era or specific theme based on the recollection information.

[0157] [Appendix 3] 3. The control device according to claim 1, further comprising: a measuring unit that measures a silent time from when the question is provided until the user starts to answer; and a determining unit that determines when the user has finished answering based on the silent time measured by the measuring unit.

[0158] [Appendix 4] The system further includes a return instruction unit that, when the user's speech deviates from an initial theme based on a scenario program during execution of the voice conversation function, returns the user to the initial theme through free conversation by the dialogue generation AI execution unit. A control device according to any one of appendices 1 to 3.

[0159] [Appendix 5] Computer, Operated as a control device according to any one of appendices 1 to 4. Digital Buddhist altar control program. [Explanation of symbols]

[0160] (First embodiment) NW Network 10 Autobiography Creation System 12 users 14 Terminal Equipment 16 Voice conversation generation device 18. Mike 20 speakers 22 Display section 24 Control section 26 Control device 28 UI (User Interface) 26A processor 26B ROM 26C RAM 26D Storage 26E Communication Interface 26F Input / Output Interface 30 processors 32 ROM 34 RAM 36 Storage 38 Timer 38 42 Internal Bus 50 Control section (recollection control section) 52 Interactive Generative AI Execution Unit 54 Personal History Writing Reception Department 56 Voice information transmission unit 58 Voice information receiving unit 60 Scenario Database 62 Recollection and summary information storage section 64 Personal History Timeline Creation Department 66 Update section 68 Thematic Return Instructions 70 User-facing information acquisition unit 72 Response end time determination section (Second embodiment) 16A Voice conversation generating device 52A Interactive Generative AI Execution Unit 74 Digital Buddhist Altar User Terminal Device 76 Digital Buddhist Altar Control Unit 78 User Identification Part 80 Audio information acquisition unit 82 Request and response analysis section 84 Answer acquisition part 86 Answer voice output section

Claims

1. The system is executed using a predetermined scenario program and an interactive generation AI execution unit that assists the progress of the scenario program, and is provided with a voice conversation function that receives answers from the user by providing the user with questions to elicit information related to a specific theme that prompts the user to reminisce about their personal history, and has a recollection control unit that accumulates multiple pieces of recollection information of the user obtained by repeating the voice conversation function, The interactive generation AI execution unit A control device having a user synthetic voice generation function for generating synthetic voice of the user, which generates a response to a question posed by a user wishing to converse with the user based on the accumulated recollection information of the user, and when outputting the response aloud, generates and outputs a synthetic voice similar to that of the user.

2. The recollection control unit The control device according to claim 1 , wherein a summary of the user's personal history is created for each period or each specific theme based on the recollection information.

3. 2. The control device according to claim 1, further comprising: a measuring unit that measures a silent time from when the question is provided until the user starts to answer; and a determining unit that determines when the user has finished answering based on the silent time measured by the measuring unit.

4. The control device according to claim 1, further comprising a regression instruction unit that, if the user's speech content deviates from an original theme based on a scenario program during execution of the voice conversation function, returns the user to the original theme through free conversation by the interactive generation AI execution unit.

5. Computer, Operated as the control device according to any one of claims 1 to 4. Digital Buddhist altar control program.

Citation Information

Patent Citations

  • Method and device for on-hook detection, and method and device for continuous voice recognition

    JP1996115093A

  • Voice management server device, conversation voice provision method, and conversation voice provision system

    JP2016110151A

  • Voice recognition system

    JP2019105735A

  • Individual reminiscence sharing system

    JP2020057350A