Information processing device, information processing method, program, and storage medium

JP7919790B1Active Publication Date: 2026-09-14SIDEPEAK CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2026089218
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2026-04-28
Filing Date
2026-05-27
Publication Date
2026-09-14
Estimated Expiration
2046-05-27

AI Technical Summary

Benefits of technology

【0007】 本実施形態の情報処理装置は、接触情報及び音声情報から推定されるユーザの居場所感を、複数の異なるレベルの居場所感として定めた判断基盤を用いて、ユーザの居場所感のレベルを推定し、ユーザの居場所感のレベルに応じた応答音声を作成する。したがって、ユーザ端末で出力される応答音声を生成するにあたり、ユーザの居場所感に対する応答音声の生成性能及び生成精度を向上できる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007919790000001_ABST
    Figure 0007919790000001_ABST
Patent Text Reader

Abstract

This information processing device uses both contact information and voice information to estimate the user's psychological state, taking advantage of the characteristics of each type of information, and creates a response voice that is adapted to the estimated "level corresponding to the user's psychological state." [Solution] A management computer 11 having a processor 47, the processor 47 is provided with an information processing device that performs the following: a first process of acquiring contact information of the user 15 with an object 13; a second process of acquiring voice information of the user 15; a third process of estimating the level of the user 15's psychological state using a judgment base that includes a database associating multiple different psychological states with the level of said psychological state, and a learning model that takes the contact information and voice information as input and outputs the level of the user 15's psychological state; and a fourth process of creating a response voice that corresponds to the level of the user 15's psychological state and can be output by the user terminal 12 used by the user 15.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an information processing apparatus, an information processing method, a program, and a storage medium that process information in which a user interacts with an object.

Background Art

[0002] An example of an entertainment apparatus that processes information in which a user interacts with an object is described in Patent Document 1. The entertainment apparatus described in Patent Document 1 comprises: a wireless receiver configured to receive, from an interactive toy, data held by the interactive toy, the data describing physical capabilities of the interactive toy during operation and defining one or more sensors included in the interactive toy and one or more actuators included in the interactive toy; a processor that generates an interactive command signal; and a wireless transmitter configured to transmit the interactive command signal to the interactive toy during operation. The processor is further configured to generate the interactive command signal in response to the data describing the physical capabilities of the interactive toy during operation.

Prior Art Literature

Patent Literature

[0003]

Patent Document 1

Summary of the Invention

Problem to be Solved by the Invention

[0004] The inventor of the present application has recognized that, with the entertainment apparatus described in Patent Document 1, the problem arises that the processing load on a computer, which integrates heterogeneous data obtained from a plurality of sensors in real time and generates an appropriate response voice promptly in response to the user's psychological state, increases, and the generation of the response voice is delayed.

[0005] The purpose of this disclosure is to provide an information processing device, information processing method, program, and storage medium that can estimate the user's psychological state by utilizing the characteristics of both contact information and voice information, and select a response voice that is appropriate to the estimated "level corresponding to the user's psychological state." [Means for solving the problem]

[0006] This embodiment discloses an information processing device having a memory and a processor coupled to the memory, wherein the processor executes program instructions stored in the memory to perform the following: a first process of acquiring contact information indicating the physical contact state of a user with an object; a second process of acquiring voice information of a user using the object; a third process of estimating the level of the user's psychological state using a judgment platform that includes a database associating a plurality of different psychological states with the level of said psychological state, and a learning model that takes the contact information and the voice information as input and outputs the level of the user's psychological state; and a fourth process of creating a response voice that corresponds to the level of the user's psychological state estimated in the third process and can be output by a user terminal used by the user, using at least one of the response voice database or the response voice generation model stored in the memory. [Effects of the Invention]

[0007] The information processing device of this embodiment estimates the user's sense of location level using a judgment base that defines the user's sense of location, estimated from contact information and voice information, as multiple different levels of sense of location, and creates a response voice corresponding to the user's sense of location level. Therefore, when generating a response voice output at the user terminal, the generation performance and generation accuracy of the response voice corresponding to the user's sense of location can be improved. [Brief explanation of the drawing]

[0008] [Figure 1]This is a conceptual diagram showing an example configuration of a management computer and user terminal included in an information processing system. [Figure 2] This is a conceptual diagram showing an example of the configuration of objects included in an information processing system. [Figure 3] This is a conceptual diagram showing the functional configuration realized by the processor of the management computer. [Figure 4] This is a flowchart showing the first specific example of an information processing method performed by an information processing system. [Figure 5] This diagram shows an example of the data structure of the database used by the management computer to estimate the user's psychological state, level of psychological state, and response voice. [Figure 6] This flowchart shows a second specific example of an information processing method performed by an information processing system. [Modes for carrying out the invention]

[0009] (Description of the information processing system) Figures 1 and 2 show the overall configuration of the information processing system 10. The information processing system 10 includes devices such as a management computer 11, a user terminal 12, and an object 13 as constituent elements. The user terminal 12 and the object 13 are used by the user 15. The user terminal 12 and the management computer 11 can communicate bidirectionally via the network 16. The user terminal 12 and the object 13 can communicate bidirectionally via the network 17.

[0010] Information is transmitted and received between the management computer 11 and the user terminal 12. Information is also transmitted and received between the user terminal 12 and the object 13. The management computer 11 may be configured to connect to an external computer 52 other than the user terminal 12 via the network 16. The external computer 52 may provide various types of information to the management computer 11 via the network 16. In this embodiment, the various types of information transmitted and received between the management computer 11 and the user terminal 12, between the object 13 and the user terminal 12, and between the external computer 52 and the management computer 11 may include information, data, and commands.

[0011] Network 16 consists of one or more communication systems, which are either wireless communication systems or wired communication systems. Wireless communication systems include radio wave communication, optical communication, infrared communication, radio wave communication, satellite communication, etc. Wired communication systems include communication circuits and communication cables. Network 16 includes one or more networks, which are the Internet, intranet, wide area network, and intranet. Network 17 includes short-range wireless communication systems. Short-range wireless communication systems are wireless communication systems that transmit information using radio waves or light, etc., with a communication distance of 5m to 15m. Short-range wireless communication systems include, for example, wireless LAN (Wi-Fi®) and Bluetooth®.

[0012] (Description of the object) The object 13 may be, for example, a pillow, cushion, doll, stuffed animal, etc. When the user 15 touches the surface of the object 13 with a part of their body, for example, their hand, the object 13 deforms due to the contact pressure. The user 15 may carry the object 13.

[0013] User 15 can confirm their sense of belonging by interacting with object 13. Interaction includes touching object 13 with a part of User 15's body, and User 15 emitting sounds to simulate a conversation with object 13. User 15's sense of belonging is the feeling that "it's okay for me to be here" and "this is where I belong." For example, User 15's sense of belonging includes feeling that their existence is acknowledged, feeling that it's okay to be themselves, feeling calm and secure, and feeling that they are useful in some way. User 15's sense of belonging may also be expressed as a level value that indicates User 15's psychological state.

[0014] The object 13 may have a structure in which the outside of the main body is covered with a cover. The main body may be made of synthetic resin, silicone, sponge, etc. The cover may be made of cloth, nonwoven fabric, synthetic resin sheet, etc. The object 13 may be provided with a processor 19, main memory 20, auxiliary memory 21, pressure sensor 22, speaker 23, communication device 25, power supply 26, and identification information display unit 71, as shown in Figure 2.

[0015] The identification information display unit 71 displays unique object information for each object 13. The object information includes an identifier for the object 13, and the identifier for the object 13 serves to identify the type and design of the object 13, and the character attached to the cover of the object 13. The object information may be displayed as, for example, a one-dimensional code, a two-dimensional code, a serial number, etc. A one-dimensional code includes a barcode. A two-dimensional code includes a QR code (registered trademark). The identification information display unit 71 may be attached to the object 13 by either printing or pasting.

[0016] The electrical components and electronic components provided on the main body function with electric power supplied from the power source 26. The processor 19 may have a structure communicatively connected to the main memory 20, the auxiliary memory 21, the pressure sensor 22, the speaker 23, and the communication device 25 via a communication bus 27. Further, the main memory 20 may have a structure built into the processor 19. As described above, the processor 19 is a processing circuit having a structure communicatively coupled to the main memory 20 and the auxiliary memory 21. The processor 19 may be provided inside the main body and configured by a central processing unit (CPU) in which an arithmetic unit (arithmetic circuit) and a control unit (control circuit) are integrated.

[0017] Further, instead of the central processing unit, the processor 19 may be implemented by an arithmetic processing circuit such as a digital signal processor, an application specific integrated circuit (ASIC), a graphics processing unit (GPU), or a micro processing unit (MPU). Image information and sound information may be processed by the GPU.

[0018] The main memory 20 is a volatile storage device, and functions as a work area and a buffer area when the processor 19 executes processing. The auxiliary memory 21 includes a non-transitory storage medium. A non-transitory program is stored in the auxiliary memory 21. Further, information processed by the processor 19 is stored in the auxiliary memory 21. The processor 19 may execute various processes by storing the non-transitory program read from the auxiliary memory 21 into the main memory 20 and executing instructions.

[0019] The processor 19 may execute processing for outputting sound information from the speaker 23, processing for a signal detected by the pressure sensor 22, processing for determining an operation of the operation unit 24, processing for transmitting information to the user terminal 12 via the communication device 25, processing for information acquired from the user terminal 12 via the communication device 25, and the like.

[0020] A plurality of pressure sensors (pressure-sensitive devices) 22 are provided in a planar region at a predetermined position of the object 13, for example, a position of the outer surface of the main body covered by the cover. The pressure sensor 22 may be of any type, for example, among a piezoelectric element type, a capacitance type, a resistive film type, an optical type, MEMS (Micro Electro-Mechanical System), and the like. The pressure sensor 22 detects pressure received when the user 15 brings a part of the body into contact with the object 13, and outputs a signal.

[0021] The speaker 23 is an electronic device that outputs sound information in response to receiving a control signal from the processor 19. The sound information may include speech, musical instrument sounds, artificial sounds, and the like. The speaker 23 may output sound information generated by the management computer 11 and acquired via the user terminal 12. The operation unit 24 may have a structure of at least one or more among various switches, buttons, levers, touch panels, and the like provided on the main body of the object 13.

[0022] The communication device 25 may have a structure of at least one or more among a communication circuit, a communication port, a communication cable, an antenna, and the like. The communication device 25 is connected to the user terminal 12 via the network 17. The power source 26 may be either a DC power source or an AC power source. The DC power source may be either a primary battery or a secondary battery.

[0023] (Description of User Terminal) The user terminal 12 is a computer used by the user 15. The user terminal 12 may be either a fixed computer fixedly placed on a table, or a portable computer that can be carried and moved by the user 15. The portable computer may be implemented by a smartphone, a tablet computer, a notebook computer, or the like.

[0024] The user terminal 12 may include a main unit (casing), a processor 37, a main memory 38, an auxiliary memory 39, an input device 40, an output device 41, a communication device 45, and a current location determination unit 62. The processor 37 is located inside the main unit and is implemented as a central processing unit (CPU) in which an arithmetic unit (arithmetic circuit) and a control device (control circuit) are integrated. The processor 37 may be connected to the main memory 38, auxiliary memory 39, input device 40, output device 41, and communication device 45 via a communication bus 46. The main memory 38 may be built into the processor 37. The processor 37 is a processing circuit with a structure that is connected to the main memory 38 and the auxiliary memory 39 in a manner that enables communication.

[0025] Furthermore, the processor 37 may be implemented by a arithmetic processing circuit such as a digital signal processor, an application-specific integrated circuit (ASIC), a GPU (Graphics Processing Unit), or an MPU (Micro Processing Unit) instead of a central processing unit. Image information and sound information may be processed by the GPU.

[0026] The main memory 38 is a volatile storage device and functions as a work area and buffer area when the processor 37 performs processing. The auxiliary memory 39 may be implemented by a main unit and a non-temporary storage medium 39A. The auxiliary memory 39 stores non-temporary programs, various information and data used by the processor 37 to perform various processes, and various information and data as a result of the processor 37 performing various processes. The programs stored in the auxiliary memory 39 may be programs downloaded and installed from the management computer 11. The auxiliary memory 39 operates according to input and output instructions from the processor 37.

[0027] The storage medium 39A can be implemented by, for example, a magnetic disk, an optical disk, flash memory, etc. An example of a magnetic disk is a hard disk drive. Examples of optical disks include compact discs, digital video discs, Blu-ray discs, etc. Flash memory is a type of semiconductor memory, and examples of flash memory include SD memory cards, USB flash drives, solid-state drives, etc. The storage medium 39A may be fixed to the main unit or can be attached to and removed from the main unit; either structure is acceptable.

[0028] User 15 can input various types of information to the user terminal 12 using the input device 40. This information may include image information, sound information, and text data. The input device 40 may be implemented using a microphone 42, a display 44, a mouse 91, a keyboard 92, an image acquisition device 70, etc. The display 44 may include at least one element from, for example, a liquid crystal display or an organic electroluminescent display.

[0029] The display 44 is controlled based on control signals from the processor 37. The display 44 may show a website, a user registration screen, a response information selection screen, a cursor, operation tabs, etc. The microphone 42 is a device that acquires the voice emitted by the user 15 and converts it into an electrical signal for output. The image acquisition device 70 captures at least a portion of the object to be photographed to generate an image and converts the generated image into an electrical signal. The image acquisition device 70 includes, for example, a camera, a scanner, a barcode reader, an optical character recognition device, etc.

[0030] The objects to be photographed by the image acquisition device 70 may include the object 13, the identification information display unit 71, the user's face 15, documents, etc. The image format may include video, still images, photographs, etc. The image content may include printed text, handwritten characters, illustrations, symbols, etc. Images acquired by the image acquisition device 70 may be stored in the auxiliary memory 39.

[0031] The output device 41 may be implemented by an audio information output device 43, a display 44, a printer, etc. The printer is a device that prints information onto paper based on control signals output from the processor 37. The audio information output device 43 may be implemented by a speaker, headphones, earphones, etc. The audio information output device 43 may be connected to the processor 37 by wired or wireless communication to the communication bus 46. The audio information output device 43 is a device that outputs audio information based on control signals output from the processor 37.

[0032] The communication device 45 includes a communication circuit, communication cable, antenna, etc. The communication device 45 is connected to the management computer 11 via the network 16, the communication device 45 is connected to the object 13 via the network 17, and the communication device 45 is connected to the object 13 via the network 17.

[0033] The current location determination unit 62 is composed of, for example, a detection circuit and a determination circuit. The current location determination unit 62 processes signals received from the satellite via the communication device 45, and signals from the gyro sensor and magnetic sensor installed in the user terminal 12. Based on these processing results and the map data stored in the auxiliary memory 39, the current location determination unit 62 determines the current location of the user terminal 12 and may also determine the user terminal 12's movement path. The user terminal 12's current location and movement path are the user terminal 12's current location and movement path in the map data. The user terminal 12's current location can be understood as the current location of the user 15 who possesses the user terminal 12. The user 15 may operate the input device 40 to display the map data on the display 44.

[0034] Furthermore, the current location determination unit 62 may process the communication status between the communication device 45 and the multiple access points constituting the short-range communication system, and determine the current location of the user terminal 12 based on the processing results. Note that the current location determination unit 62 may be integrated into the processor 37.

[0035] The processor 37 performs various processes by executing instructions for non-temporary programs stored in the auxiliary memory 39. The processes performed by the processor 37 include calculations, decisions, comparisons, control, and storage in the auxiliary memory 39. The processes performed by the processor 37 may also include reading information from the auxiliary memory 39, processing information obtained from the management computer 11, and processing information obtained from the object 13.

[0036] The information acquired by the processor 37 from the object 13 includes contact information determined from the signal of the pressure sensor 22. The processing performed by the processor 37 may include processing of various information and commands input using the input device 40, processing of the output device 41, processing of various information to be sent to the management computer 11, and processing of various information to be sent to the object 13.

[0037] (Description of the management computer) The management computer 11 may be managed and used by an administrator. The management computer 11 may have a main unit (casing), a processor 47, main memory 48, auxiliary memory 49, input device 73, output device 74, and communication device 50. The processor 47 may be located inside the main unit and consist of a central processing unit (CPU) in which an arithmetic unit (arithmetic circuit) and a control device (control circuit) are integrated.

[0038] Furthermore, the processor 47 may be implemented by a digital signal processor, an application-specific integrated circuit (ASIC), a graphics processing unit (GPU), or a microprocessing unit (MPU) in addition to the central processing unit, instead of the central processing unit. Image information and sound information may be processed by the GPU.

[0039] Furthermore, the processor 47 may be connected to the main memory 48, auxiliary memory 49, communication device 50, input device 73, and output device 74 via the communication bus 51. The main memory 48 may be built into the processor 47. The processor 47 is a processing circuit with a structure that is connected to the main memory 48 and auxiliary memory 49 in a manner that enables communication.

[0040] The processor 47 performs various processes by executing instructions for a non-temporary program. The various processes performed by the processor 47 may include calculations, decisions, comparisons, control, etc. The processes performed by the processor 47 may also include reading various information from the auxiliary memory 49, processing various information and commands obtained from the user terminal 12 and the external computer 52, adding and modifying various information stored in the auxiliary memory 49, processing various information and commands input using the input device 73, and processing signals that control the output device 74.

[0041] The main memory 48 is a volatile storage device, and when the processor 47 performs processing, the main memory 48 functions as a work area and a buffer area.

[0042] The auxiliary memory 49 is a non-volatile storage device and may be implemented by the main unit and a non-temporary storage medium 49A. The auxiliary memory 49 operates according to input and output commands from the processor 47. The storage medium 49A may be implemented by, for example, a magnetic disk, an optical disk, flash memory, etc. Examples of magnetic disks include hard disk drives. Examples of optical disks include compact discs, digital video discs, Blu-ray discs, etc. Flash memory is a type of semiconductor memory, and examples of flash memory include SD memory cards, USB flash drives, solid-state drives, etc. The storage medium 49A may be fixed to the main unit or may be removable from the main unit.

[0043] Furthermore, the storage medium 49A of the auxiliary memory 49 may be implemented as shown in Figure 3, comprising a program storage unit 81, various information storage units 82, and a model storage unit 83. The program storage unit 81 stores a non-temporary program 84. The various information storage unit 82 may store various information used by the processor 47 to execute various processes, various information before processing by the processor 47, various information after processing by the processor 47, and various information used by the artificial intelligence unit of the processor 47 to perform machine learning.

[0044] The database 82A of the various information storage unit 82 may store various information and data used by the processor 47 to perform various processes, and various information and data that are the results of the processor 47 performing various processes. For example, the various information storage unit 82 may store various information and database 82A used by the contact information processing unit 87 to determine contact information based on the signal from the pressure sensor 22. The various information storage unit 82 may also store database 82A used to estimate the psychological state (emotions) of the user 15 based on contact information. The various information storage unit 82 may also store database 82A used to estimate the psychological state (emotions) of the user 15 based on voice information. Furthermore, the various information storage unit 82 may also store database 82A used to estimate the psychological state (emotions) of the user 15 based on image information.

[0045] In other words, database 82A stores data structures that associate contact information with the psychological state (emotions) of user 15, data structures that associate voice information with the psychological state (emotions) of user 15, and data structures that associate image information with the psychological state (emotions) of user 15.

[0046] Furthermore, database 82A stores data structures for estimating the level of user 15's psychological state from user 15's voice information, data structures for estimating the level of user 15's psychological state from user 15's contact information, and data structures for estimating the level of user 15's psychological state from user 15's image information.

[0047] Furthermore, the various information storage units 82 may store, for example, various information and a database 82A for creating response voices based on the estimation results of the user's psychological state, such as the level of their sense of belonging. The model storage unit 83 may store a learning model 83A and a large-scale language model used by the artificial intelligence unit 59 for processing. The configuration and functions of the artificial intelligence unit 59 will be described later. In addition, the various information storage units 82 may store map data for the processor 47 to determine the current location of the user terminal 12.

[0048] The communication device 50 includes devices, equipment, and standards for connecting the management computer 11 to the network 16 by at least one communication system, either a wireless communication system or a wired communication system. The devices and equipment for connecting the management computer 11 to the network 16 may include one or more components, such as a communication circuit, a communication cable, an antenna, a communication port, and a repeater.

[0049] The input device 73 may be operated by an administrator. The input device 73 may be implemented as, for example, a mouse, keyboard, stylus, touch panel, image acquisition device, microphone, etc. The administrator may use the input device 73 to input various information and instructions to the management computer 11. The various information and instructions input using the input device 73 may be processed by the processor 47 and stored in the auxiliary memory 49.

[0050] The processor 47 may implement functional units such as the website management unit 54, user information processing unit 55, object information processing unit 56, artificial intelligence unit 59, sound information processing unit 60, image information processing unit 61, various information processing unit 57, contact information processing unit 87, text data processing unit 90, location sense determination unit 89, and response voice processing unit 88 shown in Figure 3 by executing instructions for non-temporary programs.

[0051] The website management unit 54 may perform processes such as providing a website at a predetermined URL (Uniform Resource Locator) on the network 16, providing various screens to computers connected to the website, providing programs to computers connected to the website, and mutually sending and receiving various information with computers connected to the website.

[0052] The user information processing unit 55 processes user information obtained from the user terminal 12 and may also store the user information in the auxiliary memory 49. User information may include items such as name, gender, date of birth, address (home address), telephone number, occupation, email address, user ID, password, object information, etc. Object information is information obtained from the object 13 owned by user 15. The user ID is for the purpose of individually identifying user 15 and may consist of numbers and symbols. The user information processing unit 55 may associate the user information with the object information of the object 13 owned by user 15 and store it in the auxiliary memory 49.

[0053] The object information processing unit 56 may process the object information displayed on the object identification information display unit 71 of the object 13, and may also store the object information in the auxiliary memory 49.

[0054] The contact information processing unit 87 may determine the user's contact information with the object 13 by processing the signal from the pressure sensor 22. The contact information may include the strength and duration of the user's grip on the object 13, the number of times the user strokes the object 13 within a predetermined time, the speed at which the user strokes the object 13, the force with which the user strikes the object 13, the number of times the user strikes the object 13 within a predetermined time, and so on. The contact information processing unit 87 may store the result of its determination of the user's contact information in the auxiliary memory 49.

[0055] The sound information processing unit 60 may process sound information acquired from the user terminal 12 and store the processing results in the auxiliary memory 49. The sound information acquired from the user terminal 12 may include the voice of the user 15. The sound information processing unit 60 may process the voice of the user 15 in cooperation with the artificial intelligence unit 59 and the text data processing unit 90 to estimate the emotions of the user 15, convert it into text data, and store that text data in the auxiliary memory 49.

[0056] Furthermore, the technology for estimating human emotions from human voice information is publicly known, as shown in, for example, Japanese Patent Publication No. 2874858, Japanese Patent Publication No. 3676969, Japanese Patent Publication No. 3676981, Japanese Patent Publication No. 4580190, Japanese Patent Publication No. 4670431, Japanese Unexamined Patent Publication No. 2012-59107, Japanese Unexamined Patent Publication No. 2015-141428, Japanese Unexamined Patent Publication No. 2021-11071, etc., so a detailed explanation will be omitted. The image information processing unit 61 may process the image information acquired from the user terminal 12 in cooperation with the artificial intelligence unit 59 and the text data processing unit 90 to estimate the emotions of the user 15, and may also store the emotion estimation result as text data in the auxiliary memory 49. Furthermore, since the technology for analyzing human emotions by processing image information of people is publicly known, as shown in, for example, Japanese Patent Publication No. 6868422, Japanese Patent Publication No. 6703893, Japanese Patent Publication No. 6042015, Japanese Patent Publication No. 3953024, Japanese Patent Publication No. 4458888, etc., a detailed explanation will be omitted.

[0057] The text data processing unit 90 may, in cooperation with the artificial intelligence unit 59, perform a process to convert the emotions of the user 15 obtained by processing image information and the emotions of the user 15 obtained by processing sound information into text data.

[0058] The sense of place determination unit 89 may, in cooperation with at least one of the artificial intelligence unit 59 or auxiliary memory 49, determine, that is, estimate, the user's psychological sense of place based on contact information, the user's voice information, and the determination base 85. The sense of place determination unit 89 may express the user's sense of place as a level corresponding to the level of the sense of place. The determination base 85 may be implemented by a learning model 83A used by the artificial intelligence unit 59 to process contact information and the user's voice information, a database 82A stored in the various information storage units 82 of the auxiliary memory 49 for the processor 47 to execute processing, instructions of the program 84, etc. The various information and database 82A used by the processor 47 for determination may be stored in the auxiliary memory 49, acquired by the management computer 11 from an external computer 52, or input to the management computer 11 from the input device 73.

[0059] User 15's sense of belonging means that, while having connections with people other than themselves, User 15 feels that "it's okay for me to be here," "I feel accepted by those around me," "I feel safe," "I can be myself," and "I feel that I'm useful here and my existence is acknowledged."

[0060] Psychological factors influencing User 15's sense of belonging may include acceptance, mental stability, freedom of action, ease of thinking and introspection, self-esteem, and freedom from others. User 15 may have a sense of belonging in various settings, such as school (class, club activities, friend groups, relationships with teachers), home (parent-child relationships, sibling relationships), workplace (department, team, relationships with superiors and colleagues), and relationships with lovers or partners. Examples of how to assess a sense of belonging will be discussed later.

[0061] The response voice processing unit 88 may, in cooperation with the location sense determination unit 89, the artificial intelligence unit 59, and the auxiliary memory 49, create a response voice that corresponds to the user 15's sense of location. The response voices created by the response voice processing unit 88 may store, for example, the voices of the user 15's family, male response voices, female response voices, anime character response voices, celebrity response voices, sports player response voices, etc. The item 80 and the QR code of the item will be described later.

[0062] The various information processing units 57 may perform processing that is not performed by the website management unit 54, user information processing unit 55, object information processing unit 56, sound information processing unit 60, image information processing unit 61, contact information processing unit 87, response voice processing unit 88, and location sense determination unit 89, and store the processing results in the auxiliary memory 49. In addition, the various information processing units 57 may process signals from the gyro sensor and magnetic sensor acquired from the user terminal 12, determine the current location of the user terminal 12 and the movement path of the user terminal 12 based on the map data stored in the auxiliary memory 49.

[0063] Furthermore, the various information processing units 57 may process at least one piece of information from among the signals from the pressure sensor 22, sound information, image information, etc., and, in cooperation with the artificial intelligence unit 59 and the auxiliary memory 49, determine (estimate) the psychological state of the user 15. The various information processing units 57 may also determine whether the psychological state of the user 15 estimated by processing the signals from the pressure sensor 22 is different from (discrepancies with) the psychological state of the user 15 estimated by processing the user's voice. In addition, if the various information processing units 57 determine that the psychological state of the user 15 estimated by processing the signals from the pressure sensor 22 is different from the psychological state of the user 15 estimated by processing the user's voice, it may store a log including the date, time (time zone), and user ID in the auxiliary memory 49.

[0064] Furthermore, the various information processing units 57 may determine a "specific tendency" in the psychological state of user 15 based on the fact that they have determined that the emotions of user 15 estimated by processing the signals from the pressure sensor 22 are different from the emotions of user 15 estimated by processing the user's voice. Here, "specific tendency" may mean that there is a period of time during which the user 15's action of contacting the object 13 and the voice that user 15 makes diverge. The predetermined period of time may be, for example, from 6:00 p.m. on the previous day to 1:00 a.m. on the following day. Furthermore, after the predetermined period of time has been stored in the auxiliary memory 49, when the various information processing units 57 estimate the psychological state of user 15 based on the period of time, they may prioritize estimating the level of the psychological state of user 15 estimated based on contact information over the psychological state of user 15 estimated based on voice information.

[0065] The artificial intelligence unit 59 may perform various processes in cooperation with the contact information processing unit 87, the response voice processing unit 88, and the location sense determination unit 89 by executing instructions of a non-temporary program 84 stored in the auxiliary memory 49.

[0066] The artificial intelligence unit 59 may perform learning and inference processes. In the learning process, for example, in the machine learning stage, the artificial intelligence unit 59 may obtain output data by processing the input data to be learned with the learning model 83A. The machine learning performed by the artificial intelligence unit 59 includes three types: supervised learning, unsupervised learning, and reinforcement learning.

[0067] In supervised learning, the learning model 83A is trained based on the error obtained by comparing the output data of the learning model 83A with the correct answer data. Specifically, the weights and biases are adjusted to minimize the error. Unsupervised learning is a method of training using training data that does not contain the correct answer, and it involves classifying the training data into groups of data with similar features (such as regularities). Reinforcement learning is a method of learning the learning model 83A through trial and error based on its output data, acquiring guidance on the optimal action to take in order to achieve the objective.

[0068] The artificial intelligence unit 59 may use a neural network as the learning model 83A for realizing machine learning. The neural network has an input layer, an intermediate layer (hidden layer), and an output layer. The artificial intelligence unit 59 may be an artificial intelligence (AI) equipped with transformers including GPT (registered trademark: Generative Pre-trained Transformer), BERT (Bidirectional Encoder Representations from Transformers), etc., and language models such as recurrent neural networks, and may perform processing as a generative artificial intelligence.

[0069] Furthermore, a portion of the learning model 83A that implements the processing of the artificial intelligence unit 59 may be composed of large language models (LLMs). Large language models are a type of generative artificial intelligence specialized in natural language processing (NLP).

[0070] Large-scale language models are language models that achieve advanced natural language generation (NLG) by being built through deep learning of vast amounts of text data. Natural language processing combines processes such as morphological analysis, syntactic analysis, semantic analysis, contextual analysis, and intent analysis to process natural language, enabling machine translation, text summarization, speech recognition, and more.

[0071] The large-scale language model may perform tasks such as parsing and analyzing text data contained in input prompts, parsing prompts and generating text data according to the analysis results, automatically generating prompts according to various types of information, and converting speech information into text.

[0072] A large-scale language model may train the learning model 83A by using a large amount of training data in the learning process and analyzing which word sequences have a high probability of occurrence. The training data consists of, for example, various information obtained from an external computer 52 and various information stored in auxiliary memory 49. The training data may also include text data, image information, sound information, book data, etc.

[0073] The learning model 83A may be configured as a multimodal neural network comprising a recurrent neural network that processes time-series data of contact information, or a network such as long-term and short-term memory, and a language model such as a transformer that processes text data of speech information, and then combines the output features of each and outputs the final sense of place level through a fully connected layer.

[0074] The response voices pre-stored in the auxiliary memory 49, or the response voices individually generated in response to the user 15's voice, are based on the user 15's sense of location score.

[0075] In this embodiment, the training data, including training target data, correct answer data, etc., for training the learning model 83A used by the artificial intelligence unit 59, may be acquired and stored in the auxiliary memory 49, for example, as follows. The training data may be created by giving a number of different questions (questionnaires) to people who constitute a specific group and obtaining their answers. The people who constitute a specific group may be a group in which all age groups, genders, living environments, etc. are the same, or a group in which all age groups, genders, living environments, etc. are different. Living environment may include items such as whether they are students or working adults, whether they live with family or alone, whether they run a company or work for a company, etc.

[0076] Furthermore, several different questions include words that correspond to feelings of affirmation or satisfaction, for example, "I feel like I can be my true self in this place." "I feel like I belong here." "I feel safe here." These may include, for example.

[0077] Furthermore, several different questions include words that correspond to negative or dissatisfied feelings in people, for example, "I don't think I can be my true self in this place." "I don't feel like I belong here." "I don't feel safe here." They are equivalent.

[0078] And in response to each question, "I don't think so at all." "I don't really think so." "I kind of think so," "I really think so." As shown above, multiple levels of answer examples may be provided. Furthermore, each level of answer examples may be pre-assigned a score (numerical value). The score (numerical value) assigned to each level of answer example will be different. For example, if the content of the question is "to include words corresponding to positive feelings or satisfaction," then the answer to the question will be assigned a higher score the more the degree of agreement or positive words is included, and a lower score the less the degree of agreement or positive words is included. Specifically, an answer of "strongly disagree" may be assigned a score of "1," an answer of "somewhat disagree" may be assigned a score of "2," an answer of "somewhat agree" may be assigned a score of "3," and an answer of "strongly agree" may be assigned a score of "5."

[0079] In contrast, if the question contains words that express negativity or dissatisfaction, the lower the degree of agreement or affirmation in the answer to the question, the higher the score will be. Specifically, an answer of "strongly disagree" may receive a score of 5, an answer of "somewhat disagree" may receive a score of 3, an answer of "somewhat agree" may receive a score of 2, and an answer of "strongly agree" may receive a score of 1.

[0080] Then, each person making up a specific group is given multiple questions, and their answers to each question are obtained. Furthermore, the sum of the points assigned to each answer, or the average of the points, is stored in database 82A as a numerical value indicating the "level of belonging" for each person, and is used to determine their sense of belonging. In addition, database 82A is linked to the level of belonging with the age group, gender, and living environment of the people making up the specific group, and this information is stored in auxiliary memory 49.

[0081] In the management computer 11, the artificial intelligence unit 59 may estimate the user's sense of location by processing the user's voice information with the learning model 83A.

[0082] The management computer 11 may then create a response voice that matches the estimated level of the user 15's sense of belonging. When the management computer 11 uses the voice of the user 15 obtained from the user terminal 12 to estimate the user 15's "level of belonging," it may match the user 15's age group, gender, and living environment with the age group, gender, and living environment of people who make up a specific group included in the database 82A, using two or more items.

[0083] The artificial intelligence unit 59 outputs an inference result by processing the data to be inferred with the learning model 83A during the inference process. In this embodiment, the artificial intelligence unit 59 may output contact information by processing the signal of the pressure sensor 22, which is an example of the data to be inferred, with the learning model. Alternatively, the artificial intelligence unit 59 may output a level indicating the user's sense of belonging by processing the contact information, which is an example of the data to be inferred, and text data indicating the user's emotions with the learning model 83A. Furthermore, the artificial intelligence unit 59 in this embodiment may output a response voice that matches the level of belonging.

[0084] Furthermore, the learning model 83A may be a pre-trained model constructed using a machine learning algorithm such as a neural network, a support vector machine (SVM), or a random forest. In the generation phase of the learning model 83A, the management computer 11 may take contact information collected from multiple test users, such as the signal strength of the pressure sensor 22, contact time, and the speed of movement of the contact position, and voice information, such as the pitch of the voices of users 15, the tone of the voice, and text data of the content of the utterances, as input data, and perform machine learning using training data in which the sense of belonging level based on the actual psychological state of the test users is used as ground truth data (labels).

[0085] Furthermore, at the same time as the collection of input data (for example, while the test user is interacting with object 13, or immediately after the interaction), several different questions (questionnaires) are administered to the test user, and the "level of belonging" calculated based on the answers is obtained as ground truth data (labels). The management computer 11 may use the pairs of input data and ground truth data obtained in this way as training data to learn the correlation between the input data and the psychological state of user 15, and perform machine learning.

[0086] Furthermore, the artificial intelligence unit 59 may, during the inference stage, input user 15's contact information and voice information into the learning model 83A and perform a first process that outputs a level indicating user 15's sense of place (for example, a score from 1 to 5) as an inference result. Furthermore, the artificial intelligence unit 59 may, during the inference stage, input user 15's level indicating their sense of place into the learning model 83A and perform a second process that outputs a response voice adapted to user 15's level indicating their sense of place as an inference result. The artificial intelligence unit 59 may also perform the first and second processes comprehensively.

[0087] Furthermore, the artificial intelligence unit 59 may perform machine learning on the learning model 83A so that it can determine the psychological state of user 15 by referring to logs corresponding to the user ID of user 15 that are the subject of inference during the inference stage. In addition, when the artificial intelligence unit 59 performs machine learning on the learning model 83A for determining the psychological state of user 15, it may autonomously correct or optimize the weighting for determining the psychological state of user 15 by time of day, based on the "specific tendencies" of user 15's psychological state.

[0088] For example, if the psychological state of user 15 estimated by processing the signal from the pressure sensor 22 differs from the psychological state of user 15 estimated by processing the user's voice during a predetermined time period, the artificial intelligence unit 59 may weight the emotions of user 15 estimated by processing the signal from the pressure sensor 22 over the psychological state of user 15 estimated by processing the user's voice, based on the "specific tendencies" of user 15's psychological state.

[0089] (First specific example of an information processing method) Figure 4 shows a first specific example of an information processing method performed by the information processing system 10. User 15 may obtain, for example, purchase, an item 80 with a QR code attached at a store in step S10. Item 80 may be, for example, a card. The QR code attached to item 80 represents the URL of a website that the management computer 11 provides to the network 16. The QR code may also be assigned a unique identifier. When user 15 reads the QR code of item 80 with the image acquisition device 70 of user terminal 12 in step S11, user terminal 12 accesses the management computer 11 via the URL. Once user terminal 12 accesses the management computer 11, the IP address of user terminal 12 is transmitted from user terminal 12 to the management computer 11.

[0090] In step S30, the management computer 11 sends the user screen to the accessed user terminal 12. In step S11, the user screen obtained from the management computer 11 is displayed on the user terminal 12's display 44. On the user screen, users can input user information such as name, gender, date of birth, telephone number, type of response voice, occupation, email address, and user account. User 15 can select one of the following response voice types: the response voice of user 15's family, a male response voice, a female response voice, an anime character's response voice, a celebrity's response voice, a sports player's response voice, etc.

[0091] Furthermore, if a response voice for user 15's family is selected, the input device 40 must be used to input the voice of user 15's family. User 15 inputs user information using the input device 40 and sends it to the management computer 11 in step S11. Also, if the voice of user 15's family is input to the user terminal 12, the voice of user 15's family is sent from the user terminal 12 to the management computer 11 in step S11.

[0092] When the management computer 11 obtains user information from the user terminal 12 in step S30, it may associate the user information with a QR code-specific identifier, perform user registration, and store it in the auxiliary memory 49. After performing user registration, the management computer 11 provides a program to the user terminal 12 in step S30.

[0093] In step S12, the user terminal 12 may download the program from the management computer 11 and install it into the auxiliary memory 39. After step S12, the user 15 may physically touch the object 13 in step S13, for example, with their hand or arm.

[0094] In step S40, the processor of the object may acquire the signal output from the pressure sensor 22 and transmit the pressure signal to the user terminal 12.

[0095] In step S14, the user terminal 12 may acquire the signal from the pressure sensor 22 from the object 13 and transmit the signal from the pressure sensor 22 to the management computer 11.

[0096] The management computer 31 may process the signal from the pressure sensor 22 in step S31 to determine contact information and estimate the user's sense of location based on the contact information. The management computer 31 may also create a response voice based on the estimated sense of location of the user 15. The management computer 11 may perform the following first or second process when creating the response voice.

[0097] In the first process, various information and a database 82A for determining the user 15's sense of location from contact information, and a database 82A of response voices corresponding to the user 15's sense of location may be pre-stored in the auxiliary memory 49. The management computer 11 selects a response voice that matches the user 15's sense of location. In the second process, the artificial intelligence unit 59 processes the contact information with a learning model 83A to determine the user 15's sense of location, and the artificial intelligence unit 59 may automatically generate a response voice corresponding to the user 15's sense of location.

[0098] The process by which the management computer 11 processes the signal from the pressure sensor 22 to determine contact information, and the process by which it estimates the user's sense of location based on the contact information, will be described later. The management computer 11 may also transmit the response voice obtained by executing the first or second process to the user terminal 12 in step S31.

[0099] The user terminal 12 may acquire a response voice from the management computer 11 in step S15 and output the response voice from the sound information output device 43.

[0100] User 15 may input their voice into user terminal 12 in step S16. User terminal 12 may transmit user 15's voice to management computer 11 in step S16.

[0101] The management computer 31 may process the user 15's voice in step S32 to estimate the user 15's sense of location. Alternatively, the management computer 31 may create a response voice based on the estimated sense of location of the user 15. When creating the response voice, the management computer 11 may perform the following third or fourth process.

[0102] In the third process, various information and a database 82A for determining user 15's sense of location from user 15's voice are stored in the auxiliary memory 49. A database 82A of response voices corresponding to user 15's sense of location may also be pre-stored in the auxiliary memory 49. The management computer 11 selects a response voice that matches user 15's sense of location.

[0103] In the fourth process, the artificial intelligence unit 59 processes the user 15's voice using the learning model 83A to determine the user 15's sense of location. The artificial intelligence unit 59 and the response voice processing unit 88 may then work together to generate a response voice corresponding to the user 15's level of location. The management computer 11 may transmit the response voice obtained by executing the third or fourth process to the user terminal 12 in step S32.

[0104] The user terminal 12 may acquire a response voice from the management computer 11 in step S17 and output the response voice from the sound information output device 43.

[0105] In the flowchart of Figure 4, the timing of the occurrence of the first process P1, which includes steps S13, S40, S14, S31, and S15, and the timing of the occurrence of the second process P2, which includes steps S16, S32, and S17, are not limited. For example, the timing of the occurrence of the first process P1 may be before the timing of the occurrence of the second process P2, or it may be after the timing of the occurrence of the second process P2. Furthermore, the timing of the occurrence of the first process P1 and the timing of the occurrence of the second process P2 may overlap at least partially.

[0106] Furthermore, the first process P1 may be repeated two or more times, and the second process P2 may be repeated two or more times. Also, the timing of the user 15 making contact with the object 13 in step S13 and the timing of the user 15 inputting voice into the user terminal 12 in step S16 may overlap at least partially. Moreover, either the first process P1 or the second process P2 may be executed by itself.

[0107] Furthermore, the user terminal 12 may acquire a facial image of user 15 using the camera included in the image acquisition device 70 in step S13, and transmit the facial image of user 15 to the management computer 11 in step S14. The management computer 11 may acquire the facial image of user 15 from the user terminal 12 in step S31, and the artificial intelligence unit 59 may process the contact information of user 15's contact with the object 13, as well as the facial image of user 15, using the learning model 83A to estimate user 15's sense of location.

[0108] Furthermore, in step S16, the user terminal 12 may acquire a facial image of user 15 using the camera included in the image acquisition device 70 and transmit the facial image of user 15 to the management computer 11. In step S32, the management computer 11 may acquire the facial image of user 15 from the user terminal 12, and the artificial intelligence unit 59 may process the voice of user 15 and the facial image of user 15 using the learning model 83A to estimate user 15's sense of location.

[0109] Furthermore, the management computer 11 may perform steps S31 and S32 simultaneously, that is, as a single first step. In other words, it may estimate the level of the user 15's sense of location based on one or more pieces of information from the user 15's contact information or voice information, and create a response voice corresponding to the level of the user 15's sense of location determined in the single first step, as a single first step. Furthermore, the management computer 11 may transmit the response voice corresponding to the level of the user 15's sense of location determined in the single first step to the user terminal 12 as a single first step.

[0110] Furthermore, the management computer 11 may execute the processes of steps S31 and S32 simultaneously in a single second step. That is, it may estimate the level of the user 15's sense of location based on one or more pieces of information from the user 15's contact information or voice information, and image information, and create a response voice corresponding to the level of the user 15's sense of location estimated in the single second step in a single second step. Furthermore, the management computer 11 may transmit the response voice created in the single second step to the user terminal 12 in a single second step.

[0111] (Example of a process where a management computer processes pressure sensor signals to determine contact information) The management computer 11 can process the signal from the pressure sensor 22 to estimate contact information. The contact information may include whether the user 15 grasped the object 13 with their hand, the duration of the grasp, whether the user 15 stroked (rubbed) the object 13 with their hand within a predetermined time, the number of strokes, the speed of the strokes, whether the user 15 struck the object 13, the number of times the object 13 was struck within a predetermined time, and so on.

[0112] The management computer 11 may assume that the user 15 is not in contact with the object 13, that is, is not grasping it, if the pressure estimated by processing the signal from the pressure sensor 22 is below a predetermined value. The management computer 11 may also assume that the user 15 has made contact with the object 13, that is, grasped it, if the pressure estimated by processing the signal from the pressure sensor 22 exceeds a predetermined value.

[0113] If the position of a pressure sensor 22 whose estimated pressure exceeds a predetermined value moves within a predetermined time, it may be estimated that the user 15 stroked the object 13 with their hand. Alternatively, the speed at which the user 15 stroked the object 13 with their hand may be estimated based on the time at which the position of a pressure sensor 22 whose estimated pressure exceeds a predetermined value moves within a predetermined time, and the time at which the pressure of each pressure sensor 22 exceeds a predetermined value.

[0114] If the estimated pressure alternates between being below a predetermined value and being above a predetermined value a predetermined number of times or more within a predetermined time, it may be presumed that the user 15 struck the object 13. Conversely, if the estimated pressure does not alternate between being below a predetermined value and being above a predetermined value a predetermined number of times or more within a predetermined time, it may be presumed that the user 15 did not strike the object 13. The number of times the object 13 was struck within the predetermined time may be estimated from the number of times the pressure alternated between being below a predetermined value and being above a predetermined number of times within the predetermined time.

[0115] (Example of a process that estimates a user's sense of location from contact information) In this embodiment, the management computer 11 may estimate the user 15's sense of place based on contact information. The management computer 11 may also represent the user 15's sense of place with a level corresponding to the level of place. This level may be represented by the relative magnitudes of integer values. For example, a level of place represented by "score 5" means that it is higher than a level of place represented by "score 1". The management computer 11's function of handling the user 15's sense of place with a level corresponding to the level of place is the same even when estimating the user 15's sense of place based on information other than contact information.

[0116] If the user 15 strokes the object 13 for a predetermined amount of time or longer, the management computer 11 may treat this as the user 15 experiencing a sense of security and safety, a bond and attachment with others and objects, a feeling of not being excluded (not being rejected), ease of self-disclosure and emotional expression, etc., and may assign a score of "5" to this sense of belonging.

[0117] Furthermore, if the user 15 holds onto the object 13 for a predetermined period of time or longer, the management computer 11 may treat this as the user 15 enduring anxiety and tension by holding onto the object 13, and may assign a score of "2" to the user's sense of belonging.

[0118] Furthermore, if user 15 is hitting object 13, the management computer 11 may treat this as user 15 experiencing frustration, anger, dissatisfaction, or isolation, and may assign a score of "1" to it as a sense of belonging.

[0119] (Example of a process that estimates a user's sense of location from their voice) If the management computer 11 finds that user 15's voice includes phrases such as, "When I pet you, I feel like telling you what I really think," "When I pet you, I feel calm," or "Today was a fulfilling day," it may treat this as user 15 experiencing a sense of security and safety, a bond and attachment with others and objects, a feeling of not being excluded (not being rejected), and ease of self-disclosure and emotional expression, and may assign a score of "5" to this sense of belonging.

[0120] If the management computer 11 detects that user 15's voice includes phrases such as, "I'm a little nervous, but maybe I'll be able to talk more once I get more used to it," or "I'll just try talking for now," it may treat this as a mixture of reassurance and anxiety and assign a score of "3" to indicate a sense of belonging.

[0121] If the management computer 11 detects that user 15's voice includes phrases such as, "I don't know what to say right now," or "I'm worried about how they'll react if I say something strange, so I don't really want to talk," it may treat this as a strong sense of alienation and isolation from those around them and assign a score of "1" to indicate a sense of belonging.

[0122] (Example of a process that estimates a user's sense of location based on their emotions inferred from image information) The management computer 11 may assign a score of "5" to the user's feeling of belonging if the emotion of the user 15, as estimated from the image information, is one of reassurance, joy, exhilaration, etc.

[0123] The management computer 11 may assign a score of "3" to the user's sense of belonging if the user's emotions, estimated from the image information, are calm, peaceful, normal, etc.

[0124] The management computer 11 may assign a score of "1" to the user's sense of belonging if the emotions of the user 15, estimated from the image information, are sadness, anger, anxiety, etc.

[0125] (Example of creating a response voice based on the user's sense of place score) The response voice generated by the management computer 11 may vary depending on the score indicating the user 15's sense of belonging. If the user 15's sense of belonging is scored "5" or "4", the corresponding response voice may be words that foster trust in the user 15, words that reflect on the relationship between the user 15 and the object 13, etc.

[0126] The words that foster trust in user 15 are: "I've become more able to clearly express my feelings, saying 'This is what I think,' than before." "Even though they are going through a tough time, their willingness to work together to figure out 'what to do' comes through." "It feels like you're trying to look at yourself as you are right now, including any doubts or uncertainties you might have." They are equivalent.

[0127] The words used to reflect on the relationship between user 15 and object 13 are: "Looking back on the consultations we've had so far, what meaning has this place held for you?" "If there's anything you'd like to change to make this a more comfortable place to talk, please let me know." "I'd like to make this a place where 'any version of yourself can be put aside at least once.' What do you think?" They are equivalent.

[0128] The response voice when User 15's sense of belonging is scored as "3" or "2" may include statements that confirm the relationship between User 15 and Object 13, respect User 15's pace or choices, carefully acknowledge User 15's changes or efforts, etc.

[0129] The response voice used to confirm the relationship between user 15 and object 13 is, for example, "Having met several times like this, what are your impressions of this place?" "After talking with me, was there even a moment when you felt, even just a little, 'Maybe I can talk here'?" "It seems there are moments of reassurance and moments of anxiety. I want to cherish both of those feelings together." They are equivalent.

[0130] Furthermore, language that respects the user's pace or choice is, "Please decide the order or content of your remarks at your own pace." "Today, let's just focus on thinking about this topic together." "If you feel like you want to end our discussion here for today, please let me know." They are equivalent.

[0131] Furthermore, the words used to carefully capture the changes or efforts of user 15 are: "I felt like you went into a little more detail than last time." "I take your honest story very seriously." They are equivalent.

[0132] When user 15's sense of belonging is rated "1," the type of response voice may include words of appreciation and welcome, words acknowledging their anxiety or tension, etc. Words of appreciation and welcome could be, for example, "I feel that the very fact that we are able to talk like this right now is a very important step." "This time is for you, so feel free to talk at your own pace." They are equivalent.

[0133] Furthermore, words that simply acknowledge anxiety and tension are, for example, "You're feeling anxious, aren't you? That's a very natural feeling." "You don't have to try to speak perfectly." "I hope we can discuss this together, including any doubts like, 'Is it okay to talk about this?'" "You don't have to talk about things you don't want to talk about." They are equivalent.

[0134] Figure 5 shows the data structure of the database 82A for estimating the user 15's psychological state, level, and response voice from contact information, audio information, and image information in this embodiment. Contact information includes types, i.e., contact patterns, such as "stroking," "grabbing," and "tapping." "Joy" may be estimated as the psychological state corresponding to "stroking." "Calmness" may be estimated as the psychological state corresponding to "grabbing." "Sadness" may be estimated as the psychological state corresponding to "tapping." The types of words included in the audio information include "happy," "no change or normal," and "painful." "Joy" may be estimated as the psychological state corresponding to "happy." "Calmness" may be estimated as the psychological state corresponding to "no change or normal." "Sadness" may be estimated as the psychological state corresponding to "painful."

[0135] The types of faces of user 15 included in the image information include "smiling," "neutral," and "crying." The psychological state corresponding to "smiling" may be estimated as "joy." The psychological state corresponding to "neutral" may be estimated as "peaceful." The psychological state corresponding to "crying" may be estimated as "sadness."

[0136] Furthermore, a "Level 5" corresponding to the psychological state of "joy" may be estimated, a "Level 3" corresponding to the psychological state of "calmness" may be estimated, and a "Level 1" corresponding to the psychological state of "sadness" may be estimated. Furthermore, a voice identifier "0005" for the response voice corresponding to "Level 5" may be estimated, a voice identifier "0003" for the response voice corresponding to the psychological state of "calmness" may be estimated, and a voice identifier "0001" for the response voice corresponding to the psychological state of "sadness" may be estimated.

[0137] Furthermore, when contact information, audio information, and image information are input as inference targets, the learning model 83A may output items such as type, psychological state, level, and response voice as inference results. The inference results output by the learning model 83A may be the same as, or different from, the data stored in the database 82A.

[0138] In this embodiment, if the level estimated from contact information differs from the level estimated from voice information, the processor 47 may determine the final level by prioritizing a predetermined priority, for example, prioritizing contact information. Alternatively, the processor 47 may calculate the average value of the levels estimated from contact information and the average value of the levels estimated from voice information, and treat the higher of the two average values ​​as the final level.

[0139] (Examples of application) User 15 may read the identification information from the identification information display unit 71 using the image acquisition device 70 of the user terminal 12 and transmit the identification information to the management computer 11. The management computer 11 may transmit a response voice corresponding to the identification information from the identification information display unit 71 to the user terminal 12. For example, the management computer 11 may transmit a response voice corresponding to a character represented by the identification information to the user terminal 12. Different characters will result in different response voices. The process by which the user terminal 12 transmits the identification information from the identification information display unit 71 to the management computer 11 may be performed in either step S14 or step S16 of Figure 4. The process by which the management computer 11 transmits a response voice corresponding to a character represented by the identification information to the user terminal 12 may be performed in either step S31 or step S32 of Figure 4.

[0140] The user terminal 12 may transmit the current location information of the user terminal 12, as determined by the current location determination unit 62, to the management computer 11. The management computer 11 may transmit a response voice to the user terminal 12 corresponding to the user terminal 12's current location.

[0141] If the current location of user terminal 12 is the home address of user 15 in the map data, the response voice will be: "Welcome back. You must be a little tired today. This is your home, so please relax and make yourself at home." "You did a great job today. Everything's alright now. You don't need to do anything here." They are equivalent.

[0142] If the current location of user terminal 12 is not the home address of user 15 in the map data, the response voice will be: "You were seeking solace while you were outdoors. It might seem a little hectic around you, but let's take this moment as a small respite." "You're working hard outside right now, aren't you? Why don't you take some time to rest your mind?" They are equivalent.

[0143] The process by which the user terminal 12 transmits its current location information to the management computer 11 may be performed in either step S14 or step S16 of Figure 4. The process by which the management computer 11 transmits a response voice corresponding to the user terminal 12's current location to the user terminal 12 may be performed in either step S31 or step S32 of Figure 4.

[0144] Furthermore, after the user terminal 12 outputs the response voice acquired from the management computer 11, the management computer 11 may process the voice of user 15 transmitted from the user terminal 12 to the management computer 11 in step S32, and the management computer 11 may estimate the user 15's satisfaction level. The management computer 11 may estimate that the user 15's satisfaction level is above a predetermined value (satisfied) if the voice of user 15 includes words such as "I'm happy," "I'm satisfied," "I'm grateful," or "That's helpful."

[0145] In response to this, the management computer 11 may estimate that user 15's satisfaction level is below a predetermined value (dissatisfied) if the user 15's voice contains words such as "I don't like it," "I'm not convinced," "I'm having trouble," "I want you to do something about it," or "I want you to fix it."

[0146] Furthermore, immediately before estimating that the user 15's satisfaction level is above a predetermined value (satisfied), the management computer 11 may store the response voice it sent to the user terminal 12 as a good response voice in the auxiliary memory 49, and train the learning model 83A using the good response voice.

[0147] (A method for estimating a user's sense of belonging from their facial expressions when using an object.) The management computer 11 may estimate the user's emotions from the user's facial image and assign a score for sense of belonging based on the estimated emotions of the user and the database 82A stored in the auxiliary memory 49. If the management computer 11 estimates that the user's emotions are secure or uplifted, it may assign a score of "5" for sense of belonging. If the management computer 11 estimates that the user's emotions are stable or neutral, it may assign a score of "3" for sense of belonging. If the management computer 11 estimates that the user's emotions are anger or sadness, it may assign a score of "1" for sense of belonging.

[0148] (Effects of this embodiment) In this embodiment, the management computer 11 estimates the user's sense of location level from contact information and voice information using a "decision base that defines multiple different levels of sense of location," and creates a response voice appropriate to the user's sense of location level. Therefore, when generating the response voice output to the user terminal 12, the performance and accuracy of generating the response voice for the user's sense of location can be improved.

[0149] Furthermore, conventional voice dialogue systems or response voice generation systems generated response voices based on the meaning of the user's utterances or a simple sentiment analysis. Therefore, it was difficult to accurately capture the user's complex and profound psychological state, such as a sense of belonging, feelings of alienation, or connection with others. Additionally, systems that relied solely on voice information faced the technical challenge of being unable to generate appropriate response voices when the user did not speak or when there was a discrepancy between their utterances and their actual psychological state.

[0150] In contrast, the management computer 11 of this embodiment comprehensively analyzes disparate data, such as the user 15's physical contact information and voice information with respect to the object 13, thereby improving the accuracy of estimating the user 15's level of "sense of being present." Furthermore, based on the estimated "level of the user 15's sense of being present," the management computer 11 of this embodiment can generate a response voice adapted to the user 15's current psychological state. Therefore, it is possible to improve the interaction technology between the user 15 and the management computer 11, thereby achieving a technical improvement in the human-machine interface that enables more natural and appropriate dialogue.

[0151] Furthermore, the decision-making base 85 used by the processor 47 to execute processing is stored in the auxiliary memory 49. Therefore, if the functionality of the processor 47 deteriorates, the deteriorated processor 47 can be removed from the communication bus 51 and replaced with a new processor 47, and the new processor 47 can use the decision-making base stored in the auxiliary memory 49 to estimate the user 15's sense of location.

[0152] Furthermore, the decision platform 85 indicates the user's sense of location, estimated from image information including the user's face image, at multiple different levels. Therefore, the processor 47 uses at least one piece of information, either contact information or voice information, along with the image information including the user's face image, to estimate the user's sense of location level. Consequently, the accuracy with which the processor 47 determines the user's sense of location level, and the performance and generation accuracy of creating response voices corresponding to the user's sense of location level, are further improved.

[0153] Furthermore, in this embodiment, the management computer 11 acquires contact information from the user terminal 12 in the first process and acquires voice information from the user terminal 12 in the second process. The management computer 11 also sends the generated voice information to the user terminal 12.

[0154] (Supplementary explanation 1) In step S31, the management computer 11 may determine the user 15's sense of belonging level using either the database 82A or the learning model 83A, or one or more of those. In step S32, the management computer 11 may also determine the user 15's sense of belonging level using either the database 82A or the learning model 83A, or one or more of those.

[0155] Furthermore, the output speed of the response voice may decrease as the user 15's sense of belonging increases. Additionally, the pitch range of the response voice may decrease as the user 15's sense of belonging increases. User 15's sense of belonging is one example of User 15's psychological state. Furthermore, a user's psychological state could also be represented by their stress level or fatigue level.

[0156] The management computer 11 may estimate that the user's stress level is higher the greater the force with which the user 15 grasps the object 13. The management computer 11 may also estimate that the user's stress level is higher the higher the tone of the user 15's voice. Furthermore, the management computer 11 may create a response voice that contains more words that promote relaxation for the user 15, or a response voice that promotes relaxation for a longer period of time, as the stress level increases.

[0157] The management computer 11 may estimate that the user's fatigue level is higher the greater the force with which the user 15 grasps the object 13. The management computer 11 may also estimate that the user's fatigue level is higher the higher the tone of the user 15's voice. Furthermore, the management computer 11 may create a response voice that contains more words that promote relaxation for the user 15, or a response voice that promotes relaxation for a longer period of time, as the fatigue level increases.

[0158] An example of the technical meaning disclosed in this embodiment is as follows: Management computer 11 is an example of an information processing device. Main memory 48 and auxiliary memory 49 are examples of memory. Processor 47 is an example of a processor. User terminal 12 is an example of a user terminal. Communication device 50 is an example of a communication device. Program 84 is an example of a program. Object 13 is an example of an object. Artificial intelligence unit 59 is an example of an artificial intelligence unit. Decision base 85 is an example of a decision base. Database 82A is an example of a response voice database and a reference database. Learning model 83A is an example of a response voice generation model and a learning model. Network 16 is an example of a network. User 15's sense of belonging, stress level, fatigue level, and concentration level are examples of the user 15's psychological state.

[0159] The processes performed by the management computer 11 in the flowchart of Figure 4 are examples of information processing methods. Step S31 in Figure 4 is an example of the first process. Step S32 is an example of the second process. Steps S31 and S32 are examples of the third and fourth processes, respectively. Steps S15 and S17 are examples of the fifth process, respectively.

[0160] The level indicating the user 15's sense of belonging may be represented by numbers, English letters, or Chinese characters. For example, it may be represented by Level A, Level B, Level C, and Level D. Level A is higher than Level B, Level B is higher than Level C, and Level C is higher than Level D. For example, it may be represented by Level A, Level B, Level C, and Level D. Level A is higher than Level B, Level B is higher than Level C, and Level C is higher than Level D. In this embodiment, the predetermined time may be any one of the following, for example, 30 seconds, 1 minute, 2 minutes, 3 minutes, etc., and the predetermined time is stored in the auxiliary memory 49 in advance.

[0161] (Second specific example of an information processing method) Furthermore, a second specific example of the information processing method performed by the information processing system 10 is shown in Figure 6. In Figure 6, the same steps as those performed in Figure 4 are assigned the same step numbers. In step S18, which is after step S13, the user terminal 12 may acquire the signal from the pressure sensor 22 and the user voice, and transmit the signal from the pressure sensor 22 and the user voice to the management computer 11. The process performed in step S18 includes the process performed in step S14 and the process performed in step S16 in Figure 4.

[0162] The management computer 11 may process the signal from the pressure sensor 22 acquired from the user terminal 12 in step S33, which is after step S30, to estimate the psychological state of the user 15. Alternatively, the management computer 11 may process the user voice acquired from the user terminal 12 and estimate the psychological state of the user 15 in step S33.

[0163] The management computer 11 may determine in step S34 whether the psychological state of user 15 estimated by processing the signal from the pressure sensor 22 is different from the psychological state of user 15 estimated by processing the user's voice. For example, if the psychological state estimated from the contact information shown in Figure 5 is different from the psychological state estimated from the voice information, the management computer 11 may determine in step S34 to be Yes. Conversely, if the psychological state estimated from the contact information shown in Figure 5 is the same from the psychological state estimated from the voice information, the management computer 11 may determine in step S34 to be No.

[0164] If the management computer 11 determines Yes in step S34, it may store a log in the auxiliary memory 49 in step S35 that includes the date, time, and user ID at which it determined that the emotion of user 15 estimated by processing the signal from the pressure sensor 22 is different from the emotion of user 15 estimated by processing the user's voice. Alternatively, the management computer 11 may perform machine learning on the learning model 83A in step S35 so that, during the inference stage performed by the learning model 83A, it refers to the log corresponding to the user ID of the user 15 being inferred to determine the emotion of user 15.

[0165] In step S36, following step S35, the management computer 11 may determine the user 15's sense of place and level of place, and create a response voice based on the user 15's level of place. Alternatively, in step S36, the management computer 11 may transmit the response voice to the user terminal 12. If the management computer 11 has executed the processing in step S35 and proceeded to step S36, it may determine the user 15's sense of place and level of place by prioritizing the user 15's psychological state estimated by processing the signal from the pressure sensor 22 over the user 15's psychological state estimated by processing the user voice.

[0166] If the management computer 11 determines No in step S34, it may skip the processing in step S35 and execute the processing in step S36. If the management computer 11 skips the processing in step S35 and proceeds to step S36, it may determine the user's sense of belonging and level of belonging based on the psychological state of the user 15 estimated by processing the signal from the pressure sensor 22 and the psychological state of the user 15 estimated by processing the user's voice.

[0167] If the management computer 11 proceeds from step S33 to step S35 and stores a predetermined time period in the auxiliary memory 49, and then again determines the user 15's psychological state within the predetermined time period in step S33, it may skip steps S34 and S35 as shown by the dashed line in Figure 6 and proceed to step S36, and determine the user 15's sense of place and level of place by prioritizing the user 15's psychological state estimated based on contact information over the user 15's psychological state estimated based on voice information.

[0168] Furthermore, if the management computer 11 proceeds from step S33 to step S35 and stores the predetermined time period in the auxiliary memory 49, and then determines in step S33 that the user's psychological state is outside of the predetermined time period, it may proceed to step S34.

[0169] (Other effects of this embodiment) The management computer 11 may determine whether the psychological state of user 15 estimated based on contact information differs from the psychological state of user 15 estimated based on voice information, as shown in the information processing method in Figure 6. If the management computer 11 determines that the psychological state of user 15 estimated based on contact information differs from the psychological state of user 15 estimated based on voice information, it can prioritize the psychological state of user 15 estimated based on contact information over the psychological state of user 15 estimated based on voice information to determine user 15's sense of belonging and level of belonging. Therefore, the accuracy of the management computer 11's judgment is improved.

[0170] Furthermore, the management computer 11 may store in the auxiliary memory 49 the time periods in which it determined that the psychological state of user 15 estimated based on contact information differs from the psychological state of user 15 estimated based on voice information. After the management computer 11 stores in the auxiliary memory 49 the time periods in which it determined that the psychological state of user 15 estimated based on contact information differs from the psychological state of user 15 estimated based on voice information, if the management computer 11 estimates the psychological state of user 15 based on time periods, it may prioritize the psychological state of user 15 estimated based on contact information over the psychological state of user 15 estimated based on voice information when estimating the sense of belonging and the level of belonging.

[0171] Therefore, even if user 15 intentionally conceals their emotions through voice during a given time period, their sense of belonging and level of belonging can be estimated with higher accuracy by reflecting their past behavioral characteristics (personal habits).

[0172] (Supplementary explanation 2) The process performed by the management computer 11 in the flowchart of Figure 6 is an example of an information processing method. Step S33 in Figure 6 is an example of the first and second processes. Steps S33, S34, S35, and S36 are examples of the third process. Step S36 is an example of the fourth process. Steps S15 and S17 are examples of the fifth process, respectively.

[0173] (Supplementary explanation 3) Furthermore, the management computer 11 may consist of a single computer or may be distributed across multiple computers. The management computer 11 includes at least one computer from among servers, supercomputers, mainframes, workstations, etc.

[0174] Furthermore, the program 84 disclosed in this embodiment may be understood as a program product. Moreover, a program product may be realized using the program 84 and the storage medium 49A. [Industrial applicability]

[0175] This embodiment can be used as an information processing device, information processing method, program, and storage medium for processing information about a user's interaction with an object. [Explanation of Symbols]

[0176] 11...Management computer, 12...User terminal, 13...Object, 16...Network, 47...Processor, 48...Main memory, 49...Auxiliary memory, 49A...Storage medium, 50...Communication device, 59...Artificial intelligence unit, 82A...Database, 83A...Learning model, 84...Program, 85...Decision-making platform

Claims

1. An information processing apparatus having a memory and a processor coupled to the memory, The processor executes the program instructions stored in the memory, A first process to acquire contact information indicating the user's physical contact state with the object, A second process to acquire voice information of the user using the object, A database that associates the aforementioned contact information and the aforementioned voice information with multiple different psychological states and the level of said psychological state. And, A learning model that takes the aforementioned contact information and the aforementioned voice information as input and outputs the level of the user's psychological state. A third process that estimates the level of the user's psychological state using a judgment basis that includes, A fourth process involves creating a response voice that corresponds to the level of the user's psychological state estimated in the third process and that can be output by the user terminal used by the user, using at least one of the response voice database or response voice generation model stored in the memory. An information processing device that performs this task.

2. An information processing apparatus according to claim 1, The information processing apparatus, wherein the fourth process performed by the processor includes the process of selecting a response voice corresponding to the estimated level of psychological state from a response voice database that associates the level of psychological state with the response voice.

3. An information processing apparatus according to claim 1, The information processing apparatus includes a fourth process performed by the processor, which involves outputting a response voice corresponding to the estimated level of the psychological state using a response voice generation model that generates a response voice with the level of the psychological state as input.

4. An information processing apparatus according to claim 1, The aforementioned decision-making basis displays the user's psychological state, estimated from image information including the user's facial image, at multiple different levels. The third process performed by the processor is an information processing device that estimates the level of the user's psychological state using at least one piece of information, either contact information or voice information, and image information, including a facial image of the user.

5. An information processing apparatus according to claim 1, A communication device is provided that is connected to the user terminal via a network. The first process performed by the processor acquires the contact information from the user terminal, The second process performed by the processor acquires the voice information from the user terminal, The processor is an information processing device that performs a fifth process to send the response voice generated in the fourth process to the user terminal.

6. An information processing apparatus according to claim 1, The third process performed by the aforementioned processor is: A process for estimating the user's psychological state based on the aforementioned contact information, A process for estimating the user's psychological state based on the aforementioned audio information, A process for determining whether the psychological state of the user estimated based on the contact information differs from the psychological state of the user estimated based on the voice information, Includes, The fourth process performed by the aforementioned processor is: If the third process determines that the psychological state of the user estimated based on the contact information is different from the psychological state of the user estimated based on the voice information, the information processing device creates the response voice prioritizing the psychological state of the user estimated based on the contact information over the psychological state of the user estimated based on the voice information.

7. The information processing apparatus according to claim 6, The third process performed by the aforementioned processor is: A process to store in memory the time period during which it is determined that the user's psychological state estimated based on the contact information differs from the user's psychological state estimated based on the voice information. An information processing device in which, after the aforementioned time period has been stored in the memory, the processor estimates the user's psychological state during that time period, prioritizing the estimation of the user's psychological state based on the contact information over the estimation of the user's psychological state based on the voice information.

8. A computer having memory and a processor coupled to the memory, the information processing method performed by the processor, The processor executes the program instructions stored in the memory, A first process to acquire contact information indicating the user's physical contact state with the object, A second process to acquire voice information of the user using the object, A database that associates the aforementioned contact information and the aforementioned voice information with multiple different psychological states and the level of said psychological state. And, A learning model that takes the aforementioned contact information and the aforementioned voice information as input and outputs the level of the user's psychological state. A third process that estimates the level of the user's psychological state using a judgment basis that includes, A fourth process involves creating a response voice that corresponds to the level of the user's psychological state estimated in the third process and that can be output by the user terminal used by the user, using at least one of the response voice database or response voice generation model stored in the memory. An information processing method that performs [this action].

9. A non-temporary program that is stored in memory and, when read and executed by a processor connected to the memory, allows the processor to perform processing, The processor executes the program instructions stored in the memory, A first process to acquire contact information indicating the user's physical contact state with the object, A second process to acquire voice information of the user using the object, A database that associates the aforementioned contact information and the aforementioned voice information with multiple different psychological states and the level of said psychological state. And, A learning model that takes the aforementioned contact information and the aforementioned voice information as input and outputs the level of the user's psychological state. A third process that estimates the level of the user's psychological state using a judgment basis that includes, A fourth process involves creating a response voice that corresponds to the level of the user's psychological state estimated in the third process and that can be output by the user terminal used by the user, using at least one of the response voice database or response voice generation model stored in the memory. A program that executes something.

10. A non-temporary storage medium that stores a non-temporary program for a processor to perform processing, The processor executes the instructions of the program, A first process to acquire contact information indicating the user's physical contact state with the object, A second process to acquire voice information of the user using the object, A database that associates the aforementioned contact information and the aforementioned voice information with multiple different psychological states and the level of said psychological state. And, A learning model that takes the aforementioned contact information and the aforementioned voice information as input and outputs the level of the user's psychological state. A third process that estimates the level of the user's psychological state using a judgment basis that includes, A fourth process involves creating a response voice that corresponds to the level of the user's psychological state estimated in the third process and that can be output by the user's terminal, using at least one of the response voice database or response voice generation model stored in the storage medium. A storage medium that performs this operation.

Citation Information

Patent Citations

  • Method of producing vibrationnproof cell

    JP1979098938A

  • Electronic toy and method for controlling same and storage medium

    JP2001327765A

  • Voice synthesizer, pseudofeeling expressing device and voice synthesizing method

    JP2002049385A

  • Communication system, communication device, program and communication control method

    JP2013202080A

  • System

    JP2025057611A